Semantic-driven node engine voice interaction method and system
By monitoring and identifying the noise characteristics and emotional information of the wake-up voice through IoT devices, combined with noise filtering processing, voice interaction is optimized, solving the problem that the user's emotional state is not taken into account in existing technologies, and improving the accuracy and response speed of the interaction.
Patent Information
- Application Number
- CN202511007745.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-22
AI Technical Summary
Existing device voice interaction technology does not fully consider the user's emotional state, resulting in inaccurate interaction processing and delayed responses, affecting the user experience.
Through IoT devices, the wake-up voice is monitored, noise characteristics and wake-up emotion information are identified, and corrections are made. In addition, noise filtering is performed in combination with noise characteristics to optimize interactive semantic recognition and device control.
It improves the accuracy and response speed of voice interaction, optimizes the user experience, and adapts to different environments and user emotional states.
Smart Images

Figure CN120808778A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a semantic-driven node engine voice interaction method and system. BACKGROUND
[0002] In the current device voice interaction technology, the device mainly relies on speech recognition and natural language processing to parse user instructions and perform corresponding operations. However, the prior art generally does not fully consider the emotional state of the user, resulting in inaccurate interaction processing, response delay and other problems in the recognition and response process, thereby affecting the user experience. Therefore, there is a technical problem of low voice interaction recognition accuracy, slow response, and poor user experience in the prior art. SUMMARY
[0003] The present application provides a semantic-driven node engine voice interaction method and system to solve the technical problem of lack of efficient real-time data acquisition and intelligent analysis means in the traditional heat exchange system operation fault analysis method, and the technical problem of insufficient timeliness and accuracy in fault analysis.
[0004] The technical solution of the present application to solve the above technical problems is as follows: In a first aspect, the present application provides a semantic-driven node engine voice interaction method, comprising: obtaining a wake-up voice issued by a user through an Internet of Things device, performing semantic recognition, and obtaining noise feature information and wake-up emotion information when a wake-up node is triggered, and correcting to obtain corrected wake-up emotion information, wherein the node engine further includes a collection node, a processing node and a control node; Collecting the interactive voice of the user, performing noise filtering processing according to the noise feature information, performing interactive semantic recognition, obtaining a recognition result and a recognition repetition rate; According to the corrected wake-up emotion information and the repeated recognition rate, generating a device control result and feedback information according to the recognition result, and performing feedback interaction and device control.
[0005] In a second aspect, the present application provides a semantic-driven node engine voice interaction system, comprising: a wake-up emotion recognition module for obtaining a wake-up voice issued by a user through an Internet of Things device, performing semantic recognition, and obtaining noise feature information and wake-up emotion information when a wake-up node is triggered, and correcting to obtain corrected wake-up emotion information, wherein the node engine further includes a collection node, a processing node and a control node; An interactive semantic recognition module for collecting the interactive voice of the user, performing noise filtering processing according to the noise feature information, performing interactive semantic recognition, obtaining a recognition result and a recognition repetition rate; The control feedback interaction module is used for generating device control results and feedback information according to the recognition results, performing feedback interaction and device control according to the corrected wake-up emotion information and the repeated recognition rate.
[0006] The present application has the advantages that the present application can effectively improve the accuracy, response speed and user experience of voice interaction. The present application can monitor and obtain the wake-up voice of the user through the Internet of Things device, identify noise characteristic information and wake-up emotion information in combination with semantic recognition when triggering the wake-up node, and correct them, so as to identify the accurate emotion information of the user and adjust the interaction strategy according to the user emotion state to optimize the interaction experience. In the voice interaction process, the user interaction voice is obtained through the collection node, and noise filtering is performed based on the noise characteristic information, so as to reduce the interference of environmental noise on voice recognition and improve the accuracy of interaction semantic recognition. The control node generates the corresponding control feedback analysis strategy according to the corrected wake-up emotion information and the repeated recognition rate, and analyzes and generates the corresponding device control results and feedback information, which can adapt to the user emotion state and optimize the response strategy, so as to reduce unnecessary interaction delay, improve the accuracy of interaction control, and further achieve the technical effects of improving the intelligent degree of interaction and user satisfaction. BRIEF DESCRIPTION OF DRAWINGS
[0007] Figure 1 A flowchart of a semantic-driven node engine voice interaction method provided by the present application is shown in the figure. Figure 2 A structural diagram of a semantic-driven node engine voice interaction system provided by the present application is shown in the figure.
[0008] In the figure, the components represented by the respective numbers are described as follows: The wake-up emotion recognition module 11, the interaction semantic recognition module 12, and the control feedback interaction module 13. DETAILED DESCRIPTION
[0009] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0010] In the description of the present application, the terms "first" and "second" are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.
[0011] In the description of the present application, the term "for example" is used to mean "serving as an example, instance, or illustration." Any embodiment described as "for example" in this patent is not necessarily to be construed as preferred or advantageous over other embodiments. The following description is presented to enable any person skilled in the art to make and use the application. In the following description, for purposes of explanation, specific details are set forth to provide a thorough understanding of the present application. It will be apparent to one skilled in the art, however, that the present application can be practiced without using these specific details. In other instances, well-known structures and processes are not elaborated in order not to obscure the description of the present application with unnecessary details. Thus, the present application is not intended to be limited by the embodiments shown, but is to be accorded with the widest scope consistent with the principles and features disclosed.
[0012] Embodiment one, as shown in the present application, provides a semantic-driven node engine voice interaction method, which specifically includes the following steps: Figure 1 S10: Acquire the wake-up voice issued by the user through the Internet of Things device monitoring, perform semantic recognition, identify the noise feature information and wake-up emotion information when triggering the wake-up node, and correct the corrected wake-up emotion information, wherein the node engine further includes a collection node, a processing node and a control node. S10: Acquire the wake-up voice issued by the user through the Internet of Things device monitoring, perform semantic recognition, identify the noise feature information and wake-up emotion information when triggering the wake-up node, and correct the corrected wake-up emotion information, wherein the node engine further includes a collection node, a processing node and a control node.
[0013] In the embodiments of the present application, the wake-up voice issued by the user is monitored and acquired through the Internet of Things device, and semantic recognition is performed thereon to determine whether to trigger the wake-up node.
[0014] At the same time of triggering the wake-up node, the noise feature information and the wake-up emotion information are further identified, and the wake-up emotion information is corrected to acquire the current emotion state of the user as the data basis for subsequent device interaction control, so as to improve the interaction accuracy under different environments and user emotion states.
[0015] Among them, the node engine includes a wake-up node, a collection node, a processing node and a control node to process the data of different steps in the voice interaction device control, realize voice interaction, the wake-up node is used to wake up the device through the wake-up voice, the collection node is used to collect the interactive voice of the user, the processing node is used to identify the semantics in the interactive voice, and the control node is used to control the device according to the semantics and feedback the user, so as to realize continuous Internet of Things device interaction control.
[0016] The step S10 in the method provided by the embodiments of the present application includes: Acquire the wake-up voice issued by the user through the Internet of Things device monitoring, input the semantic recognizer, and obtain the recognized semantics; determining whether the recognized semantic is a wake-up, if yes, triggering a wake-up node, if no, not triggering the wake-up node, and the Internet of Things device continues standby; When the wake-up node is triggered, noise feature recognition and emotion recognition are performed on the wake-up voice to obtain noise feature information and wake-up emotion information.
[0017] In the embodiments of the present application, the Internet of Things device, such as a smart speaker, a smart phone, smart furniture (such as a smart lamp), etc., is configured with a microphone array to listen to the user's voice input in real time to perform device interaction control through voice recognition, such as controlling the switch of the smart lamp.
[0018] The microphone array in the Internet of Things device listens to the user's voice input in real time as a wake-up voice, inputs the semantic recognizer, obtains the recognized semantic, and further determines whether the recognized semantic is a wake-up.
[0019] Exemplarily, the semantic recognizer is constructed based on natural language processing (NLP) and deep learning models (such as BERT, Transformer) in the prior art to perform semantic recognition, analyze the input voice, and identify the core semantic. For example, the Internet of Things device monitors and obtains the wake-up voice "hello" issued by the user, and the semantic recognizer identifies the recognized semantic as "hello". The wake-up voice can be sent to the cloud for recognition through the Internet of Things, and then the recognized semantic obtained by recognition is transmitted back.
[0020] Further, whether the recognized semantic is a wake-up is determined, specifically whether it is consistent with a preset wake-up word. There can be multiple preset wake-up words, such as "hello", "hello", "hello", etc. If the recognized semantic is consistent with any one of the preset wake-up words, it is determined that the recognized semantic is a wake-up, the wake-up node is triggered, and the Internet of Things device switches from the standby state to the wake-up state. If the recognized semantic is not consistent with any one of the preset wake-up words, it is determined that the recognized semantic is not a wake-up, which can be only a conversation between the user and others, the wake-up node is not triggered, and the Internet of Things device continues standby and continues to monitor the wake-up voice issued by the user.
[0021] When the wake-up node is triggered, further noise feature recognition and emotion recognition are performed on the wake-up voice to obtain noise feature information and wake-up emotion information, to analyze the user's emotion and the noise in the environment, correct the user's emotion, and serve as a noise filtering reference for subsequent recognition of the user's voice.
[0022] The step S10 in the method provided in the embodiments of the present application further includes: According to the wake-up voice data in the historical time, a sample wake-up voice set is collected, and noise feature extraction and wake-up emotion level identification are performed on each sample wake-up voice to obtain a sample noise feature information set and a sample wake-up emotion information set. The sample wake-up voice set, the sample noise feature information set and the sample wake-up emotion information set are used as input features and output features, and a wake-up voice recognizer is trained based on deep learning. The wake-up voice is input into the wake-up voice recognizer, and noise feature information and wake-up emotion information are obtained.
[0023] In the embodiments of the present application, voice samples are extracted from wake-up voice data in historical time and converted into spectrum graphs to form a sample wake-up voice set. These samples include wake-up voices of different users in different environments, such as a quiet room (30 dB or less), an office (50-60 dB), a car (70-80 dB) and a noisy outdoor environment (85 dB or more). Through diversified data collection, the wake-up voice recognizer is ensured to have generalization ability for different scenes.
[0024] Further, noise feature extraction and wake-up emotion level identification are performed on each sample wake-up voice. Specifically, noise frequency, noise decibel value, signal-to-noise ratio and the like in the sample wake-up voice are extracted as noise feature information, and the user's emotional state when issuing the wake-up noise is obtained through questionnaire survey and the like, for example, divided into calm (0 level), slight excitement (1 level), anxious (2 level), angry (3 level) and the like, as sample wake-up emotion information. In this way, the sample noise feature information set and the sample wake-up emotion information set are collected and identified.
[0025] Further, deep learning is used, and the sample wake-up voice set, the sample noise feature information set and the sample wake-up emotion information set are used as input features and output features to train the wake-up voice recognizer to identify the environmental noise features and user emotion levels in the wake-up voice.
[0026] Exemplarily, a convolutional neural network is used to construct and train the wake-up voice recognizer, which includes an input layer, a convolutional layer, a fully connected layer and an output layer. The convolutional layer has two layers, each layer containing 64 3x3 convolutional kernels, using a ReLU activation function, and the fully connected layer includes 256 neurons. In the training process, the sample wake-up voice is input, the output environmental noise features and user emotion levels are obtained, the error with the corresponding sample noise feature information and sample wake-up emotion information is calculated as a loss, the network parameters are optimized using Adam to reduce the loss until the requirement is met, for example, less than 1%. The iterative training is performed, for example, 50 rounds of iterative training, and then tested. If the test loss meets the requirements, the training is completed, otherwise the iterative training is continued until the training is completed.
[0027] Based on the trained wake-up voice recognizer, the spectrum graph of the current collected wake-up voice is input, and corresponding noise feature information and wake-up emotion information are obtained by identification output. The wake-up emotion information includes the user's emotion level.
[0028] The embodiment of the present application trains the wake-up voice recognizer by analyzing historical voice data to improve the accuracy and adaptability of voice interaction in complex environments and different user emotional states.
[0029] The step S10 in the method provided by the embodiment of the present application further includes: Extracting the noise decibel in the noise feature information; According to the ratio of the noise decibel to the average noise decibel, the emotion level in the wake-up emotion information is corrected and calculated to obtain a corrected wake-up emotion level as the corrected wake-up emotion information.
[0030] In the embodiment of the present application, the environmental noise will further affect the user's emotion. Therefore, in order to improve the adaptability of subsequent voice interaction control and user emotion, after obtaining the noise feature information and the wake-up emotion information, the wake-up emotion information is corrected according to the noise feature information to improve the accuracy of user emotion analysis.
[0031] Specifically, the noise decibel in the noise feature information is extracted, and the greater the noise decibel, the greater the impact on the user's emotion. For example, the noise decibel is 55 dB.
[0032] Further, the average noise decibel of the previously identified user wake-up voice is obtained, which reflects the average noise decibel size in the environment of the Internet of Things device.
[0033] Further, the ratio of the current noise decibel to the average noise decibel is calculated, and the emotion level in the wake-up emotion information is corrected and calculated, for example, the ratio is multiplied by the emotion level and rounded to obtain a corrected wake-up emotion level, which is further used as the corrected wake-up emotion information. For example, the average noise decibel is 40 dB, and the emotion level in the wake-up emotion information is 2, then the corrected wake-up emotion level = 55 dB / 40 dB*2 = 2.75. The integer part is 3, and the corrected wake-up emotion level is 3, which is used as the corrected wake-up emotion information.
[0034] In this way, the greater the current noise decibel, the greater the noise impact on the corrected wake-up emotion information, the greater the impact of noise on the user's emotion, and the more negative the emotion. In this way, the corrected wake-up emotion information is obtained.
[0035] The embodiment of the present application improves the accuracy and adaptability of voice interaction emotion analysis by noise feature analysis and emotion correction calculation, and can more intelligently adapt to different environments and user states, thereby optimizing user experience in subsequent voice interaction processing.
[0036] S20: collect the interactive voice of the user, perform noise filtering processing according to the noise feature information, perform interactive semantic recognition, and obtain a recognition result and a recognition repetition rate.
[0037] In the embodiments of the present application, after the wake-up node of the Internet of Things device triggers wake-up, the device performs interactive voice collection and recognition processing of the collection node and the processing node, collects the interactive voice of the user for voice interaction, and then performs interactive semantic recognition to obtain the intention of the user for voice interaction control of the device and subsequent device control.
[0038] After the interactive voice is recognized, a recognition result and a recognition repetition rate of the recognition result are obtained. The greater the recognition repetition rate is, the greater the accuracy of the current recognition and control is, which serves as a data basis for subsequent generation of device control results.
[0039] The step S20 in the method provided by the embodiments of the present application includes: collecting the interactive voice of the user, performing noise filtering processing according to the noise feature information, and obtaining filtered interactive voice; performing interactive semantic recognition on the filtered interactive voice to obtain a recognition result; calculating the ratio of the number of the same recognition results to the number of all recognition results in a historical time to obtain a recognition repetition rate.
[0040] In the embodiments of the present application, after the Internet of Things device is woken up, the collection node is entered, the interactive voice is collected, and the interactive voice is filtered based on the noise feature information obtained by the wake-up node recognition. Since the wake-up voice and the interactive voice are collected in the same environment, the noise features in the environment are also the same. Therefore, noise filtering is performed to improve the accuracy of voice recognition.
[0041] For example, the interactive voice is filtered based on the noise frequency and the signal-to-noise ratio in the noise feature information, for example, noise filtering processing is used for filtering, and filtered interactive voice after filtering out noise is obtained.
[0042] Further, the filtered interactive voice after filtering is subjected to interactive semantic recognition. For example, the filtered interactive voice is recognized by using natural language processing (NLP) in the prior art to obtain a recognition result. The recognition accuracy can be improved by performing recognition after filtering according to the noise feature information.
[0043] For example, the recognition result can be “play music”, “turn off the light”, “turn on the light”, etc., which is related to the type of the Internet of Things device.
[0044] Further, a ratio of the same recognition result quantity of the recognition result to a total recognition result quantity in a historical time is calculated to obtain a recognition repetition rate, to reflect a proportion of repeated appearance of the current recognition result, and the greater the ratio, the greater the accuracy of the recognition result.
[0045] For example, the Internet of Things device has performed 1000 times of interactive voice recognition in a historical time, the total recognition result quantity is 1000, and the same recognition result "playing music" appears 620 times, so the recognition repetition rate is 620 / 1000=62%. Among them, the user may have different recognition results for the same interactive control purpose of the Internet of Things device, for example, "playing music", "randomly playing a song", and the like, but the recognition result appearing more times is more in line with the user's habits and has a higher accuracy.
[0046] The recognition repetition rate is obtained by calculation, which can reflect the accuracy of the current recognition result, and then serve as a data basis for subsequent device control feedback, to improve the adaptability of device interactive control feedback.
[0047] S30: generating device control results and feedback information according to the recognition result according to the corrected wake-up emotion information and the repetition recognition rate, and performing feedback interaction and device control.
[0048] In the embodiments of the present application, after the collection node and the processing node, the Internet of Things device enters the control node to perform device control and feedback according to the semantic recognition result of the user, to complete the complete voice interactive control. Specifically, according to the corrected wake-up emotion information and the repetition recognition rate, the device control results and the feedback information are generated according to the recognition result. Among them, the computing power configuration for generating the device control results and the feedback information is considered according to the corrected wake-up emotion information of the user and the repetition recognition rate of the recognition result, to improve the accuracy and timeliness of the device control feedback, to fit the current user emotion and the recognition result, and to further improve the user experience.
[0049] The method provided in the embodiments of the present application includes the following steps S30: An integrated learning is used to train an interactive control feedback channel including M interactive control feedback paths, and M is a positive integer; According to the corrected wake-up emotion information and the repetition recognition rate, the emotion recognition quantity and the accuracy recognition quantity are calculated in combination with M, and a composite recognition quantity N is calculated, and N is a positive integer less than or equal to M; Randomly selecting N interactive control feedback paths, inputting the recognition result, outputting N path device control results and path feedback information, and screening the device control results and the feedback information with the highest appearance proportion to perform feedback interaction and device control.
[0050] In the voice interaction, the Internet of Things device needs to generate reasonable device control results and feedback information according to the voice instruction recognition result of the user in the embodiments of the present application. In order to improve the accuracy of device control and feedback, integrated learning is used to train M interactive control feedback paths to form an interactive control feedback channel. M is a positive integer, for example, 10.
[0051] The method provided in the embodiments of the present application includes the step of "using integrated learning to train an interactive control feedback channel including M interactive control feedback paths". According to the voice interaction processing data in the historical time, a sample recognition result set is collected, and the correct device control results and feedback information corresponding to different sample recognition results are collected and labeled as a sample device control result set and a sample feedback information set. The sample recognition result set, the sample device control result set, and the sample feedback information set are divided to obtain M groups of interactive control feedback training data. M interactive control feedback paths are trained based on machine learning using the M groups of interactive control feedback training data respectively to obtain an interactive control feedback channel.
[0052] In the embodiments of the present application, in order to improve the accuracy of device control and user experience, multiple interactive control feedback paths need to be trained through historical interaction data to form an interactive control feedback channel.
[0053] Specifically, according to the voice interaction processing data in the historical time, the recognition result after interactive voice semantic recognition is collected to obtain a sample recognition result set, for example, sample recognition results such as "turn on the light" and "play music". Then the correct device control results and feedback information corresponding to different sample recognition results are collected, for example, the device control results of turning on the light or playing music for the Internet of Things device, and the corresponding voice feedback information such as "the light has been turned on" and "playing your favorite music", as a sample device control result set and a sample feedback information set. Based on different Internet of Things devices, Table 1 shows part of the sample recognition result set, the sample device control result set, and the sample feedback information set.
[0054] Sample recognition result Sample device control result Sample feedback information "Turn on the light" Turn on the light "Light turned on" "Turn up the volume" Volume +10% "Volume increased by 10%" "Turn off the air conditioner" Turn off the air conditioner "Air conditioner turned off" "Play music" Start playing music "Started playing music" "Turn down the volume" Volume -10% "Volume decreased by 10%" Table 1 Further, the sample recognition result set, the sample device control result set, and the sample feedback information set are divided, and each time a certain percentage is divided with replacement, for example, 60% of the data is randomly divided, and the division is performed M times to obtain M groups of interactive control feedback training data.
[0055] Then M interactive control feedback paths are trained based on machine learning using the M groups of interactive control feedback training data respectively to obtain an interactive control feedback channel.
[0056] The training step of the single interaction control feedback path is taken as an example for illustration. For example, based on a feedforward neural network, an interaction control feedback path is constructed, which includes an input layer, a hidden layer, and an output layer, the input feature of the input layer is the recognition result, and the output feature of the output layer is the device control result and the feedback information. In the training process, the sample recognition result set is input, the output device control result set and the feedback information set are obtained, it is judged whether they are consistent with the corresponding device control result set and the sample feedback information set, the proportion of consistency is calculated, the accuracy rate is obtained, it is judged whether it meets the requirement, for example, the accuracy rate is greater than 95%, if not, the difference from the accuracy rate is calculated as the loss, the Adam optimization is used to adjust the network parameters of the interaction control feedback path, so that the loss is reduced and the accuracy rate is improved, until the accuracy rate meets the requirement, the training is completed. Based on the same way, the training of M interaction control feedback paths is completed, and the training processes of the M interaction control feedback paths are the same, and the training data are different.
[0057] After the training is completed, the M interaction control feedback paths are combined to obtain an interaction control feedback channel. Since the training data of the M interaction control feedback paths are different, the M interaction control feedback paths have different performances. By integrating the outputs of the M interaction control feedback paths, the accuracy of the device control feedback can be improved.
[0058] In the embodiments of the present application, according to the current analyzed user's correction wake-up emotion information and the repeated recognition rate of the recognition result, the emotion recognition quantity and the accuracy recognition quantity are calculated and obtained in combination with M, and the composite recognition quantity N is calculated, N is a positive integer less than or equal to M.
[0059] Among them, the greater the correction emotion level in the user's correction wake-up emotion information, the more negative the user's emotion, the more accurate the device control result and the feedback information are needed to avoid the user's emotion deterioration and affect the user experience, the more interaction control feedback paths are needed for device control and feedback information generation to improve the accuracy. In addition, the greater the repeated recognition rate, the more accurate the current recognition result, and the more the same recognition result device control and interaction feedback have been performed, the smaller the probability of generating an error device control result and feedback information, and the faster the control feedback response speed is needed to improve the user's use experience, and the fewer the interaction control feedback paths are needed for device control and feedback information generation to reduce the data processing amount and improve the control feedback efficiency.
[0060] The method provided in the embodiments of the present application includes the step of "according to the correction wake-up emotion information and the repeated recognition rate, combining M to calculate and obtain the emotion recognition quantity and the accuracy recognition quantity, and calculating the composite recognition quantity N". obtaining the maximum emotion level; calculating a ratio of the corrected wake-up emotion level in the corrected wake-up emotion information to the maximum emotion level, multiplying M and taking an integer to obtain an emotion recognition number; According to the repeated recognition rate, a control feedback recognition coefficient is calculated; The control feedback recognition coefficient is multiplied by M and taken as an integer to obtain an accuracy recognition number; The average of the emotion recognition number and the accuracy recognition number is calculated to obtain a composite recognition number N.
[0061] In the embodiments of the present application, the maximum emotion level, i.e. the strongest negative emotion level of the user, is first obtained, for example, 3.
[0062] Further, the ratio of the corrected wake-up emotion level in the corrected wake-up emotion information of the current user to the maximum emotion level is calculated, and then multiplied by M and taken as an integer to obtain an emotion recognition number. For example, the corrected wake-up emotion level in the corrected wake-up emotion information of the current user is 2, the maximum emotion level is 3, and M is 10, then the emotion recognition number is 2 / 3*10 and taken as an integer to obtain 7. In this way, the larger the corrected wake-up emotion level is, the larger the emotion recognition number is, and the more accurate the generation of the device control result and the feedback information is, the probability of error is reduced, the user's emotion is further affected, and the user experience is improved.
[0063] Further, according to the repeated recognition rate, a control feedback recognition coefficient is calculated. For example, the control feedback recognition coefficient is calculated by subtracting the repeated recognition rate from 1. For example, the repeated recognition rate is 62%, and the control feedback recognition coefficient is 38%.
[0064] Further, the control feedback recognition coefficient is multiplied by M and taken as an integer to obtain an accuracy recognition number. For example, M is 10, and the accuracy recognition number is 10*38% and taken as an integer to obtain 4.
[0065] Wherein, the larger the repeated recognition rate is, the greater the probability of generating accurate device control results and feedback information is, and fewer interactive control feedback paths can also achieve accurate generation of device control results and feedback information, therefore, the number of calls of the interactive control feedback path is reduced, the larger the repeated recognition rate is, the smaller the accuracy recognition number is, the control feedback timeliness and efficiency are improved, and the user experience is further improved.
[0066] Further, the average of the emotion recognition number and the accuracy recognition number is calculated to obtain a composite recognition number N. For example, the emotion recognition number is 7, and the accuracy recognition number is 4, then the composite recognition number N is the average of the two and taken as an integer to obtain 6.
[0067] Therefore, the embodiment of the application comprehensively considers the user's emotion and the accuracy of the current speech interaction recognition result, and makes a calculation decision on the number of interactive control feedback path invocations in the device control feedback generation analysis, more interactive control feedback paths are invoked when the negative level of the user's emotion is higher, the accuracy is improved, and the user experience is improved, fewer interactive control feedback paths are invoked when the accuracy of the recognition result is higher, the processing efficiency is improved, and the user experience is improved, and the accuracy and timeliness of the control feedback can be comprehensively ensured in combination with the two dimensions.
[0068] In the embodiment of the application, N interactive control feedback paths are randomly selected in the interactive control feedback channel according to the number N of composite recognitions, the current recognition result is input, N path device control results and path feedback information respectively output are obtained, and then the device control result and the feedback information with the highest occurrence proportion (the most number of occurrences) are screened out as the final device control result and the feedback information for feedback interaction and device control.
[0069] The semantic-driven node engine speech interaction method provided by the embodiment of the application has at least the following technical effects: The embodiment of the application can effectively improve the accuracy, response speed and user experience of speech interaction. The embodiment of the application monitors and acquires the wake-up voice issued by the user through the Internet of Things device, combines semantic recognition, identifies noise characteristic information and wake-up emotion information when triggering the wake-up node, and corrects them, so as to identify the accurate emotion information of the user, adjust the interaction strategy according to the user emotion state, and optimize the interaction experience. In the speech interaction process, the user interaction voice is acquired through the acquisition node, and noise filtering processing is performed based on the noise characteristic information, so as to reduce the interference of environmental noise on speech recognition and improve the accuracy of interaction semantic recognition. The control node generates the strategy of corresponding control feedback analysis according to the corrected wake-up emotion information and the recognition repetition rate, and analyzes and generates corresponding device control results and feedback information, which can adapt to the user emotion state and optimize the response strategy, thereby reducing unnecessary interaction delay, improving the accuracy of interaction control, and further achieving the technical effects of improving the intelligent degree of interaction and user satisfaction.
[0070] Embodiment two, as Figure 2 shown, based on the same inventive concept of the semantic-driven node engine speech interaction method provided in embodiment one, the embodiment of the application further provides a semantic-driven node engine speech interaction system, which comprises: The wake-up emotion recognition module 11 is configured to monitor and acquire the wake-up voice issued by the user through the Internet of Things device, perform semantic recognition, identify and obtain noise characteristic information and wake-up emotion information when triggering the wake-up node, and correct and obtain corrected wake-up emotion information, wherein the node engine further comprises an acquisition node, a processing node and a control node. The interactive semantic recognition module 12 is configured to collect the interactive voice of the user, perform noise filtering processing according to the noise feature information, perform interactive semantic recognition, and obtain a recognition result and a recognition repetition rate. The control feedback interaction module 13 is configured to generate a device control result and feedback information according to the recognition result, and perform feedback interaction and device control according to the correction wake-up emotion information and the recognition repetition rate.
[0071] Further, the semantic-driven node engine voice interaction system is further configured to: acquire a wake-up voice issued by a user through an Internet of Things device, input the semantic recognizer, and obtain a recognized semantic. Determine whether the recognized semantic is a wake-up, if yes, trigger a wake-up node, and if not, do not trigger the wake-up node, and the Internet of Things device continues to standby. When the wake-up node is triggered, perform noise feature recognition and emotion recognition on the wake-up voice, and obtain noise feature information and wake-up emotion information.
[0072] Further, the semantic-driven node engine voice interaction system is further configured to: acquire a sample wake-up voice set according to wake-up voice data in a historical time, and perform noise feature extraction and wake-up emotion level identification on each sample wake-up voice, to obtain a sample noise feature information set and a sample wake-up emotion information set. Use the sample wake-up voice set, the sample noise feature information set and the sample wake-up emotion information set as input features and output features, train a wake-up voice recognizer based on deep learning. Input the wake-up voice into the wake-up voice recognizer, and output the noise feature information and the wake-up emotion information.
[0073] Further, the semantic-driven node engine voice interaction system is further configured to: extract a noise decibel in the noise feature information. According to a ratio of the noise decibel to an average noise decibel, correct and calculate an emotion level in the wake-up emotion information, to obtain a corrected wake-up emotion level as the corrected wake-up emotion information.
[0074] Further, the semantic-driven node engine voice interaction system is further configured to: collect an interactive voice of a user, perform noise filtering processing according to the noise feature information, and obtain a noise-filtered interactive voice. Perform interactive semantic recognition on the noise-filtered interactive voice, and obtain a recognition result. Calculate a ratio of a same recognition result quantity of the recognition result to a total recognition result quantity in a historical time, to obtain a recognition repetition rate.
[0075] Further, the semantic-driven node engine voice interaction system is further used for: training an interaction control feedback channel including M interaction control feedback paths by using ensemble learning, M being a positive integer; According to the corrected wake-up emotion information and the repeated recognition rate, the number of emotion recognitions and the number of accuracy recognitions are calculated in combination with M, and a composite recognition number N is calculated, N being a positive integer less than or equal to M; N interaction control feedback paths are randomly selected, the recognition result is input, N path device control results and path feedback information are output, the device control result and the feedback information with the highest occurrence ratio are screened, and feedback interaction and device control are performed.
[0076] Further, the semantic-driven node engine voice interaction system is further used for: collecting a sample recognition result set according to voice interaction processing data in a historical time, and collecting correct device control results and feedback information corresponding to different sample recognition results, and labeling as a sample device control result set and a sample feedback information set; The sample recognition result set, the sample device control result set and the sample feedback information set are divided to obtain M sets of interaction control feedback training data; M interaction control feedback paths are trained based on machine learning by using the M sets of interaction control feedback training data respectively, and an interaction control feedback channel is obtained.
[0077] Further, the semantic-driven node engine voice interaction system is further used for: Obtaining a maximum emotion level; Calculating the ratio of the corrected wake-up emotion level in the corrected wake-up emotion information to the maximum emotion level, multiplying M and taking an integer to obtain a number of emotion recognitions; According to the repeated recognition rate, a control feedback recognition coefficient is calculated; The control feedback recognition coefficient is multiplied by M and an integer is taken to obtain a number of accuracy recognitions; The mean of the number of emotion recognitions and the number of accuracy recognitions is calculated to obtain a composite recognition number N.
[0078] Although the preferred embodiments of the present application have been described, those skilled in the art can make further changes and modifications to these embodiments once they know the basic inventive concept.
[0079] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the present application and its equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A semantic-driven node engine voice interaction method, characterized in that: The method comprises: The user's wake-up voice is monitored by IoT devices and semantic recognition is performed. When the wake-up node is triggered, noise feature information and wake-up emotion information are identified and corrected to obtain the corrected wake-up emotion information. The node engine also includes an acquisition node, a processing node, and a control node. Collecting the user's interactive speech, performing noise filtering according to the noise feature information, performing interactive semantic recognition, and obtaining a recognition result and a recognition repetition rate; According to the corrected awakening emotion information and the repeated recognition rate, device control results and feedback information are generated according to the recognition results, and feedback interaction and device control are performed.
2. The semantic-driven node engine voice interaction method according to claim 1, characterized in that: The IoT device monitors the user's wake-up voice and performs semantic recognition. When the wake-up node is triggered, noise feature information and wake-up emotion information are identified, including: The user's wake-up voice is acquired through IoT device monitoring, input into the semantic recognizer, and the recognition semantics are obtained; Determine whether the recognition semantics is wake-up, if so, trigger the wake-up node; if not, do not trigger the wake-up node, and the IoT device continues to standby; When the wake-up node is triggered, noise feature recognition and emotion recognition are performed on the wake-up speech to obtain noise feature information and wake-up emotion information.
3. The semantic-driven node engine voice interaction method according to claim 1, characterized in that: When the wake-up node is triggered, noise feature recognition and emotion recognition are performed on the wake-up speech, including: Based on the historical wake-up speech data, a sample wake-up speech set is collected, and noise features are extracted and the wake-up emotion level is marked for each sample wake-up speech to obtain a sample noise feature information set and a sample wake-up emotion information set; Using the sample wake-up speech set, the sample noise feature information set, and the sample wake-up emotion information set as input features and output features, and training a wake-up speech recognizer based on deep learning; The wake-up speech is input into the wake-up speech recognizer, and noise feature information and wake-up emotion information are obtained by output.
4. The semantic-driven node engine voice interaction method according to claim 1, characterized in that: Correction obtains correction arousal emotional information, including: Extracting noise decibels from the noise characteristic information; According to the ratio of the noise decibel to the average noise decibel, a correction calculation is performed on the emotion level in the awakening emotion information to obtain a corrected awakening emotion level as the corrected awakening emotion information.
5. The semantic-driven node engine voice interaction method according to claim 1, characterized in that: Collecting the user's interactive speech, performing noise filtering according to the noise feature information, performing interactive semantic recognition, and obtaining recognition results and recognition repetition rates, including: Collecting the user's interactive voice, performing noise filtering according to the noise characteristic information, and obtaining the noise-filtered interactive voice; Performing interactive semantic recognition on the noise-filtered interactive speech to obtain a recognition result; The ratio of the number of identical recognition results of the recognition results to the number of all recognition results in the historical time is calculated to obtain the recognition repetition rate.
6. The semantic-driven node engine voice interaction method according to claim 1, characterized in that: According to the corrected awakening emotion information and the repeated recognition rate, generating a device control result and feedback information according to the recognition result, and performing feedback interaction and device control, including: Ensemble learning is used to train an interactive control feedback channel including M interactive control feedback paths, where M is a positive integer; According to the corrected awakening emotion information and the repeated recognition rate, combined with M, the number of emotion recognitions and the number of accurate recognitions are calculated, and the number of composite recognitions N is calculated, where N is a positive integer less than or equal to M; Randomly select N interactive control feedback paths, input the identification results, output N path device control results and path feedback information, screen and obtain the device control results and feedback information with the highest occurrence ratio, and perform feedback interaction and device control.
7. The semantic-driven node engine voice interaction method according to claim 1, characterized in that: Ensemble learning is used to train an interactive control feedback channel comprising M interactive control feedback paths, including: Based on the historical voice interaction processing data, a sample recognition result set is collected, and the correct device control results and feedback information corresponding to different sample recognition results are collected, and marked as a sample device control result set and a sample feedback information set; Dividing the sample recognition result set, the sample device control result set, and the sample feedback information set to obtain M groups of interactive control feedback training data; The M groups of interactive control feedback training data are respectively used to train M interactive control feedback paths based on machine learning to obtain interactive control feedback channels.
8. The semantic-driven node engine voice interaction method according to claim 7, characterized in that: According to the corrected awakening emotion information and the repeated recognition rate, combined with M, the number of emotion recognitions and the number of accurate recognitions are calculated, and the number of composite recognitions N is calculated as follows: Get the maximum emotion level; Calculating a ratio of the corrected arousal emotion level in the corrected arousal emotion information to the maximum emotion level, multiplying the ratio by M and rounding the ratio to obtain the number of emotion recognitions; Calculating a control feedback recognition coefficient based on the repeated recognition rate; The control feedback recognition coefficient is multiplied by M and rounded to obtain the accuracy recognition number; The average of the emotion recognition number and the accuracy recognition number is calculated to obtain the composite recognition number N.
9. A semantically driven node engine voice interaction system, characterized in that: The steps for implementing the semantic-driven node engine voice interaction method according to any one of claims 1 to 8 include: The wake-up emotion recognition module is used to monitor and obtain the wake-up speech issued by the user through the Internet of Things device, perform semantic recognition, identify and obtain noise feature information and wake-up emotion information when the wake-up node is triggered, and correct and obtain the corrected wake-up emotion information. The node engine also includes an acquisition node, a processing node, and a control node; An interactive semantic recognition module is used to collect the user's interactive speech, perform noise filtering according to the noise feature information, perform interactive semantic recognition, and obtain recognition results and recognition repetition rate; The control feedback interaction module is used to generate device control results and feedback information according to the correction awakening emotion information and the repeated recognition rate, and perform feedback interaction and device control according to the recognition result.
Citation Information
Patent Citations
Highly anthropomorphic voice interaction algorithm and emotion interaction algorithm for robot and robot
CN109741746A
Voice interaction device and method and computer readable storage medium
CN110827821A
Voice emotion recognition method and device
CN116259307A
Virtual interaction method and system based on voice recognition
CN117995222A
Emotion understanding and feedback method for full-duplex interactive question and answer digital human
CN119311119A