A semantic-driven node engine voice interaction method and system

By monitoring and identifying noise characteristics and emotional information in wake-up voices through IoT devices, and combining deep learning and ensemble learning, voice interaction processing is optimized. This solves the problem that the user's emotional state is not taken into account in device voice interaction, and improves the accuracy of interaction and user experience.

CN120808778BActive Publication Date: 2026-03-17JIANGSU SHANXIN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing voice interaction technologies do not fully consider the user's emotional state, resulting in inaccurate interaction processing, response delays, and negatively impacting the user experience.

Method used

By monitoring wake-up voice through IoT devices, identifying noise characteristics and emotional information during wake-up, and making corrections, the wake-up voice recognizer is trained using deep learning. Interactive voice is collected and noise is filtered, and device control results and feedback information are generated using ensemble learning.

Benefits of technology

It improves the accuracy and response speed of voice interaction, optimizes the user experience, reduces environmental noise interference, adapts to the user's emotional state, and reduces interaction delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808778B_ABST
    Figure CN120808778B_ABST
Patent Text Reader

Abstract

The application relates to a semantic-driven node engine voice interaction method and system, and relates to the technical field of data processing.The method comprises the following steps: obtaining a wake-up voice issued by a user through Internet of Things equipment monitoring, performing semantic recognition, obtaining noise characteristic information and wake-up emotion information when a wake-up node is triggered, and correcting the wake-up emotion information to obtain corrected wake-up emotion information, wherein the node engine further comprises a collection node, a processing node and a control node; collecting interactive voice of the user, performing noise filtering according to the noise characteristic information, performing interactive semantic recognition, obtaining a recognition result and a recognition repetition rate; according to the corrected wake-up emotion information and the recognition repetition rate, generating device control results and feedback information according to the recognition result, and performing feedback interaction and device control.The application considers user emotions and semantic analysis accuracy to perform device control feedback, and achieves the technical effects of improving interaction accuracy, timeliness and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a semantically driven node engine voice interaction method and system. Background Technology

[0002] Current device voice interaction technologies primarily rely on speech recognition and natural language processing to parse user commands and execute corresponding operations. However, existing technologies generally do not adequately consider the user's emotional state, leading to inaccurate interaction processing and response delays during recognition and response, thus impacting user experience. Therefore, existing technologies suffer from low accuracy in voice interaction recognition, slow response times, and poor user experience. Summary of the Invention

[0003] This invention addresses the technical problem of traditional heat exchange system operation fault analysis methods lacking efficient real-time data acquisition and intelligent analysis means, resulting in insufficient timeliness and accuracy of fault analysis. It provides a semantic-driven node engine voice interaction method and system to solve this problem.

[0004] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:

[0005] In a first aspect, the present invention provides a semantically driven node engine voice interaction method, comprising: monitoring and acquiring the wake-up voice issued by the user through an Internet of Things device, performing semantic recognition, and when the wake-up node is triggered, identifying and acquiring noise feature information and wake-up emotion information, and correcting and acquiring corrected wake-up emotion information, wherein the node engine further comprises a collection node, a processing node and a control node.

[0006] Collect user's interactive voice, perform noise filtering according to the noise feature information, perform interactive semantic recognition, and obtain recognition results and recognition repetition rate;

[0007] Based on the corrected arousal emotion information and repetition rate, device control results and feedback information are generated according to the recognition results, and feedback interaction and device control are performed.

[0008] Secondly, the present invention provides a semantically driven node engine voice interaction system, including: a wake-up emotion recognition module, used to monitor and acquire the wake-up voice issued by the user through an Internet of Things device, perform semantic recognition, and when the wake-up node is triggered, identify and obtain noise feature information and wake-up emotion information, and correct and obtain corrected wake-up emotion information. The node engine also includes a collection node, a processing node and a control node.

[0009] The interactive semantic recognition module is used to collect the user's interactive voice, perform noise filtering according to the noise feature information, perform interactive semantic recognition, and obtain the recognition result and recognition repetition rate.

[0010] The control feedback interaction module is used to generate device control results and feedback information based on the recognition results according to the corrected arousal emotion information and repetition recognition rate, and to perform feedback interaction and device control.

[0011] The beneficial effects of this invention are: it effectively improves the accuracy, response speed, and user experience of voice interaction. This invention monitors and acquires the user's wake-up voice through IoT devices, and combines this with semantic recognition. Simultaneously with triggering the wake-up node, it identifies noise feature information and wake-up emotion information, and performs corrections. This allows for accurate identification of the user's emotional information, and the interaction strategy is adjusted according to the user's emotional state, optimizing the interactive experience. During voice interaction, the user's interactive voice is acquired through a data acquisition node, and noise filtering is performed based on noise feature information, thereby reducing the interference of environmental noise on voice recognition and improving the accuracy of interactive semantic recognition. The control node generates corresponding control feedback analysis strategies based on the corrected wake-up emotion information and recognition repetition rate, and analyzes and generates corresponding device control results and feedback information. This adapts to the user's emotional state and optimizes the response strategy, thereby reducing unnecessary interaction delays, improving the accuracy of interactive control, and ultimately achieving the technical effect of improving the intelligence level of interaction and user satisfaction. Attached Figure Description

[0012] Figure 1 A flowchart illustrating a semantically driven node engine voice interaction method provided by the present invention;

[0013] Figure 2 This is a schematic diagram of the structure of a semantically driven node engine voice interaction system provided by the present invention.

[0014] The components represented by each number in the attached diagram are explained below:

[0015] The system includes an emotion recognition module 11, an interaction semantic recognition module 12, and a control feedback interaction module 13. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0018] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0019] Example 1, as Figure 1 As shown, this embodiment of the invention provides a semantically driven node engine voice interaction method, which specifically includes the following steps:

[0020] S10: The system monitors and acquires the wake-up voice issued by the user through IoT devices, performs semantic recognition, and identifies and obtains noise feature information and wake-up emotion information when the wake-up node is triggered, and corrects and obtains corrected wake-up emotion information. The node engine also includes acquisition nodes, processing nodes and control nodes.

[0021] In this embodiment of the application, the wake-up voice issued by the user is monitored and acquired through an Internet of Things device, and semantic recognition is performed on it to determine whether to trigger the wake-up node.

[0022] While triggering the wake-up node, noise feature information and wake-up emotion information are further identified, and the wake-up emotion information is corrected to obtain the user's current emotional state as the data basis for subsequent device interaction control, so as to improve the accuracy of interaction in different environments and user emotional states.

[0023] The node engine includes wake-up nodes, acquisition nodes, processing nodes, and control nodes to perform data processing for different steps in the control of voice interaction devices, thereby enabling voice interaction. The wake-up node is used to wake up the device by using voice commands, the acquisition node is used to collect the user's interactive voice, the processing node is used to identify the semantics within the interactive voice, and the control node is used to control the device based on the semantics and provide feedback to the user, thus realizing continuous interactive control of IoT devices.

[0024] Step S10 in the method provided in this application embodiment includes:

[0025] The system monitors and acquires the user's wake-up voice through IoT devices, inputs it into a semantic recognizer, and obtains the recognized semantics.

[0026] Determine whether the identified semantic is wake-up. If yes, trigger the wake-up node; otherwise, do not trigger the wake-up node and the IoT device remains in standby mode.

[0027] When the wake-up node is triggered, noise feature recognition and emotion recognition are performed on the wake-up voice to obtain noise feature information and wake-up emotion information.

[0028] In this embodiment of the application, IoT devices such as smart speakers, smartphones, and smart furniture (such as smart lights) are equipped with microphone arrays to listen to the user's voice input in real time for voice recognition-based device interaction control, such as controlling the on / off switch of smart lights.

[0029] The device uses a microphone array within the IoT device to monitor the user's voice input in real time. This voice input is then used as a wake-up command and input into a semantic recognizer to obtain the recognized semantics. The system then determines whether the recognized semantics indicate that the user is ready to wake up.

[0030] For example, the semantic recognizer is built based on existing technologies such as Natural Language Processing (NLP) and deep learning models (e.g., BERT, Transformer) to perform semantic recognition, parse the input speech, and identify the core semantics. For instance, if an IoT device detects that a user's wake-up voice is "Hello," the semantic recognizer will identify its semantic meaning as "Hello." The wake-up voice can be sent to the cloud via the IoT for recognition, and then the obtained semantic meaning will be transmitted back.

[0031] Further, it is determined whether the recognized semantics is a wake-up. Specifically, it is determined whether it is consistent with a preset wake-up word. There can be multiple preset wake-up words, such as "Hello", "你好", "你好啊", etc. If the recognized semantics is consistent with any one of the preset wake-up words, it is determined that the recognized semantics is a wake-up, triggering a wake-up node, and the IoT device switches from the standby state to the wake-up state. If the recognized semantics is not consistent with any one of the preset wake-up words, it is determined that the recognized semantics is not a wake-up, which may be only a conversation between the user and others, and the wake-up node is not triggered. The IoT device continues to standby and continues to monitor the wake-up voice issued by the user.

[0032] When the wake-up node is triggered, noise feature recognition and emotion recognition are further performed on the wake-up voice to obtain noise feature information and wake-up emotion information, so as to analyze the user's emotion and the noise in the environment, perform correction of the user's emotion, and use it as a noise filtering reference for subsequent recognition of the user's voice.

[0033] Step S10 in the method provided by the embodiment of the present application further includes:

[0034] According to the wake-up voice data within the historical time, a sample wake-up voice set is collected, and noise feature extraction and wake-up emotion level identification are performed on each sample wake-up voice to obtain a sample noise feature information set and a sample wake-up emotion information set;

[0035] Using the sample wake-up voice set, the sample noise feature information set and the sample wake-up emotion information set as input features and output features, a wake-up voice recognizer is trained based on deep learning;

[0036] Input the wake-up voice into the wake-up voice recognizer, and output to obtain noise feature information and wake-up emotion information.

[0037] In the embodiment of the present application, voice samples are extracted from the wake-up voice data within the historical time and converted into spectrograms to form a sample wake-up voice set. These samples include wake-up voices of different users in different environments, such as a quiet room (below 30dB), an office (50 - 60dB), a car (70 - 80dB), and an outdoor noisy environment (above 85dB). Through diversified data collection, it is ensured that the wake-up voice recognizer has the generalization ability for different scenarios.

[0038] Furthermore, noise features and emotional level labels are extracted for each sample wake-up speech. Specifically, the noise frequency, decibel value, and signal-to-noise ratio within the sample wake-up speech are extracted as noise feature information. Additionally, the user's emotional state when emitting the wake-up noise is obtained through methods such as questionnaires, categorized into multiple emotional levels such as calm (level 0), slightly excited (level 1), anxious (level 2), and angry (level 3), serving as sample wake-up emotional information. In this way, the collection and labeling yield a set of sample noise feature information and a set of sample wake-up emotional information.

[0039] Furthermore, deep learning is used to train a wake-up speech recognizer by employing a set of sample wake-up speech, a set of sample noise features, and a set of sample wake-up emotion information as input and output features, in order to recognize the environmental noise features and the user's emotion level within the wake-up speech.

[0040] For example, a convolutional neural network is used to construct and train a wake-up speech recognition system. This system includes an input layer, convolutional layers, fully connected layers, and an output layer. The convolutional layers consist of two layers, each containing 64 3×3 convolutional kernels using the ReLU activation function. The fully connected layers contain 256 neurons. During training, sample wake-up speech is input, and the environmental noise features and user emotion level of the output are obtained. The error between the output and the corresponding sample noise feature information and sample wake-up emotion information is calculated as the loss. The network parameters are optimized using Adam to reduce the loss until a requirement is met, such as less than 1%. This process is iteratively trained, for example, 50 rounds. Then, a test is performed. If the test loss meets the requirement, training is complete; otherwise, iterative training continues until training is complete.

[0041] Based on the trained wake-up speech recognizer, the spectrogram of the currently collected wake-up speech is input, and the corresponding noise feature information and wake-up emotion information are obtained from the recognition output. The wake-up emotion information includes the user's emotion level.

[0042] This invention analyzes historical voice data to train a wake-up voice recognizer, thereby improving the accuracy and adaptability of voice interaction in complex environments and under different user emotional states.

[0043] Step S10 in the method provided in this application embodiment further includes:

[0044] Extract the noise decibels from the noise feature information;

[0045] Based on the ratio of the noise decibels to the average noise decibels, the emotional level within the arousal emotional information is corrected and calculated to obtain the corrected arousal emotional level, which is used as the corrected arousal emotional information.

[0046] In this embodiment, environmental noise can further affect the user's emotions. Therefore, to improve the adaptability of subsequent voice interaction control to the user's emotions, after identifying and obtaining noise feature information and arousal emotion information, the arousal emotion information is corrected based on the noise feature information to improve the accuracy of user emotion analysis.

[0047] Specifically, the noise level (in decibels) is extracted from the noise feature information. The higher the noise level (in decibels), the greater the impact on the user's mood. For example, the noise level (in decibels) is 55 dB.

[0048] Furthermore, the average noise level in decibels of the previously identified user wake-up voice is obtained, which reflects the average noise level in the environment in which the IoT device is located.

[0049] Furthermore, the ratio of the current noise level to the average noise level is calculated, and the emotional level within the arousal emotional information is corrected accordingly. For example, this ratio is multiplied by the emotional level and rounded to obtain the corrected arousal emotional level, which is then used as the corrected arousal emotional information. For instance, if the average noise level is 40 dB and the emotional level within the arousal emotional information is level 2, then the corrected arousal emotional level = 55 dB / 40 dB * 2 = 2.75. Rounded to 3, the corrected arousal emotional level is 3, which is used as the corrected arousal emotional information.

[0050] Thus, the higher the current noise level, the greater the impact of noise on the corrected arousal emotional information, the more the user's emotions are affected by noise, and the more negative the emotions become. In this way, the correction obtains the corrected arousal emotional information.

[0051] This application embodiment effectively improves the accuracy and adaptability of voice interaction emotion analysis through noise feature analysis and emotion correction calculation, enabling it to more intelligently adapt to different environments and user states, thereby optimizing the user experience in subsequent voice interaction processing.

[0052] S20: Collect the user's interactive voice, perform noise filtering according to the noise feature information, perform interactive semantic recognition, and obtain the recognition result and recognition repetition rate.

[0053] In this embodiment of the application, after the wake-up node of the IoT device triggers wake-up, the device performs interactive voice acquisition and recognition processing between the acquisition node and the processing node, acquires the interactive voice of the user's voice interaction, and then performs interactive semantic recognition to obtain the user's intention to control the device through voice interaction, and then performs subsequent device control.

[0054] The process involves recognizing the interactive voice, obtaining the recognition result, and the recognition repetition rate of the result. The higher the recognition repetition rate, the greater the accuracy of the current recognition and control, which serves as the data basis for generating subsequent device control results.

[0055] Step S20 in the method provided in this application embodiment includes:

[0056] Collect the user's interactive voice, perform noise filtering processing according to the noise feature information, and obtain noise-filtered interactive voice;

[0057] Perform interactive semantic recognition on the noise-filtered interactive speech to obtain the recognition result;

[0058] The recognition duplication rate is obtained by calculating the ratio of the number of identical recognition results to the total number of recognition results within the historical period.

[0059] In this embodiment, after the IoT device is woken up, it enters the acquisition node to collect interactive voice. Based on the noise feature information obtained by the wake-up node, the interactive voice is filtered. Since the wake-up voice and the interactive voice are collected in the same environment, their noise features are the same. This filtering improves the accuracy of voice recognition.

[0060] For example, based on the noise frequency and signal-to-noise ratio within the noise feature information, noise filtering processing is performed on the interactive voice, such as using noise filtering processing to obtain noise-filtered interactive voice.

[0061] Furthermore, interactive semantic recognition is performed on the noise-filtered interactive speech. For example, Natural Language Processing (NLP) technology is used to recognize the noise-filtered interactive speech to obtain recognition results. Specifically, by performing recognition based on noise feature information after noise filtering, the recognition accuracy can be improved.

[0062] For example, the recognition results may be: "Play music", "Turn off the lights", "Turn on the lights", etc., which are related to the type of IoT device.

[0063] Furthermore, the ratio of the number of identical identification results to the total number of identification results in the historical period is calculated to obtain the identification repetition rate, which reflects the proportion of repeated occurrences of the current identification result. The higher the rate, the greater the accuracy of the identification result.

[0064] For example, if an IoT device performs 1000 interactive voice recognition operations over a historical period, and the total number of recognition results is 1000, and the same recognition result "play music" appears 620 times, then the recognition repetition rate is 620 / 1000 = 62%. Furthermore, the same interactive control purpose of a user on an IoT device may result in different recognition results, such as "play music" or "play any song," etc. However, the more frequently a recognition result appears, the more it matches the user's habits, and the higher the accuracy.

[0065] The accuracy of the current recognition result can be reflected by calculating the recognition repetition rate, which can then serve as the data basis for subsequent equipment control feedback and improve the adaptability of equipment interactive control feedback.

[0066] S30: Based on the corrected arousal emotion information and repetition rate, generate device control results and feedback information according to the recognition results, and perform feedback interaction and device control.

[0067] In this embodiment, after the acquisition node and processing node, the IoT device enters the control node to control and provide feedback based on the user's semantic recognition results, completing the full voice interaction control. Specifically, based on the corrected wake-up emotion information and the repetition rate of the recognition results, device control results and feedback information are generated. The computing power configuration for generating device control results and feedback information takes into account the user's corrected wake-up emotion information and the repetition rate of the recognition results to improve the accuracy and timeliness of device control feedback, thus aligning it with the current user emotion and recognition results, and ultimately enhancing the user experience.

[0068] Step S30 in the method provided in this application embodiment includes:

[0069] Ensemble learning is used to train an interactive control feedback channel with M interactive control feedback paths, where M is a positive integer;

[0070] Based on the corrected arousal emotion information and the repetition rate, combined with M, the number of emotion recognitions and the number of accurate recognitions are calculated, and the number of composite recognitions N is calculated, where N is a positive integer less than or equal to M.

[0071] N interactive control feedback paths are randomly selected. The recognition results are input, and the device control results and path feedback information of the N paths are output. The device control results and feedback information with the highest occurrence rate are selected and feedback interaction and device control are performed.

[0072] In this embodiment, during voice interaction, the IoT device needs to generate reasonable device control results and feedback information based on the user's voice command recognition results. To improve the accuracy of device control and feedback, ensemble learning is used to train M interactive control feedback paths, forming an interactive control feedback channel. M is a positive integer, for example, 10.

[0073] The step "using ensemble learning to train an interactive control feedback channel including M interactive control feedback paths" in the method provided in this application embodiment includes:

[0074] Based on voice interaction processing data over a historical period, a set of sample recognition results is collected, along with the correct device control results and feedback information corresponding to different sample recognition results, which are labeled as the sample device control result set and the sample feedback information set.

[0075] The sample identification result set, sample device control result set, and sample feedback information set are divided to obtain M sets of interactive control feedback training data.

[0076] Using the M sets of interactive control feedback training data respectively, M interactive control feedback paths are trained based on machine learning to obtain interactive control feedback channels.

[0077] In this embodiment of the application, in order to improve the accuracy of device control and user experience, it is necessary to train multiple interactive control feedback paths through historical interaction data to form an interactive control feedback channel.

[0078] Specifically, based on historical voice interaction processing data, the recognition results after interactive voice semantic recognition are collected to obtain a sample recognition result set, such as sample recognition results like "turn on the light" or "play music". Then, the correct device control results and feedback information corresponding to different sample recognition results are collected, such as the device control results for turning on the light or playing music for IoT devices, and the corresponding voice feedback information, such as "the light is on" or "play your favorite music", as sample device control result sets and sample feedback information sets. Based on different IoT devices, Table 1 shows a partial sample recognition result set, sample device control result set, and sample feedback information set.

[0079] Sample identification results Sample equipment control results Sample feedback information Turn on the light Turn on the lights The lights are on. Turn up the volume Volume +10% "Volume has been increased by 10%" Turn off the air conditioner Air conditioner off "Air conditioning is off" Play music Start playing music Start playing music Turn down the volume Volume -10% "Volume has been reduced by 10%"

[0080] Table 1

[0081] Furthermore, the sample identification result set, sample device control result set, and sample feedback information set are divided. Specifically, each division is performed with replacement according to a certain proportion, for example, a random division at a proportion of 60%, and the division is performed M times to obtain M sets of interactive control feedback training data.

[0082] Then, using M sets of interactive control feedback training data, based on machine learning, M interactive control feedback paths are trained to obtain interactive control feedback channels.

[0083] The training steps for a single interactive control feedback path are illustrated using an example. For instance, an interactive control feedback path is constructed based on a feedforward neural network. This path includes an input layer, a hidden layer, and an output layer. The input features of the input layer are the recognition results, and the output features of the output layer are the device control results and feedback information. During training, the set of sample recognition results is input, and the set of output device control results and feedback information is obtained. It is then determined whether these match the corresponding set of device control results and sample feedback information, and the percentage of consistency is calculated to obtain the accuracy. The accuracy is then assessed to determine if it meets the requirements, for example, an accuracy greater than 95%. If not, the difference between the accuracy and the target accuracy is calculated as the loss. Adam optimization is then used to adjust the network parameters of the interactive control feedback path to reduce the loss and improve the accuracy until the accuracy meets the requirements, at which point training is complete. Using the same method, M interactive control feedback paths are trained, with the same training process but different training data.

[0084] After training, the M interactive control feedback paths are combined to obtain the interactive control feedback channel. Since the training data of the M interactive control feedback paths are different, they have different performance. By integrating the outputs of the M interactive control feedback paths, the accuracy of device control feedback can be improved.

[0085] In this embodiment of the application, based on the corrected arousal emotion information of the user being analyzed and the repetition rate of the recognition results, combined with M, the number of emotion recognitions and the number of accurate recognitions are calculated, and the number of composite recognitions N is calculated, where N is a positive integer less than or equal to M.

[0086] Specifically, the higher the corrected emotional level within the user's corrected arousal emotional information, the more negative the user's emotion. This necessitates more accurate device control results and feedback information to prevent the user's emotional state from worsening and impacting the user experience. Consequently, more interactive control feedback paths are needed for device control and feedback information generation to improve accuracy. Conversely, a higher repetition rate indicates a more accurate current recognition result. Furthermore, since more device control and interactive feedback processes have previously yielded the same results, the probability of generating incorrect device control results and feedback information is lower. In this case, a faster control feedback response speed is needed to improve the user experience. Therefore, fewer interactive control feedback paths are required for device control and feedback information generation to reduce data processing volume and improve control feedback efficiency.

[0087] The step "calculating the number of emotion recognitions and the number of accurate recognitions based on the corrected arousal emotion information and the repetition rate, combined with M, and calculating the number of composite recognitions N" in the method provided in this application embodiment includes:

[0088] Get the highest emotion level;

[0089] Calculate the ratio of the corrected arousal emotion level to the maximum emotion level within the corrected arousal emotion information, multiply it by M and round it to obtain the number of emotions recognized.

[0090] Based on the repeated recognition rate, the control feedback recognition coefficient is calculated.

[0091] The accurate recognition count is obtained by multiplying the control feedback recognition coefficient by M and rounding it down.

[0092] Calculate the average of the number of emotion recognitions and the number of accuracy recognitions to obtain the composite recognition count N.

[0093] In this embodiment of the application, the maximum emotion level is first obtained, that is, the emotion level of the user's strongest negative emotion, for example, 3.

[0094] Furthermore, the ratio of the corrected arousal emotion level to the maximum emotion level within the current user's corrected arousal emotion information is calculated, then multiplied by M and rounded up to obtain the number of emotions recognized. For example, if the corrected arousal emotion level within the current user's corrected arousal emotion information is 2, the maximum emotion level is 3, and M is 10, then the number of emotions recognized is 2 / 3 * 10, rounded up to 7. Thus, the higher the corrected arousal emotion level, the greater the number of emotions recognized, resulting in more accurate device control results and feedback information generation, reducing the probability of errors, preventing further impact on user emotions, and improving user experience.

[0095] Furthermore, the control feedback identification coefficient is calculated based on the duplicate identification rate. For example, 1 minus the duplicate identification rate is used as the control feedback identification coefficient. For instance, if the duplicate identification rate is 62%, the control feedback identification coefficient is 38%.

[0096] Furthermore, the accuracy recognition count is obtained by multiplying the control feedback recognition coefficient by M and rounding it up. For example, if M is 10, the accuracy recognition count is 10 * 38% and rounded up to 4.

[0097] The higher the repetition rate, the greater the probability of generating accurate device control results and feedback information. The fewer interactive control feedback paths required, the more accurate device control results and feedback information can be generated. Therefore, reducing the number of interactive control feedback path calls results in a higher repetition rate and a smaller number of accurate recognitions, improving the timeliness and efficiency of control feedback, and thus enhancing the user experience.

[0098] Furthermore, the average of the number of emotion recognitions and the number of accuracy recognitions is calculated to obtain the composite recognition count N. For example, if the number of emotion recognitions is 7 and the number of accuracy recognitions is 4, then the composite recognition count N is the average of the two, rounded up to 6.

[0099] Thus, this embodiment of the application comprehensively considers the user's emotions and the accuracy of the current voice interaction recognition results to calculate and decide the number of interactive control feedback paths to be called in the device control feedback generation analysis. When the user's negative emotion level is higher, more interactive control feedback paths are called to improve accuracy and user experience. When the recognition result accuracy is higher, fewer interactive control feedback paths are called to improve processing efficiency and user experience. Combining these two dimensions can comprehensively ensure the accuracy and timeliness of control feedback.

[0100] In this embodiment of the application, according to the composite recognition quantity N, N interactive control feedback paths are randomly selected in the interactive control feedback channel. The current recognition result is input, and the N path device control results and path feedback information are output respectively. Then, the device control result and feedback information with the highest occurrence ratio (most occurrences) are selected as the final device control result and feedback information for feedback interaction and device control.

[0101] The semantically driven node engine voice interaction method provided in this embodiment of the invention has at least the following technical effects:

[0102] This invention effectively improves the accuracy, response speed, and user experience of voice interaction. By monitoring and acquiring the user's wake-up voice through IoT devices, and combining this with semantic recognition, noise feature information and emotional information are identified and corrected upon triggering the wake-up node. This allows for accurate identification of the user's emotional information, and the interaction strategy is adjusted based on the user's emotional state to optimize the interactive experience. During voice interaction, the user's interactive voice is acquired through a data acquisition node, and noise filtering is performed based on noise feature information to reduce the interference of environmental noise on voice recognition and improve the accuracy of interactive semantic recognition. The control node generates corresponding control feedback analysis strategies based on the corrected wake-up emotional information and recognition repetition rate, and analyzes and generates corresponding device control results and feedback information. This adapts to the user's emotional state and optimizes the response strategy, thereby reducing unnecessary interaction delays, improving the accuracy of interactive control, and ultimately enhancing the intelligence of the interaction and user satisfaction.

[0103] Example 2, as Figure 2 As shown, based on the same inventive concept as the semantic-driven node engine voice interaction method provided in Embodiment 1, this embodiment of the invention also provides a semantic-driven node engine voice interaction system, including:

[0104] The wake-up emotion recognition module 11 is used to monitor and acquire the wake-up voice issued by the user through IoT devices, perform semantic recognition, and when the wake-up node is triggered, identify and obtain noise feature information and wake-up emotion information, and correct and obtain corrected wake-up emotion information. The node engine also includes a collection node, a processing node and a control node.

[0105] Interactive semantic recognition module 12 is used to collect the user's interactive voice, perform noise filtering according to the noise feature information, perform interactive semantic recognition, and obtain recognition results and recognition repetition rate;

[0106] The control feedback interaction module 13 is used to generate device control results and feedback information based on the recognition results according to the corrected arousal emotion information and repetition recognition rate, and to perform feedback interaction and device control.

[0107] Furthermore, the semantic-driven node engine voice interaction system is also used to: monitor and acquire the wake-up voice issued by the user through IoT devices, input it into a semantic recognizer, and obtain the recognized semantics;

[0108] Determine whether the identified semantic is wake-up. If yes, trigger the wake-up node; otherwise, do not trigger the wake-up node and the IoT device remains in standby mode.

[0109] When the wake-up node is triggered, noise feature recognition and emotion recognition are performed on the wake-up voice to obtain noise feature information and wake-up emotion information.

[0110] Furthermore, the semantically driven node engine voice interaction system is also used to: collect a set of sample wake-up voices based on wake-up voice data within a historical time period, and extract noise features and identify the wake-up emotion level for each sample wake-up voice to obtain a set of sample noise feature information and a set of sample wake-up emotion information.

[0111] Using the sample wake-up speech set, sample noise feature information set, and sample wake-up emotion information set as input and output features, a wake-up speech recognizer is trained based on deep learning;

[0112] The wake-up voice is input into the wake-up voice recognizer, and noise feature information and wake-up emotion information are output.

[0113] Furthermore, the semantically driven node engine voice interaction system is also used to: extract noise decibels from the noise feature information;

[0114] Based on the ratio of the noise decibels to the average noise decibels, the emotional level within the arousal emotional information is corrected and calculated to obtain the corrected arousal emotional level, which is used as the corrected arousal emotional information.

[0115] Furthermore, the semantically driven node engine voice interaction system is also used to: collect the user's interactive voice, perform noise filtering processing according to the noise feature information, and obtain noise-filtered interactive voice;

[0116] Perform interactive semantic recognition on the noise-filtered interactive speech to obtain the recognition result;

[0117] The recognition duplication rate is obtained by calculating the ratio of the number of identical recognition results to the total number of recognition results within the historical period.

[0118] Furthermore, the semantically driven node engine voice interaction system is also used to: train an interaction control feedback channel including M interaction control feedback paths using ensemble learning, where M is a positive integer;

[0119] Based on the corrected arousal emotion information and the repetition rate, combined with M, the number of emotion recognitions and the number of accurate recognitions are calculated, and the number of composite recognitions N is calculated, where N is a positive integer less than or equal to M.

[0120] N interactive control feedback paths are randomly selected. The recognition results are input, and the device control results and path feedback information of the N paths are output. The device control results and feedback information with the highest occurrence rate are selected and feedback interaction and device control are performed.

[0121] Furthermore, the semantically driven node engine voice interaction system is also used to: collect a set of sample recognition results based on voice interaction processing data within a historical time period, and collect the correct device control results and feedback information corresponding to different sample recognition results, and label them as a set of sample device control results and a set of sample feedback information.

[0122] The sample identification result set, sample device control result set, and sample feedback information set are divided to obtain M sets of interactive control feedback training data.

[0123] Using the M sets of interactive control feedback training data respectively, M interactive control feedback paths are trained based on machine learning to obtain interactive control feedback channels.

[0124] Furthermore, the semantically driven node engine voice interaction system is also used for:

[0125] Get the highest emotion level;

[0126] Calculate the ratio of the corrected arousal emotion level to the maximum emotion level within the corrected arousal emotion information, multiply it by M and round it to obtain the number of emotions recognized.

[0127] Based on the repeated recognition rate, the control feedback recognition coefficient is calculated.

[0128] The accurate recognition count is obtained by multiplying the control feedback recognition coefficient by M and rounding it down.

[0129] Calculate the average of the number of emotion recognitions and the number of accuracy recognitions to obtain the composite recognition count N.

[0130] Although preferred embodiments of the invention have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.

[0131] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for semantic driven node engine voice interaction, the method comprising: The method includes: The system monitors and acquires the wake-up voice issued by the user through IoT devices, performs semantic recognition, and identifies and obtains noise feature information and wake-up emotion information when the wake-up node is triggered, and corrects and obtains corrected wake-up emotion information. The node engine also includes acquisition nodes, processing nodes and control nodes. Collect user's interactive voice, perform noise filtering according to the noise feature information, perform interactive semantic recognition, and obtain recognition results and recognition repetition rate, including: Collect the user's interactive voice, perform noise filtering processing according to the noise feature information, and obtain noise-filtered interactive voice; Perform interactive semantic recognition on the noise-filtered interactive speech to obtain the recognition result; The recognition duplication rate is obtained by calculating the ratio of the number of identical recognition results to the total number of recognition results within the historical period. Based on the corrected arousal emotion information and recognition repetition rate, device control results and feedback information are generated according to the recognition results, and feedback interaction and device control are performed, including: Ensemble learning is used to train an interactive control feedback channel with M interactive control feedback paths, where M is a positive integer; Based on the corrected arousal emotion information and recognition repetition rate, combined with M, the number of emotion recognitions and the number of accurate recognitions are calculated, and the number of composite recognitions N is calculated, where N is a positive integer less than or equal to M, including: Get the highest emotion level; Calculate the ratio of the corrected arousal emotion level to the maximum emotion level within the corrected arousal emotion information, multiply it by M and round it to obtain the number of emotions recognized. Based on the recognition repetition rate, the control feedback recognition coefficient is calculated. The accurate recognition count is obtained by multiplying the control feedback recognition coefficient by M and rounding it down. Calculate the average of the number of emotion recognitions and the number of accuracy recognitions to obtain the composite recognition count N; N interactive control feedback paths are randomly selected. The recognition results are input, and the device control results and path feedback information of the N paths are output. The device control results and feedback information with the highest occurrence rate are selected and feedback interaction and device control are performed.

2. The semantic driven node engine voice interaction method of claim 1, wherein, By monitoring and acquiring the user's wake-up voice through IoT devices, semantic recognition is performed. When the wake-up node is triggered, noise feature information and wake-up emotion information are identified, including: The system monitors and acquires the user's wake-up voice through IoT devices, inputs it into a semantic recognizer, and obtains the recognized semantics. Determine whether the identified semantic is wake-up. If yes, trigger the wake-up node; otherwise, do not trigger the wake-up node and the IoT device remains in standby mode. When the wake-up node is triggered, noise feature recognition and emotion recognition are performed on the wake-up voice to obtain noise feature information and wake-up emotion information.

3. The semantic driven node engine voice interaction method of claim 1, wherein, When the wake-up node is triggered, noise feature recognition and emotion recognition are performed on the wake-up voice, including: Based on wake-up voice data from a historical period, a set of sample wake-up voices is collected, and noise features are extracted and wake-up emotion level is labeled for each sample wake-up voice, to obtain a set of sample noise feature information and a set of sample wake-up emotion information. Adopt the sample wake-up voice set, sample noise feature information set and sample wake-up emotion information set as input features and output features, train the wake-up voice recognizer based on deep learning; Input the wake-up voice into the wake-up voice recognizer, and output the obtained noise feature information and wake-up emotion information.

4. The semantic driven node engine voice interaction method of claim 1, wherein, Correct the obtained correction wake-up emotion information, including: Extract the noise decibel in the noise feature information; According to the ratio of the noise decibel to the average noise decibel, the emotion level in the wake-up emotion information is corrected and calculated to obtain the corrected wake-up emotion level as the corrected wake-up emotion information.

5. The semantic driven node engine voice interaction method of claim 1, wherein, Adopt ensemble learning to train the interactive control feedback channel including M interactive control feedback paths, including: According to the voice interaction processing data in the historical time, a sample recognition result set is collected, and the correct device control result and feedback information corresponding to different sample recognition results are collected and labeled as a sample device control result set and a sample feedback information set; Divide the sample recognition result set, sample device control result set and sample feedback information set to obtain M groups of interactive control feedback training data; Respectively adopt the M groups of interactive control feedback training data to train M interactive control feedback paths based on machine learning to obtain the interactive control feedback channel.

6. A semantically driven node engine voice interaction system, characterized by The steps of a semantic-driven node engine voice interaction method for implementing any one of claims 1-5, including: The wake-up emotion recognition module is used to monitor and obtain the wake-up voice issued by the user through the Internet of Things device, perform semantic recognition, and identify and obtain the noise feature information and the wake-up emotion information when triggering the wake-up node, and correct the obtained correction wake-up emotion information, wherein the node engine further includes a collection node, a processing node and a control node; The interactive semantic recognition module is used to collect the interactive voice of the user, perform noise filtering processing according to the noise feature information, perform interactive semantic recognition, and obtain the recognition result and the recognition repetition rate; The control feedback interaction module is used to generate the device control result and the feedback information according to the recognition result according to the corrected wake-up emotion information and the recognition repetition rate, and perform feedback interaction and device control.

Citation Information

Patent Citations

  • Highly anthropomorphic voice interaction algorithm and emotion interaction algorithm for robot and robot

    CN109741746A

  • Voice interaction device and method and computer readable storage medium

    CN110827821A