Smart Home Voice Interaction Control Method and System

By integrating voice and non-voice information, understanding the user's emotional state, and dynamically adjusting the device control strategy based on the user's historical operating habits and preferences, the problem that the existing smart home voice interaction system cannot accurately understand the user's intentions is solved, and higher speech recognition accuracy and user satisfaction are achieved.

CN119601010BActive Publication Date: 2025-06-24XIAMEN OCEAN VOCATIONAL & TECH COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510141656.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-24
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The existing smart home voice interaction system cannot accurately understand the user's emotional state and intentions, especially in noisy environments, voice recognition accuracy is low, and device control strategies lack personalization and real-time optimization.

Method used

By integrating voice and non-voice information, users' voice input signals and non-voice information (such as facial expressions and hand movements), synchronous processing is carried out to understand the user's emotional state, and dynamically adjust the device control strategy based on the user's historical operating habits and preferences, provide personalized services, and continuously optimize the control strategy through user feedback.

Benefits of technology

It improves the accuracy of recognition of voice commands, fully understands users' emotional state and needs, improves user satisfaction, provides more humane services, and improves the intelligence level and user experience of smart home systems through continuous optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119601010B_ABST
    Figure CN119601010B_ABST
Patent Text Reader

Abstract

The present invention discloses a smart home voice interaction control method and system, which relates to the technical field of smart homes. In order to solve the problems that the prior art cannot accurately understand the user's intention through non-verbal information such as the user's mood and expression, the control strategy is rigid and lacks personalization, and it cannot adapt to the user's needs and mood in real time; the present invention enhances the accurate recognition of the user's intention by integrating voice and non-verbal information, improves the recognition accuracy of voice commands, and more comprehensively understands the user's emotional state and needs. Especially in a noisy environment or a complex situation, the recognition and response effect of the system is significantly improved. By using the user's sentiment analysis and context understanding technology, the interaction method is dynamically adjusted according to the user's emotional state, which can flexibly adapt to the emotional needs of different users, provide more user-friendly services, and continuously optimize the control strategy according to the user's feedback to improve user satisfaction, thereby improving the intelligence level and user experience of the smart home system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart home, and particularly to a method and system for voice interaction control in smart home. Background Art

[0002] Existing smart home voice interaction systems often have problems such as response latency, low speech recognition accuracy, and inability to fully understand the user's emotional state. Especially in a noisy environment, the speech recognition accuracy is prone to decline. For example, a Chinese patent application with the publication number CN116758914A discloses a smart home voice interaction control method applied to a central control gateway, which is characterized by including: receiving text information sent by a voice panel, where the text information is obtained by converting voice information collected by the voice panel; performing semantic parsing on the received text information to determine the semantic content; controlling smart home devices to perform corresponding operations according to the semantic content, and replying the corresponding text information to the voice panel to play the corresponding voice content through the voice panel.

[0003] Although the above patent uses the combination of a central control gateway and a voice panel to replace multiple gateway devices with screens, saving costs, the voice interaction method often cannot accurately understand the user's intention through non-verbal information such as the user's mood and expression. In addition, the device control strategy of the existing system is relatively fixed, lacking personalized adjustment and optimization, and unable to make effective feedback and adjustment according to the user's real-time needs and emotional state. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for smart home voice interaction control, which improves the recognition accuracy of voice commands by integrating voice and non-verbal information, dynamically adjusts the interaction method according to the user's emotional state, flexibly adapts to the emotional needs of different users, provides a more user-friendly service, and continuously optimizes the control strategy according to the user's feedback to improve user satisfaction, so as to solve the problems raised in the above background art.

[0005] To achieve the above purpose, the present invention provides the following technical solutions:

[0006] A smart home voice interaction control method, including the following steps:

[0007] Step 1: Collection and preprocessing of voice input: Collect the user's voice input signal through a smart voice collection device, and preprocess the voice input signal;

[0008] Step 2: Recognition and parsing of voice commands: Perform text recognition on the preprocessed voice input signal, convert the recognition result into a text command, and perform semantic parsing on the text command to determine the user's intention and generate a user command;

[0009] Step 3: Generation of device control strategy: Match the generated user instructions with the corresponding device control strategies to determine the operation logic between the user requirements and the target device. Meanwhile, adjust the device control strategies according to the user's historical operation habits and preferences.

[0010] Step 4: Execution and real-time monitoring of device operations: Establish a communication link with the target device according to the device control strategy, send operation instructions to the target device, and monitor the response of the device in real time.

[0011] Step 5: Feedback of operation results and continuous optimization: After the device operation is completed, feedback the execution result to the user by voice and record the user operation data.

[0012] Furthermore, the said Step 1: Acquisition and preprocessing of voice input further includes:

[0013] Match the obtained voice input signal with the preset wake-up words and phrases to trigger the wake-up monitoring data acquisition device.

[0014] When collecting the voice input signal, collect the user's non-voice information based on the monitoring data acquisition device to obtain user monitoring data.

[0015] Identify the user's facial expressions and hand movements from the obtained user monitoring data, analyze the data characteristics of the non-voice information, and understand the user's emotional state and intentions.

[0016] Meanwhile, synchronize the voice input signal and the non-voice information based on the time stamp, establish the correlation between the voice input signal and the non-voice information, and supplement the semantics of the voice input signal based on the non-voice information.

[0017] Furthermore, the said Step 2: Recognition and parsing of voice instructions specifically includes:

[0018] Effective voice extraction: After collecting the voice input signal, remove the invalid silent parts before and after the voice through endpoint detection technology, retain the effective voice information, and split the long voice input signal into multiple logical segments.

[0019] Voice quality assessment: Monitor and evaluate the quality of the voice input signal in real time. If the voice quality does not meet the standard, actively optimize the parameters or prompt the user to re-enter. Meanwhile, dynamically adjust the acquisition parameters of the microphone according to the environmental factors.

[0020] Text instruction generation: Perform voice recognition processing on the qualified voice input signal, convert the recognition result into text information, and correct the noise, accent or unclear pronunciation in the voice input signal.

[0021] Semantic analysis: Perform word segmentation on the converted text information, split the text instructions into words or phrases, identify the part of speech of each word based on the semantic recognition of the supplementary voice input signal, determine the grammatical relationship between words, determine the user's operation intention according to the parsing result, and identify the key entities in the text;

[0022] Context understanding: After each round of conversation ends, record the relevant information and status of the user input, perform context understanding on continuous voice instructions, maintain the coherence of the conversation, and infer the user's further needs based on the previous instructions and operation results;

[0023] User instruction generation: Generate structured user instructions according to the recognition result of semantic analysis and the inference result of context understanding.

[0024] Furthermore, monitor and evaluate the quality of the voice input signal in real time, including:

[0025] Monitor the signal characteristic parameters of the voice input signal in real time, where the signal characteristic parameters include the signal-to-noise ratio, the signal difference ratio corresponding to the distortion degree, the jitter parameter, and the drift parameter;

[0026] Obtain the first voice quality evaluation coefficient by using the signal-to-noise ratio and the signal difference ratio corresponding to the distortion degree;

[0027] Among them, the first voice quality evaluation coefficient is obtained through the following formula:

[0028]

[0029] Among them, Q 01 represents the first voice quality evaluation coefficient; n represents the number of unit times included in the voice input signal; S i+1 represents the signal-to-noise ratio of the voice input signal corresponding to the i + 1th unit time; S i represents the signal-to-noise ratio of the voice input signal corresponding to the ith unit time; P i represents the signal difference ratio corresponding to the distortion degree of the voice input signal corresponding to the ith unit time; P b represents the standard deviation of the signal difference ratio corresponding to the distortion degree corresponding to n unit times; S b represents the standard deviation of the signal-to-noise ratio of the voice input signal corresponding to n unit times;

[0030] Obtain the second voice quality evaluation coefficient by using the jitter parameter and the drift parameter;

[0031] Among them, the second voice quality evaluation coefficient is obtained through the following formula:

[0032]

[0033] Among them, Q 02 represents the second voice quality evaluation coefficient; n represents the number of unit times included in the voice input signal; J i and Y i represent the jitter parameter and drift parameter of the voice input signal corresponding to the i-th unit time; J i+1 and Y i+1 represent the jitter parameter and drift parameter of the voice input signal corresponding to the (i + 1)-th unit time; J b and Y b represent the standard deviation of the jitter parameter and the standard deviation of the drift parameter of the voice input signal corresponding to n unit times;

[0034] Use the first voice quality evaluation coefficient and the second voice quality evaluation coefficient to determine the quality of the voice input signal.

[0035] Furthermore, using the first voice quality evaluation coefficient and the second voice quality evaluation coefficient to determine the quality of the voice input signal includes:

[0036] Use the preset first evaluation coefficient reference value and second evaluation coefficient reference value to perform normalization processing on the first voice quality evaluation coefficient and the second voice quality evaluation coefficient respectively, and obtain the normalized first voice quality evaluation coefficient and second voice quality evaluation coefficient;

[0037] Compare the normalized first voice quality evaluation coefficient and the second voice quality evaluation coefficient;

[0038] When the first voice quality evaluation coefficient is lower than the second voice quality evaluation coefficient, the first comprehensive evaluation model is retrieved; when the first voice quality evaluation coefficient is higher than or equal to the second voice quality evaluation coefficient, the second comprehensive evaluation model is retrieved;

[0039] Use the first comprehensive evaluation model to combine the first voice quality evaluation coefficient and the second voice quality evaluation coefficient to obtain the first comprehensive evaluation coefficient, and compare the first comprehensive evaluation coefficient with the preset first comprehensive coefficient threshold; when the first comprehensive evaluation coefficient exceeds the preset first comprehensive coefficient threshold, it is determined that the voice quality does not meet the standard;

[0040] Among them, the first comprehensive evaluation coefficient is obtained through the following formula:

[0041]

[0042] Among them, Z 01 represents the first comprehensive evaluation coefficient; Q x01 represents the normalized first voice quality evaluation coefficient; Q x02Represents the second speech quality evaluation coefficient after normalization;

[0043] Use the second comprehensive evaluation model to obtain the second comprehensive evaluation coefficient by combining the first speech quality evaluation coefficient and the second speech quality evaluation coefficient, and compare the second comprehensive evaluation coefficient with a preset second comprehensive coefficient threshold; when the second comprehensive evaluation coefficient exceeds the preset second comprehensive coefficient threshold, it is determined that the speech quality does not meet the standard;

[0044] Among them, the second comprehensive evaluation coefficient is obtained through the following formula:

[0045]

[0046] Among them, Z 02 Represents the second comprehensive evaluation coefficient; Q x01 Represents the first speech quality evaluation coefficient after normalization; Q x02 Represents the second speech quality evaluation coefficient after normalization.

[0047] Furthermore, in the step two: recognition and parsing of voice commands, it also includes emotional perception and context understanding of text information. Based on multi-modal data, by analyzing the intonation and speech rate characteristics of the user's voice, combined with the data characteristics of non-speech information obtained synchronously, the user's emotional state is recognized, and the device control strategy is adjusted according to the user's emotional state.

[0048] Furthermore, in the step three, before executing the control strategy, it also includes confirming and interacting with the user through voice to ensure the accuracy of the control operation. The specific steps are as follows:

[0049] Voice feedback and confirmation: After generating the control strategy, the expected operation is fed back to the user through voice. The feedback content will include a summary of the understood user needs to help the user quickly understand the upcoming operation;

[0050] User confirmation and adjustment: Receive the user's confirmation or adjustment instruction through voice. If the confirmation content of the user is not clear, it will automatically prompt and ask for clarification;

[0051] Multi-round interaction and confirmation: When the user's instruction involves multiple devices and operation steps, ensure that each operation conforms to the user's intention through step-by-step confirmation, and optimize the user's operation experience based on the confirmation feedback after each step is completed;

[0052] Personalized confirmation: Dynamically adjust the confirmation process based on the user's emotional state;

[0053] Confirmation information feedback: Once the user-confirmed control strategy is received, the specific content and operation result to be executed will be clearly informed to the user through voice before the operation. After the operation is completed, the final result will be confirmed through voice continuously, and operation review will be provided.

[0054] Furthermore, the personalized confirmation specifically includes:

[0055] Obtain the intonation and speech rate characteristics of the user's voice, respond with intonation and tone matching the user's emotional state, and match corresponding template statements in the pre-stored voice database according to the user's emotional state for response;

[0056] Classify the user's emotional state, determine the user's emotional category, match corresponding interaction styles for users with different emotional states for interaction, and adjust the execution priority of tasks according to the user's emotional state.

[0057] Furthermore, Step Five: Feedback and continuous optimization of operation results also includes:

[0058] Regularly ask the user about their satisfaction with the smart home system through voice, summarize the satisfaction evaluation report, and help the user understand their usage habits and improvement process;

[0059] Conduct correlation analysis on the user's feedback data, the user's specific operations, and the system response data, evaluate the key indicators of each specific operation, and generate a user satisfaction optimization strategy according to the user's usage habits and preferences;

[0060] Identify the deficiencies of certain functions based on the user's long-term feedback and needs, and promote the iterative update of the function.

[0061] The present invention provides another technical solution, a smart home voice interaction control system, including:

[0062] A data acquisition module, used to collect the user's voice input signal and monitoring video data, and preprocess the collected voice input signal and monitoring video data;

[0063] A voice recognition and parsing module, used to perform text conversion on the preprocessed voice input signal, perform semantic parsing on the text based on natural language processing technology to determine the operation instruction. At the same time, verify the recognition result of the voice input signal based on the monitoring video data, and generate prompt or clarification statements for fuzzy or incomplete instructions according to the recognition result;

[0064] A control strategy generation module, used to match the operation logic of the target device according to the parsing result, generate a basic control strategy, and adjust the control strategy according to the user's historical operation data;

[0065] The device control module is used to establish a communication link with the target device, send operation instructions through the home Internet of Things protocol, and monitor the response status of the device in real time. If the device is abnormal, it will trigger an exception handling mechanism and feedback to the user;

[0066] The feedback optimization module is used to obtain user instructions and device response data, construct a user behavior data set, generate shortcut operation logic based on the user's frequently used instructions, collect evaluation feedback from the user at preset time intervals, and optimize the control instructions based on the evaluation feedback.

[0067] Compared with the prior art, the beneficial effects of the present invention are:

[0068] By integrating voice and non-voice information, the accurate recognition of user intentions is enhanced. Not only the recognition accuracy of voice instructions is improved, but also the user's emotional state and needs can be more comprehensively understood. Especially in noisy environments or complex situations, the recognition and response effects of the system are significantly improved. Using the user's sentiment analysis and context understanding technology, the interaction method is dynamically adjusted according to the user's emotional state, which can flexibly adapt to the emotional needs of different users, provide more user-friendly services, and continuously optimize the control strategy according to the user's feedback to improve user satisfaction, thereby enhancing the intelligence level and user experience of the smart home system. Description of the Drawings

[0069] Figure 1 It is a flowchart of the smart home voice interaction control method of the present invention. Detailed Embodiments

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0071] In order to solve the technical problems that the prior art cannot accurately understand the user's intentions through non-voice information such as the user's emotions and expressions, the control strategy is fixed and lacks personalization, and cannot adapt to the user's needs and emotions in real time, please refer to Figure 1 This embodiment provides the following technical solutions:

[0072] The smart home voice interaction control method includes the following steps:

[0073] Step 1: Collection and Preprocessing of Voice Input: Collect the voice input signal of the user through intelligent voice collection devices (such as microphone arrays or intelligent voice terminals), and preprocess the voice input signal, including noise suppression, voice signal enhancement, echo cancellation, and reverberation suppression, etc., to ensure the quality of voice data;

[0074] Step 2: Recognition and Parsing of Voice Commands: Perform text recognition on the preprocessed voice input signal, convert the recognition result into a text command, and perform semantic parsing on the text command to determine the user's intention, generate a user command, perform emotional perception and context understanding on the text information, analyze the intonation and speech rate characteristics of the user's voice based on multi-modal data, combine the data characteristics of the non-speech information obtained synchronously, recognize the user's emotional state, such as anxiety, pleasure, tension, etc., and adjust the device control strategy according to the user's emotional state. For example, when the user says "turn off the light" with a tired expression, the system can infer that the user is in a rest state and automatically adjust the states of other home appliances (such as air conditioners, curtains) to provide a more comfortable rest environment and more user-friendly services;

[0075] Step 3: Generation of Device Control Strategy: Match the generated user command with the corresponding device control strategy to determine the operation logic of the user's needs and the target device. At the same time, adjust the device control strategy according to the user's historical operation habits and preferences to achieve personalized control;

[0076] Step 4: Execution and Real-time Monitoring of Device Operations: Establish a communication link with the target device according to the device control strategy, send an operation command to the target device, and monitor the response of the device in real time;

[0077] Step 5: Feedback of Operation Results and Continuous Optimization: After the device operation is completed, feedback the execution result to the user through voice, and record the user operation data to optimize the subsequent experience.

[0078] In this embodiment, by integrating voice and non-voice information, the accurate recognition of the user's intention is enhanced. It not only improves the recognition accuracy of voice commands, but also can more comprehensively understand the user's emotional state and needs. Especially in a noisy environment or a complex situation, the recognition and response effect of the system is significantly improved. Using the user's emotional analysis and context understanding technology, dynamically adjust the interaction method according to the user's emotional state, can flexibly adapt to the emotional needs of different users, provide more user-friendly services, and continuously optimize the control strategy according to the user's feedback to improve user satisfaction, thereby improving the intelligence level and user experience of the smart home system.

[0079] In this embodiment, the above-mentioned Step 1: Collection and Preprocessing of Voice Input, further includes:

[0080] Match the obtained voice input signal with a preset wake-up word and phrase to trigger the wake-up monitoring data acquisition device, such as "Hello, smart home" or "Turn on the home assistant";

[0081] When collecting the voice input signal, collect non-voice information such as the user's facial expressions and gestures based on the monitoring data acquisition device to obtain user monitoring data;

[0082] Identify the user's facial expressions and hand movements from the obtained user monitoring data, analyze the data characteristics of the non-voice information, understand the user's emotional state and intentions, and thus adjust the subsequent interaction strategy. For example, if the user shows a tense or eager expression, the system will speed up the response speed and give priority to processing emergency operation requests;

[0083] At the same time, synchronize the voice input signal and the non-voice information based on the time stamp, establish the association relationship between the voice input signal and the non-voice information, and supplement the semantics of the voice input signal based on the non-voice information. For example, when the user says "Open the curtain" and makes a gesture pointing to the curtain, the system will give priority to recognizing the gesture and performing the relevant operation;

[0084] In this embodiment, by fusing multi-modal data, the accuracy of speech recognition and the comprehensiveness of user intention understanding are improved, and the multi-modal data is synchronously processed. With the assistance of non-voice information such as facial expressions and gestures, the accuracy of speech recognition is further enhanced, the depth of user intention understanding is increased, a reliable data basis is provided for subsequent speech command recognition and parsing, and it is ensured that the system can operate smoothly in complex environments and various interaction scenarios.

[0085] In this embodiment, step two: the recognition and parsing of voice commands specifically includes:

[0086] Effective voice extraction: After collecting the voice input signal, remove the invalid silent parts before and after the voice through endpoint detection technology, retain the effective voice information, and split the long voice input signal into multiple logical segments;

[0087] Voice quality assessment: Real-time monitor and evaluate the quality of the voice input signal. If the voice quality does not meet the standard (such as signal distortion, excessive noise, etc.), actively optimize the parameters or prompt the user to re-enter to ensure that the voice data meets the requirements of subsequent processing. At the same time, dynamically adjust the acquisition parameters such as the gain and sampling frequency of the microphone according to environmental factors such as environmental noise level, user distance, and device performance to optimize the voice acquisition effect;

[0088] Text command generation: Perform speech recognition processing on the voice input signal that passes the evaluation, convert the recognition result into text information, and correct the noise, accent, or fuzzy pronunciation in the voice input signal to improve the recognition accuracy;

[0089] Semantic parsing: Perform word segmentation on the transformed text information, split the text instructions into words or phrases, identify the part of speech of each word based on the semantic recognition of the supplementary voice input signal, such as nouns, verbs, adjectives, etc., determine the grammatical relationship between words, and determine the user's operation intention according to the parsing result, such as turning on / off the light, adjusting the temperature, etc. Identify the key entities in the text, such as device type, device name, location, time, operation instructions and their values, etc. For example, if the user says "Increase the temperature in the living room by two degrees", the system will recognize "Increase" as the operation instruction, "living room" as the target device, and "temperature" as the adjustment attribute, and finally extract the "Increase temperature" instruction and determine the target temperature control device;

[0090] Context understanding: After each round of conversation ends, record the relevant information and status of the user input, perform context understanding on consecutive voice instructions, maintain the coherence of the conversation, and infer the user's further needs based on the previous instructions and operation results. The system can identify and process multi-round conversations, realize the logical association across instructions, and avoid repeated recognition or invalid operations; for example, when the user says "Lower the living room lights" and "Increase the air conditioner temperature" in succession, the system not only recognizes the two independent instructions, but also understands the time and space relationship between the two instructions, and coordinates according to the current device status in the living room;

[0091] User instruction generation: Generate structured user instructions according to the recognition result of semantic parsing and the inference result of context understanding.

[0092] In this embodiment, through effective voice extraction, voice quality assessment, text instruction generation, semantic parsing, context understanding, and user instruction generation, the interaction ability and user experience of the smart home system are significantly improved, the accuracy of voice data and the comprehensiveness of instruction parsing are ensured, so that the system can more accurately understand the user's intention, generate operation instructions accordingly, improve the accuracy and efficiency of voice recognition, enhance the interaction coherence between the user and the system, realize more natural and intelligent voice interaction, and thus greatly improve the user's satisfaction and the practicality of the smart home system.

[0093] Specifically, the quality of the voice input signal is monitored and evaluated in real time, including:

[0094] The signal characteristic parameters of the voice input signal are monitored in real time, where the signal characteristic parameters include signal-to-noise ratio, signal difference ratio corresponding to distortion degree, jitter parameter, and drift parameter;

[0095] Obtain the first voice quality evaluation coefficient by using the signal-to-noise ratio and the signal difference ratio corresponding to the distortion degree;

[0096] Among them, the first voice quality evaluation coefficient is obtained through the following formula:

[0097]

[0098] Among them, Q 01 represents the first voice quality evaluation coefficient; n represents the number of unit time intervals included in the voice input signal; S i+1 represents the signal-to-noise ratio of the voice input signal corresponding to the (i + 1)-th unit time interval; S i represents the signal-to-noise ratio of the voice input signal corresponding to the i-th unit time interval; P i represents the signal difference ratio corresponding to the distortion degree of the voice input signal corresponding to the i-th unit time interval; P b represents the standard deviation of the signal difference ratios corresponding to the distortion degrees for n unit time intervals; S b represents the standard deviation of the signal-to-noise ratios of the voice input signals for n unit time intervals;

[0099] Obtain a second voice quality evaluation coefficient by using the jitter parameter and the drift parameter;

[0100] Among them, the second voice quality evaluation coefficient is obtained through the following formula:

[0101]

[0102] Among them, Q 02 represents the second voice quality evaluation coefficient; n represents the number of unit time intervals included in the voice input signal; J i and Y i represent the jitter parameter and the drift parameter of the voice input signal corresponding to the i-th unit time interval; J i+1 and Y i+1 represent the jitter parameter and the drift parameter of the voice input signal corresponding to the (i + 1)-th unit time interval; J b and Y b represent the standard deviations of the jitter parameter and the drift parameter of the voice input signals for n unit time intervals;

[0103] Determine the quality of the voice input signal by using the first voice quality evaluation coefficient and the second voice quality evaluation coefficient.

[0104] The technical effects of the above technical solution are as follows: By real-time monitoring multiple signal characteristic parameters of the voice input signal (signal-to-noise ratio, signal difference ratio corresponding to distortion degree, jitter parameter, and drift parameter), this technical solution can comprehensively and meticulously evaluate the quality of the voice signal. These parameters cover the main quality characteristics of the voice signal, thus ensuring the accuracy and reliability of the evaluation results. This technical solution uses unit time (such as seconds, milliseconds, etc.) as the basic unit of evaluation. By monitoring and calculating the signal characteristic parameters within each unit time, it can reflect the change of the voice signal quality in real time. This dynamically adaptive evaluation method helps to promptly detect and handle quality problems in the voice signal. By calculating the first voice quality assessment coefficient (Q 01 ) and the second voice quality assessment coefficient (Q 02 ) through specific formulas, this technical solution converts the quality of the voice signal into specific numerical indicators. This quantitative evaluation standard not only facilitates the understanding and comparison of the quality differences of different voice signals but also provides a clear reference basis for subsequent voice signal processing. This technical solution uses statistical quantities such as standard deviation to reflect the distribution of signal characteristic parameters, thus simplifying the calculation process and improving the evaluation efficiency. At the same time, through the formulaic calculation method, this technical solution can quickly obtain the voice quality assessment coefficient, providing strong support for application scenarios such as real-time voice communication. The formulas and parameters in this technical solution can be adjusted and optimized according to actual application requirements. For example, different weight coefficients or thresholds can be set according to different voice coding standards, transmission conditions, or application scenarios to adapt to different evaluation needs. This technical solution mainly relies on the monitoring and calculation of signal characteristic parameters and does not require complex hardware support or additional signal processing steps. Therefore, it is relatively simple to implement and has a low cost, and is applicable to voice communication systems of various scales.

[0105] In summary, by real-time monitoring multiple signal characteristic parameters of the voice input signal and calculating the corresponding voice quality assessment coefficients, this technical solution can comprehensively and accurately evaluate the quality of the voice signal and has the advantages of high dynamic adaptability, quantitative evaluation standard, high efficiency, flexibility, and easy implementation. These technical effects provide strong technical support and guarantee for application scenarios such as real-time voice communication.

[0106] Specifically, determining the quality of the voice input signal by using the first voice quality assessment coefficient and the second voice quality assessment coefficient includes:

[0107] Normalizing the first voice quality assessment coefficient and the second voice quality assessment coefficient respectively by using a preset first assessment coefficient reference value and second assessment coefficient reference value to obtain the normalized first voice quality assessment coefficient and second voice quality assessment coefficient;

[0108] Compare the first voice quality evaluation coefficient and the second voice quality evaluation coefficient after the normalization process;

[0109] When the first voice quality evaluation coefficient is lower than the second voice quality evaluation coefficient, retrieve the first comprehensive evaluation model; when the first voice quality evaluation coefficient is higher than or equal to the second voice quality evaluation coefficient, retrieve the second comprehensive evaluation model;

[0110] Use the first comprehensive evaluation model to combine the first voice quality evaluation coefficient and the second voice quality evaluation coefficient to obtain a first comprehensive evaluation coefficient, and compare the first comprehensive evaluation coefficient with a preset first comprehensive coefficient threshold; when the first comprehensive evaluation coefficient exceeds the preset first comprehensive coefficient threshold, it is determined that the voice quality does not meet the standard;

[0111] Among them, the first comprehensive evaluation coefficient is obtained through the following formula:

[0112]

[0113] Among them, Z 01 represents the first comprehensive evaluation coefficient; Q x01 represents the first voice quality evaluation coefficient after the normalization process; Q x02 represents the second voice quality evaluation coefficient after the normalization process;

[0114] Use the second comprehensive evaluation model to combine the first voice quality evaluation coefficient and the second voice quality evaluation coefficient to obtain a second comprehensive evaluation coefficient, and compare the second comprehensive evaluation coefficient with a preset second comprehensive coefficient threshold; when the second comprehensive evaluation coefficient exceeds the preset second comprehensive coefficient threshold, it is determined that the voice quality does not meet the standard;

[0115] Among them, the second comprehensive evaluation coefficient is obtained through the following formula:

[0116]

[0117] Among them, Z 02 represents the second comprehensive evaluation coefficient; Q x01 represents the first voice quality evaluation coefficient after the normalization process; Q x02 represents the second voice quality evaluation coefficient after the normalization process.

[0118] The technical effects of the above technical solution are as follows: By normalizing the first speech quality evaluation coefficient and the second speech quality evaluation coefficient with the preset first evaluation coefficient reference value and second evaluation coefficient reference value, these two coefficients can be compared on the same scale. This normalization process helps to eliminate the comparison difficulties caused by different dimensions or value ranges between different evaluation coefficients, thereby improving the accuracy and effectiveness of comparison. This technical solution introduces the first comprehensive evaluation model and the second comprehensive evaluation model, which can calculate the first comprehensive evaluation coefficient and the second comprehensive evaluation coefficient respectively by combining the normalized first speech quality evaluation coefficient and the second speech quality evaluation coefficient. This comprehensive evaluation method not only considers the information of individual evaluation coefficients but also integrates the comprehensive effects of multiple evaluation coefficients, thus enhancing the comprehensiveness and accuracy of the evaluation. By comparing the first comprehensive evaluation coefficient and the second comprehensive evaluation coefficient with the preset first comprehensive coefficient threshold and second comprehensive coefficient threshold respectively, this technical solution can standardly determine whether the quality of the speech input signal meets the standard. This threshold comparison method not only simplifies the determination process but also improves the objectivity and consistency of the determination. This technical solution can comprehensively evaluate the speech input signal according to different quality characteristics (such as signal-to-noise ratio, distortion degree, jitter parameter, drift parameter, etc.). By adjusting the reference values in the normalization process, the parameters in the comprehensive evaluation model, and the thresholds, this technical solution can flexibly adapt to different application scenarios and evaluation requirements. This technical solution improves the efficiency and accuracy of speech quality evaluation through a formulaic calculation method and a standardized determination process. At the same time, since the comprehensive evaluation model can integrate the information of multiple evaluation coefficients, this technical solution has higher accuracy and reliability when evaluating the quality of complex speech signals. This technical solution mainly relies on formulaic calculations and a standardized determination process, without the need for complex hardware support or additional signal processing steps. Therefore, it is relatively simple to implement and has a low cost. At the same time, due to the good scalability of this technical solution, more evaluation coefficients or comprehensive evaluation models can be added according to actual needs to further improve the accuracy and comprehensiveness of the evaluation.

[0119] In summary, this technical solution realizes a comprehensive, accurate, and efficient evaluation of the quality of the speech input signal through methods such as normalization, comprehensive evaluation model, and threshold comparison. These technical effects provide strong technical support and guarantee for application fields such as voice communication and speech recognition.

[0120] In this embodiment, before executing the control strategy in step three, it further includes confirming and interacting with the user through voice to ensure the accuracy of the control operation. The specific steps are as follows:

[0121] Voice Feedback and Confirmation: After generating the control strategy, the expected operations are fed back to the user via voice, including key parameters such as the type of device, operation instructions, and target device location. For example: "The air conditioner in the living room is about to be set to a temperature two degrees higher. Do you confirm?" The feedback content will include a summary of the understood user requirements to help the user quickly understand the upcoming operations;

[0122] User Confirmation and Adjustment: Receive the user's confirmation or adjustment instructions via voice. For example, the user can confirm with a simple "yes" or "no". If the user wishes to adjust the settings, they can further indicate the modification content, such as "lower it a bit" or "raise the temperature by five degrees". If the user's confirmation content is unclear, it will be automatically prompted and a clarification will be requested. For example: "May I ask how many degrees you want to raise the air conditioner temperature?";

[0123] Multi-round Interaction and Confirmation: When the user's instructions involve multiple devices and operation steps, step-by-step confirmation is used to ensure that each operation conforms to the user's intention. For example: "The instruction 'turn off the lights in the living room' is being executed. Next is the 'adjust temperature' operation. Do you confirm?" Optimize the user's operation experience based on the confirmation feedback after each step to ensure that each instruction is clear and accurate;

[0124] Personalized Confirmation: Dynamically adjust the confirmation process based on the user's emotional state. When it is recognized that the user's mood is relatively tense or anxious, accelerate the response through a faster and more concise confirmation interaction method. Conversely, when the user shows a relaxed or happy mood, a more gentle and slow interaction method is adopted to enhance the personalization of the user experience. For example, if the user's voice is accompanied by anxiety, the system may simplify the confirmation process and quickly execute the instruction: "I have raised the air conditioner temperature by two degrees to ensure your comfort." Specifically, it includes:

[0125] Obtain the intonation and speech rate characteristics of the user's voice, respond with a tone and mood that match the user's emotional state, and match the corresponding template sentences in the pre-stored voice database according to the user's emotional state for response. For example, provide encouragement or comfort when the user is in a low mood;

[0126] Classify the user's emotional state, determine the user's emotional category, match the corresponding interaction styles for different emotional state users for interaction, such as more lively, formal or warm, etc.; and adjust the execution priority of tasks according to the user's emotional state. For example, prioritize relevant tasks when the user is urgent or anxious;

[0127] Confirmation information feedback: Once the user's confirmed control strategy is received, the specific content and operation result to be executed will be clearly informed to the user through voice before the operation, further increasing the user's sense of transparency and control. For example, "The temperature of your air conditioner has been raised by two degrees, and then the system will adjust the brightness of the lights." After the operation is completed, continue to confirm the final result through voice and provide an operation review. For example, "The temperature of the living room air conditioner has been adjusted to the target value, and the operation is completed."

[0128] In this embodiment, the confirmation interaction with the user through voice ensures the accurate understanding of the user's intention, reduces misoperations, and at the same time takes into account the user's emotional state, provides a more user-friendly service, enhances the user's trust and satisfaction with the smart home system, improves the transparency of operations, enables the user to control the home environment more reassuringly and conveniently, and thus enhances the overall quality of life and the intelligent living experience.

[0129] In this embodiment, step five: feedback and continuous optimization of operation results, further includes:

[0130] Regularly ask the user about their satisfaction with the smart home system through voice, such as: "Are you satisfied with the recent smart home operation experience?" "Which functions do you hope to improve?" Summarize the satisfaction evaluation report to help the user understand their usage habits and improvement progress. For example, the report generated by the system every month can include content such as "In the past month, your satisfaction with the home temperature control system has increased by 15%", showing the effect of continuous improvement to the user;

[0131] Perform correlation analysis on the user's feedback data with the user's specific operations and system response data, and evaluate key indicators such as the success rate, accuracy rate, and response time of each specific operation, and generate a user satisfaction optimization strategy according to the user's usage habits and preferences;

[0132] Identify the deficiencies of certain functions based on the user's long-term feedback and needs, and promote the iterative update of the function. For example, if the user frequently feedbacks "inaccurate voice recognition" or "long control delay", then enhance the accuracy of voice recognition and the response speed through software updates.

[0133] In this embodiment, the continuous optimization mechanism ensures the synchronization of the system with the user's needs, effectively improves the user experience and performance of the smart home system, enhances user satisfaction and loyalty, and promotes the continuous improvement and development of smart home technology.

[0134] To better implement the smart home voice interaction control method, the present invention provides a smart home voice interaction control system, including:

[0135] A data acquisition module, which is used to collect the user's voice input signal and monitoring video data, and preprocess the collected voice input signal and monitoring video data;

[0136] A voice recognition and parsing module, which is used to perform text conversion on the preprocessed voice input signal, perform semantic parsing on the text based on natural language processing technology to determine operation instructions. At the same time, it verifies the recognition result of the voice input signal based on the monitoring video data, and generates prompt or clarification statements for fuzzy or incomplete instructions according to the recognition result, such as "Please specify which room's device to control?";

[0137] A control strategy generation module, which is used to match the operation logic of the target device according to the parsing result, generate a basic control strategy, and adjust the control strategy according to the user's historical operation data to conform to the user's habit preferences;

[0138] A device control module, which is used to establish a communication link with the target device, send operation instructions through the home Internet of Things protocol, and monitor the response status of the device in real time to ensure that the device operates as expected. If the device is abnormal (such as unable to respond to instructions), it triggers an exception handling mechanism and feedbacks to the user, such as "The air conditioner did not respond. Please check the connection.";

[0139] A feedback and optimization module, which is used to obtain user instruction and device response data, construct a user behavior data set, generate shortcut operation logic based on the user's frequently used instructions, such as "One-key enable the away mode", and collect evaluation feedback (such as operation satisfaction) from the user at a preset time interval. Optimize the control instructions based on the evaluation feedback to ensure that the system continuously learns and adapts to the user's needs, thus significantly improving the user experience, enhancing the intelligence of home automation and the quality of user interaction, and creating a more comfortable, convenient and intelligent living environment for the user.

[0140] In this embodiment, through accurate voice recognition and video data verification, the accuracy and reliability of the instructions are improved. At the same time, considering the user's habits and preferences, the personalization and convenience of the operation are enhanced.

[0141] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A smart home voice interaction control method, characterized in that: The following steps are involved: Step 1: Voice input acquisition and preprocessing: The user's voice input signal is acquired through an intelligent voice acquisition device, and the voice input signal is preprocessed, which also includes: Based on the matching of the acquired voice input signal with the preset wake-up words and phrases, triggering the wake-up monitoring data collection device; When collecting voice input signals, the user's non-voice information is collected based on the monitoring data collection device to obtain user monitoring data; Identify the user's facial expressions and hand movements based on the acquired user monitoring data, analyze the data features of non-voice information, and understand the user's emotional state and intentions; At the same time, the voice input signal and the non-voice information are synchronized based on the timestamp, the association relationship between the voice input signal and the non-voice information is established, and the semantics of the voice input signal is supplemented based on the non-voice information; Step 2: Recognition and analysis of voice commands: After collecting the voice input signal, monitor and evaluate the quality of the voice input signal in real time, perform text recognition on the pre-processed voice input signal, convert the recognition result into a text command, and perform semantic analysis on the text command to determine the user's intention and generate a user command; Among them, real-time monitoring and evaluation of the quality of voice input signals include: Real-time monitoring of signal characteristic parameters of the voice input signal, wherein the signal characteristic parameters include a signal-to-noise ratio, a signal difference ratio corresponding to a distortion, a jitter parameter, and a drift parameter; Obtaining a first speech quality assessment coefficient using a signal difference ratio corresponding to the signal-to-noise ratio and the distortion; Acquire a second voice quality assessment coefficient using the jitter parameter and the drift parameter; The first speech quality assessment coefficient is obtained by the following formula: Among them, Q 01 represents the first speech quality assessment coefficient; n represents the number of unit times contained in the speech input signal; S i+1 represents the signal-to-noise ratio of the speech input signal corresponding to the i+1th unit time; S i represents the signal-to-noise ratio of the speech input signal corresponding to the i-th unit time; P i represents the signal difference ratio corresponding to the distortion of the speech input signal corresponding to the i-th unit time; P b S represents the standard deviation of the signal difference ratio corresponding to the distortion corresponding to n unit time; b Represents the standard deviation of the signal-to-noise ratio of the speech input signal corresponding to n unit time; The second voice quality assessment coefficient is obtained by the following formula: Among them, Q 02 represents the second speech quality assessment coefficient; n represents the number of unit time contained in the speech input signal; J i and Y i J represents the jitter parameter and drift parameter of the speech input signal corresponding to the i-th unit time; i+1 and Y i+1 represents the jitter parameter and drift parameter of the speech input signal corresponding to the i+1th unit time; J b and Y b Indicates the jitter parameter standard deviation and drift parameter standard deviation of the speech input signal corresponding to n unit time; Determining the quality of a speech input signal using the first speech quality assessment coefficient and the second speech quality assessment coefficient; Step 3: Generation of device control strategy: Match the corresponding device control strategy according to the generated user instructions, determine the user needs and the operation logic of the target device, and adjust the device control strategy according to the user's historical operation habits and preferences; Step 4: Execution and real-time monitoring of equipment operations: Establish a communication link with the target equipment according to the equipment control strategy, send operation instructions to the target equipment, and monitor the response of the equipment in real time; Step 5: Feedback and continuous optimization of operation results: After the device operation is completed, the execution results will be fed back to the user through voice, and the user operation data will be recorded.

2. The smart home voice interaction control method according to claim 1, characterized in that: The step 2: recognition and analysis of voice commands, specifically includes: Effective speech extraction: After collecting the speech input signal, the endpoint detection technology is used to remove the invalid silent parts before and after the speech, retain the effective speech information, and split the long speech input signal into multiple logical segments; Voice quality assessment: monitor and evaluate the quality of voice input signals in real time. If the voice quality does not meet the standards, optimize the parameters or prompt the user to re-enter the voice. At the same time, dynamically adjust the microphone acquisition parameters according to environmental factors. Text instruction generation: Perform speech recognition processing on qualified speech input signals, convert the recognition results into text information, and correct the noise, accent or ambiguous pronunciation in the speech input signals; Semantic parsing: Perform word segmentation on the converted text information, split the text instructions into words or phrases, identify the part of speech of each word based on the semantics of the supplementary voice input signal, determine the grammatical relationship between words, determine the user's operation intention based on the parsing results, and identify key entities in the text; Contextual understanding: After each round of conversation, the system records the relevant information and status of the user's input, understands the context of continuous voice commands, maintains the continuity of the conversation, and infers the user's further needs based on the previous commands and operation results; User instruction generation: Generate structured user instructions based on the recognition results of semantic parsing and the inference results of context understanding.

3. The smart home voice interaction control method according to claim 2, characterized in that: Determining the quality of a speech input signal by using the first speech quality assessment coefficient and the second speech quality assessment coefficient comprises: Using a preset first evaluation coefficient reference value and a preset second evaluation coefficient reference value to respectively perform normalization processing on the first speech quality evaluation coefficient and the second speech quality evaluation coefficient, and obtain the normalized first speech quality evaluation coefficient and the second speech quality evaluation coefficient; Comparing the normalized first speech quality assessment coefficient with the second speech quality assessment coefficient; When the first speech quality assessment coefficient is lower than the second speech quality assessment coefficient, the first comprehensive assessment model is retrieved; when the first speech quality assessment coefficient is higher than or equal to the second speech quality assessment coefficient, the second comprehensive assessment model is retrieved; Using the first comprehensive evaluation model in combination with the first speech quality evaluation coefficient and the second speech quality evaluation coefficient to obtain a first comprehensive evaluation coefficient, and comparing the first comprehensive evaluation coefficient with a preset first comprehensive coefficient threshold; when the first comprehensive evaluation coefficient exceeds the preset first comprehensive coefficient threshold, determining that the speech quality does not meet the standard; The first comprehensive evaluation coefficient is obtained by the following formula: Among them, Z 01 Indicates the first comprehensive evaluation coefficient; Q x01 represents the first speech quality assessment coefficient after normalization; Q x02 represents the second speech quality assessment coefficient after normalization; Using the second comprehensive evaluation model in combination with the first speech quality evaluation coefficient and the second speech quality evaluation coefficient to obtain a second comprehensive evaluation coefficient, and comparing the second comprehensive evaluation coefficient with a preset second comprehensive coefficient threshold; when the second comprehensive evaluation coefficient exceeds the preset second comprehensive coefficient threshold, determining that the speech quality does not meet the standard; The second comprehensive evaluation coefficient is obtained by the following formula: Among them, Z 02 Indicates the second comprehensive evaluation coefficient; Q x01 represents the first speech quality assessment coefficient after normalization; Q x02 It represents the second speech quality assessment coefficient after normalization.

4. The smart home voice interaction control method according to claim 3, characterized in that: The step 2: recognition and analysis of voice commands also includes emotional perception and contextual understanding of text information, identifying the user's emotional state by analyzing the intonation and speed characteristics of the user's voice based on multimodal data, combining the data characteristics of non-voice information obtained simultaneously, and adjusting the device control strategy according to the user's emotional state.

5. The smart home voice interaction control method according to claim 4, characterized in that: The step three, before executing the control strategy, also includes confirming interaction with the user through voice to ensure the accuracy of the control operation. The specific steps are as follows: Voice feedback and confirmation: After the control strategy is generated, the expected operation is fed back to the user through voice. The feedback content includes a summary of the understood user needs, helping the user to quickly understand the operation to be performed; User confirmation and adjustment: Receive the user's confirmation or adjustment instructions through voice. If the user's confirmation content is unclear, it will automatically prompt and ask for clarification; Multi-round interaction and confirmation: If the user's instructions involve multiple devices and operation steps, ensure that each operation is in line with the user's intention through step-by-step confirmation, and optimize the user's operation experience based on the confirmation feedback after each step is completed; Personalized confirmation: Dynamically adjust the confirmation process based on the user’s emotional state; Confirmation information feedback: Once the control strategy is confirmed by the user, the user will be clearly informed of the specific content and operation results to be performed by voice before the operation is performed. After the operation is completed, the final result will continue to be confirmed by voice and an operation review will be provided.

6. The smart home voice interaction control method according to claim 5, characterized in that: The personalized confirmation specifically includes: Acquire the intonation and speech speed features of the user's voice, respond with the intonation and tone that matches the user's emotional state, and respond by matching the corresponding template sentence in the pre-stored voice database according to the user's emotional state; Classify the user's emotional state, determine the user's emotional category, match the corresponding interaction style to users in different emotional states, and adjust the execution priority of the task according to the user's emotional state.

7. The smart home voice interaction control method according to claim 6, characterized in that: The step 5: feedback and continuous optimization of operation results also includes: Regularly ask users about their satisfaction with the smart home system through voice, summarize satisfaction evaluation reports, and help users understand their usage habits and improvement progress; Conduct correlation analysis based on user feedback data, user specific operations and system response data, evaluate key indicators of each specific operation, and generate user satisfaction optimization strategies based on user usage habits and preferences; Identify the shortcomings of certain functions based on users' long-term feedback and needs, and promote iterative updates of the functions.

8. A smart home voice interaction control system, applied in the smart home voice interaction control method according to claim 7, characterized in that: include: The data acquisition module is used to collect the user's voice input signal and monitoring video data, and pre-process the collected voice input signal and monitoring video data; The speech recognition and analysis module is used to convert the pre-processed speech input signal into text, perform semantic analysis on the text based on natural language processing technology, determine the operation instructions, and verify the recognition results of the speech input signal based on the monitoring video data, and generate prompts or clarification statements for ambiguous or incomplete instructions based on the recognition results; The control strategy generation module is used to match the operation logic of the target device according to the analysis results, generate the basic control strategy, and adjust the control strategy according to the user's historical operation data; The device control module is used to establish a communication link with the target device, send operation instructions through the home IoT protocol, and monitor the response status of the device in real time. If the device is abnormal, the exception handling mechanism is triggered and feedback is given to the user; The feedback optimization module is used to obtain user instructions and device response data, build a user behavior data set, generate shortcut operation logic based on the user's high-frequency usage instructions, collect evaluation feedback from the user at preset time intervals, and optimize the control instructions based on the evaluation feedback.

Citation Information

Patent Citations

  • Smart home voice interaction control method and system

    CN116758914A

  • Sound box interaction method and sound box system

    CN118629380A