Display voice interaction system and method

Through the multimodal fusion technology of microphone array and RNN-T model combining touch and gesture, the problem of insufficient context association and cross-device collaboration in smart home interaction is solved, and more accurate voice data conversion and personalized control are achieved, improving user experience.

CN120340481AActive Publication Date: 2025-07-18SHENZHEN AO MIHOO ELECTRONICS

Patent Information

Application Number
CN202510406214.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The existing smart home interaction methods lack the ability to integrate complex multi-round dialogue, contextual association and multi-modal inputs, and lack cross-device collaboration, resulting in poor user experience.

Method used

The microphone array is used to obtain speech data, use the improved RNN-T speech recognition model to convert text, and combine touch and gestures to perform multi-modal fusion. The context weight is adjusted through incremental learning algorithms, and a semantic protocol is established to work in collaboration with smart home devices.

Benefits of technology

It realizes more accurate voice data conversion and multi-modal control, generates target control instructions that meet user needs, and improves the efficiency of equipment collaborative work and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340481A_ABST
    Figure CN120340481A_ABST
Patent Text Reader

Abstract

The invention relates to the field of display voice interaction, in particular to a display voice interaction system. The method comprises the following steps: acquiring voice data in a real-time environment through a microphone array, converting the voice data into text data by using an improved RNN-T voice recognition model, and marking a timestamp and confidence to obtain an initial control instruction; acquiring historical dialogue data, associating the initial control instruction with the historical dialogue data, dynamically adjusting a context weight through an incremental learning algorithm, taking a display as a main node, establishing a semantic protocol with the smart home equipment, analyzing the target control instruction, distributing the target control instruction to associated equipment, and synchronizing a context; and analyzing feedback logs in the display and the smart home equipment. The setting of the equipment can be automatically adjusted, personalized intelligent control of the display and the equipment is realized, and the use experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent interaction, and particularly to a display voice interaction system and method. Background Art

[0002] In the prior art, users' requirements for the interaction experience of smart homes are also getting higher and higher. The traditional smart home interaction methods mainly rely on single remote control or simple touch operations, which have many limitations. For example, most display voice interaction systems are based on single-round conversations or simple instruction responses, lacking the ability to dynamically integrate complex multi-round conversations, context associations, and multi-modal inputs (voice + touch / gesture). Although the multi-round conversation ability is improved through historical conversation data in the prior art, the problem of real-time scene dynamic adaptability has not been solved; and it relies on Bluetooth and Wi-Fi transmissions, and the semantic coherence of cross-device collaboration has not been achieved. Summary of the Invention

[0003] The purpose of the present invention is to solve the above problems and design a display voice interaction system and method.

[0004] Further, in the above-mentioned display voice interaction system, the display voice interaction system includes the following modules:

[0005] A voice data acquisition module, configured to obtain voice data in the real-time environment through a microphone array, convert the voice data into text data by using an improved RNN-T speech recognition model, and mark the time stamp and confidence level to obtain a voice command;

[0006] A multi-modal fusion module, configured to judge the priorities of the voice command, the display touch operation, and the user gesture, where the voice command has the highest weight, and the touch and gesture are auxiliary, to obtain an initial control command;

[0007] A dynamic perception module, configured to obtain historical conversation data, associate the initial control command with the historical conversation data, and dynamically adjust the context weight through an incremental learning algorithm to obtain a target control command;

[0008] A device collaboration execution module, configured to use the display as the main node, establish a semantic protocol with smart home devices, parse the target control command, distribute it to associated devices, and synchronize the context;

[0009] An execution feedback module, configured to analyze the feedback logs in the display and smart home devices, record the high-frequency instruction periods and response preferences of users, predict potential demands through a collaborative filtering algorithm, and perform interactive display in the display.

[0010] Further, in the above display voice interaction system, the voice data acquisition module includes the following sub-modules:

[0011] The data acquisition sub-module is used to acquire voice data in the real-time environment through a microphone array and extract Mel Frequency Cepstral Coefficient (MFCC) features from the collected voice signals;

[0012] The feature extraction sub-module is used to set the voice data as x m (t), where m = 1, 2, ···, M, t = 1, 2, ···, T, M represents the number of microphones in the microphone array, T represents the total acquisition duration, and the nth frame and the kth MFCC feature extracted from x m (t) is f m,n,k , where n = 1, 2, ···, N, k = 1, 2, ···, K, N represents the total number of frames, and k represents the dimension of the MFCC feature;

[0013] The model optimization sub-module is used to add a CNN (Convolutional Neural Network) to the RNN-T (Recurrent Neural Network - Transducer) speech recognition model to perform local feature extraction on the voice data; Let the convolutional kernel of the CNN layer be W c , and the bias be b c , and the output after the convolution operation is h c,n , and the calculation formula is:

[0014] b c,n = ReLU(W c f m,n,k + b c )

[0015] where ReLU(x) = max(0, x) is the activation function.

[0016] Further, in the above display voice interaction system, the voice data acquisition module further includes the following sub-modules:

[0017] The model optimization sub-module is used to input the output h c,n of the CNN convolutional layer into the RNN layer in the RNN-T speech recognition model, and the final output of the encoder is H = [h1, h2, ···, h n , where h n represents the hidden state of the LSTM in the RNN layer at the nth moment;

[0018] The data generation sub-module is used to generate the first text data using the prediction network in the RNN layer. Let the input of the prediction network at time u be y u-1 , and the hidden state be g u , and the calculation formula is:

[0019] g u = LSTM(yu-1 , g u-1 )

[0020] A feature fusion sub-module, which is used to fuse the output h of the encoder and the output g of the prediction network based on the joint network to obtain a joint feature j. The calculation formula is: n and the output g of the prediction network u to obtain a joint feature j n,u , and the calculation formula is:

[0021] j n,u = ReLU(W j [h n ; g u + b j )

[0022] where [h n ; g u means concatenating h n and g u together, W j represents the weight matrix, b j represents the bias vector, and the probability distribution of each time step and each vocabulary is calculated through the softmax function;

[0023] An instruction generation sub-module, which is used to find the most likely text sequence from the probability distribution using Beam Search beam search to obtain the target text data, record the encoder time step corresponding to each vocabulary in the target text data, convert the time step into an actual timestamp according to the frame length and frame shift during feature extraction, and calculate the confidence through the product of the probabilities of the decoding path to obtain the voice instruction.

[0024] Furthermore, in the above display voice interaction system, the multi-modal fusion module includes the following sub-modules:

[0025] An operation classification sub-module, which is used to classify the display touch operations according to the touch position, duration, and sliding direction features, analyze the user gesture images collected by the image depth sensor using a computer image recognition algorithm, identify the gesture types, and map the touch operations and gesture types to manipulation intentions;

[0026] A priority judgment sub-module, which is used to judge the priorities of the voice instruction, the display touch operation, and the user gesture. If there is a conflict between the voice instruction and the manipulation intention within the same time period, it is determined as a conflict situation;

[0027] An instruction obtaining sub-module, which is used to, in case of a conflict situation, preferentially adopt the voice instruction. If there are more than 3 conflicts in the voice instruction, it is judged in combination with historical instructions and context information. If the voice instruction conflicts with other input methods, the voice instruction is adopted to obtain the initial manipulation instruction.

[0028] Furthermore, in the above display voice interaction system, the dynamic perception module includes the following units:

[0029] An instruction receiving unit, which is configured to, after receiving an initial control instruction, query relevant historical conversation data from a database according to the user ID and time range, perform word segmentation on the initial control instruction and the historical conversation content, split the sentence into individual words, label the part-of-speech and entity recognition for each word segmentation, and obtain a word segmentation result;

[0030] A similarity calculation unit, which is configured to, according to the word segmentation result, obtain common keywords in the initial control instruction and the historical conversation, and use cosine similarity to calculate the semantic similarity between the initial control instruction and the historical conversation;

[0031] An instruction screening unit, which is configured to screen out historical conversation data with a high degree of association with the initial control instruction according to a set association threshold to obtain a prediction result;

[0032] A weighted fusion unit, which is configured to calculate the loss between the prediction result and the actual result according to the current context weight and features, update the context weight according to the gradient of the loss function, and perform weighted fusion on the associated historical conversation data and the initial control instruction according to the updated context weight to obtain a target control instruction.

[0033] Furthermore, in the above display voice interaction system, the device collaborative execution module includes the following units:

[0034] A protocol establishment unit, which is configured to use the display as the main node to design a hierarchical semantic protocol architecture with a smart home device, including at least an application layer, a transport layer, and a physical layer;

[0035] An instruction parsing unit, which is configured to define the data format of the target control specification and the status information fed back by the smart home device to the display, and parse the target control instruction to obtain a device identification field, an action parsing field, and a parameter extraction field;

[0036] An instruction execution unit, which is configured to obtain a smart home device according to the device identification field, perform corresponding operations on the smart home device according to the action parsing field and the parameter extraction field, and record the execution situation of the target control instruction, including instruction content, execution time, and execution result.

[0037] Furthermore, in the above display voice interaction system, the execution feedback module includes the following units:

[0038] A feedback data collection unit, which is configured to collect feedback logs in the display and the smart home device, including at least instruction execution time, instruction content, execution result, and device status;

[0039] A similarity calculation unit, which is used to assume that there are u users and v different instructions. Then, the usage frequency of instruction j by user i is represented by R i,j and calculate the similarity between users by using cosine similarity. For user i and user k, the calculation formula for their similarity sim(i,k) is:

[0040]

[0041] where R k,j represents the usage frequency of instruction j by user k, represents the sum of the products of the usage frequencies of user i and user k on all instructions, and respectively represent the square roots of the sum of the squares of the usage frequencies of user i and user k on all instructions;

[0042] A prediction calculation unit, which is used to find the N users with the highest similarity to the target user i, denoted as the neighbor user set N(i). For an unused instruction j, predict the potential demand score P of the target user i for instruction j ij , and the calculation formula is:

[0043]

[0044] where ∑ k∈N(i) sim(i,k)R i,j represents the weighted sum of the potential demands of the target user i based on the neighbor user set N(i), and the weight is the similarity sim(i,k) between users; ∑ k∈N(i) sim(i,k) represents the sum of the similarities between the users in the neighbor user set N(i) and the target user i;

[0045] A demand display unit, which is used to convert the analysis result and the predicted potential demand into natural language text, use the TTS speech synthesis technology to convert the natural language text into voice output, and display the analysis result and the potential demand prediction in a visual manner on the display.

[0046] Furthermore, in the above display voice interaction method, the display voice interaction method includes the following steps:

[0047] Obtain voice data in the real-time environment through a microphone array, convert the voice data into text data by using an improved RNN-T speech recognition model, and mark the time stamp and confidence level to obtain voice instructions;

[0048] Judge the priorities of the voice instructions, display touch operations and user gestures, where the voice instructions have the highest weight and the touch and gestures are auxiliary to obtain initial manipulation instructions;

[0049] Obtain historical dialogue data, associate the initial manipulation instruction with the historical dialogue data, and dynamically adjust the context weight through an incremental learning algorithm to obtain a target manipulation instruction;

[0050] Use the display as the main node, establish a semantic protocol with the smart home device, parse the target manipulation instruction, distribute it to the associated device, and synchronize the context;

[0051] Analyze the feedback logs in the display and the smart home device, record the high-frequency instruction time period and response preference of the user, predict potential needs through a collaborative filtering algorithm, and perform interactive display in the display.

[0052] Furthermore, in the above display voice interaction method, the obtaining of voice data in the real-time environment through the microphone array includes:

[0053] Obtain voice data in the real-time environment through the microphone array, and extract Mel Frequency Cepstral Coefficient (MFCC) features from the collected voice signal;

[0054] Let the voice data be x m (t), where m = 1, 2, ···, M, t = 1, 2, ···, T, M represents the number of microphones in the microphone array, T represents the total acquisition duration, and the nth frame and the kth MFCC feature extracted from x m (t) is f m,n,k , where n = 1, 2, ···, N, k = 1, 2, ···, K, N represents the total number of frames, and k represents the dimension of the MFCC feature;

[0055] Add a CNN (Convolutional Neural Network) to the RNN-T (Recurrent Neural Network - Transducer) speech recognition model to perform local feature extraction on the voice data; Let the convolutional kernel of the CNN layer be W c , and the bias be b c , and the output after the convolution operation is h c,n , and the calculation formula is:

[0056] h c,n = ReLU(W c f m,n,k + b c )

[0057] where ReLU(x) = max(0, x) is the activation function.

[0058] Furthermore, in the above display voice interaction method, it is characterized in that the converting of the voice data into text data by using an improved RNN-T speech recognition model, and marking the timestamp and confidence to obtain a voice instruction includes:

[0059] The output h of the CNN convolutional layerc,n Input into the RNN layer of the RNN-T speech recognition model, the final output of the encoder is H = [h1, h2, ···, h n , where h n represents the hidden state of the LSTM in the RNN layer at the nth moment;

[0060] Use the prediction network in the RNN layer to generate the first text data. Let the input of the prediction network at time u be y u-1 , and the hidden state be g u , and the calculation formula is:

[0061] g u = LSTM(y u-1 , g u-1 )

[0062] Based on the joint network, fuse the output h n of the encoder and the output g u of the prediction network to obtain the joint feature j n,u , and the calculation formula is:

[0063] j n,u = ReLU(W j [h n ; g u · + b j )

[0064] Among them, [h n ; g u means concatenating h n and g u together, W j represents the weight matrix, b j represents the bias vector, and calculate the probability distribution of each time step and each vocabulary through the softmax function;

[0065] Use Beam Search to find the most likely text sequence from the probability distribution to obtain the target text data. Record the encoder time step corresponding to each vocabulary in the target text data. According to the frame length and frame shift during feature extraction, convert the time step to the actual timestamp, and calculate the confidence through the product of the probabilities of the decoding path to obtain the voice command.

[0066] Its beneficial effects are as follows: 1. Provide more precise control through touch and gestures. Among them, voice commands have the highest weight, ensuring that users can quickly issue commands through voice in most cases, improving operation efficiency; touch and gestures, as supplements, enrich the interaction methods and meet the diverse needs of users. 2. Can more accurately convert voice data in the real-time environment into text data. At the same time, the operations of marking timestamps and confidence levels enable the system to more precisely evaluate the source and accuracy of voice commands. 3. Can generate target manipulation commands that better meet the user's needs based on the user's historical operation habits and current context information. 4. Can accurately parse the target manipulation commands and distribute them to associated devices, while synchronizing context information. This enables better cooperation between different devices and improves the overall operation efficiency of the display system. 5. Can anticipate the user's needs in advance and provide more personalized services. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention.

[0068] Figure 1 Schematic diagram of the first embodiment of a display voice interaction system in an embodiment of the present invention;

[0069] Figure 2 Schematic diagram of the second embodiment of a display voice interaction system in an embodiment of the present invention;

[0070] Figure 3 Schematic diagram of the third embodiment of a display voice interaction system in an embodiment of the present invention;

[0071] Figure 4 Schematic diagram of the first embodiment of a display voice interaction method in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0072] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not used to limit the present invention.

[0073] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0074] The present invention will be specifically described below with reference to the accompanying drawings. As Figure 1 shown, a display voice interaction system, the display voice interaction system includes the following modules:

[0075] A voice data acquisition module, configured to obtain voice data in the real-time environment through a microphone array, convert the voice data into text data by using an improved RNN-T voice recognition model, and mark time stamps and confidence levels to obtain a voice command;

[0076] Specifically, this embodiment further includes

[0077] A data acquisition sub-module, configured to obtain voice data in the real-time environment through a microphone array, and perform Mel-frequency cepstral coefficient feature extraction on the collected voice signal;

[0078] A feature extraction sub-module, configured to set the voice data as x m (t), where m = 1, 2, ···, M, t = 1, 2, ···, T, M represents the number of microphones in the microphone array, T represents the total acquisition duration, and the nth frame and the kth MFCC feature extracted from x m (t) are f m,n,k , where n = 1, 2, ···, N, k = 1, 2, ···, K, N represents the total number of frames, and k represents the dimension of the MFCC feature;

[0079] A model optimization sub-module, configured to add a CNN convolutional neural network to the RNN-T voice recognition model to perform local feature extraction on the voice data; set the convolutional kernel of the CNN layer as W c , and the bias as b c , and the output after the convolutional operation is h c,n , and the calculation formula is:

[0080] h c,n = ReLU(W c f m,n,k + b c )

[0081] Where ReLU(x) = max(0, x) is the activation function.

[0082] The model optimization sub-module is used to input the output h of the CNN convolutional layer c,n into the RNN layer in the RNN-T speech recognition model. The final output of the encoder is H = [h1, h2, ···, h n , where h n represents the hidden state of the LSTM in the RNN layer at the nth moment;

[0083] The data generation sub-module is used to generate the first text data by using the prediction network in the RNN layer. Let the input of the prediction network at time u be y u-1 , and the hidden state be g u . The calculation formula is:

[0084] g u = LSTM(y u-1 , g u-1 )

[0085] The feature fusion sub-module is used to fuse the output h of the encoder n and the output g of the prediction network u based on the joint network to obtain the joint feature j n,u . The calculation formula is:

[0086] j n,u = ReLU(W j [h n ; g u +b j )

[0087] where, [h n ; g u means concatenating h n and g u together. W j represents the weight matrix, and b j represents the bias vector. The probability distribution of each time step and each vocabulary is calculated through the softmax function;

[0088] The instruction generation sub-module is used to find the most likely text sequence from the probability distribution by using Beam Search beam search to obtain the target text data, record the encoder time step corresponding to each vocabulary in the target text data, convert the time step into the actual timestamp according to the frame length and frame shift during feature extraction, and calculate the confidence through the product of the probabilities of the decoding path to obtain the voice instruction.

[0089] The multi-modal fusion module is used to judge the priorities of the voice instruction, the display touch operation, and the user gesture. Among them, the voice instruction has the highest weight, and the touch and gesture are auxiliary to obtain the initial control instruction;

[0090] Specifically, this embodiment also includes

[0091] An operation classification sub-module, which is used to classify the touch operations on the display according to the touch position, duration, and sliding direction features, analyze the user gesture images collected by the image depth sensor using a computer image recognition algorithm, identify the gesture types, and map the touch operations and gesture types to manipulation intentions;

[0092] A priority judgment sub-module, which is used to judge the priorities of the voice commands, display touch operations, and user gestures. In the same time period, if a voice command conflicts with the manipulation intention, it is determined as a conflict situation;

[0093] An instruction obtaining sub-module, which is used to, in case of a conflict situation, preferentially adopt the voice command. If there are more than 3 mutually conflicting voice commands, it combines historical commands and context information for judgment. If the voice command conflicts with other input methods, it adopts the voice command to obtain an initial manipulation command.

[0094] A dynamic perception module, which is used to obtain historical conversation data, associate the initial manipulation command with the historical conversation data, and dynamically adjust the context weight through an incremental learning algorithm to obtain a target manipulation command;

[0095] Specifically, this embodiment further includes

[0096] An instruction receiving unit, which is used to, after receiving the initial manipulation command, query relevant historical conversation data from the database according to the user ID and time range, perform word segmentation on the initial manipulation command and the historical conversation content, split the sentence into individual words, label the part of speech and entity recognition for each word segmentation, and obtain a word segmentation result;

[0097] A similarity calculation unit, which is used to, according to the word segmentation result, obtain the common keywords between the initial manipulation command and the historical conversation, and use cosine similarity to calculate the semantic similarity between the initial manipulation command and the historical conversation;

[0098] An instruction screening unit, which is used to screen out the historical conversation data with a high degree of association with the initial manipulation command according to a set association threshold to obtain a prediction result;

[0099] A weighted fusion unit, which is used to calculate the loss between the prediction result and the actual result according to the current context weight and features, update the context weight according to the gradient of the loss function, and perform weighted fusion on the associated historical conversation data and the initial manipulation command according to the updated context weight to obtain a target manipulation command.

[0100] A device collaborative execution module, which is used to use the display as the main node, establish a semantic protocol with smart home devices, parse the target manipulation command, distribute it to associated devices, and synchronize the context;

[0101] Specifically, this embodiment further includes

[0102] A protocol establishment unit, which takes the display as the master node and designs a hierarchical semantic protocol architecture with the smart home device, including at least an application layer, a transport layer, and a physical layer;

[0103] An instruction parsing unit, which defines the data format of the target control instruction and the status information feedback from the smart home device to the display, parses the target control instruction, and obtains a device identification field, an action parsing field, and a parameter extraction field;

[0104] An instruction execution unit, which obtains the smart home device according to the device identification field, performs corresponding operations on the smart home device according to the action parsing field and the parameter extraction field, and records the execution situation of the target control instruction, including the instruction content, the execution time, and the execution result.

[0105] An execution feedback module, which analyzes the feedback logs in the display and the smart home device, records the user's high-frequency instruction time period and response preference, predicts potential demands through a collaborative filtering algorithm, and performs interactive display in the display.

[0106] Specifically, this embodiment further includes

[0107] A feedback data collection unit, which collects the feedback logs in the display and the smart home device, including at least the instruction execution time, the instruction content, the execution result, and the device status;

[0108] A similarity calculation unit, which assumes that there are u users and v different instructions, and the usage frequency of user i for instruction j is represented by R i,j It is used to calculate the similarity between users by using the cosine similarity. For users i and k, the calculation formula for their similarity sim(i,k) is:

[0109]

[0110] Among them, R k,j represents the usage frequency of user k for instruction j, represents the sum of the products of the usage frequencies of user i and user k for all instructions, and respectively represent the square root of the sum of the squares of the usage frequencies of user i and user k for all instructions;

[0111] A prediction calculation unit, which finds the N users with the highest similarity to the target user i, denoted as the neighbor user set N(i), and for the instruction j that has not been used, predicts the potential demand score P of the target user i for instruction j ij, the calculation formula is:

[0112]

[0113] Among them, ∑ k∈N(i) sim(i,k)R i,j represents the weighted sum of the potential demands of the target user i based on the neighbor user set N(i), and the weight is the similarity sim(i,k) between users; ∑ k∈N(i) sim(i,k) represents the sum of the similarities between the users in the neighbor user set N(i) and the target user i;

[0114] A demand display unit, configured to convert the analysis result and the predicted potential demand into natural language text, use the TTS speech synthesis technology to convert the natural language text into voice output, and visually display the analysis result and the potential demand prediction on a display.

[0115] Its beneficial effects are as follows: 1. Provide more refined control through touch and gestures. Among them, the voice command has the highest weight, ensuring that in most cases, users can quickly issue commands through voice, improving the operation efficiency; touch and gestures are used as auxiliary means to enrich the interaction method and meet the diverse needs of users. 2. Can more accurately convert the voice data in the real-time environment into text data. At the same time, the operation of marking the timestamp and confidence level enables the system to more precisely evaluate the source and accuracy of the voice command. 3. Can generate target manipulation commands that more meet the user's needs according to the user's historical operation habits and current context information. 4. Can accurately parse the target manipulation command and distribute it to the associated device, and synchronize the context information at the same time. Enable better cooperation between different devices and improve the overall operation efficiency of the smart home system. 5. Can understand the user's needs in advance and provide more personalized services for the user.

[0116] In this embodiment, please refer to Figure 2 , the second embodiment of a display voice interaction system in the embodiment of the present invention, the voice data acquisition module includes the following sub-modules:

[0117] A model optimization sub-module, configured to input the output h c,n of the CNN convolutional layer into the RNN layer in the RNN-T speech recognition model, and the final output of the encoder is H = [h1, h2, ···, h n , where h n represents the hidden state of the LSTM in the RNN layer at the nth moment;

[0118] A data generation sub-module, configured to generate first text data by using the prediction network in the RNN layer. Let the input of the prediction network at time u be y u-1 , and the hidden state be gu , the calculation formula is:

[0119] g u = LSTM(y u-1 , g u-1 )

[0120] A feature fusion sub-module, configured to fuse the output h n of the encoder and the output g u of the prediction network based on the joint network to obtain a joint feature j n,u , the calculation formula is:

[0121] j n,u = ReLU(W j [h n ; g u · + b j )

[0122] Among them, [h n ; g u means concatenating h n and g u together, W j represents a weight matrix, and b j represents a bias vector, and the probability distribution of each time step and each word is calculated through the softmax function;

[0123] An instruction generation sub-module, configured to use Beam Search to find the most likely text sequence from the probability distribution to obtain target text data, record the encoder time step corresponding to each word in the target text data, convert the time step into an actual timestamp according to the frame length and frame shift during feature extraction, and calculate the confidence through the probability product of the decoding path to obtain a voice command.

[0124] Its beneficial effect is that the system can more accurately evaluate the source and accuracy of voice commands. The order of commands can be understood through timestamps to better handle context relationships; while the confidence helps the system judge the reliability of commands, and corresponding measures can be taken when the confidence is low.

[0125] In this embodiment, please refer to Figure 3 , the third embodiment of a display voice interaction system in the embodiments of the present invention, the dynamic perception module includes the following sub-units:

[0126] An instruction receiving unit, configured to, after receiving an initial control instruction, query relevant historical conversation data from a database according to the user ID and time range, perform word segmentation on the initial control instruction and the historical conversation content, split the sentence into individual words, and label the part of speech and entity recognition for each word segmentation to obtain a word segmentation result;

[0127] A similarity calculation unit, which is used to obtain the common keywords between the initial control instruction and the historical conversation according to the word segmentation result, and calculate the semantic similarity between the initial control instruction and the historical conversation using cosine similarity;

[0128] An instruction screening unit, which is used to screen out the historical conversation data with a high degree of association with the initial control instruction according to the set association threshold to obtain a prediction result;

[0129] A weighted fusion unit, which is used to calculate the loss between the prediction result and the actual result according to the current context weight and features, update the context weight according to the gradient of the loss function, and perform weighted fusion on the associated historical conversation data and the initial control instruction according to the updated context weight to obtain a target control instruction.

[0130] Its beneficial effect is that, according to the user's previous operation records of smart home devices at different times, the settings of the devices are automatically adjusted to achieve personalized intelligent control of the display and the devices. It not only improves the user experience, but also better meets the user's needs in different scenarios, making the smart home system more intelligent and considerate.

[0131] The above describes a display voice interaction system provided by an embodiment of the present invention. Next, a display voice interaction method of an embodiment of the present invention will be described. Please refer to Figure 4 , an embodiment of the display voice interaction method in the embodiment of the present invention includes:

[0132] Step 401: Obtain voice data in the real-time environment through a microphone array, convert the voice data into text data using an improved RNN-T speech recognition model, and mark the time stamp and confidence level to obtain a voice instruction;

[0133] Step 402: Judge the priorities of the voice instruction, the display touch operation, and the user gesture, where the voice instruction has the highest weight and the touch and gesture are auxiliary to obtain an initial control instruction;

[0134] Step 403: Obtain historical conversation data, associate the initial control instruction with the historical conversation data, and dynamically adjust the context weight through an incremental learning algorithm to obtain a target control instruction;

[0135] Step 404: Use the display as the main node to establish a semantic protocol with the smart home device, parse the target control instruction, distribute it to the associated device, and synchronize the context;

[0136] Step 405: Analyze the feedback logs in the display and the smart home devices, record the high-frequency instruction periods and response preferences of the users, predict potential demands through a collaborative filtering algorithm, and perform interactive display in the display.

[0137] The above has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A display voice interaction system, characterized in that, The described display voice interaction system includes the following modules: A voice data acquisition module, which is used to obtain voice data in the real-time environment through a microphone array, convert the voice data into text data by using an improved RNN-T voice recognition model, and mark the timestamp and confidence level to obtain a voice command; A multi-modal fusion module, which is used to judge the priorities of the voice command, the display touch operation, and the user gesture. Among them, the weight of the voice command is the highest, and the touch and gesture are auxiliary, so as to obtain an initial control command; A dynamic perception module, which is used to obtain historical conversation data, associate the initial control command with the historical conversation data, and dynamically adjust the context weight through an incremental learning algorithm to obtain a target control command; A device collaborative execution module, which is used to use the display as the main node, establish a semantic protocol with the smart home device, parse the target control command, distribute it to the associated device, and synchronize the context; An execution feedback module, which is used to analyze the feedback logs in the display and the smart home device, record the high-frequency command period and response preference of the user, predict potential demands through a collaborative filtering algorithm, and perform interactive display in the display.

2. The voice interaction system for a display according to claim 1, characterized in that, The voice data acquisition module includes the following sub-modules: A data acquisition sub-module, which is used to obtain voice data in the real-time environment through a microphone array, and extract the Mel-frequency cepstral coefficient features of the collected voice signal; A feature extraction sub-module for setting the speech data as x m (t), where m = 1, 2, ···, M, t = 1, 2, ···, T, M represents the number of microphones in the microphone array, T represents the total acquisition duration, and the nth frame extracted from x m (t), the kth MFCC feature is f m,n,k , where n = 1, 2, ···, N, k = 1, 2, ···, K, N represents the total number of frames, and k represents the dimension of the MFCC feature; A model optimization sub-module, which is used to add a CNN (Convolutional Neural Network) to the RNN-T (Recurrent Neural Network - Transducer) speech recognition model to perform local feature extraction on the speech data; assume that the convolution kernel of the CNN layer is W c , and the bias is b c , and the output after the convolution operation is h c,n , and the calculation formula is: h c,n = ReLU(W c f m,n,k + b c ) Among them, ReLU(x) = max(0, x) is the activation function.

3. The voice interaction system of a display according to claim 1, wherein, The voice data acquisition module also includes the following sub-modules: A model optimization sub-module for inputting the output h of the CNN convolutional layer c,n into the RNN layer in the RNN-T speech recognition model, and the final output of the encoder is H = [h1, h2, ···, h n , where h n represents the hidden state of the LSTM in the RNN layer at the nth moment; A data generation sub-module, which is used to generate first text data by using a prediction network in an RNN layer. Let the input of the prediction network at time u be y u-1 , and the hidden state be g u . The calculation formula is as follows: g u = LSTM(y u-1 , g u-1 ) A feature fusion sub-module, which is used to fuse the output h of the encoder and the output g of the prediction network based on the joint network to obtain the joint feature j, and the calculation formula is: n and the output g of the prediction network u for fusion to obtain the joint feature j n,u , the calculation formula is: j n,u = RELU(W j [h n ; g u + b j ) Among them, [h n ; g u means concatenating h n and g u together, W j represents the weight matrix, b j represents the bias vector, and the probability distribution of each time step and each word is calculated through the softmax function; A command generation sub-module, which is used to use Beam Search to find the most likely text sequence from the probability distribution to obtain the target text data, record the encoder time step corresponding to each vocabulary in the target text data, convert the time step into an actual timestamp according to the frame length and frame shift during feature extraction, and calculate the confidence level through the probability product of the decoding path to obtain a voice command.

4. A display voice interaction system according to claim 1, characterized in that, The multi-modal fusion module includes the following sub-modules: An operation classification sub-module, which is used to classify the display touch operation according to the touch position, duration, and sliding direction features, analyze the user gesture image collected by the image depth sensor by using a computer image recognition algorithm, identify the gesture type, and map the touch operation and the gesture type to a control intention; A priority judgment sub-module, which is used to judge the priorities of the voice command, the display touch operation, and the user gesture. If a voice command conflicts with the control intention during the same time period, it is determined as a conflict situation; A command obtaining sub-module, which is used to, in case of a conflict situation, preferentially adopt the voice command. If there are more than 3 mutually conflicting voice commands, it is judged by combining the historical commands and context information. If the voice command conflicts with other input methods, the voice command is adopted to obtain an initial control command.

5. The voice interaction system for a display according to claim 1, wherein The dynamic perception module includes the following units: An instruction receiving unit, which is configured to, after receiving an initial control instruction, query relevant historical conversation data from a database according to the user ID and time range, perform word segmentation on the initial control instruction and the historical conversation content, split the sentence into individual words, label the part of speech and entity recognition for each word segment, and obtain a word segmentation result; A similarity calculation unit, which is configured to, according to the word segmentation result, obtain common keywords in the initial control instruction and the historical conversation, and use cosine similarity to calculate the semantic similarity between the initial control instruction and the historical conversation; An instruction screening unit, which is configured to screen out historical conversation data with a high degree of association with the initial control instruction according to a set association threshold to obtain a prediction result; A weighted fusion unit, which is configured to calculate the loss between the prediction result and the actual result according to the current context weight and features, update the context weight according to the gradient of the loss function, and perform weighted fusion on the associated historical conversation data and the initial control instruction according to the updated context weight to obtain a target control instruction.

6. The voice interaction system of a display according to claim 1, wherein The device collaborative execution module includes the following units: A protocol establishment unit, which is configured to use the display as the main node to design a hierarchical semantic protocol architecture with smart home devices, including at least an application layer, a transport layer, and a physical layer; An instruction parsing unit, which is configured to define the data format of the target control specification and the status information fed back by the smart home device to the display, and parse the target control instruction to obtain a device identification field, an action parsing field, and a parameter extraction field; An instruction execution unit, which is configured to obtain a smart home device according to the device identification field, perform corresponding operations on the smart home device according to the action parsing field and the parameter extraction field, and record the execution situation of the target control instruction, including the instruction content, the execution time, and the execution result.

7. The voice interaction system for a display according to claim 1, wherein The execution feedback module includes the following units: A feedback data collection unit, which is configured to collect feedback logs in the display and smart home devices, including at least the instruction execution time, the instruction content, the execution result, and the device status; Similarity calculation unit, which is used to assume that there are u users and v different instructions, and the usage frequency of instruction j by user i is represented by R i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j i,j Among them, R k,j represents the usage frequency of instruction j by user k, represents the sum of the products of the usage frequencies of all instructions by user i and user k, and respectively represent the square roots of the sum of the squares of the usage frequencies of all instructions by user i and user k; A prediction calculation unit is used to find the N users with the highest similarity to the target user i, denoted as the neighbor user set N(i). For an instruction j that has not been used, it predicts the potential demand score P of the target user i for the instruction j ij , and the calculation formula is: Among them, ∑ k∈N(i) sim(i,k)R i,j represents the weighted sum of the potential demands of the target user i based on the set of neighbor users N(i), where the weight is the similarity sim(i,k) between users; ∑ k∈N(i) sim(i,k) represents the sum of the similarities between the users in the set of neighbor users N(i) and the target user i; A requirement display unit, which is configured to convert the analysis result and the predicted potential requirements into natural language text, use TTS speech synthesis technology to convert the natural language text into voice output, and visually display the analysis result and the potential requirement prediction on the display.

8. A method for voice interaction of a display, characterized in that, The display voice interaction method includes the following steps: Obtain voice data in the real-time environment through a microphone array, use an improved RNN-T speech recognition model to convert the voice data into text data, and mark the time stamp and confidence level to obtain a voice instruction; Judge the priorities of the voice instruction, the display touch operation, and the user gesture, where the voice instruction has the highest weight, and the touch and gesture are auxiliary, to obtain an initial control instruction; Obtain historical conversation data, associate the initial control instruction with the historical conversation data, and dynamically adjust the context weight through an incremental learning algorithm to obtain a target control instruction; Use the display as the main node to establish a semantic protocol with smart home devices, parse the target control instruction, distribute it to the associated devices, and synchronize the context; Analyze the feedback logs in the display and smart home devices, record the high-frequency instruction periods and response preferences of users, predict potential demands through the collaborative filtering algorithm, and perform interactive display on the display.

9. The method for voice interaction of a display according to claim 8, wherein The acquisition of voice data in the real-time environment through the microphone array includes: Acquire voice data in the real-time environment through the microphone array, and extract the Mel Frequency Cepstral Coefficient (MFCC) features of the collected voice signals. Let the speech data be x m (t), where m = 1, 2, ···, M, t = 1, 2, ···, T, M represents the number of microphones in the microphone array, T represents the total acquisition duration, the nth frame extracted from x m (t), and the kth MFCC feature is f m,n,k , where n = 1, 2, ···, N, k = 1, 2, ···, K, N represents the total number of frames, and k represents the dimension of the MFCC feature; In the RNN-T speech recognition model, a CNN convolutional neural network is added to extract local features of the speech data; let the convolutional kernel of the CNN layer be W c , and the bias be b c , and the output after the convolution operation is h c,n , and the calculation formula is: h c,n = ReLU(W c f m,n,k + b c ) Among them, RELU(x) = max(0, x) is the activation function.

10. The method for voice interaction of a display according to claim 8, wherein, The conversion of the voice data into text data by using the improved RNN-T speech recognition model, and marking the timestamp and confidence level to obtain voice instructions includes: Input the output h of the CNN convolutional layer c,n into the RNN layer in the RNN-T speech recognition model. The final output of the encoder is H = [h1, h2, ···, h n , where h n represents the hidden state of the LSTM in the RNN layer at the nth moment; Generate the first text data using the prediction network in the RNN layer. Let the input of the prediction network at time u be y u-1 , and the hidden state be g u . The calculation formula is as follows: g u = LSTM(y u-1 , g u-1 ) Based on the joint network, fuse the output h of the encoder n and the output g of the prediction network u to obtain the joint feature j n,u , and the calculation formula is: j n,u = ReLU(W j [h n ; g u + b j ) Among them, [h n ; g u means concatenating h n and g u , W j represents the weight matrix, b j represents the bias vector, and the probability distribution of each time step and each word is calculated through the softmax function; Use Beam Search to find the most likely text sequence from the probability distribution to obtain the target text data, record the encoder time steps corresponding to each vocabulary in the target text data, convert the time steps into actual timestamps according to the frame length and frame shift during feature extraction, and calculate the confidence level through the product of the probabilities of the decoding paths to obtain voice instructions.

Citation Information

Patent Citations

  • Multi-modal man-machine interaction system and control method thereof

    CN106569613A

  • Holographic interactive teleconference method and system

    CN116760942A

  • Multi-mode composite man-machine interaction system and method

    CN117891341A

  • Multi-dimensional perception interaction method and system

    CN117930984A

  • Exhibition hall interface design and man-machine interaction method based on voice instruction

    CN118069090A

Cited By

  • Audio and video player control method based on voice instruction

    CN121053987A