Voice interaction method and device, electronic equipment and readable storage medium
By detecting user area changes in the smart home system, combining voice commands and dialogue data, seamless interaction across devices is achieved, and the problem of breakage between devices is solved when users are talking across regions is improved, and the usability and user experience of the system are improved.
Patent Information
- Application Number
- CN202510359666.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
AI Technical Summary
When users have cross-regional conversations, it is impossible to seamlessly switch between different smart home devices to complete complex interactive tasks.
By detecting that when a user moves from one area to another area, the target user identification is determined based on the received voice command. If the identification is the same, the dialogue data corresponding to the user is obtained, and the operation is determined in combination with the voice command to achieve continuous interaction across devices.
It realizes seamless interaction across devices, improving the availability and user satisfaction of smart home systems.
Smart Images

Figure CN120220677A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of smart home, and specifically relates to a voice interaction method, device, electronic device and readable storage medium. Background Art
[0002] With the progress of technology and the change of market demand, voice control technology has been widely used in fields such as smart home and smart office. In a smart home environment, users usually interact with multiple devices through voice control technology to meet different needs.
[0003] In the related art, when a smart home device receives a voice command sent by a user, the smart home device closest to the user executes the corresponding control operation.
[0004] However, when the user has a cross-region conversation, seamless switching between different smart home devices cannot be achieved to complete complex interaction tasks. Summary of the Invention
[0005] This application aims to provide a voice interaction method, device, electronic device and readable storage medium, which at least solves the problem that seamless switching between different smart home devices cannot be achieved to complete complex interaction tasks when the user has a cross-region conversation in the prior art.
[0006] In a first aspect, an embodiment of this application discloses a voice interaction method, and the method includes:
[0007] When it is detected that the user moves from the first area to the second area, determine a target user identifier according to the received first voice command;
[0008] If the target user identifier is the same as the first user identifier, obtain the conversation data corresponding to the first user, determine an execution operation according to the conversation data and the first voice command, and execute the execution operation; the conversation data includes the received second voice command other than the first voice command.
[0009] In a second aspect, an embodiment of this application discloses a voice interaction device, and the device includes:
[0010] A first determination module, configured to determine a target user identifier according to the received first voice command when it is detected that the user moves from the first area to the second area;
[0011] A second determination module, configured to, if the target user identifier is the same as the first user identifier, obtain the conversation data corresponding to the first user, determine an execution operation according to the conversation data and the first voice command, and execute the execution operation; the conversation data includes the received second voice command other than the first voice command.
[0012] In a third aspect, an embodiment of the present application also discloses an electronic device, including a processor and a memory. The memory stores a program or instructions that can run on the processor. When the program or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
[0013] In a fourth aspect, an embodiment of the present application also discloses a readable storage medium, on which a program or instructions are stored. When the program or instructions are executed by a processor, the steps of the method described in the first aspect are implemented.
[0014] In summary, in the embodiment of the present application, when it is detected that the user moves from the first area to the second area, the target user identifier is determined according to the received first voice instruction. If the target user identifier is the same as the first user identifier, the operation to be executed can be determined according to the conversation data corresponding to the first user and the first voice instruction. By considering the conversation data when executing the first voice instruction, the context can be made coherent, continuous interaction across devices can be achieved, and the usability and user satisfaction of the smart home system can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In the drawings:
[0016] Figure 1 is a flowchart of the steps of a voice interaction method provided by an embodiment of the present application;
[0017] Figure 2 is a flowchart of the steps of another voice interaction method provided by an embodiment of the present application;
[0018] Figure 3 is a flowchart of the steps of yet another voice interaction method provided by an embodiment of the present application;
[0019] Figure 4 is a detailed flowchart of the steps of yet another voice interaction method provided by an embodiment of the present application;
[0020] Figure 5 is a block diagram of a voice interaction device provided by an embodiment of the present application;
[0021] Figure 6 is a block diagram of an electronic device of an embodiment provided by an embodiment of the present application;
[0022] Figure 7 is a block diagram of an electronic device of another embodiment provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0024] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0025] With the continuous progress of technology, voice control technology has been widely used in fields such as smart home and smart office. In a smart home environment, users usually need to interact with multiple devices to meet different needs. Existing interaction systems often rely on close-range voice commands or physical operations, which limits the convenience of users. For example, when a user needs to have a cross-region conversation, after the existing interaction system interacts with the devices in area A, if the user moves to area B, the existing interaction system cannot perform subsequent interactions with the devices in area B.
[0026] In the embodiments of the present application, when it is detected that the user moves from the first area to the second area, the target user identifier is determined according to the received first voice command. If the target user identifier is the same as the first user identifier, then the operation to be executed can be determined according to the conversation data corresponding to the first user and the first voice command. By considering the conversation data when executing the first voice command, the context can be made coherent, and continuous interaction across devices can be achieved, which can improve the usability of the smart home system and user satisfaction.
[0027] Figure 1 is a flowchart of the steps of a voice interaction method provided by an embodiment of the present application. Refer to Figure 1 The method may include the following steps:
[0028] Step 101, when it is detected that the user moves from the first area to the second area, determine the target user identifier according to the received first voice command.
[0029] Exemplarily, a voice command is an interaction method that uses voice recognition technology to control various smart devices with voice commands. This technology allows users to operate home appliances, entertainment devices, transportation vehicles, etc. through voice commands, greatly improving the convenience and safety of life. For example, users can control these devices with simple voice commands such as "Turn on the TV", "Turn off the air conditioner", or "Play my music list".
[0030] Exemplarily, the first area can be the first room, the second area can be the second room, and the first room and the second room can be different rooms.
[0031] Exemplarily, after receiving the first voice command sent by the user, the first voice command can be input into a deep neural network to generate the first voiceprint information corresponding to the first voice command. Thus, the conversation data generated by the same user as the first voiceprint information can be obtained based on the first voiceprint information, which can meet the personalized needs of the user.
[0032] Step 102: If the target user identifier is the same as the first user identifier, obtain the conversation data corresponding to the first user, determine the execution operation according to the conversation data and the first voice command, and execute the execution operation; the conversation data includes the second voice command received except the first voice command.
[0033] Exemplarily, the user identifier is an identifier pre-stored in a database or the cloud. The user identifier can be the user's account. Before the user interacts with the smart home system, the user first registers, and the system collects the user's account and stores it in the database or the cloud. The first user identifier is one of the pre-stored user identifiers.
[0034] Exemplarily, the second voice command can be "Turn on the air conditioner". When the first user says "Turn on the air conditioner", the system will reply "There are multiple units, please select". The conversation data is: "Turn on the air conditioner" and "There are multiple units, please select".
[0035] Exemplarily, since all devices share the same cloud service platform, Device B can obtain the latest conversation status and history. If the first user moves to Area B and issues a subsequent first voice command to a device in Area B, for example, "Select the nth one". Then, the conversation data corresponding to the first user, "Turn on the air conditioner" and "There are multiple units, please select", are obtained, and the operation to be executed is determined based on "Turn on the air conditioner", "There are multiple units, please select", and "Select the nth one", that is, turn on the nth air conditioner and execute the operation. Device B knows that the user is selecting an air conditioner and performs the corresponding operation according to the new command "Select the nth one". Device B responds to the user's request, selects the specified air conditioner and turns it on. The system can identify the user's voiceprint information and associate this with the previous interaction session to maintain context coherence, thus achieving seamless cross-device interaction.
[0036] In an embodiment of the present application, when it is detected that the user moves from the first area to the second area, the target user identifier is determined according to the received first voice command. If the target user identifier is the same as the first user identifier, then the operation to be executed can be determined according to the conversation data corresponding to the first user and the first voice command. By considering the conversation data when executing the first voice command, context coherence can be achieved, and continuous cross-device interaction can be realized, which can improve the usability and user satisfaction of the smart home system.
[0037] Figure 2 It is a flowchart of the steps of another voice interaction method provided by the present application. Refer to Figure 2 , the method may include the following steps:
[0038] Step 201, when it is detected that the user moves from the first area to the second area, input the received first voice command into a deep neural network to generate first voiceprint information corresponding to the first voice command.
[0039] Exemplarily, deep neural networks have significant advantages in speech feature extraction, and can automatically learn discriminative features in speech signals, thereby improving the performance of speech-related tasks, such as speech recognition, voiceprint recognition, speech synthesis, etc. Deep neural networks may include: convolutional neural networks, recurrent neural networks, as well as long short-term memory artificial neural networks, gated recurrent units, etc. By using the above deep neural networks to automatically learn voiceprint features, multi-level speech features can be extracted: low-level features: such as spectrograms, time-frequency characteristics; middle-level features: such as phonemes, syllables; high-level features: such as semantics, speaker identity. Multi-level feature extraction helps to capture global and local information in speech signals.
[0040] Exemplarily, taking a convolutional neural network as an example, through operations such as convolutional layers and pooling layers, features such as the spectrogram of the speech signal corresponding to the first voice command are extracted and abstracted to learn more representative voiceprint features.
[0041] Optionally, step 201 may specifically include:
[0042] Sub-step 2011: Convert the first voice command into a spectrogram, and perform normalization processing on the spectrogram to obtain a normalized spectrogram.
[0043] Sub-step 2012: Input the normalized spectrogram into a deep neural network, and use the information output by the deep neural network as the first voiceprint information.
[0044] Regarding sub-step 2011 - sub-step 2012, the first voice command is converted into a spectrogram through operations such as convolutional layers and pooling layers, and normalization processing is performed on the spectrogram to obtain a normalized spectrogram. By extracting and abstracting the spectrogram features, more representative voiceprint features are learned. Specifically, the normalized spectrogram is input into a deep neural network, and the information output by the deep neural network is used as the first voiceprint information.
[0045] Step 202: Calculate the similarity between the first voiceprint information and the stored voiceprint features, and based on the similarity, find the target voiceprint features from the stored voiceprint features.
[0046] Exemplarily, similarity refers to the degree of similarity between two or more things in terms of certain features, attributes, or structures, etc. It is a relative concept used to measure the similarity relationship between things.
[0047] Exemplarily, the similarity can be the cosine similarity. The cosine similarity is a method for measuring the similarity between two vectors in a vector space. It determines their similarity by calculating the cosine value of the angle between the two vectors. The closer the value is to 1, the more similar the two vectors are, and the closer the value is to -1, the less similar the two vectors are.
[0048] Exemplarily, the first voiceprint information and the stored voiceprint features can be converted into word vectors, and the cosine similarity between the two word vectors is calculated. Based on the cosine similarity, the target voiceprint features are found from the stored voiceprint features.
[0049] Optionally, step 202 may specifically include:
[0050] Sub-step 2021: Calculate the similarity between the first voiceprint information and each stored voiceprint feature using cosine similarity.
[0051] Exemplarily, data preprocessing can be performed on the first voiceprint information and each stored voiceprint feature respectively. For example, normalization. In the first voiceprint information and each stored voiceprint feature, each dimensional value of the vector is scaled to a certain range, such as [0,1] or [-1,1], to improve the accuracy and stability of subsequent calculations.
[0052] Exemplarily, cosine similarity is a method for measuring the similarity between two vectors. In the field of voiceprint recognition, we convert voiceprint information into vector form and judge the similarity of voiceprints by calculating the cosine value between vectors. Calculate the similarity between the feature vector corresponding to the first voiceprint information and the vectors of each stored voiceprint feature according to cosine similarity.
[0053] Sub-step 2022: Use the voiceprint feature with the highest similarity among all similarity values as the target voiceprint feature.
[0054] Exemplarily, after obtaining the similarity values between the first voiceprint information and each stored voiceprint feature, further analysis and applications can be carried out based on these values. For example, set a similarity threshold. When the similarity value is greater than this threshold, it is considered that the first voiceprint information and the corresponding stored voiceprint feature belong to the same speaker. Another example is to sort the stored voiceprint features according to the similarity values and find the voiceprint feature most similar to the first voiceprint information.
[0055] Step 203: According to the pre-established correspondence between voiceprint features and user identifiers, obtain the target user identifier that matches the target voiceprint feature.
[0056] Exemplarily, the pre-established correspondence between voiceprint features and user identifiers can be established by collecting the user's voice samples through the microphone of the intelligent device and extracting the voiceprint features. These voice samples are created by the user reading specific sentences or words when using the device for the first time and are used for subsequent identity verification, ensuring the personalization and security of the interaction.
[0057] Exemplarily, after obtaining the target voiceprint information, according to the pre-established correspondence between voiceprint features and user identifiers stored in the cloud, the target voiceprint information can be determined from the stored voiceprint features, and its corresponding user identifier is used as the target user identifier.
[0058] Exemplarily, the user issues a first voice command to device B, "Select the nth one". Device B collects the user's first voice command and uploads it to the cloud service for voiceprint comparison. The cloud platform recognizes that this is the voice of the same user and automatically associates the previous conversation session started on device A. In this way, the system can understand that what the user wants to complete now is the previous unfinished operation, that is, to select the air conditioner with a specific number.
[0059] Optionally, before step 203, the method further includes:
[0060] Step A1: Collect the user identification and the user's voice information;
[0061] Step A2: Input the user's voice information into a deep neural network to obtain the voiceprint feature corresponding to the user's voice information;
[0062] Step A3: Establish the correspondence between the voiceprint feature and the user identification, and store it in the cloud.
[0063] For steps A1 - A3, creating the user's voiceprint feature mainly extracts the voiceprint feature from the voice information and generates representative features based on these features for subsequent voiceprint recognition. For example, after the deep neural network is trained, part of the output is extracted as the voiceprint feature. In the deep neural network, the final output can be used as the voiceprint feature, which contains the unique feature representation of the speaker learned by the model. Finally, establish the correspondence between the voiceprint feature and the user identification, and store it in the cloud database for the user's subsequent voiceprint recognition.
[0064] Optionally, before step A3, the method further includes:
[0065] Step A4: Encrypt the voiceprint feature using an encryption algorithm to obtain the encrypted voiceprint feature;
[0066] Step A3 may specifically include:
[0067] Sub - step A31: Establish the correspondence between the encrypted voiceprint feature and the user identification.
[0068] For steps A4 and sub - step A31, voiceprint information, as a type of data with high personal distinctiveness, its security cannot be ignored. Applying encryption technology to the storage of voiceprint information is a key measure to ensure data security. Encryption technology can use algorithms such as symmetric encryption, asymmetric encryption, or national encryption algorithms to encrypt the voiceprint information and store it in the cloud database. The system uses encryption technology to protect the user's voiceprint information, ensuring the security of data throughout the interaction process. This protection measure prevents unauthorized access and data leakage, enhancing the user's trust in the smart home system.
[0069] Symmetric encryption algorithms use the same key for both encryption and decryption operations. When encrypting voiceprint information, a secure key needs to be generated first. This key can be a randomly generated string of binary digits of a specific length. For example, the Advanced Encryption Standard (AES) algorithm typically supports key lengths of 128 bits, 192 bits, or 256 bits. Then, using the selected symmetric encryption algorithm, the generated key is applied to the voiceprint information data. Taking the AES algorithm as an example, through a series of operations such as byte substitution, row shift, column mixing, and round key addition, the original voiceprint information is transformed into ciphertext. The encrypted voiceprint information ciphertext is stored in the cloud database. When these voiceprint information are needed, the ciphertext is retrieved from the cloud database and decrypted using the same key. The decryption process is the inverse operation of the encryption process, and the ciphertext is restored to the original voiceprint information through corresponding algorithm steps.
[0070] Asymmetric encryption algorithms use a pair of keys, namely the public key and the private key. The public key can be publicly distributed and used to encrypt data; the private key is properly kept by the owner and used to decrypt data. For voiceprint information encryption, first, the receiving party generates a pair of public key and private key. The receiving party makes the public key public. After the sending party obtains the public key, it uses the public key to encrypt the voiceprint information. During the encryption process, specific mathematical operations are performed on the voiceprint information using the public key to transform it into ciphertext.
[0071] The national cryptographic algorithms are a series of independently developed cryptographic algorithms, such as block cipher algorithms. When encrypting voiceprint information, first, an encryption key is generated according to the specifications of the block cipher algorithm. The key generation process follows strict algorithm rules to ensure the randomness and security of the key. Then, the generated key is used to perform the encryption operation on the voiceprint information. The block cipher algorithm uses 32 rounds of iterative encryption, and through a series of operations such as byte transformation and linear transformation, the original voiceprint information is converted into ciphertext.
[0072] Step 204: If the target user identifier is the same as the first user identifier, obtain the conversation data corresponding to the first user, determine the execution operation according to the conversation data and the first voice command, and execute the execution operation; the conversation data includes the received second voice command other than the first voice command.
[0073] This step can specifically refer to the above step 102 and will not be elaborated here.
[0074] Optionally, step 204 may specifically include:
[0075] Sub-step 2041: Analyze the conversation data and extract the context information related to the first voice command;
[0076] Sub-step 2042: Determine the execution operation according to the context information and the first voice command.
[0077] For sub-steps 2041 - 2042, the dialogue data includes the interaction records between the user and the system, which may involve multiple dialogue rounds. For example, "Turn on the air conditioner" and "There are multiple ones, please select". Analyze the dialogue data to extract the context information related to the first voice command. At this time, the system already knows that the user wants to turn on the air conditioner. Based on the first voice command, the execution operation can be determined. For example, if the first voice command is "Select the nth one", it can be determined that the user wants to turn on the nth air conditioner. Thus, according to this intention, the nth air conditioner can be turned on.
[0078] Optionally, after step 204, the method further includes:
[0079] Step 205: In response to a click operation on the feedback button, display the items to be evaluated on the feedback interface;
[0080] Step 206: Obtain the updated dialogue data from the cloud and determine the target items to be evaluated related to the updated dialogue data among the items to be evaluated; the updated dialogue data further includes: the first voice command;
[0081] Step 207: In the input box corresponding to the target item to be evaluated, input the evaluation content corresponding to the target item to be evaluated;
[0082] Step 208: After inputting all the evaluation content, click the submit button to store the evaluation content in the cloud.
[0083] Regarding steps 205 - 208, when the user browses or uses various functions within the application (APP, Application), if they find that they need to provide feedback on certain aspects, they can click the feedback button. In response to this click operation, the APP quickly loads the feedback interface. In this interface, a series of items to be evaluated are clearly displayed in a list form, and each item to be evaluated is described in concise and clear text, enabling the user to quickly understand its meaning. At the same time as the user opens the feedback interface, the APP background quickly sends a data request to the cloud to obtain the updated dialogue data. This dialogue data includes the text content converted from the voice conversations between the user and the APP during the recent interaction process, text chat records, etc. After receiving the request, the cloud quickly transmits the updated dialogue data to the APP according to an efficient data transmission protocol.
[0084] After the APP receives the data, it performs word segmentation on the dialogue data to extract key information and keywords. Then, it precisely compares these keywords with the keywords in the item to be evaluated, and determines the target item to be evaluated related to the latest dialogue data in the item to be evaluated by calculating the keyword matching degree and semantic similarity. In the feedback interface, the user can, according to their own usage experience, enter the detailed evaluation content corresponding to the target item to be evaluated in the input box. After entering all the evaluation content, click the submit button to store the evaluation content in the cloud. After the storage is completed, the cloud returns a successful storage response message to the APP. After the APP receives the response, it pops up a prompt box to inform the user that the evaluation content has been successfully submitted, and thanks the user for the feedback. The user feedback mechanism provided by the system allows users to evaluate the interaction results. These feedbacks are used to continuously optimize the voiceprint recognition and interaction algorithms, and improve the accuracy and personalization of the interaction. In this way, the system can continuously learn and adapt to the user's behavior patterns, making the interaction experience continuously optimized over time.
[0085] Optionally, the dialogue data further includes: first context information related to the first voice command, second context information related to the second voice command, and step 206 may specifically include:
[0086] Sub-step 2061: Use a semantic understanding model to perform semantic encoding on the first voice command, the first context information, the second voice command, and the second context information in the updated dialogue data to obtain first semantic information, and perform semantic encoding on the item to be evaluated to obtain a plurality of second semantic information.
[0087] Sub-step 2062: Calculate the semantic similarity corresponding to each of the second semantic information and the second semantic information in the first semantic information respectively. If the semantic similarity corresponding to the second semantic information is higher than a preset threshold, determine the item to be evaluated corresponding to the second semantic information as the target item to be evaluated.
[0088] For sub-step 2061 - sub-step 2062, the semantic understanding model can be a Bidirectional Encoder Representations from Transformers (BERT) model. The updated dialogue data is input into the BERT model, and the model will perform word embedding processing on each word in the dialogue data, converting it into a low-dimensional dense vector. Then, through multiple layers of encoders, these vectors are deeply semantically fused and feature extracted to finally obtain the first semantic information that can comprehensively represent the dialogue meaning. Similarly, similar operations are performed on the items to be evaluated. Each item to be evaluated is sequentially input into the BERT model, and after the same word embedding and multiple-layer encoder processing process, the second semantic information corresponding to each item to be evaluated is obtained. These second semantic information can accurately reflect the core semantic connotation of each item to be evaluated. After obtaining the first semantic information and each second semantic information, the cosine similarity algorithm or other professional semantic similarity measurement methods are used to calculate the semantic similarity between the first semantic information and each second semantic information. If the calculated semantic similarity is higher than the preset threshold, it indicates that the item to be evaluated has a high semantic correlation with the updated dialogue data. At this time, the item to be evaluated corresponding to the second semantic information is determined as the target item to be evaluated. For example, if the preset threshold is 0.7, when it is calculated that the similarity between the second semantic information of a certain item to be evaluated and the first semantic information reaches 0.75, then this item to be evaluated will be identified as the target item to be evaluated.
[0089] Optionally, the method further includes:
[0090] Step 209, when receiving the second voice command, input the second voice command into a deep neural network to generate second voiceprint information corresponding to the second voice command, and determine a user identifier according to the second voiceprint information;
[0091] Step 210, if the user identifier exists in the user identifiers stored in the cloud, then perform interaction with the smart home device.
[0092] Regarding steps 209 - 210, when the user interacts with device A in area A, the system identifies the user's identity by comparing the real-time voice information with the stored voiceprint template and starts a multi-round interaction session.
[0093] When the user is in area A and wants to interact with the smart home device A, they can directly issue a second voice command. For example, the user says "turn on the air conditioner". At this time, device A captures the user's second voice command through the built-in microphone and sends the real-time second voice command to the cloud for processing. The cloud inputs the second voice command into a pre-trained deep neural network to obtain the second voiceprint information corresponding to this voice command. This voiceprint information is a vector that can uniquely represent the user's voice characteristics and contains information such as the user's pronunciation habits, tone color, and intonation.
[0094] After the cloud successfully determines the user identification, it checks whether this user identification exists in the list of user identifications stored in the cloud. If the user identification exists in the list of legitimate user identifications, it indicates that the user has been registered in the system and the identity verification has passed. At this time, the system successfully identifies the user and starts to initiate a personalized multi-round conversation session for this user to interact with the corresponding smart home device.
[0095] The deep neural network can accurately extract the voiceprint information from the second voice command to determine the user identification. This process realizes the accurate identification of the user's identity. Only when the user identification matches the identification stored in the cloud is the interaction with the smart home device allowed. This mechanism provides a reliable security guarantee for the smart home system, preventing unauthorized access and operations, and protecting the user's privacy and device security.
[0096] See Figure 3 , which shows the step flowchart of another voice interaction method provided by the embodiment of the present application. The steps include:
[0097] Step M1, voiceprint information collection and registration;
[0098] Step M2, voiceprint recognition and user identity confirmation;
[0099] Step M3, context maintenance and cross-device interaction;
[0100] Step M4, execution of multi-round interaction;
[0101] Step M5, security guarantee;
[0102] Step M6, user feedback and continuous learning.
[0103] See Figure 4 , which shows the detailed step flowchart of another voice interaction method provided by the embodiment of the present application. The steps include:
[0104] Step S1, collect the user's second voice command and wake up the device;
[0105] Step S2: Extract features from the second voice command to obtain second voiceprint information;
[0106] Step S3: Identify the user's identity based on the second voiceprint information and the stored encrypted voiceprint features, start a multi-round interaction session, and record the conversation data;
[0107] Step S4: Collect the user's first voice command;
[0108] Step S5: Extract features from the first voice command to obtain first voiceprint information;
[0109] Step S6: Obtain the target user identifier that matches the first voiceprint information according to the pre-established correspondence between voiceprint features and user identifiers;
[0110] Step S7: Obtain the conversation data of the target user, determine the operation to be performed based on the conversation data and the first voice command, and the system performs the operation;
[0111] Step S8: Obtain user evaluations and optimize the voiceprint recognition and interaction algorithms.
[0112] In the embodiment of the present application, when it is detected that the user moves from the first area to the second area, the target user identifier is determined according to the received first voice command. If the target user identifier is the same as the first user identifier, then the operation to be performed can be determined according to the conversation data corresponding to the first user and the first voice command. By considering the conversation data when executing the first voice command, context coherence can be achieved, and continuous interaction across devices can be realized, which can improve the usability and user satisfaction of the smart home system.
[0113] See Figure 5 , which shows a voice interaction device 30 provided by an embodiment of the present application. The voice interaction device 30 includes:
[0114] The first determination module 301 is configured to determine the target user identifier according to the received first voice command when it is detected that the user moves from the first area to the second area;
[0115] The second determination module 302 is configured to, if the target user identifier is the same as the first user identifier, obtain the conversation data corresponding to the first user, determine the operation to be performed according to the conversation data and the first voice command, and execute the operation to be performed; the conversation data includes the received second voice command other than the first voice command.
[0116] Optionally, the first determination module includes:
[0117] The generation sub-module is configured to input the first voice command into a deep neural network to generate first voiceprint information corresponding to the first voice command;
[0118] A calculation sub-module, configured to calculate the similarity between the first voiceprint information and the stored voiceprint features, and find the target voiceprint features from the stored voiceprint features according to the similarity;
[0119] An identification determination sub-module, configured to obtain the target user identification that matches the target voiceprint features according to the pre-established correspondence between the voiceprint features and the user identifications;
[0120] Optionally, the generation sub-module includes:
[0121] A normalization processing unit, configured to convert the first voice command into a spectrogram, and perform normalization processing on the spectrogram to obtain a normalized spectrogram;
[0122] A first determination unit, configured to input the normalized spectrogram into a deep neural network, and use the information output by the deep neural network as the first voiceprint information.
[0123] Optionally, the calculation sub-module includes:
[0124] A calculation unit, configured to calculate the similarity between the first voiceprint information and each of the stored voiceprint features by using cosine similarity;
[0125] A second determination unit, configured to use the voiceprint feature with the highest similarity among all the similarity values as the target voiceprint feature.
[0126] Optionally, the apparatus further includes:
[0127] A collection module, configured to collect user identifications and the voice information of users;
[0128] A third determination module, configured to input the voice information of the user into a deep neural network to obtain the voiceprint features corresponding to the voice information of the user;
[0129] A first establishment module, configured to establish the correspondence between the voiceprint features and the user identifications, and store same in the cloud.
[0130] Optionally, the apparatus further includes:
[0131] An encryption module, configured to encrypt the voiceprint features by using an encryption algorithm to obtain encrypted voiceprint features;
[0132] A second establishment module, configured to establish the correspondence between the encrypted voiceprint features and the user identifications.
[0133] Optionally, the second determination module includes:
[0134] An analysis sub-module, configured to analyze the conversation data and extract context information related to the first voice command;
[0135] A first determination sub-module, configured to determine the execution operation according to the context information and the first voice command.
[0136] Optionally, the apparatus further includes:
[0137] A display module, configured to display items to be evaluated on a feedback interface in response to a click operation on a feedback button;
[0138] A fourth determination module, configured to obtain updated conversation data from the cloud and determine target items to be evaluated related to the updated conversation data; the updated conversation data further includes: the first voice command;
[0139] An input module, configured to input evaluation content corresponding to the target item to be evaluated in an input box corresponding to the target item to be evaluated;
[0140] A storage module, configured to click a submit button after inputting all the evaluation content and store the evaluation content in the cloud.
[0141] Optionally, the conversation data further includes: first context information related to the first voice command, and second context information related to the second voice command; the fourth determination module includes:
[0142] An encoding sub-module, configured to perform semantic encoding on the first voice command, the first context information, the second voice command, and the second context information in the updated conversation data by using a semantic understanding model to obtain first semantic information, and perform semantic encoding on the items to be evaluated to obtain a plurality of second semantic information;
[0143] A second determination sub-module, configured to calculate the semantic similarity between each second semantic information and the second semantic information corresponding in the first semantic information respectively. If the semantic similarity corresponding to the second semantic information is higher than a preset threshold, the item to be evaluated corresponding to the second semantic information is determined as the target item to be evaluated.
[0144] Optionally, the apparatus further includes:
[0145] A generation module, configured to input the second voice command into a deep neural network when receiving the second voice command, generate second voiceprint information corresponding to the second voice command, and determine a user identifier according to the second voiceprint information;
[0146] An interaction module, which is used to interact with the smart home device if the user identifier exists in the user identifiers stored in the cloud.
[0147] In an embodiment of the present application, when it is detected that the user moves from the first area to the second area, the target user identifier is determined according to the received first voice command. If the target user identifier is the same as the first user identifier, the operation to be executed can be determined according to the conversation data corresponding to the first user and the first voice command. By considering the conversation data when executing the first voice command, the context can be made coherent, continuous interaction across devices can be achieved, and the usability and user satisfaction of the smart home system can be improved.
[0148] Figure 6 It is a block diagram of an electronic device 400 according to an embodiment provided by an embodiment of the present application. Refer to Figure 6 , the electronic device 400 may include one or more of the following components: a processing component 402, a memory 404, a power supply component 406, a multimedia component 408, an audio component 410, an input / output (I / O) interface 412, a sensor component 414, and a communication component 416.
[0149] The processing component 402 generally controls the overall operation of the electronic device 400, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 402 may include one or more processors 420 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 402 may include one or more modules to facilitate the interaction between the processing component 402 and other components. For example, the processing component 402 may include a multimedia module to facilitate the interaction between the multimedia component 408 and the processing component 402.
[0150] The memory 404 is used to store various types of data to support the operation of the electronic device 400. Examples of these data include instructions for any application or method operating on the electronic device 400, contact data, phone book data, messages, pictures, multimedia, etc. The memory 404 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0151] The power supply component 406 provides power for various components of the electronic device 400. The power supply component 406 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 400.
[0152] The multimedia component 408 includes an interface that provides an output interface between the electronic device 400 and the user. In some embodiments, the interface may include a liquid crystal display (LCD) and a touch panel (TP). If the interface includes a touch panel, the interface can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 408 includes a front camera and / or a rear camera. When the electronic device 400 is in an operating mode, such as a shooting mode or a multimedia mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0153] The audio component 410 is used to output and / or input audio signals. For example, the audio component 410 includes a microphone (MIC) that is used to receive external audio signals when the electronic device 400 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 404 or transmitted via the communication component 416. In some embodiments, the audio component 410 further includes a speaker for outputting audio signals.
[0154] The input / output (I / O) interface 412 provides an interface between the processing component 402 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0155] The sensor component 414 includes one or more sensors for providing a status assessment of various aspects of the electronic device 400. For example, the sensor component 414 can detect the on / off state of the electronic device 400, the relative positioning of components, such as the display and the keypad of the electronic device 400. The sensor component 414 can also detect a change in the position of the electronic device 400 or a component of the electronic device 400, the presence or absence of user contact with the electronic device 400, the orientation or acceleration / deceleration of the electronic device 400, and the temperature change of the electronic device 400. The sensor component 414 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 414 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 414 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0156] The communication component 416 is used to facilitate communication between the electronic device 400 and other devices in a wired or wireless manner. The electronic device 400 can access a communication standard-based wireless network, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 416 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 416 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0157] In an exemplary embodiment, the electronic device 400 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to implement a voice interaction method provided by an embodiment of the present application.
[0158] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 404 including instructions. The above instructions can be executed by a processor 420 of the electronic device 400 to complete the above method. For example, the non-transitory storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0159] Figure 7 It is a block diagram of an electronic device 500 according to another embodiment provided by an embodiment of the present application. For example, the electronic device 500 can be provided as a server. Refer to Figure 7 , the electronic device 500 includes a processing component 522, which further includes one or more processors, and memory resources represented by a memory 532 for storing instructions executable by the processing component 522, such as application programs. The application programs stored in the memory 532 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 522 is configured to execute instructions to perform a voice interaction method provided by an embodiment of the present application.
[0160] The electronic device 500 may further include a power supply component 526 configured to perform power management of the electronic device 500, a wired or wireless network interface 550 configured to connect the electronic device 500 to a network, and an input / output (I / O) interface 558. The electronic device 500 may operate based on an operating system stored in the memory 532, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSD TM or the like.
[0161] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0162] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A voice interaction method, characterized in that: The method comprises: When detecting that the user moves from the first area to the second area, determining the target user identifier according to the received first voice instruction; If the target user identifier is the same as the first user identifier, the conversation data corresponding to the first user is obtained, the execution operation is determined according to the conversation data and the first voice instruction, and the execution operation is performed; the conversation data includes a second voice instruction received in addition to the first voice instruction.
2. The method according to claim 1, characterized in that The step of determining the target user identifier according to the received first voice instruction includes: Inputting the first voice command into a deep neural network to generate first voiceprint information corresponding to the first voice command; Calculating the similarity between the first voiceprint information and the stored voiceprint features, and finding the target voiceprint features from the stored voiceprint features according to the similarity; According to the pre-established correspondence between the voiceprint feature and the user identification, a target user identification matching the target voiceprint feature is obtained.
3. The method according to claim 2, characterized in that The step of inputting the first voice command into a deep neural network to generate first voiceprint information corresponding to the first voice command includes: Converting the first voice command into a spectrogram, and normalizing the spectrogram to obtain a normalized spectrogram; The normalized spectrum graph is input into a deep neural network, and the information output by the deep neural network is used as the first voiceprint information.
4. The method according to claim 2, characterized in that: The calculating the similarity between the first voiceprint information and the stored voiceprint features, and finding the target voiceprint features from the stored voiceprint features according to the similarity, includes: Calculating the similarity between the first voiceprint information and each stored voiceprint feature using cosine similarity; The voiceprint feature with the highest similarity among all similarity values is used as the target voiceprint feature.
5. The method according to claim 2, characterized in that: Before obtaining the target user identification matching the target voiceprint feature according to the pre-established correspondence between the voiceprint feature and the user identification, the method further includes: Collect user ID and user voice information; Inputting the user's voice information into a deep neural network to obtain a voiceprint feature corresponding to the user's voice information; A corresponding relationship between the voiceprint feature and the user identification is established and stored in the cloud.
6. The method according to claim 5, characterized in that Before establishing the correspondence between the voiceprint feature and the user identifier, the method further includes: Encrypting the voiceprint feature using an encryption algorithm to obtain an encrypted voiceprint feature; The establishing of the correspondence between the voiceprint feature and the user identifier includes: A corresponding relationship between the encrypted voiceprint feature and the user identification is established.
7. The method according to claim 1, characterized in that The determining the execution operation according to the conversation data and the first voice instruction includes: Analyzing the conversation data to extract context information related to the first voice command; The execution operation is determined according to the context information and the first voice instruction.
8. The method according to claim 1, characterized in that After determining an execution operation according to the dialogue data and the first voice instruction, and executing the execution operation, the method further includes: In response to a click operation on a feedback button, displaying the items to be evaluated on a feedback interface; Acquire updated conversation data from the cloud, and determine a target item to be evaluated that is related to the updated conversation data among the items to be evaluated; the updated conversation data also includes: the first voice instruction; In the input box corresponding to the target item to be evaluated, input the evaluation content corresponding to the target item to be evaluated; After entering all the evaluation content, click the Submit button to store the evaluation content in the cloud.
9. The method according to claim 8, characterized in that The dialogue data further includes: first context information related to the first voice instruction, and second context information related to the second voice instruction; and determining a target item to be evaluated related to the updated dialogue data among the items to be evaluated includes: Using a semantic understanding model, semantically encoding the first voice instruction, the first context information, the second voice instruction, and the second context information in the updated dialogue data to obtain first semantic information, and semantically encoding the item to be evaluated to obtain multiple second semantic information; The semantic similarity between each of the second semantic information and the second semantic information in the first semantic information is calculated respectively. If the semantic similarity corresponding to the second semantic information is higher than a preset threshold, the item to be evaluated corresponding to the second semantic information is determined as the target item to be evaluated.
10. The method according to claim 1, characterized in that The method further comprises: When receiving the second voice instruction, inputting the second voice instruction into a deep neural network, generating second voiceprint information corresponding to the second voice instruction, and determining a user identification according to the second voiceprint information; If the user identifier exists in the user identifiers stored in the cloud, interaction with the smart home device is performed.
11. A voice interaction device, characterized in that: The device comprises: A first determination module, configured to determine a target user identifier according to a received first voice instruction when detecting that a user moves from a first area to a second area; A second determination module is used to obtain conversation data corresponding to the first user if the target user identifier is the same as the first user identifier, determine an execution operation based on the conversation data and the first voice instruction, and execute the execution operation; the conversation data includes a second voice instruction received in addition to the first voice instruction.
12. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 10 are implemented.
13. A readable storage medium, characterized in that: The readable storage medium stores a program or an instruction, and when the program or the instruction is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
14. An air conditioner, characterized in that: Comprising the electronic device as claimed in claim 12.