Voice interaction method, device, equipment and system
Through image data, the redundancy and false wake-up problems of existing voice wake-up methods are solved, wake-up-free voice interaction is realized, and user experience and device efficiency are improved.
Patent Information
- Application Number
- CN202110314068.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-24
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2041-03-24
AI Technical Summary
The existing voice wake-up method requires users to repeatedly speak wake-up words, resulting in redundant user experience and increasing the transmission pressure of cloud links, and a high false wake-up rate.
By acquiring image data, determine whether there is a user in the scene who is expected to interact with the device with a voice, and detect whether the user has voice or has a pronunciation intention, and initiate the interactive service only when the user has voice or has a pronunciation intention.
It enables the activation of device interactive services without the user speaking wake-up words, reducing the false wake-up rate and reducing the pressure of cloud data transmission.
Smart Images

Figure CN115206306B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of voice interaction, and in particular to a voice interaction method, apparatus, device, and system. Background Art
[0002] Voice interaction is a cutting-edge form of human-computer interaction. It's the process by which users give commands to machines through natural language to achieve their goals.
[0003] In the existing technology, voice wake-up is a prerequisite for realizing voice interaction.
[0004] Voice wake-up occurs when the user speaks a specific wake-up word. Upon detecting the user speaking the wake-up word, the device switches from sleep mode to active mode. Only after waking up does the device begin providing voice interaction services, such as uploading the voice to the cloud for speech recognition and semantic understanding, and then providing feedback based on the recognition results.
[0005] The wake-up method based on a wake-up word goes against the user's language expression habits. Moreover, in multi-round dialogue interaction scenarios, the user is often required to repeat the wake-up word to wake the device, which is redundant and cumbersome, affecting the user interaction experience.
[0006] Setting a wake-up delay after waking up can theoretically reduce the number of times the user needs to say the wake-up word. However, this isn't truly wake-up-free. This is because the user has no way of knowing the specific duration of the wake-up delay and, therefore, no idea when they can speak directly. In practice, users will continue to say the wake-up word. Furthermore, the device's continuous microphone reception will cause a large amount of voice to be uploaded to the cloud, increasing the transmission pressure on the cloud link and, in turn, increasing false wake-ups.
[0007] Therefore, an effective wake-up-free voice interaction solution is needed. Summary of the Invention
[0008] A technical problem to be solved by the present disclosure is to provide an effective wake-up-free voice interaction solution.
[0009] According to a first aspect of the present disclosure, a voice interaction method is provided for interacting with a device, comprising: acquiring image data, the image data containing information representing a scene in which the device is located; determining, based on the image data, whether there is a user in the scene who desires to perform voice interaction with the device; when it is determined that there is a user in the scene, detecting whether the user utters a voice and / or whether the user has an intention to pronounce a voice; when it is detected that the user utters a voice or has an intention to pronounce a voice, initiating an interactive service for the device.
[0010] According to a second aspect of the present disclosure, a voice interaction method is provided, comprising: receiving audio data uploaded by a device, the audio data being uploaded by the device when it is determined, based on image data used to characterize the scene in which the device is located, that there is a user who desires to perform voice interaction with the device, and when it is detected that the user utters a voice or has an intention to pronounce a voice; performing voice recognition on the audio data; judging, based on the voice recognition result, whether the user performs voice interaction with the device; and when it is determined that the user does not perform voice interaction with the device, sending a sound pickup termination instruction to the device.
[0011] According to a third aspect of the present disclosure, a voice interaction method is provided, comprising: receiving audio data uploaded by a first device, the audio data being sent by the first device when it is determined, based on image data used to characterize the scene in which the second device is located, that there is a user who desires to perform voice interaction with the second device, and when it is detected that the user utters a voice or has an intention to pronounce a voice; performing voice recognition on the audio data; judging, based on the voice recognition result, whether the user performs voice interaction with the second device; and when it is determined that the user does not perform voice interaction with the second device, sending a pickup termination instruction to the first device.
[0012] According to a fourth aspect of the present disclosure, a voice interaction method is provided, comprising: judging whether there is a user who desires to perform voice interaction with the device in the scene based on image data used to characterize the scene in which the device is located; and starting an interactive service for the device when it is determined that there is a user in the scene.
[0013] According to a fifth aspect of the present disclosure, a voice interaction system is provided, comprising: a first device, configured to image a scene in which the first device is located, and determine, based on the obtained image data, whether there is a user in the scene who desires to perform voice interaction with the first device; when it is determined that there is a user in the scene, detecting whether the user utters a voice and / or whether the user has an intention to pronounce a voice; when it is detected that the user utters a voice or the user has an intention to pronounce a voice, sending the collected audio data to a server; a server, configured to receive the audio data sent by the first device, perform voice recognition on the audio data, and determine, based on the voice recognition result, whether the user performs voice interaction with the first device; when it is determined that the user does not perform voice interaction with the first device, sending a sound pickup termination instruction to the first device; the first device is further configured to stop collecting audio data in response to receiving the sound pickup termination instruction.
[0014] According to the sixth aspect of the present disclosure, a voice interaction system is provided, comprising: a second device; a first device, which is arranged in the same scene as the second device, the first device being used to image the scene, and to determine whether there is a user in the scene who desires to perform voice interaction with the second device based on the obtained image data; when it is determined that the user exists in the scene, detecting whether the user utters a voice and / or whether the user has an intention to pronounce a voice, and when it is detected that the user utters a voice or the user has an intention to pronounce a voice, uploading the collected audio data to a server; a server being used to receive the audio data sent by the first device, perform voice recognition on the audio data, and determine whether the user performs voice interaction with the second device based on the voice recognition result; when it is determined that the user does not perform voice interaction with the second device, sending a sound pickup termination instruction to the first device, and the first device being further used to stop collecting audio data in response to receiving the sound pickup termination instruction.
[0015] According to a seventh aspect of the present disclosure, a smart device is provided, comprising: a communication module; a sound pickup module; an imaging module for imaging a scene in which the smart device is located to obtain image data; and a processor for determining, based on the image data, whether there is a user in the scene who desires to perform voice interaction with the smart device; and when it is determined that the user exists in the scene, detecting whether the user utters a voice and / or whether the user has an intention to pronounce a voice; and when it is detected that the user utters a voice or has an intention to pronounce a voice, uploading the audio data collected by the sound pickup module to a server via the communication module, and the server performing voice recognition on the audio data.
[0016] According to an eighth aspect of the present disclosure, a smart device is provided, which is suitable for being arranged in the same scene with an Internet of Things device, and the smart device includes: a communication module; a sound pickup module; an imaging module for imaging the scene to obtain image data; and a processor for determining, based on the image data, whether there is a user in the scene who desires to perform voice interaction with the Internet of Things device; when it is determined that the user exists in the scene, detecting whether the user utters a voice and / or whether the user has an intention to pronounce a voice; when it is detected that the user utters a voice or the user has an intention to pronounce a voice, uploading the audio data collected by the sound pickup module to a server through the communication module, and the server performing voice recognition on the audio data.
[0017] According to the ninth aspect of the present disclosure, a voice interaction device is provided, comprising: an acquisition module for acquiring image data, the image data including information representing a scene in which a device is located; a judgment module for judging, based on the image data, whether there is a user in the scene who desires to perform voice interaction with the device; a detection module for detecting, when it is determined that the user exists in the scene, whether the user utters a voice and / or whether the user has an intention to pronounce a voice; and a startup module for starting an interactive service for the device when it is detected that the user utters a voice or the user has an intention to pronounce a voice.
[0018] According to the tenth aspect of the present disclosure, a voice interaction device is provided, comprising: a receiving module for receiving audio data uploaded by a device, the audio data being uploaded by the device when it is determined, based on image data used to characterize the scene in which the device is located, that there is a user who desires to perform voice interaction with the device, and when it is detected that the user utters a voice or has an intention to pronounce a voice; a voice recognition module for performing voice recognition on the audio data; a judgment module for judging, based on the voice recognition result, whether the user performs voice interaction with the device; and a sending module for sending a sound pickup termination instruction to the device when it is determined that the user does not perform voice interaction with the device.
[0019] According to the eleventh aspect of the present disclosure, a voice interaction device is provided, including: a receiving module for receiving audio data uploaded by a first device, the audio data being sent by the first device when it is determined, based on image data used to characterize the scene in which the second device is located, that there is a user who desires to perform voice interaction with the second device in the scene, and when it is detected that the user utters a voice or has an intention to pronounce a voice; a voice recognition module for performing voice recognition on the audio data; a judgment module for judging whether the user performs voice interaction with the second device based on the voice recognition result; and a sending module for sending a sound pickup termination instruction to the first device when it is determined that the user does not perform voice interaction with the second device.
[0020] According to the twelfth aspect of the present disclosure, a voice interaction device is provided, including: a judgment module for judging whether there is a user who desires to perform voice interaction with the device in the scene based on image data used to characterize the scene in which the device is located; and a starting module for starting an interactive service for the device when it is determined that the user exists in the scene.
[0021] According to the thirteenth aspect of the present disclosure, a computing device is provided, comprising: a processor; and a memory on which executable code is stored, and when the executable code is executed by the processor, the processor executes the method described in any one of the first to fourth aspects above.
[0022] According to the fourteenth aspect of the present disclosure, a non-temporary machine-readable storage medium is provided, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor executes the method described in any one of the first to fourth aspects above.
[0023] Therefore, according to an exemplary embodiment of the present disclosure, the interactive service for the device is started only when it is determined based on image data whether there is a user in the scene who desires to interact with the device by voice, and it is detected that the user makes a voice or has an intention to pronounce. In this way, the interactive service of the device can be activated without the user saying the wake-up word, and the false wake-up rate can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings, wherein like reference numerals generally represent like components in the exemplary embodiments of the present disclosure.
[0025] Figure 1 A schematic flowchart of a voice interaction method according to an embodiment of the present disclosure is shown.
[0026] Figure 2 A schematic diagram of an application scenario according to an embodiment of the present disclosure is shown.
[0027] Figure 3 A schematic flowchart of a voice interaction method according to another embodiment of the present disclosure is shown.
[0028] Figure 4 A schematic flowchart of a voice interaction method according to another embodiment of the present disclosure is shown.
[0029] Figure 5 Shown Figure 4 A schematic diagram of an application scenario of the voice interaction method shown.
[0030] Figure 6 A structural diagram of a voice interaction system according to an embodiment of the present disclosure is shown.
[0031] Figure 7 A structural diagram of a voice interaction system according to another embodiment of the present disclosure is shown.
[0032] Figure 8 A schematic structural diagram of a smart device according to an embodiment of the present disclosure is shown.
[0033] Figure 9 A structural diagram of a voice interaction device according to an embodiment of the present disclosure is shown.
[0034] Figure 10A structural diagram of a voice interaction device according to another embodiment of the present disclosure is shown.
[0035] Figure 11 A structural diagram of a voice interaction device according to another embodiment of the present disclosure is shown.
[0036] Figure 12 A schematic structural diagram of a computing device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0037] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although preferred embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0038] This disclosure proposes a wake-up-free voice interaction solution using image data, preferably multimodal data including image data and audio data. Wake-up-free means that the user can interact directly with the device without saying a wake-up word. That is, this disclosure activates the device's voice interaction process without the user saying a wake-up word.
[0039] It should be noted that the acquisition and use of the data involved in this disclosure (such as image data, audio data, and other data obtained by analyzing image data / audio data) are all carried out after authorization.
[0040] Figure 1 A schematic flow chart of a voice interaction method according to an embodiment of the present disclosure is shown. The voice interaction method of the present disclosure is used to interact with a device, which can be a device with an imaging function (i.e., the first device described below) or a device without an imaging function (i.e., the second device described below).
[0041] Figure 1 The method shown can be performed by a device with an imaging function (i.e., an image acquisition function) (particularly a multimodal device with both imaging and sound pickup functions). For ease of distinction, this device can be referred to as a first device. The first device can be an audio and video acquisition device with a camera and a microphone, such as a smart speaker with a camera.
[0042] The first device may execute Figure 1 The method shown implements self-wake-up-free voice interaction.
[0043] The first device can also execute Figure 1The method shown implements wake-up-free voice interaction for one or more second devices (that do not support the multimodal data collection function) by leveraging the multimodal data collection function of the first device.
[0044] The second device may be a device without an imaging function, such as a single-mode device with only a sound pickup function, or a device without even a sound pickup function. For example, the second device may include, but is not limited to, IoT devices such as smart sockets, smart buttons, smart kitchen appliances, smart light bulbs, and smart TVs.
[0045] Figure 1 The method shown can also be performed by a device that does not have an imaging function (i.e., the second device). The second device can use the imaging function of the first device (optionally also the sound pickup function of the first device) to achieve its own wake-up-free voice interaction.
[0046] Reference below Figure 1 The implementation principle of the voice interaction method disclosed in the present invention is exemplified.
[0047] See also Figure 1 In step S110, image data is acquired, where the image data includes information representing the scene in which the device is located.
[0048] The device referred to herein refers to a device capable of providing voice interaction services to users, that is, a device that can perform corresponding operations in response to user voice commands. The device can be the first device with imaging capabilities described above, or the second device without imaging capabilities described above. When the device has imaging capabilities, the image data can be captured by the device; when the device does not have imaging capabilities, the image data can be captured by another device in the same scene as the device.
[0049] Image data used to represent the scene in which the device is located may refer to a scene image obtained by imaging the scene in which the device is located. The scene in which the device is located may refer to a spatial area near the device, or a spatial area in a specific direction of the device (such as in front of the device).
[0050] Alternatively, the image data may refer to image data of a portion of the scene in which the device is located that corresponds to the user's activity range. That is, the image data may be obtained by imaging the portion of the scene in which the device is located that corresponds to the user's activity range.
[0051] In step S120 , it is determined based on the image data whether there is a user in the scene who desires to perform voice interaction with the device.
[0052] By analyzing the image data, we can obtain multi-dimensional feature information such as the facial image, expression, distance from the device, and whether the device is in contact with the user in the scene. Based on this feature information, we can determine whether some or all users in the scene expect to interact with the device through voice, and then determine whether there are specific users in the scene who expect to interact with the device through voice.
[0053] For example, body language is a way people express their thoughts, and compared to speaking, it's even more consistent with people's communication habits. When users interact with devices through voice, they often unconsciously use body language to indicate their intentions. For example, when a user tells a smart speaker to play a song, they typically look at the speaker or speak in its direction. When a user asks a smart rice cooker if the food is ready, they subconsciously turn their head or gaze toward the cooker.
[0054] Therefore, the present disclosure can identify the body movements of users in a scene based on image data, and based on the recognition results of the body movements of users in the scene, determine whether there is a user in the scene who desires voice interaction with the device. The judgment method can be that the body movements conform to a preset pattern. For example, based on the recognition results of the body movements of users in the scene, it can be determined whether there is a user in the scene who is pointing the device with their body movements. The user who is pointing the device with their body movements is the user who desires voice interaction with the device.
[0055] Body movements may include but are not limited to: one or more of the following: the direction of the user's eye gaze (i.e., line of sight), facial orientation, head twisting movements, the distance between the user and the device and / or distance change information, the user's body movements (such as the direction of the arm / palm / fingers), etc.
[0056] Body movements indicating the device may include, but are not limited to: the user's eyes pointing towards the device; the user's face facing the device; the user's fingers / palm / arm pointing towards the device; the user moving towards the device, such as the distance between the user and the device becoming smaller, etc., one or more of the following items.
[0057] As an example, before executing step S120, you can first determine based on the image data whether there is a user in the scene whose distance from the device is less than a threshold (such as 1m). If it is determined that there is a user in the scene whose distance from the device is less than the threshold, execute step S120 again, such as determining whether the user's body movements indicate the device.
[0058] Alternatively, other methods may be used to determine whether a user who desires voice interaction with the device exists in the scene. For example, the user's identity information may be recognized from the user's facial image in the image data, and based on this identity information, it may be determined whether the user has previously used the device, such as whether the user has registered on the device. If the determination result is that the user has previously used the device, then the user is determined to be a user who desires voice interaction with the device.
[0059] If it is determined based on the image data that there is no user in the scene who desires to perform voice interaction with the device, the process may return to step S120 and continue to determine based on the real-time collected image data whether there is a user in the scene who desires to perform voice interaction with the device.
[0060] The user who expects to interact with the device by voice as determined in step S120 may not be the user who actually expects to interact with the device by voice. Therefore, the present disclosure proposes that when it is determined in step S120 that there is a user who expects to interact with the device by voice in the scene, step S130 can be executed to detect whether the user (i.e., the user who expects to interact with the device by voice as determined in step S120) has uttered a voice and / or whether the user has an intention to pronounce a voice, thereby further determining whether the user expects to interact with the device by voice, thereby reducing the false wake-up rate.
[0061] Taking step S120 as an example, determining whether there is a user in the scene who desires to interact with the device by voice by identifying the user's body movements, if the user's body movements indicate the device and the user makes a voice or has an intention to pronounce, then it can be determined with a high probability that the user desires to interact with the device by voice, thereby reducing the false wake-up rate.
[0062] When detecting whether the user has uttered a voice and / or whether the user has an intention to pronounce a voice, the detection can be performed based only on image data, or based only on audio data obtained by collecting audio from the scene in which the device is located, or the detection can be performed in combination with image data and audio data.
[0063] That is, based on the audio data and / or image data, it can be determined whether the user has uttered a voice and / or whether the user has an intention to pronounce a voice. Thus, the present disclosure can also obtain audio data. Wherein, when the device targeted by the voice interaction method has a sound pickup function, the above-mentioned audio data can be obtained by the device through audio collection. When the device does not have a sound pickup function, the above-mentioned audio data can also be obtained by audio collection by other devices in the same scene as the device.
[0064] As an example, Figure 1The voice interaction method shown can be executed by a first device with a sound pickup function. Initially, the sound pickup function of the first device can be turned off. When it is determined based on step S120 that there is a user who wishes to interact with the device through voice in the scene, the sound pickup function of the first device can be turned on to collect audio data so as to determine whether the user has made a voice and / or whether the user has a pronunciation intention based on the audio data, thereby reducing device consumption. Alternatively, the first device can be executed Figure 1 The image acquisition function and sound pickup function of the first device of the voice interaction method shown can also be turned on at the same time, wherein only audio data near the acquisition time of the image data used by the user who is determined to desire to interact with the device by voice (such as the time difference from the acquisition time is within a predetermined threshold time range) can be saved. This can reduce the data storage pressure on the device side while meeting the needs.
[0065] Taking the example of determining whether a user has uttered speech and / or intended to pronounce speech based on image data, it is possible to determine whether the user has uttered speech and / or intended to pronounce speech by identifying the user's facial movements or gestures in image data (e.g., a sequence of image data including multiple frames of image data). For example, it is possible to detect the user's facial image in the multiple frames of image data and analyze the user's lip movement information to determine whether the user has uttered speech and / or intended to pronounce speech. For example, if the user's lips show signs of movement, it can be determined that the user has the intention to speak by lip movement. Furthermore, if continuous signs of lip movement are detected based on the multiple frames of image data, it can be determined that the user is speaking. Optionally, the sound pickup function of the device used to collect audio data can be turned off by default. When the user's lips show signs of movement (e.g., lips are open) for the first time, it can be determined that the user has the intention to speak, and the device's sound pickup function can be turned on again to collect as complete a user's pronunciation data as possible while reducing device power consumption.
[0066] Taking the example of determining whether the user has uttered a voice and / or whether the user has an utterance intention based on audio data, it is possible to determine whether the user has uttered a voice by detecting whether voice data exists in the audio data. For example, if no voice data is detected, it can be determined that the user has not uttered a voice.
[0067] Taking the example of detecting whether the user has uttered a voice based on image data and audio data, the attribute information used to identify the identity of the user (i.e., the user who desires to perform voice interaction with the device as determined in step S120) can be determined based on the image data, and the voice features of the voice data in the audio data can be determined. It is then determined whether the user identity represented by the voice features is consistent with the user identity represented by the attribute information. When the determination result is consistent, it can be determined that the user has uttered a voice.
[0068] The voice features mentioned here may be, but are not limited to, voice features such as timbre and voiceprint that can characterize the identity of the user to a certain extent. Attribute information refers to information that is determined based on image data and can characterize the identity of the user to a certain extent, such as information used to characterize the group to which the user belongs (such as whether he is male or female, a child or a young person, etc.). It is possible to determine whether the identity represented by the attribute information is consistent with the identity represented by the voice features (or whether there is a conflict), to determine whether the user has made a voice and / or whether the user has the intention to pronounce. For example, if the user who desires to interact with the device is determined to be a child based on the image data in step S120, but the voice features of the acquired audio data represent the voice of an adult male, it can be determined that the user (i.e., the child) has not made a voice.
[0069] In addition, the attribute information may also be the voice features of the user further determined after the identity of the user is determined based on the image attributes. In this way, the voice features can be directly compared with the voice features extracted from the acquired audio data, and whether the user has made a voice can be determined by comparing whether the two are similar.
[0070] Determining whether the voice features determined based on the audio data are consistent with the attribute information determined based on the image data for identifying the user's identity can be regarded as an alignment detection of the audio data and the image data, which can reduce the false awakening rate.
[0071] For example, identity information such as facial images, timbre, and voiceprints of users who have used the device can be pre-stored. When a user who desires to interact with the device through voice is present in a scene, the image data can be used to determine whether the user has previously used the device. If the user has previously used the device, the pre-stored timbre and / or voiceprint of the user can be extracted and the timbre and / or voiceprint determined based on the currently collected audio data can be compared with the extracted timbre and / or voiceprint. If the result is consistent, it can be determined that the user has uttered a voice. If the user has not previously used the device, it can be determined whether the voice features, such as the timbre and / or voiceprint, determined based on the collected audio data, match the attribute information used to identify the user determined based on the image data. In other words, it can be determined whether the user identity represented by the voice features is consistent with the user identity represented by the attribute information. Based on this, it can be determined whether the user has uttered a voice. If it is determined that the user has uttered a voice, the user's identity information, such as the facial image, timbre, and voiceprint, can also be stored for subsequent determination.
[0072] As an example, (spectral) feature extraction can be performed on the speech data in the audio data to obtain speech (spectral) feature data; feature extraction can be performed on the user's facial image in the image data to obtain facial image feature data; the speech (spectral) feature data and facial image feature data are input into a pre-trained machine learning model to obtain a recognition result output by the machine learning model that indicates whether the user has uttered speech. The machine learning model can be, but is not limited to, an LSTM time series network.
[0073] The present disclosure may adopt one or more of the above-mentioned methods to detect whether the user has uttered a voice and / or whether the user has an intention to pronounce a voice. When adopting the above-mentioned multiple methods to detect whether the user has uttered a voice and / or whether the user has an intention to pronounce a voice, a weight value for characterizing the credibility of the detection result may be set for the detection results of different detection methods, and then the weight values of the same detection results of all detection methods are accumulated, and the detection result with the larger value is used as the final detection result. It is also possible to determine that the user has uttered a voice or that the user has an intention to pronounce a voice only when the detection results of the above-mentioned multiple detection methods are all that the user has uttered a voice or that the user has an intention to pronounce a voice.
[0074] When no user speech is detected or the user has an intention to pronounce, the process returns to step S120 and continues to determine whether there is a user in the scene who desires to interact with the device by voice based on the real-time collected image data.
[0075] In step S140, when it is detected that the user has made a speech or the user has a speech intention, the interactive service for the device is started. Starting the interactive service for the device means waking up or activating the interactive process of the device to provide the user with the interactive service.
[0076] Interactive services can be based solely on voice data, i.e., traditional voice interactive services, or they can be based on multimodal data including voice data, i.e., multimodal interactive services.
[0077] Taking the interactive service based solely on voice data (i.e., voice interaction service) as an example, the voice interaction process generally includes steps such as voice recognition, semantic analysis, command generation, and command execution.
[0078] Taking the execution of the entire interaction process on the device side as an example, starting the interactive service for the device may include: the device performing voice recognition on the audio data (or voice data) obtained by picking up sound, performing semantic analysis on the voice recognition results, identifying the user's operation intention, and controlling the device to perform operations corresponding to the user's operation intention.
[0079] Taking the interactive process executed collaboratively by the device and the cloud as an example, starting the interactive service for the device may include: uploading the audio data (or voice data) obtained by picking up sound to the server, having the server perform voice recognition, performing semantic analysis on the voice recognition results, identifying the user's operation intention, generating interactive instructions corresponding to the user's operation intention, and sending the interactive instructions to the device, so that the device executes the operation corresponding to the interactive instructions.
[0080] In order to further reduce the false awakening rate, it is also possible to determine whether the user is performing voice interaction with the device before performing semantic analysis on the speech recognition results. Determining whether the user is performing voice interaction with the device is also to determine whether the user is having a conversation with the device, that is, to perform human-computer conversation intention recognition on the user. Among them, it is possible to determine whether the user is currently performing voice interaction with the device based on the ASR text result obtained by speech recognition. When it is determined that the user is not performing voice interaction with the device, the interactive service can be terminated, that is, the device can return to the initial state (unawakened state), and at this time, it can return to step S110.
[0081] Taking the initiation of a multimodal interactive service for a device as an example, the multimodal interaction process may include steps such as speech recognition, user attribute analysis, determination of human-computer dialogue intent, semantic parsing, command generation, and command execution. If it is determined that the user is not communicating with the device, the interactive service may be terminated, i.e., the device may return to its initial state (not awakened state), and the process may return to step S110.
[0082] Taking into account the impact of noise in the voice recording environment on the judgment of human-computer dialogue intention, for example, the recorded voice data may contain the sound of TV dramas, phone conversations or noise signals. Pure speech recognition (ASR / NLU) in a noisy environment will be greatly disturbed, resulting in a decrease in online recognition rate. In order to further reduce false wake-ups in noisy environments, the present disclosure proposes that user attribute analysis can be used to obtain multimodal information related to the user (i.e., the actual speaker), such as the speaker's orientation (such as face orientation, sound source direction), the distance between the speaker and the device, facial attributes, lip movement features and other multimodal information. Combined with the multimodal information of the actual speaker, false wake-ups in noisy environments can be reduced, and the recognition effect of human-computer dialogue intention can be enhanced.
[0083] Taking the execution of the entire interaction process on the device side as an example, starting a multimodal interaction service for the device may include: the device performs voice recognition on the audio data (or voice data) obtained by picking up sound, and determines attribute information related to any one of the user's identity, action, and posture based on the image data (such as user identity attributes, eye direction, sound source direction, distance between the user and the device, facial attributes, lip movement characteristics, and other multimodal information); based on the voice recognition results (i.e., the ASR text results obtained based on voice recognition) and attribute information, determines whether the user is performing voice interaction with the device; when the result of the judgment is that the user is performing voice interaction with the device, the user's operation intention is identified by semantically parsing the voice recognition results, and an interaction instruction corresponding to the user's operation intention is generated; finally, the device executes the operation corresponding to the interaction instruction. Among them, before performing voice recognition on the audio data, the device can also process the audio data based on the determined attribute information to enhance the audio portion corresponding to the user in the audio data and filter out other noise. For example, based on the direction and distance of the face, the voice signal near the face can be enhanced while the sound in other directions can be filtered out; the voice signal corresponding to the speaker can also be enhanced based on the user identity attributes, that is, the identity of the speaker can be determined based on the user identity attributes, so that the audio signal corresponding to the speaker's identity can be enhanced and the noise that does not belong to the speaker can be suppressed. In this way, speech recognition can be performed on the audio data that has been processed based on the user attribute information (that is, multimodal information), thereby reducing the impact of noise on speech recognition and improving the accuracy of the ASR text results.
[0084] Taking the interactive process executed collaboratively by the device and the cloud as an example, initiating a multimodal interactive service for the device may include: the device uploads the audio data (or voice data) obtained by sound pickup and the attribute information related to the user determined based on the image data (such as user identity attributes, eye direction, and the distance between the user and the device, etc.) to the server; the server performs voice recognition and determines whether the user is performing voice interaction with the device based on the voice recognition result (i.e., the ASR text result obtained based on voice recognition) and the attribute information; when the result of the judgment is that the user is performing voice interaction with the device, the user's operation intention is identified by semantically parsing the voice recognition result, and an interaction instruction corresponding to the user's operation intention is generated. The interaction instruction is then sent to the device, and the device executes the operation corresponding to the interaction instruction. Among them, before performing voice recognition on the audio data, the server can also process the audio data based on the determined attribute information to enhance the audio portion corresponding to the user in the audio data and filter out other noise. For example, based on the direction and distance of the face, the voice signal near the face can be enhanced while the sound in other directions can be filtered out; the voice signal corresponding to the speaker can also be enhanced based on the user identity attributes, that is, the identity of the speaker can be determined based on the user identity attributes, so that the audio signal corresponding to the speaker's identity can be enhanced and the noise that does not belong to the speaker can be suppressed. In this way, speech recognition can be performed on the audio data that has been processed based on the user attribute information (that is, multimodal information), thereby reducing the impact of noise on speech recognition and improving the accuracy of the ASR text results.
[0085] Therefore, the multimodal information obtained through user attribute analysis can, on the one hand, be used to process the audio data obtained by picking up sound, reduce the impact of noise on speech recognition, and improve the accuracy of ASR text results; on the other hand, it can be combined with the ASR text results obtained by speech recognition for human-computer dialogue intention recognition, thereby enhancing the effect of human-computer dialogue intention recognition.
[0086] Taking the example of audio data collected by a device and audio data uploaded by the device being received by a server, when it is determined that the user has not interacted with the device by voice, the server can also send a pickup termination instruction to the device to reduce device resource consumption and reduce the amount of data transmission between the device and the server.
[0087] As an example, the present disclosure may also determine user-related attribute information such as facial expression, information used to identify the user, eye direction, and the distance between the user and the device based on the image data. During step S140, this attribute information may be uploaded to a server along with the audio data. The server then determines whether the user has engaged in voice interaction with the device based on the speech recognition results and attribute information from the audio data. For example, the server may input the speech recognition results and attribute information into a pre-trained human-computer dialogue intent recognition model, which then determines whether the user has engaged in voice interaction with the device.
[0088] Optionally, if a user who wishes to interact with the device through voice is determined to be present in the scenario, the user's identity information can be identified based on image information, such as through face detection, fingerprint recognition, retinal recognition, etc. The user's identity information is then verified, and only after the identity is authenticated can the interactive service for the device be started. This ensures that only users who meet specific conditions (such as adults, or users who have registered on the device, with registration content including voiceprint ID and face ID) can wake up the device.
[0089] Optionally, before performing semantic parsing on the speech recognition results, or before performing speech recognition on the uploaded audio (speech) data, or before uploading the audio data, it is also possible to determine whether the acquisition time of the image data (which can be a time point or a time period) and the acquisition time of the audio data (which can also be a time point or a time period) used to determine the user's desired voice interaction with the device are synchronized, that is, to determine whether the two acquisition times are close, such as whether the time difference between the two acquisition times is within a predetermined threshold time range, or whether there is an overlapping period between the two acquisition times. When the time difference is within the predetermined threshold time range, or there is an overlapping period, the subsequent voice interaction process is executed.
[0090] Figure 2 A schematic diagram of an application scenario according to an embodiment of the present disclosure is shown.
[0091] like Figure 2 As shown, in scenario 1, user A is talking to user B, and neither user A nor user B is looking at the smart speaker. Based on the voice interaction method disclosed in this disclosure, the smart speaker will be in a dormant state, that is, the voice interaction service of the smart speaker will not be activated.
[0092] In scenario 2, user B points his gaze toward the smart speaker and speaks a voice. Based on the voice interaction method disclosed herein, the smart speaker is activated, i.e., the voice interaction service of the smart speaker is started.
[0093] As a result, users do not need to say the wake-up word, but only need to make body movements to instruct the smart speaker when saying the voice interaction command, such as looking at the smart speaker, to directly interact with the smart speaker by voice.
[0094] Figure 3 A schematic flowchart of a voice interaction method according to another embodiment of the present disclosure is shown. Figure 3 The method shown can be performed collaboratively by a device and a server. The device can be a multimodal device with image acquisition and audio acquisition functions, that is, a first device, such as a smart speaker.
[0095] In this embodiment, the status of the device may include an initial state, a listening state, and an activated state.
[0096] 1. Initial state
[0097] In the initial state, the device only enables image acquisition. It then performs face detection on the captured image signals and identifies the user's gaze direction to determine whether the user is looking at the device. Optionally, face detection, such as face and eye recognition, can be performed when the user approaches the device (e.g., when the distance between the user and the device is less than a threshold).
[0098] 2. Listening attitude
[0099] When the device detects that the user is looking at the device, it enters a listening state. Once in the listening state, the device activates its sound pickup function to collect audio signals. In the initial or listening state, it can also perform operations to determine attribute information used to identify the user based on the image signal.
[0100] In the listening state, speech spectrum features can be extracted from the audio signal, and facial features can be extracted from the image signal. The feature extraction results are then fed into an LSTM time series network for speech recognition. Optionally, the network can identify whether the current environment is noisy to improve the performance in noisy environments.
[0101] 3. Activation state
[0102] When the device recognizes the user speaking, it enters the activated state, which is also the awakened state.
[0103] After the device enters the activated state, it uploads user attributes and voice to the cloud (i.e., server), which uses the voice recognition algorithm to recognize the voice content, i.e., ASR text.
[0104] The cloud can use original speech spectrum features, ASR text, facial expressions and other attribute information as input to the human-computer dialogue intention module algorithm to comprehensively determine whether the user is chatting with the device.
[0105] If the process determines that the user is not chatting with the device (i.e., the user is not interacting with the device), a command is sent to the client to mute the microphone and stop picking up audio. At this point, the device can return to its initial state.
[0106] When the judgment result of this process is that the user is chatting with the device, the semantic understanding algorithm can be used to identify the user's operation intention, generate instructions corresponding to the user's operation intention, and send them to the device to execute the instructions.
[0107] Therefore, the present disclosure can start switching the device state only when it recognizes that a face is looking at the device, and then use multi-modal algorithm solutions such as face recognition, eye contact, audio and video alignment to comprehensively determine whether to start the interactive service, thereby reducing the amount of cloud data transmission. In addition, during the voice interaction implementation process, the human-computer intention recognition algorithm is used to determine whether the user is speaking to the device, which can greatly reduce false wake-ups. The user does not need to learn or perceive the entire process. They only need to indicate the device with body movements, such as looking at the device, and they can chat with the device at any time, which is more in line with the way natural people interact.
[0108] Figure 4 A schematic flow chart of a voice interaction method according to another embodiment of the present disclosure is shown. In this embodiment, wake-up-free voice interaction with a second device can be achieved with the help of a first device. For the first device and the second device, please refer to the relevant description above and will not be repeated here.
[0109] It should be noted that the first device and the second device can be placed in different locations within the same scene. Alternatively, the second device can be a generally fixed-position smart home appliance, such as a smart TV or smart kitchen appliance. The first device can be a mobile device, such as a smart speaker. After the first and second devices are placed in the same scene, the first device can determine the relative positional relationship between the first and second devices based on its own position and image data obtained by imaging the second device. In other words, the position of the second device is known to the first device.
[0110] See also Figure 4 In step S310, the first device may collect image data for representing the scene in which the second device is located. For example, the first device may image a spatial area near the second device to obtain image data.
[0111] In step S320, it is determined whether there is a user in the image data who indicates the second device through body movements. The user who indicates the second device through body movements is a user who desires to interact with the second device. The specific determination process can be found in the above description.
[0112] In step S330, if it is determined that there is a body movement indicating the user of the second device, it is detected whether the user speaks, that is, whether the user pronounces or intends to pronounce. The specific detection method can be found in the relevant description above.
[0113] When it is detected that the user is speaking, step S340 is executed to upload the collected audio data to the server.
[0114] In step S350, the server may perform speech recognition on the audio data uploaded by the first device.
[0115] In step S360, it is determined whether the user is having a conversation with the second device based on the voice recognition result.
[0116] When the determination result is that the user is not communicating with the second device, the server may execute step S370 to send a sound pickup termination instruction to the first device. In response to receiving the sound pickup termination instruction, the first device may execute step S380 to stop sound pickup and then return to step S320.
[0117] When the determination result is that the user is having a conversation with the second device, the server may execute step S390 to perform semantic analysis on the speech recognition result to identify the user's operation intention.
[0118] In step S393, the server may send an interaction instruction corresponding to the user's operation intention to the second device, and the second device may execute step S395 to perform the operation corresponding to the interaction instruction. Alternatively, the server may send an interaction instruction corresponding to the user's operation intention to the first device, and the first device may send the interaction instruction to the second device.
[0119] Figure 5 Shown Figure 4 A schematic diagram of an application scenario of the voice interaction method shown.
[0120] like Figure 5 As shown, smart speakers can provide wake-up detection services for IoT devices that do not have multimodal data collection capabilities, such as smart sockets and smart buttons, by executing the wake-up-free voice interaction method disclosed herein. Wake-up detection refers to detecting, based on image data, whether there is a user in the scene who wishes to interact with the IoT device through voice, and whether the user has spoken and / or intended to speak. The specific detection process can be found in the relevant description above.
[0121] If the wake-up detection passes, interactive services for the IoT device can be initiated, uploading voice, user attribute, and other data to the server. The server then performs conversation intent recognition and determines whether the user is in a conversation with the IoT device. If the user is in a conversation with the IoT device, semantic parsing, command generation, and command issuance are performed to control the IoT device to perform the operation corresponding to the user's operation intention.
[0122] Therefore, for IoT devices that do not have multimodal data collection capabilities, wake-up-free voice interaction can also be achieved for such devices based on the present disclosure.
[0123] The present disclosure also proposes a voice interaction method, including: judging whether there is a user who desires to interact with the device by voice in the scene based on image data used to represent the scene in which the device is located; when it is determined that there is a user who desires to interact with the device by voice in the scene, starting an interactive service for the device. Figure 1 The difference between the voice interaction method shown in the embodiment is that the voice interaction method of this embodiment may not include Figure 1 Step S130 shown in FIG. 1 , the details of the solution can be found in the above combined Figure 1 The relevant description will not be repeated here.
[0124] The present disclosure may also be implemented as a voice interaction system.
[0125] Figure 6 A structural diagram of a voice interaction system according to an embodiment of the present disclosure is shown. Figure 6 The voice interaction system shown is Figure 3 The voice interaction method shown corresponds to Figure 6 The voice interaction system shown can be used to perform Figure 3 The voice interaction method shown.
[0126] See also Figure 6 , the voice interaction system 600 includes a first device 610 and a server 620.
[0127] The first device 610 is used to image the scene in which the first device is located, and determine whether there is a user in the scene who desires to interact with the first device through voice based on the obtained image data. When it is determined that there is a user in the scene who desires to interact with the first device through voice, it is detected whether the user makes a voice and / or whether the user has an intention to pronounce. When it is detected that the user makes a voice or the user has an intention to pronounce, the collected audio data is sent to the server.
[0128] Optionally, when it is determined that there is a user who desires to perform voice interaction with the first device in the scene, the audio collection function (ie, sound pickup function) of the first device 610 may be activated to start collecting audio data.
[0129] The server 620 is used to receive the audio data sent by the first device, perform voice recognition on the audio data, and determine whether the user has performed voice interaction with the first device based on the voice recognition result. When it is determined that the user has not performed voice interaction with the first device, a sound pickup termination instruction is sent to the first device 610. The first device 610 is also used to stop collecting audio data in response to receiving the sound pickup termination instruction, that is, to turn off sound pickup.
[0130] For the operations that can be performed by the first device 610 and the server 620 and related details, please refer to the relevant description above and will not be repeated here.
[0131] Figure 7 A structural diagram of a voice interaction system according to another embodiment of the present disclosure is shown. Figure 7 The voice interaction system shown is Figure 4 The voice interaction method shown corresponds to Figure 7 The voice interaction system shown can be used to perform Figure 4 The voice interaction method shown.
[0132] See also Figure 7 , the voice interaction system 700 includes a first device 710, a second device 720, and a server 730. The first device 710 and the second device 720 can be arranged in the same scene.
[0133] The first device 710 is used to image the scene and determine whether there is a user in the scene who desires to interact with the second device 720 by voice based on the obtained image data. When it is determined that there is a user in the scene who desires to interact with the second device 720 by voice, it is detected whether the user makes a voice and / or whether the user has an intention to pronounce. When it is detected that the user makes a voice or the user has an intention to pronounce, the collected audio data is uploaded to the server.
[0134] Optionally, when it is determined that there is a user who desires to perform voice interaction with the first device in the scene, the audio collection function (ie, sound pickup function) of the first device 710 may be activated to start collecting audio data.
[0135] The server 730 is used to receive the audio data sent by the first device 710, perform voice recognition on the audio data, and determine whether the user has performed voice interaction with the second device 720 based on the voice recognition result. When it is determined that the user has not performed voice interaction with the second device 720, a pickup termination instruction is sent to the first device 710.
[0136] The server 730 is further configured to, when determining that a user is engaging in voice interaction with the second device 720, perform semantic analysis on the voice recognition results to identify the user's intended operation and send an interaction instruction corresponding to the user's intended operation to the second device 720. The second device 720 is configured to receive the interaction instruction sent by the server and perform the operation corresponding to the interaction instruction. For details on the operations that can be performed by the first device 710 and the server 720, please refer to the relevant description above and will not be repeated here.
[0137] The present disclosure may also be implemented as a smart device. Figure 8 A schematic diagram of the structure of a smart device according to an embodiment of the present disclosure is shown. Figure 8 The illustrated smart device corresponds to the first device described above and refers to a device with multimodal data acquisition capabilities, such as a smart speaker with a camera. The following is a brief description of the functional modules that can be included in the smart device and the operations performed by the functional modules. For details on the operations that can be performed by the smart device and the operations performed by each functional module in the smart device, please refer to the relevant description above.
[0138] See also Figure 8 The smart device 800 includes a communication module 810 , a sound pickup module 820 , an imaging module 830 and a processor 840 .
[0139] In one embodiment of the present disclosure, the smart device 800 implements wake-up-free voice interaction.
[0140] Specifically, the imaging module 830 is used to image the scene in which the smart device is located to obtain image data. The processor 840 is used to determine whether there is a user in the scene who desires to interact with the smart device through voice based on the image data. When it is determined that there is a user in the scene who desires to interact with the smart device through voice, it detects whether the user has uttered a voice and / or whether the user has an intention to utter a voice. When it is detected that the user has uttered a voice or has an intention to utter a voice, the audio data collected by the sound pickup module 820 is uploaded to the server through the communication module 810, and the server performs voice recognition on the audio data.
[0141] Optionally, the sound pickup function of the sound pickup module 820 is in an off state in the initial state. When it is determined that there is a user who desires to interact with the smart device through voice in the scene, the sound pickup module 820 is controlled to turn on the sound pickup function to collect audio data.
[0142] The communication module 810 is also configured to receive an interaction instruction or a sound pickup termination instruction issued by the server. The server also determines, based on the speech recognition results, whether the user is engaging in voice interaction with the smart device 800. The interaction instruction is sent by the server when it determines that the user is engaging in voice interaction with the smart device 800, and the sound pickup termination instruction is sent by the server when it determines that the user is not engaging in voice interaction with the smart device 800. In response to receiving the sound pickup termination instruction, the processor 840 may control the sound pickup module 820 to switch to an off state, i.e., to disable the sound pickup function and stop collecting audio data.
[0143] In another embodiment of the present disclosure, the smart device 800 can implement wake-up-free voice interaction for an IoT device that does not have a multimodal data collection function. The smart device 800 can be adapted to be deployed in the same scene as the IoT device.
[0144] The imaging module 830 can be used to image the scene in which the IoT device is located to obtain image data. The processor 840 can be used to determine whether there is a user in the scene who desires to interact with the IoT device through voice based on the image data. When it is determined that there is a user in the scene who desires to interact with the IoT device through voice, the processor 840 detects whether the user has uttered a voice and / or whether the user has an intention to utter a voice. When it is detected that the user has uttered a voice or has an intention to utter a voice, the processor uploads the audio data collected by the sound pickup module 820 to the server via the communication module 810, and the server performs voice recognition on the audio data.
[0145] Optionally, the sound pickup function of the sound pickup module 820 is in an off state in the initial state. When it is determined that there is a user who desires to interact with the IoT device through voice in the scene, the sound pickup module 820 is controlled to turn on the sound pickup function to collect audio data.
[0146] Communication module 810 is also configured to receive a sound pickup termination instruction from the server. The server also determines, based on the voice recognition results, whether the user is engaging in voice interaction with the IoT device. The sound pickup termination instruction is sent by the server if the server determines that the user is not engaging in voice interaction with the IoT device. In response to receiving the sound pickup termination instruction, processor 840 can control sound pickup module 820 to switch to an off state, i.e., disabling the sound pickup function and stopping audio data collection.
[0147] The voice interaction method disclosed herein can also be implemented as a voice interaction device.
[0148] Figure 9 The following is a schematic diagram showing the structure of a voice interaction device according to an embodiment of the present disclosure. Figure 9The voice interaction device 900 shown can be set on a device that supports multimodal data collection. The functional modules of the voice interaction device 900 can be implemented by hardware, software, or a combination of hardware and software that implements the principles of the present disclosure. It can be understood by those skilled in the art that Figure 9 The functional modules described herein may be combined or divided into sub-modules to implement the principles of the invention. Therefore, the description herein may support any possible combination, division, or further limitation of the functional modules described herein.
[0149] The following is a brief description of the functional modules that the voice interaction device 900 may have and the operations that each functional module may perform. For the details involved, please refer to the relevant description above and will not be repeated here.
[0150] See also Figure 9 The voice interaction device 900 includes an acquisition module 910 , a judgment module 920 , a detection module 930 and a start module 940 .
[0151] The acquisition module 910 is used to acquire image data, which includes information representing the scene in which the device is located. The judgment module 920 is used to determine, based on the image data, whether there is a user in the scene who desires voice interaction with the device. The detection module 930 is used to detect whether the user has uttered a voice and / or whether the user has an intention to utter a voice, if it is determined that there is a user in the scene who desires voice interaction with the device. The activation module 940 is used to activate the interactive service for the device when it is detected that the user has uttered a voice or has an intention to utter a voice.
[0152] The determination module 920 can identify the body movements of users in the scene based on the image data, and determine whether there is a user in the scene who desires to interact with the device through voice based on the results of the body movement identification of the user in the scene. For example, the determination module 920 can determine whether there is a user in the scene who is pointing the device through body movement based on the results of the body movement identification of the user in the scene, where the user who is pointing the device through body movement is the user who desires to interact with the device through voice.
[0153] The detection module 930 can determine whether the user has made a voice and / or whether the user has a pronunciation intention based on audio data and / or image data, wherein the audio data is obtained by collecting audio of the scene in which the device is located.
[0154] As an example, the detection module 930 can determine attribute information used to identify the identity of the user based on the image data; determine the voice features of the voice data in the audio data; and determine whether the user identity represented by the voice features is consistent with the user identity represented by the attribute information. When the judgment result is consistent, it is determined that the user has made a voice.
[0155] As an example, the detection module 930 may also determine whether the user has uttered a voice and / or whether the user has an intention to pronounce a voice by recognizing the facial image of the user in the image data.
[0156] As an example, the detection module 930 can also perform feature extraction on the voice data in the audio data to obtain voice feature data; perform feature extraction on the facial image of the user in the image data to obtain facial image feature data; input the voice feature data and the facial image feature data into a pre-trained machine learning model to obtain a recognition result output by the machine learning model to characterize whether the user has uttered a voice.
[0157] The startup module 940 may include an upload module and a receiving module. The upload module is used to upload the collected audio data to the server, which performs voice recognition on the audio data and determines whether the user has engaged in voice interaction with the device based on the voice recognition results. The receiving module is used to receive interaction instructions or audio pickup termination instructions issued by the server. The interaction instruction is sent by the server when it determines that the user has engaged in voice interaction with the device, and the audio pickup termination instruction is sent by the server when it determines that the user has not engaged in voice interaction with the device.
[0158] The voice interaction device 900 may also include an attribute determination module for determining attribute information related to any one of the user's identity, action, and posture based on the image data. The upload module may also upload the attribute information to a server, and the server determines whether the user performs voice interaction with the device based on the voice recognition result and the attribute information.
[0159] Figure 10 The following is a schematic diagram showing the structure of a voice interaction device according to an embodiment of the present disclosure. Figure 10 The voice interaction device 1000 shown can be set up on the server side. The functional modules of the voice interaction device 1000 can be implemented by hardware, software or a combination of hardware and software that implements the principles of the present disclosure. It can be understood by those skilled in the art that Figure 10 The functional modules described herein may be combined or divided into sub-modules to implement the principles of the invention. Therefore, the description herein may support any possible combination, division, or further limitation of the functional modules described herein.
[0160] The following is a brief description of the functional modules that the voice interaction device 1000 may have and the operations that each functional module may perform. For the details involved, please refer to the relevant description above and will not be repeated here.
[0161] See also Figure 10The voice interaction device 1000 includes a receiving module 1010 , a voice recognition module 1020 , a judgment module 1030 and a sending module 1040 .
[0162] In one embodiment of the present disclosure, the receiving module 1010 is used to receive audio data uploaded by a device. The audio data is uploaded when the device determines, based on image data representing the scene in which the device is located, that there is a user who desires to interact with the device by voice, and detects that the user utters a voice or has an intention to pronounce a voice. The speech recognition module 1020 is used to perform speech recognition on the audio data. The judgment module 1030 is used to determine whether the user has interacted with the device by voice based on the speech recognition results. The sending module 1040 is used to send a sound pickup termination instruction to the device when it is determined that the user has not interacted with the device by voice.
[0163] The voice interaction device 1000 may also include a semantic parsing module. The semantic parsing module is configured to, upon determining that a user is engaging in voice interaction with the device, perform semantic parsing on the voice recognition results to identify the user's intended operation. The sending module 1040 is further configured to send interaction instructions corresponding to the user's intended operation to the device.
[0164] Receiving module 1010 may also receive attribute information uploaded by the device. Attribute information is information related to any of the user's identity, actions, and posture, determined based on image data. Determination module 1030 may specifically determine whether the user has engaged in voice interaction with the device based on the speech recognition results and attribute information. For example, determination module 1030 may input the speech recognition results and attribute information into a pre-trained human-computer dialogue intent recognition model, which may then determine whether the user has engaged in voice interaction with the device.
[0165] In another embodiment of the present disclosure, the receiving module 1010 is used to receive audio data uploaded by the first device. The audio data is sent by the first device when it determines that there is a user who desires to perform voice interaction with the second device in the scene based on image data used to characterize the scene in which the second device is located, and detects that the user has uttered a voice or the user has an intention to pronounce a voice. The speech recognition module 1020 is used to perform speech recognition on the audio data. The judgment module 1030 is used to determine whether the user has performed voice interaction with the second device based on the speech recognition result. The sending module 1040 is used to send a pickup termination instruction to the first device when it is determined that the user has not performed voice interaction with the second device.
[0166] The voice interaction device 1000 may further include a semantic parsing module. The semantic parsing module is configured to, upon determining that the user is engaging in voice interaction with the second device, perform semantic parsing on the voice recognition results to identify the user's intended operation. The sending module 1040 is further configured to send an interaction instruction corresponding to the user's intended operation to the second device.
[0167] Figure 11 The following is a schematic diagram showing the structure of a voice interaction device according to an embodiment of the present disclosure. Figure 11 The voice interaction device 1100 shown can be set at the device end. The functional modules of the voice interaction device 1100 can be implemented by hardware, software or a combination of hardware and software that implements the principles of the present disclosure. It can be understood by those skilled in the art that Figure 11 The functional modules described herein may be combined or divided into sub-modules to implement the principles of the invention. Therefore, the description herein may support any possible combination, division, or further limitation of the functional modules described herein.
[0168] The following is a brief description of the functional modules that the voice interaction device 1100 may have and the operations that each functional module may perform. For the details involved, please refer to the relevant description above and will not be repeated here.
[0169] See also Figure 11 The voice interaction device 1100 includes a judgment module 1110 and a starting module 1120 .
[0170] The determination module 1110 is configured to determine, based on image data representing the scene in which the device is located, whether there is a user who desires to perform voice interaction with the device in the scene.
[0171] The judgment module 1110 can identify the body movements of users in the scene based on image data; according to the recognition results of the body movements of users in the scene, it can be judged whether there is a user with a body movement indication device in the scene, and the user with a body movement indication device is the user who expects to interact with the device through voice.
[0172] The starting module 1120 is used to start the interactive service for the device when it is determined that there is a user who desires to perform voice interaction with the device in the scene.
[0173] The startup module 1120 may include an upload module and a receiving module. The upload module is used to upload the collected audio data to the server, which performs voice recognition on the audio data and determines whether the user has performed voice interaction with the device based on the voice recognition results. The receiving module is used to receive interaction instructions or audio pickup termination instructions issued by the server. The interaction instruction is sent by the server when it determines that the user has performed voice interaction with the device, and the audio pickup termination instruction is sent by the server when it determines that the user has not performed voice interaction with the device.
[0174] The voice interaction device 1100 may also include an attribute determination module for determining attribute information related to the user based on image data. The upload module may also be used to upload the attribute information to a server, which determines whether the user is performing voice interaction with the device based on the voice recognition results and the attribute information.
[0175] Figure 12 A schematic diagram of the structure of a computing device that can be used to implement the above-mentioned voice interaction method according to an embodiment of the present invention is shown.
[0176] See also Figure 12 , the computing device 1200 includes a memory 1210 and a processor 1220 .
[0177] Processor 1220 may be a multi-core processor or may include multiple processors. In some embodiments, processor 1220 may include a general-purpose main processor and one or more special coprocessors, such as a graphics processing unit (GPU) or a digital signal processor (DSP). In some embodiments, processor 1220 may be implemented using customized circuits, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs).
[0178] The memory 1210 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. ROM may store static data or instructions required by the processor 1220 or other modules of the computer. The permanent storage device may be a readable and writable storage device. The permanent storage device may be a non-volatile storage device that retains stored instructions and data even after the computer loses power. In some embodiments, the permanent storage device uses a large-capacity storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In other embodiments, the permanent storage device may be a removable storage device (such as a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory may store some or all instructions and data required by the processor during operation. In addition, the memory 1210 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks may also be used. In some embodiments, the memory 1210 may include a readable and / or writable removable storage device, such as a compact disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and transient electronic signals transmitted wirelessly or wired.
[0179] The memory 1210 stores executable code, which, when processed by the processor 1220 , enables the processor 1220 to execute the voice interaction method described above.
[0180] The voice interaction method, voice interaction apparatus, voice interaction system, intelligent device, and computing device according to the present invention have been described above in detail with reference to the accompanying drawings.
[0181] In addition, the method according to the present invention may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing the above steps defined in the above method of the present invention.
[0182] Alternatively, the present invention can also be implemented as a non-transitory machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium) on which executable code (or computer program, or computer instruction code) is stored. When the executable code (or computer program, or computer instruction code) is executed by a processor of an electronic device (or computing device, server, etc.), the processor executes the various steps of the above-mentioned method according to the present invention.
[0183] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or combinations of both.
[0184] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architecture, functions and operations of the systems and methods according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0185] While various embodiments of the present invention have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A voice interaction method for interacting with a device, comprising: Acquire image data, the image data including information representing a scene in which the device is located; determining, based on the image data, whether there is a user in the scene who desires to perform voice interaction with the device; When it is determined that the user exists in the scene, detecting whether the user has uttered a voice and / or whether the user has an intention to pronounce a voice; When it is detected that the user has made a speech or the user has a pronunciation intention, the interactive service for the device is started, The method further comprises: Store the identity information of users who have used the device; Get audio data, Among them, the step of detecting whether the user has uttered a voice and / or whether the user has an intention to pronounce a voice includes: if the user is not a user who has used the device before, determining whether the user group represented by the voice features determined based on the collected audio data is consistent with the user group represented by the attribute information determined based on the image data for identifying the user; when the judgment result is consistent, determining that the user has uttered a voice.
2. The voice interaction method according to claim 1, wherein: The step of determining whether there is a user who desires to perform voice interaction with the device in the scene based on the image data includes: recognizing body movements of a user in the scene based on the image data; According to the recognition result of the user's body movement in the scene, it is determined whether there is a user in the scene who desires to perform voice interaction with the device.
3. The voice interaction method according to claim 1, wherein: The step of detecting whether the user has uttered a voice and / or whether the user has an intention to pronounce a voice further includes: Extracting features from the speech data in the audio data to obtain speech feature data; performing feature extraction on the facial image of the user in the image data to obtain facial image feature data; The voice feature data and the facial image feature data are input into a pre-trained machine learning model to obtain a recognition result output by the machine learning model for characterizing whether the user has uttered a voice.
4. The voice interaction method according to claim 1, wherein: The step of detecting whether the user has uttered a voice and / or whether the user has an intention to pronounce a voice comprises: By recognizing the facial movements or gestures of the user in the image data, it is determined whether the user has uttered a voice and / or whether the user has an intention to pronounce a voice.
5. The voice interaction method according to claim 1, wherein: The steps of starting the interactive service for the device include: Uploading the collected audio data to a server, which performs voice recognition on the audio data and determines whether the user has performed voice interaction with the device based on the voice recognition result; Receive an interaction instruction or a sound pickup termination instruction sent by the server, wherein the interaction instruction is sent by the server when it is determined that the user has performed voice interaction with the device, and the sound pickup termination instruction is sent by the server when it is determined that the user has not performed voice interaction with the device.
6. The voice interaction method according to claim 5, further comprising: determining attribute information related to any one of the identity, action, and posture of the user based on the image data; Among them, the step of starting the interactive service for the device also includes: uploading the attribute information to the server, and the server processing the audio data based on the attribute information to enhance the audio part of the audio data corresponding to the user, and / or the server determining whether the user performs voice interaction with the device based on the voice recognition result and the attribute information.
7. The voice interaction method according to claim 1, wherein: The image data is captured by the device, or the image data is captured by another device in the same scene as the device.
8. A voice interaction system, comprising: a first device, configured to image a scene in which the first device is located, determine, based on the obtained image data, whether a user who desires to perform voice interaction with the first device exists in the scene, and upon determining that the user exists in the scene, detect whether the user utters a voice and / or whether the user intends to utter a voice, and upon detecting that the user utters a voice or intends to utter a voice, send the collected audio data to a server; The server is configured to receive the audio data sent by the first device, perform voice recognition on the audio data, determine whether the user has performed voice interaction with the first device based on the voice recognition result, and send a voice pickup termination instruction to the first device when it is determined that the user has not performed voice interaction with the first device. The first device is further configured to stop collecting audio data in response to receiving the sound pickup termination instruction. Among them, the first device stores the identity information of the user who has used the device. If the user is not a user who has used the device before, the first device determines whether the group to which the user belongs, represented by the voice features determined based on the collected audio data, is consistent with the group to which the user belongs, represented by the attribute information used to identify the user determined based on the image data. When the judgment result is consistent, it is determined that the user has made a voice.
9. The voice interaction system according to claim 8, wherein: The server is further configured to, when determining that the user is performing voice interaction with the device, identify the user's operation intention by semantically parsing the voice recognition result, and send an interaction instruction corresponding to the user's operation intention to the first device. The first device is further configured to, in response to receiving the interaction instruction, execute an operation corresponding to the interaction instruction.
10. A voice interaction system, comprising: Second device; a first device, arranged in the same scene as the second device, the first device being configured to image the scene, determine, based on the obtained image data, whether a user who desires to perform voice interaction with the second device exists in the scene, and upon determining that the user exists in the scene, detect whether the user utters a voice and / or whether the user intends to utter a voice, and upon detecting that the user utters a voice or intends to utter a voice, upload the collected audio data to a server; The server is configured to receive audio data sent by the first device, perform voice recognition on the audio data, determine whether the user has performed voice interaction with the second device based on the voice recognition result, and send a voice pickup termination instruction to the first device when it is determined that the user has not performed voice interaction with the second device. The first device is further configured to stop collecting audio data in response to receiving the sound pickup termination instruction. In which, the first device stores the identity information of the user who has used the second device. If the user is not a user who has used the device before, the first device determines whether the group to which the user belongs, represented by the voice features determined based on the collected audio data, is consistent with the group to which the user belongs, represented by the attribute information used to identify the user determined based on the image data. When the judgment result is consistent, it is determined that the user has made a voice.
11. The voice interaction system according to claim 10, wherein: The server is further configured to, when determining that the user is performing voice interaction with the second device, identify the user's operation intention by semantically parsing the voice recognition result, and send an interaction instruction corresponding to the user's operation intention to the second device. The second device is used to receive the interaction instruction and perform an operation corresponding to the interaction instruction.
12. A smart device comprising: Communication module; Pickup module; An imaging module, configured to image the scene in which the smart device is located to obtain image data; The processor is configured to determine, based on the image data, whether there is a user in the scene who desires to interact with the smart device through voice; if it is determined that the user exists in the scene, detect whether the user has uttered a voice and / or whether the user has an intention to pronounce a voice; if it is detected that the user has uttered a voice or has an intention to pronounce a voice, upload the audio data collected by the sound pickup module to the server via the communication module, and have the server perform voice recognition on the audio data; In which, the smart device stores the identity information of users who have used the device. If the user is not a user who has used the device before, the processor determines whether the group to which the user belongs, represented by the voice features determined based on the collected audio data, is consistent with the group to which the user belongs, represented by the attribute information used to identify the user determined based on the image data. When the judgment result is consistent, it is determined that the user has made a voice.
13. The smart device according to claim 12, wherein: The communication module is also used to receive an interaction instruction or a sound pickup termination instruction issued by the server, wherein the server also determines whether the user performs voice interaction with the smart device based on the voice recognition result. The interaction instruction is sent by the server when it is determined that the user performs voice interaction with the smart device, and the sound pickup termination instruction is sent by the server when it is determined that the user does not perform voice interaction with the smart device.
14. A smart device, suitable for being deployed in the same scene as an IoT device, comprising: Communication module; Pickup module; An imaging module, configured to image the scene and obtain image data; The processor is configured to determine, based on the image data, whether there is a user in the scene who desires to perform voice interaction with the IoT device; if it is determined that the user exists in the scene, detect whether the user has uttered a voice and / or whether the user has an intention to pronounce a voice; if it is detected that the user has uttered a voice or has an intention to pronounce a voice, upload the audio data collected by the sound pickup module to the server via the communication module, and the server performs voice recognition on the audio data; In which, the smart device stores the identity information of users who have used the device. If the user is not a user whose identity information has been pre-stored, the processor determines whether the group to which the user belongs, represented by the voice features determined based on the collected audio data, is consistent with the group to which the user belongs, represented by the attribute information used to identify the user determined based on the image data. When the judgment result is consistent, it is determined that the user has made a voice.
15. The smart device according to claim 14, wherein: The communication module is also used to receive a voice pickup termination instruction issued by the server, wherein the server also determines whether the user has performed voice interaction with the Internet of Things device based on the voice recognition result, and the voice pickup termination instruction is sent by the server when it is determined that the user has not performed voice interaction with the Internet of Things device.
16. A voice interaction device, comprising: An acquisition module, configured to acquire image data, the image data including information representing a scene in which the device is located; a determination module, configured to determine, based on the image data, whether there is a user in the scene who desires to perform voice interaction with the device; a detection module, configured to detect, when determining that the user exists in the scene, whether the user has uttered a voice and / or whether the user has an intention to pronounce a voice; A startup module is used to start the interactive service for the device when it is detected that the user has made a voice or the user has a pronunciation intention, The voice interaction device also stores the identity information of users who have used the device and obtains audio data. If the user is not a user who has used the device before, the detection module determines whether the group to which the user belongs, represented by the voice features determined based on the collected audio data, is consistent with the group to which the user belongs, represented by the attribute information used to identify the user determined based on the image data. When the judgment result is consistent, it is determined that the user has made a voice.
17. A computing device comprising: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to perform the method according to any one of claims 1 to 7.
18. A non-transitory machine-readable storage medium having executable codes stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Voice interaction feedback method and system for intelligent television and computer readable medium
CN108683937A
Interactive method and equipment
CN109767774A
Management control method and device for interactive equipment
CN111968633A
Methods, apparatuses, storage mediums and terminal devices for authentication
US20210166241A1
Speech operation method for device, apparatus, and electronic device
WO2022179253A1