Voice interaction method, device and system
Through the multimodal interaction judgment model of image and sound information, the recording start and end time are dynamically adjusted, which solves the problem of repeated wake-up operations of voice interaction devices, realizes natural and direct voice interaction, and improves user experience and interaction efficiency.
Patent Information
- Application Number
- CN202010690864.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-17
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2040-07-17
AI Technical Summary
Existing voice interaction devices require repeated wake-up operations, which makes users feel that the operation is cumbersome and boring. At the same time, there are problems such as irrelevant sound recording or increased waiting time.
Through artificial intelligence based on image and sound information, using interaction judgment models, intent recognition models and dynamic threshold models, the user's interaction intention is determined, and the recording start and end times are dynamically adjusted to achieve natural voice interaction without wake-up words or operations.
It achieves accurate collection of user interaction information without additional operations, avoids the collection of irrelevant sounds, and improves interaction efficiency and user experience.
Smart Images

Figure CN113948076B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of human-computer interaction, and in particular to a voice interaction method, device, and system. Background Art
[0002] With the development of voice recognition technology and wireless networks, devices with voice interaction functions, such as smart speakers, have become popular. Among these devices with voice interaction functions, dedicated voice interaction devices (such as smart speakers) usually require the use of a wake-up word to enter the power-on state and record audio for recognition and feedback. For non-dedicated voice interaction devices (such as smartphones, car systems, smart home appliances, etc.), users usually need to perform special operations (for example, clicking a physical or virtual button) to enter the power-on state.
[0003] Regardless of the aforementioned wake-up method, each use requires repeated wake-up operations (e.g., speaking the wake-up word or clicking an interactive button), which can make users feel tedious and boring. Furthermore, since the microphone is turned on and off at a fixed time or requires the user to wait for a long period of silence, irrelevant sounds such as other people's voices may be recorded, or the user's waiting time may be unnecessarily increased.
[0004] Therefore, a solution is needed to more intelligently record voice interaction information. Summary of the Invention
[0005] A technical problem to be solved by the present disclosure is to provide a voice interaction solution that can determine the user's interaction intention based on collected image information or preferred audio and video information through artificial intelligence, and intelligently collect the interaction information accordingly.
[0006] According to a first aspect of the present disclosure, a voice interaction method is provided, including: turning on a camera to obtain image information and simultaneously turning on a microphone to obtain sound information; inputting the obtained image information and sound information into an interaction determination model; and based on the output of the interaction determination model, using a microphone to obtain sound information for voice interaction.
[0007] According to a second aspect of the present disclosure, a voice interaction method is provided, comprising: determining that someone is approaching and obtaining image information; inputting the image information into an interaction determination model; and obtaining sound information for voice interaction based on an output of the interaction determination model.
[0008] According to a third aspect of the present disclosure, a voice interaction method is provided, comprising: acquiring image information; inputting the image information into an interaction determination model; and acquiring sound information for voice interaction based on an output of the interaction determination model.
[0009] According to a fourth aspect of the present disclosure, a voice interaction method is provided, comprising: obtaining multimodal information, the multimodal information including at least two paths of information obtained simultaneously; inputting the multimodal information into an interaction determination model; and obtaining sound information for voice interaction based on an output of the interaction determination model.
[0010] According to a fifth aspect of the present disclosure, a voice interaction device is provided, including: a camera for acquiring image information, a microphone for acquiring sound information, and a processor for: inputting the image information acquired by the camera into an interaction determination model; and acquiring sound information for voice interaction via the microphone based on the output of the interaction determination model.
[0011] According to the sixth aspect of the present disclosure, a voice interaction system is provided, comprising: a voice interaction device according to the fourth aspect above; and a computing node communicating with the voice interaction device, the computing node storing the model and providing model output for the voice interaction device.
[0012] According to the seventh aspect of the present disclosure, a voice interaction model training method is provided, including: using an image of a person speaking as a positive label and an image of a person not speaking as a negative label to train an interaction determination model, so that the interaction determination model determines the recording start and recording reception time of sound information used for voice interaction based on input image information based on a recording start threshold and a recording end threshold.
[0013] According to an eighth aspect of the present disclosure, a computing device is provided, comprising: a processor; and a memory on which executable code is stored, and when the executable code is executed by the processor, the processor executes the method described in the first to fourth aspects above.
[0014] According to a ninth aspect of the present disclosure, a non-transitory machine-readable storage medium is provided, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor executes the method described in the first to fourth aspects above.
[0015] Therefore, the voice interaction solution of the present invention can determine the user's interaction intention based on image information, or preferably audio and video information, through artificial intelligence, and directly perform intelligent collection of interaction information based on this. Specifically, the interaction determination model can be used to determine the start and end time of recording, and the intention recognition model can be further used to dynamically adjust the recording start and end determination thresholds of the interaction determination model. The dynamic threshold model can also be used to adjust the threshold based on dynamic learning. Thus, natural and direct voice interaction can be achieved without saying the wake-up word or turning on the voice function. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings, wherein like reference numerals generally represent like components in the exemplary embodiments of the present disclosure.
[0017] Figure 1 A schematic flow chart of a voice interaction method according to an embodiment of the present invention is shown.
[0018] Figure 2 A determination example of the interactive determination model according to the present invention is shown.
[0019] Figure 3 An example of recognition using the intention recognition model according to the present invention is shown.
[0020] Figure 4 An output example of the dynamic threshold model according to the present invention is shown.
[0021] Figure 5 An example of three-model interaction of the multimodal self-adjusting input system according to the present invention is shown.
[0022] Figure 6 A schematic flow chart of a voice interaction method according to another embodiment of the present invention is shown.
[0023] Figure 7 A schematic flow chart of a voice interaction method according to another embodiment of the present invention is shown.
[0024] Figure 8 A block diagram showing the composition of a voice interaction device according to an embodiment of the present invention is shown.
[0025] Figure 9 A schematic structural diagram of a computing device that can be used to implement the above-mentioned voice interaction method according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0026] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although preferred embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0027] With the development of voice recognition technology and wireless networks, devices with voice interaction functions, such as smart speakers, have become popular. Among these devices with voice interaction functions, dedicated voice interaction devices (such as smart speakers) usually require the use of a wake-up word to enter the power-on state and record audio for recognition and feedback. For non-dedicated voice interaction devices (such as smartphones, car systems, smart home appliances, etc.), users usually need to perform special operations (for example, clicking a physical or virtual button) to enter the power-on state.
[0028] Regardless of the aforementioned wake-up method, each use requires repeated wake-up operations (e.g., speaking the wake-up word or clicking an interactive button), which can make users feel tedious and boring. Furthermore, since the microphone is turned on and off at a fixed time or requires the user to wait for a long period of silence, irrelevant sounds such as other people's voices may be recorded, or the user's waiting time may be unnecessarily increased.
[0029] To this end, the present application provides a voice interaction solution that can determine the user's interaction intention based on image information, or preferably audio and video information, through artificial intelligence, and directly perform intelligent collection of interaction information based on this. Specifically, an interaction determination model can be used to determine the start and end time of recording, and the intention recognition model can be further used to dynamically adjust the recording start and end determination thresholds of the interaction determination model. A dynamic threshold model can also be used to adjust the threshold based on dynamic learning. In this way, natural and direct voice interaction can be achieved without saying the wake-up word or turning on the voice function.
[0030] Figure 1 A schematic flow chart of a voice interaction method according to an embodiment of the present invention is shown. The method can be executed by a voice interaction device, in particular a voice interaction device including an image acquisition function (eg, equipped with a camera).
[0031] In step S110, the camera is turned on to capture image information, and the microphone is turned on to capture sound information. In step S120, the captured image information and sound information are input into an interaction determination model. Subsequently, in step S130, based on the output of the interaction determination model, the microphone is used to capture sound information for voice interaction.
[0032] Therefore, by using the trained model to process multimodal data including image and sound information, it is possible to directly start collecting user interaction information based on the output results of the model, thereby avoiding the need for the user to say the wake-up word additionally or manually turn on the voice interaction function.
[0033] For example, a voice interaction device can turn on a camera and microphone to acquire audio and video information at predetermined intervals, or it can turn on a camera and microphone to acquire audio and video information when a proximity sensor or other mechanism notifies it of a person approaching. The above audio and video information is fed into the interaction determination model in real time, and the interaction determination model can process and determine based on the input audio and video information. For example, when it is determined that a person's face is facing the device and is about to speak (for example, the input image frame is determined to be a person speaking, and the input audio frame is also determined to be a person speaking, or about to speak), the acquisition of audio information for interaction is started. Subsequently, the audio and video information acquired by the camera and microphone can be continuously fed into the interaction determination model, and the end time of the audio information acquisition is determined based on the output of the model. The audio information thus acquired (for example, the audio information recorded between the start of acquisition of interactive audio and the end of acquisition) can be semantically parsed locally or by a server or edge computing node, and feedback is provided based on the processing results.
[0034] The interaction determination model can include or be implemented as a supervised learning deep neural network model. Specifically, the positive labels used to train the deep neural network model include images of people speaking, and the negative labels include images of people not speaking. This allows the trained interaction determination model to initiate recording of interactions during subsequent inference based on the captured image representation of a person speaking.
[0035] It is understandable that in order to dynamically determine the end time of the sound recording, the camera can be used to continue to acquire image information while the microphone is used to acquire sound information for voice interaction. This facilitates determining the recording start time and / or recording end time of the sound information for voice interaction using the microphone based on the output of the interaction determination model. For example, at time t0 (for example, the 0th second), the camera and microphone start collecting image and sound information, and continuously feed the above information into the interaction determination model. At time t1 (for example, the 1st second), the interaction determination model determines that the user has an intention to interact (for example, based on the image frame of the user intending to speak acquired at time t1, and the audio frame with human voice), then the user voice interaction of sound information can be acquired based on the output of the interaction determination model. At time t2 (for example, the 12th second), the interaction determination model determines that the user's intention to interact ends (for example, based on the image frame of the user turning his head acquired at time t2, and the audio frame without human voice), then the acquisition of sound information user voice interaction can be ended.
[0036] In other words, during the 12 seconds from time t0 to time t2, the camera and microphone always collect image and sound information, and continuously feed the above information into the interaction judgment model. During the time period when the user is judged to have the intention to interact (that is, the 11 seconds from time t1 to time t2), the sound information collected by the microphone can also be reused as sound information for voice interaction. For example, the voice interaction device can process or upload the above 11 seconds of sound information locally, perform natural language processing (NLP) on the above sound, extract semantics, analyze intent and give corresponding feedback. The above 11 seconds of sound information can be transmitted and processed in segments, or it can be transmitted and processed all at once after the recording is completed.
[0037] Specifically, the interaction decision model may make a decision based on a threshold. Specifically, the interaction decision model may start recording when the output meets a recording start time threshold and / or end recording when the output meets a recording end time threshold based on the current image information and the sound information input.
[0038] Figure 2 FIG. 4 shows an example of a decision made by the interactive decision model according to the present invention. Figure 2 For example, the audio and video data collected by the camera and microphone equipped by the voice interaction device, for example, video frames with audio information (such as images including human faces as shown in the figure), can be input into the interaction determination model of the present invention as image sequences (for example, image frame sequences) and audio (for example, audio frames) respectively. The above-mentioned interaction determination model can calculate the recording start score based on the input image and audio, and start audio and video recording for voice interaction when the start score is greater than threshold 1 (recording start threshold). In the process of starting recording, the interaction determination model can continuously acquire images and audio, and calculate the recording end score. When the calculated recording end score is greater than threshold 2 (recording end threshold), the audio and video recording for voice interaction is stopped.
[0039] It should be understood that, although an interactive decision model is shown that can calculate both the start score and the end score based on audio and video input, in other embodiments, different models or different sub-models within the model can be used to calculate the start score and the end score respectively. In addition, in different embodiments, the above threshold can be a fixed threshold, a threshold adjusted based on experience, or a threshold that is dynamically self-adjusted based on the output of other models as described below.
[0040] Because it uses both image and audio modal information, the recording function is only enabled when, for example, both the sound (speaking) and the image (when the mouth starts moving) are triggered simultaneously, and there is a rough voice frequency and semantic similarity, the recording function will be enabled. If one of the two modalities is not met, the sound system will automatically shut down. This can accurately judge the user's interaction intent, allowing users to interact with voice without additional operations (such as saying the wake-up word or manually turning on the voice interaction function), while avoiding the inclusion of a large amount of meaningless video and sound.
[0041] In addition, in order to improve the accuracy of judging the user's interaction intention, the voice interaction solution of the present invention can further introduce other models to enhance the solution's ability to cope with different environments and user states.
[0042] To this end, in one embodiment, the voice interaction method may further include inputting the acquired sound information into an intent recognition model and adjusting the recording start time threshold and / or recording end time threshold based on the output of the intent recognition model. The intent recognition model may identify the user's intent based on the input sound information and, for example, lower the recording start threshold when it determines that the user intends to interact with the voice device, and change the recording end threshold when, for example, background noise is present. It should be understood that, in different embodiments, the higher the threshold, the higher the start and end conditions that must be met, or the lower the threshold, the higher the start and end conditions that must be met. The present invention does not limit the direction of the threshold. For example, if the intent recognition model determines that the environment is noisy based on the sound, the recording start threshold may be increased, and the interaction judgment model may only start recording when it receives more specific interactive audio and video data, and only end recording when it receives more specific end intention audio and video data. For example, the intent recognition model may also identify the semantics of the input sound information and increase the recording start threshold when the semantics are not relevant to the interaction.
[0043] In a preferred embodiment, the intent recognition model can also include image input, thereby determining intent based on the multimodal combination of audio and image information. For example, if the user has an interaction intent based on semantic judgment, but the image indicates that the user is on a phone call, the recording start threshold should still be raised.
[0044] Figure 3An example of recognition of the intent recognition model according to the present invention is shown. As shown in the figure, the intent recognition model can also use image sequences and audio as input, and output corresponding feature vectors (Embedding1), which can be compared with existing feature vectors (Embedding2) in the database. For example, a context-related function is fed into the function to determine the correlation between the two vectors. For example, the correlation of the vectors is determined based on Cosin (Embedding1, Embedding2), and "yes" and "no" are output based on the calculated result. Here, for example, noise intensity, noise frequency, human behavior or posture can be used as labels to train the intent recognition model. In the subsequent reasoning stage, the input image and audio can be processed by the intent recognition model and output as a feature vector (Embedding1) with multiple dimensions. The above vector can be compared with one or more context feature vectors (Embedding2) stored in the database for representing the intent to determine the correlation.
[0045] Here, if the output is "yes," it can indicate that the user has interaction intent, and accordingly, the threshold of the interaction determination model can be lowered, thereby making it easier for the interaction determination model to determine that the user wants to interact. If the output is "no," it can indicate that the user does not have interaction intent (although the interaction determination model may perceive the user as having intent), and accordingly, the threshold of the interaction determination model can be raised.
[0046] Although threshold adjustment can be performed directly based on the output of the intent recognition model, in a preferred embodiment, a third model, namely a dynamic threshold model, can be introduced. This model can obtain the output of the intent recognition model and dynamically adjust the recording start threshold and / or the recording end threshold.
[0047] Figure 4 An output example of the dynamic threshold model according to the present invention is shown. As shown in the figure, the dynamic threshold model in the present invention can be a reinforcement learning model.
[0048] Machine learning is a key research area in artificial intelligence. Depending on whether feedback is obtained from the system, machine learning can be categorized into three main types: supervised, unsupervised, and reinforcement learning. Supervised learning, also known as supervised learning, involves learning within a known input-output dataset, where the system is given a set of inputs and a corresponding set of outputs. Both the interaction determination model and the intent recognition model of the present invention can be implemented through supervised learning.
[0049] The opposite of supervised learning is unsupervised learning, also known as unsupervised learning. In unsupervised learning, only a set of outputs is given, without the corresponding outputs. The system automatically learns based on the internal structure of the given inputs. Supervised and unsupervised machine learning models can solve the vast majority of machine learning problems, but these two machine learning models differ significantly from the processes of human learning and biological evolution. Biological evolution is a learning process in which organisms actively explore their environment and evaluate and summarize the feedback from the environment to improve and adjust their behavior. The environment then responds to these new behaviors with new feedback, continuously adjusting. The learning model that embodies this concept is called reinforcement learning (RL) in the field of machine learning. Therefore, reinforcement learning is a machine learning model that stands alongside supervised and unsupervised learning.
[0050] The entire reinforcement learning system consists of five parts: agent, state, reward, action, and environment.
[0051] The agent is the core of the entire reinforcement learning system. It perceives the state of the environment and, based on the reinforcement signals (Reward Si) provided by the environment, learns to select an appropriate action to maximize the long-term reward value. In short, the agent uses the rewards provided by the environment as feedback to learn a series of mappings from environmental states (State) to actions (Action). The principle of action selection is to maximize the probability of future cumulative rewards. The selected action affects not only the reward at the current moment, but also the rewards at the next moment and even in the future. Therefore, the basic rule of the agent's learning process is: if an action (Action) brings a positive reward (Reward) from the environment, then this action will be strengthened; otherwise, it will be gradually weakened, similar to the principle of conditioned reflex in physics.
[0052] In the present invention, the reinforcement learning model used as a dynamic threshold model uses image information, sound information and the recognition result of the speech intention as input (state), and adjusts the value of the recording start threshold and / or the recording end threshold as behavior (action) in real time based on the correctness of the sound information obtained for voice interaction (reward). Figure 4As shown, the dynamic threshold model can have n behaviors to correspond to different recording start and end threshold values. The dynamic threshold model can give a corresponding set of values based on the currently input image and audio data, and the intention output of the intention recognition model (for example, "yes" or "no" corresponding to intention and no intention). The above values are given to the interaction judgment model, and the voice data used for interaction is obtained accordingly. Subsequently, the smoothness of the interaction of the above voice data is used to evaluate the correctness of the selection of the above behavior, and to modify the selection of the threshold in real time, or even the behavior itself.
[0053] Figure 5 An example of three-model interaction of the multimodal self-adjusting input system according to the present invention is shown.
[0054] This system consists of three major modules, namely three models: interaction judgment model, dynamic threshold model, and intention recognition model. These three modules preferably use image and sound data obtained by cameras and microphones as input.
[0055] The interaction judgment model is the most important module of the entire system. It extracts and fuses features of images and sounds at the same time, determines the start time and end time, and thus outputs an audio clip (preferably an audio and video clip). Among them, the start time and end time are the time points when the user starts and stops speaking in an audio or video segment, and can be stored in the form of a video segment. The output of the start time and end time is determined by two threshold parameters: the recording start threshold and the recording end threshold. When the start time score output by the interaction judgment model is greater than the recording start threshold, the output starts from this moment. Similarly, when the end time score output by the interaction judgment model is greater than the recording end threshold, the output ends from this moment.
[0056] The dynamic threshold model is a parameter adjustment module that takes as input the image and sound, along with the output of the intent recognition model (such as noise intensity, frequency, and user behavior). It outputs threshold parameters: the recording start threshold and the recording end threshold. These two parameters interact to determine whether the model should output the start and end times, thereby determining whether to save the audio or video clips.
[0057] The intent recognition model is a content / intent-related module. It uses simultaneous image and audio input (e.g., the semantics of the audio, whether the user is looking at the screen for an extended period of time in the image), to determine whether the user is intentionally operating the device. If so, it outputs "yes"; otherwise, it outputs "no."
[0058] Because we use both image and audio modal information, recording only occurs when both audio (speaking) and image (mouth movement) are triggered simultaneously, and there is a rough similarity in speech frequency and semantics. If one of the two modalities fails, the audio system automatically shuts down, eliminating the need for additional user input (such as saying the wake-up word or clicking the voice interaction button) while avoiding the capture of meaningless video and audio.
[0059] In addition, by adding an intention recognition model and a dynamic threshold adjustment model, the model parameters can be dynamically adjusted according to changes in the environment and the user's current state.
[0060] For example, when the ambient noise is very loud, the device can easily be mistakenly awakened due to the influence of the environment. At this time, the dynamic threshold model will automatically increase the threshold for activating the interaction determination model, thereby increasing the difficulty of activating the interaction determination model to record video. When the user is on a phone call, the intent recognition model determines that the user is not operating the device at this time based on the user holding the phone and the sound semantics of the conversation. Therefore, it passes the irrelevant result to the dynamic threshold model, causing the threshold of the interaction determination model to increase further, making it even more difficult to wake up the interaction determination model. The self-adjusting design of this framework ensures that the device can be fully activated when the user actually wants to use the device, while minimizing environmental interference (such as background sound) that increases the probability of false triggering of recording.
[0061] The voice interaction method of the present invention may also include a step of at least partially blurring the image information. This prevents the complete user's facial information from being extracted from the blurred image information, thereby protecting personal privacy. For example, portions that are meaningless for determining the user's interaction intent, such as the nose and ears, may be blurred. Alternatively, an algorithm may be used to blur the entire image while retaining information that is useful for determining the user's intent.
[0062] In one embodiment, image blurring can be performed directly during the image acquisition process by the camera. In another embodiment, blurring can be performed before the image information is fed into the model. In other embodiments, blurring can be performed on the information fed into the cloud, and the local image information can be deleted if necessary.
[0063] As above combined Figure 1-5 A voice interaction method according to the present invention and its preferred implementation are described. Figure 6 A schematic flow chart of a voice interaction method according to another embodiment of the present invention is shown.
[0064] In step S610, it is determined that a person is approaching and image information is obtained. In step S620, the image information is input into an interaction determination model. In step S630, based on the output of the interaction determination model, sound information is obtained for voice interaction.
[0065] Here, the input data to the interaction determination model can be made based on the user's approach to the voice interaction device. In different embodiments, the approach of someone can be determined based on different mechanisms. For example, the camera can be turned on from time to time to obtain image information and recognize the face based on relatively simple calculations such as key point extraction technology. In other embodiments, proximity information can also be received based on the networking unit. For example, in the home Internet of Things, judgments are made based on the proximity information of other devices. Proximity sensors can also be used to sense proximity. The above-mentioned proximity sensor can be installed on the voice interaction device, or it can be an Internet of Things device that can communicate with the voice device.
[0066] After the face is recognized, the screen can be lit up and the interactive content can be displayed. When the interactive content is displayed, image information for inputting the interaction determination model is obtained. For example, the voice interaction device can determine that someone is approaching based on key point extraction technology. At this time, the display screen of the device (for example, a touch screen) can be automatically lit up, and the acquired image information can be input into the interaction determination model. At this time, as long as the interaction determination model determines from the image that the user has the intention to speak, or is looking in the direction of the screen, the recording for voice interaction can be started. To this end, the image of the person speaking can be used as a positive label to train the interaction determination model; and / or the image of the person looking in the shooting direction can be used as a positive label to train the interaction determination model.
[0067] Here, the interaction determination model can use only image information for determination. In a preferred embodiment, multimodal information (e.g., also including audio information) can also be used for determination as described above. In addition, this embodiment can also use an intent recognition model and / or a dynamic threshold model to more accurately determine user intent.
[0068] Figure 7 A schematic flow chart of a voice interaction method according to another embodiment of the present invention is shown. Compared with the previous voice interaction methods, this method has a wider range of application scenarios.
[0069] In step S710, image information is acquired. In step S720, the image information is input into an interaction determination model. In step S730, based on the output of the interaction determination model, sound information is acquired for voice interaction. Thus, a trained deep learning model is used to determine the recording based on the input image. Furthermore, the method further includes: acquiring sound information simultaneously with the image information; and inputting the sound information into the interaction determination model. Thus, the deep learning model can jointly determine the start and end of a recording based on both the image and the sound.
[0070] In one embodiment, based on the output of the interaction determination model, obtaining sound information for voice interaction may include recording the sound information for voice interaction when the output of the interaction determination model is greater than a recording start threshold. Furthermore, recording of the sound information for voice interaction may be terminated when the output of the interaction determination model is greater than a recording end threshold. Thus, by introducing thresholds, the conditions required to start and end recording are adjusted.
[0071] In one embodiment, the value of the recording start threshold and / or the recording end threshold can also be adjusted based on the recognition of the speaking intention. The above-mentioned intention recognition can be implemented by a machine learning model, and for this purpose, the recognition of the speaking intention includes: inputting the acquired sound information into the intention recognition model; and obtaining the output of the intention recognition model. Correspondingly, based on the recognition of the speaking intention, adjusting the value of the recording start threshold and / or the recording end threshold includes: lowering the recording start threshold and / or the recording end threshold when a conscious interactive operation is recognized. Furthermore, a third model, a dynamic threshold model, can be introduced, which can obtain the recognition result for the speaking intention and dynamically adjust the value of the recording start threshold and / or the recording end threshold.
[0072] In addition, in addition to obtaining sound information for voice interaction based on the output of the interaction determination model, image information can also be obtained to help improve the accuracy of the interaction or to determine the end time of the recording.
[0073] The voice interaction solution of the present invention can also be implemented as a voice interaction device. Figure 8 FIG. 8 is a block diagram illustrating a voice interaction device according to an embodiment of the present invention. The device 800 may include a camera 810 , a microphone 820 , and a processor 830 .
[0074] The camera 810 is used to obtain image information, and the microphone 820 is used to obtain sound information. The processor 830 can be used to input the image information obtained by the camera 810 into the interaction determination model; and based on the output of the interaction determination model, obtain sound information for voice interaction via the microphone 820.
[0075] The processor 830 may control the camera and the microphone to be turned on and off, for example, to turn on the microphone and input the sound information acquired by the microphone into the interaction determination model.
[0076] In different embodiments, the above-mentioned model can be stored locally, or stored in a network, or both. To this end, the device may further include: a networking unit for sending the acquired image information and / or sound information, and receiving the processing results for the image information and / or sound information. Thus, it is possible to use a model stored in a cloud server, an edge computing device or other central computing node to make judgments such as the start and end of recording. As an alternative or in addition, the device may further include a storage unit for storing a model for processing the acquired image information and / or sound information.
[0077] Furthermore, the device may also include a screen for interacting with the user. For example, it may include a touch screen that lights up when a user is detected approaching. For example, a small figure may be displayed to greet the user, and recording may be initiated when the interaction determination model determines that the user is looking at the screen.
[0078] Furthermore, the device may further include: a voice output unit for outputting voice feedback of the voice interaction. The voice output unit may include a speaker, a wired or wireless headset, or a speaker.
[0079] The device may also include a proximity determination unit for determining the approach of a person. In various implementations, the proximity determination unit may include: the camera, configured to activate the camera to acquire image information and recognize a face based on key point extraction technology; a networking unit, configured to receive proximity information; and a proximity sensor, configured to sense the approach of a person.
[0080] The voice interaction device of the present invention can be implemented as a smart speaker, for example, a smart speaker with a screen and a camera. This smart speaker can implement the voice interaction method described above, implement voice interaction without the need for a wake-up word, determine the user's interaction intent when the user is in front of the smart speaker and is about to speak, and record the corresponding information.
[0081] In combination with the above-mentioned voice interaction device, the present invention may also include a voice interaction system, which may include the voice interaction device as described above; and a computing node that communicates with the voice interaction device, the computing node stores the model and provides model output for the voice interaction device.
[0082] In different implementations, computing nodes can have different identities. For example, if the voice interaction system is a locally implemented IoT system, the computing node can be a local computing node, such as a smart speaker serving as the IoT central node, or a higher-performance computing device for commercial use. In this case, in addition to connecting the voice interaction device and the central computing node, the IoT system can also connect other IoT devices. Information can be shared between these devices to meet the needs of one or more devices executing the voice interaction method of the present invention.
[0083] In a larger-scale implementation, the computing node can be an edge computing device. The edge computing device can support a larger network, such as an industrial park network or a campus network, and serve as a storage and processing server for the above model in each voice interaction device within the network.
[0084] In a larger implementation, the computing node may be a server located in the cloud, which may provide the above-mentioned voice interaction service based on user intent recognition to a large number of voice interaction devices.
[0085] The computing node may subsequently obtain sound information for voice interaction; generate and issue feedback of the voice interaction.
[0086] The present invention can also be implemented as a voice interaction model training method, including: using images of people who speak as positive labels and images of people who do not speak as negative labels to train an interaction determination model, so that the interaction determination model determines the recording start and recording reception time of sound information used for voice interaction based on the input image information based on the recording start threshold and the recording end threshold.
[0087] Furthermore, the sound information (preferably including image information) can also be used to train the intention recognition model, wherein the recording start threshold and the recording end threshold can be dynamically adjusted based on the output of the intention recognition model.
[0088] Furthermore, a dynamic threshold model can be constructed as a reinforcement learning model, which uses image information, sound information and the recognition results of the speaking intention as input, and adjusts the value of the recording start threshold and / or the recording end threshold as the behavior in real time based on the correctness of the sound information obtained for voice interaction.
[0089] Figure 9 A schematic structural diagram of a computing device that can be used to implement the above-mentioned voice interaction method according to an embodiment of the present invention is shown.
[0090] See also Figure 9 , the computing device 900 includes a memory 910 and a processor 920 .
[0091] The processor 920 may be a multi-core processor or may include multiple processors. In some embodiments, the processor 920 may include a general-purpose main processor and one or more special coprocessors, such as a graphics processing unit (GPU) or a digital signal processor (DSP). In some embodiments, the processor 920 may be implemented using customized circuits, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs).
[0092] The memory 910 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. ROM may store static data or instructions required by the processor 920 or other modules of the computer. The permanent storage device may be a readable and writable storage device. The permanent storage device may be a non-volatile storage device that does not lose stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a large-capacity storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In other embodiments, the permanent storage device may be a removable storage device (such as a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory may store some or all instructions and data required by the processor during operation. In addition, the memory 910 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks may also be used. In some embodiments, the memory 910 may include a readable and / or writable removable storage device, such as a compact disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and transient electronic signals transmitted wirelessly or wired.
[0093] The memory 910 stores executable code. When the executable code is processed by the processor 920, the processor 920 can execute the voice interaction method described above.
[0094] The voice interaction scheme according to the present invention has been described in detail above with reference to the accompanying drawings. The voice interaction scheme of the present invention can determine the user's interaction intention based on image information, or preferably audio and video information, through artificial intelligence, and directly perform intelligent collection of interaction information based on this. Specifically, an interaction determination model can be used to determine the start and end time of recording, and the intention recognition model can be further used to dynamically adjust the recording start and end determination thresholds of the interaction determination model, and a dynamic threshold model can be used to perform threshold adjustment based on dynamic learning. In this way, natural and direct voice interaction can be achieved without saying the wake-up word or turning on the voice function. By using both image and sound modalities to complete the video recording process, and designing a dynamic threshold model and an intention recognition model, the recording process can be made more personalized and intelligent.
[0095] In a broader implementation, the present invention can utilize a combination of information other than audio and visual information, namely, multimodal information, to determine a user's interaction intent. Thus, the present invention can be implemented as a voice interaction method, comprising: obtaining multimodal information, wherein the multimodal information includes at least two pieces of information obtained simultaneously; inputting the multimodal information into an interaction determination model; and, based on the output of the interaction determination model, obtaining sound information for voice interaction.
[0096] In one embodiment, one channel of the multimodal information includes acquired sound information. In another embodiment, one channel of the multimodal information includes user status information acquired via a sensor. For example, the audio and video information described above can be used to determine the model's intent, or location information acquired by a device worn by the user (e.g., a smartwatch) can be used in conjunction with sound information to determine the model's intent. In other embodiments, other sensors capable of acquiring information reflecting the user's communication intent can also be used to determine the model's intent.
[0097] In addition, the method according to the present invention may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing the above steps defined in the above method of the present invention.
[0098] Alternatively, the present invention can also be implemented as a non-transitory machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium) on which executable code (or computer program, or computer instruction code) is stored. When the executable code (or computer program, or computer instruction code) is executed by a processor of an electronic device (or computing device, server, etc.), the processor executes the various steps of the above-mentioned method according to the present invention.
[0099] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or combinations of both.
[0100] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architecture, functions and operations of the systems and methods according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0101] While various embodiments of the present invention have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A voice interaction method, comprising: Turn on the camera to obtain image information, and turn on the microphone to obtain sound information; inputting the acquired image information and sound information into an interaction determination model; as well as Based on the output of the interaction determination model, a microphone is used to obtain sound information for voice interaction. The interaction determination model simultaneously extracts and fuses features of the image information and the sound information. When the input image frame is determined to be a person speaking and the input audio frame is determined to be a person speaking or preparing to speak, the model starts acquiring sound information used for interaction and determines the recording start and end times. The image information and the sound information are simultaneously input into the intention recognition model to determine whether the current user has consciously operated the device, and the image information, the sound information and the output of the intention recognition model are simultaneously input into the dynamic threshold model to output threshold parameters: recording start threshold and recording end threshold. The output of the recording start time and the recording end time is determined by the recording start threshold and the recording end threshold. Among them, the dynamic threshold model is a reinforcement learning model, which uses the recognition results of image information, sound information and speaking intention as input, and adjusts the values of the recording start threshold and the recording end threshold as behaviors in real time based on the correctness of the sound information obtained for voice interaction.
2. The method of claim 1, further comprising: Performing speech recognition on the acquired sound information for voice interaction; as well as Based on the result of the speech recognition, feedback of the speech interaction is output.
3. The method according to claim 1, wherein The interaction determination model includes: Deep neural network models for supervised learning.
4. The method according to claim 3, wherein: The positive labels for training the deep neural network model include images of people who are speaking, and the negative labels include images of people who are not speaking.
5. The method of claim 1 , further comprising: While the microphone is used to obtain sound information for voice interaction, the camera continues to obtain image information.
6. The method according to claim 5, wherein: Based on the output of the interaction determination model, using a microphone to obtain sound information for voice interaction includes: Based on the output of the interaction determination model, the recording start time and the recording end time of recording sound information for voice interaction using a microphone are determined.
7. The method according to claim 6, wherein: Determining, based on an output of the interaction determination model, the recording start time and the recording end time of the sound information for voice interaction recorded using a microphone includes: The interaction decision model starts recording when the output meets the recording start threshold and / or ends recording when the output meets the recording end threshold based on the current image information and the sound information input.
8. The method of claim 7, wherein: The acquired sound information is input into the intention recognition model, and the value of the recording start threshold and / or the recording end threshold is adjusted based on the output of the intention recognition model.
9. The method of claim 8, wherein: Adjusting the value of the recording start threshold and / or the recording end threshold includes: The dynamic threshold model obtains the output of the intent recognition model and dynamically adjusts the value of the recording start threshold and / or the recording end threshold.
10. The method of claim 1, further comprising: The image information is at least partially blurred.
11. A computing device comprising: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to execute the method according to any one of claims 1 to 10.
12. A non-transitory machine-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to execute the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Artificial intelligence-based control method and control system for intelligent interaction equipment
CN105159111A
Wake-up method for intelligent product, intelligent product, and computer-readable storage medium
CN107679506A
Multi-mode interaction method applied to television scene
CN110335603A
Human-machine interaction method and device
CN111063354A