Methods and systems for human-computer voice interaction

By collecting and analyzing the voice and lip movement video of the priority operator and selecting an appropriate wake-up word recognition mode, the problem of chaotic wake-up word recognition among multiple users in the vehicle is solved, enabling accurate wake-up and control command execution for the priority operator and improving the user experience.

CN116364075BActive Publication Date: 2026-04-03LINGYUE DIGITAL INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-12
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, when multiple users in a vehicle simultaneously utter the wake word, the human-machine interaction system cannot distinguish between users, leading to confusion and unnecessary wake-ups. Furthermore, it cannot differentiate between wake words and control commands in normal conversation.

Method used

By collecting the voice and lip movement video of the priority operator, analyzing the lip movement pattern, selecting wake word recognition modes with different levels of accuracy, recognizing the wake word of the priority operator, and distinguishing users based on voice components to execute specific control commands.

Benefits of technology

It achieves a distinction in the accuracy of wake word recognition for the priority operator, avoiding unnecessary wake-ups and confusion, and providing a better user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116364075B_ABST
    Figure CN116364075B_ABST
Patent Text Reader

Abstract

This disclosure relates to methods and systems for human-computer voice interaction. According to one aspect of this disclosure, a method for human-computer voice interaction is provided, comprising: acquiring voice input via a voice acquisition device; simultaneously acquiring the voice input via the voice acquisition device and acquiring a video of at least a priority operator via a video acquisition device, the video of the priority operator including at least a video of the priority operator's lips; analyzing the video of the priority operator to determine a lip movement pattern of the priority operator during the voice input; and selecting a corresponding wake-word recognition mode from a plurality of wake-word recognition modes based on the lip movement pattern of the priority operator, wherein different wake-word recognition modes correspond to different wake-word recognition accuracies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to human-computer interaction, and more specifically to a method, system, device, and medium for human-computer voice interaction. Background Technology

[0002] Currently, an increasing number of vehicles offer human-machine voice interaction functions. Users can wake up / activate the human-machine interaction function by saying a preset wake-up word, such as "Hello, car". However, when there are multiple users in the vehicle, anyone saying "Hello, car" will activate the human-machine interaction function. That is, in the existing technology, it does not distinguish which user said the wake-up word. When multiple users say the same wake-up word at the same time, followed by different specific control commands (e.g., please play music; please stop playing; please search for nearby restaurants, etc.), it can lead to confusion.

[0003] Furthermore, the human-computer interaction interface will activate the wake-up function as soon as it recognizes a wake-up word, regardless of whether the wake-up word is actually used to issue control commands or is just part of normal conversation (e.g., the wake-up word is used during a chat). Summary of the Invention

[0004] This disclosure relates to methods, systems, devices, and media for human-computer voice interaction.

[0005] According to one aspect of this disclosure, a method for human-computer voice interaction is provided, comprising: acquiring voice input via a voice acquisition device; simultaneously acquiring the voice input via the voice acquisition device and acquiring a video of at least a priority operator via a video acquisition device, the video of the priority operator including at least a video of the priority operator's lips; analyzing the video of the priority operator to determine a lip movement pattern of the priority operator during the voice input; and selecting a corresponding wake word recognition mode from a plurality of wake word recognition modes based on the lip movement pattern of the priority operator, wherein different wake word recognition modes correspond to different wake word recognition accuracies.

[0006] According to some embodiments, the operation of analyzing the video of the priority operator to determine the lip movement pattern of the priority operator during the speech input further includes: comparing the acquired video of the priority operator's lips with a previously acquired set of images / video segments associated with the lip movement pattern of the spoken wake word to determine whether the acquired video of the priority operator's lips contains a video segment that matches the previously acquired set of images / video segments associated with the lip movement pattern of the spoken wake word; and in response to determining that the acquired video of the priority operator's lips contains a video segment that matches the previously acquired set of images / video segments associated with the lip movement pattern of the spoken wake word, further determining whether there is continuous lip movement near the matching video segment.

[0007] According to some embodiments, the operation of selecting a corresponding wake word recognition mode from multiple wake word recognition modes based on the lip movement pattern of the preferred operator further includes: in response to determining that there is no continuous lip movement near the matched video segment, selecting a first wake word recognition mode from the multiple wake word recognition modes; in response to determining that there is continuous lip movement near the matched video segment, selecting a second wake word recognition mode from the multiple wake word recognition modes; wherein the wake word recognition accuracy of the first wake word recognition mode is lower than the wake word recognition accuracy of the second wake word recognition mode.

[0008] According to some embodiments, the method further includes identifying a wake word based on a selected wake word recognition pattern.

[0009] According to some embodiments, the method further includes: not processing the speech input when the lip movement pattern of the priority operator indicates that the priority operator has no lip movement during the speech input.

[0010] According to some embodiments, the voice input includes at least one voice component from at least one operator, and the method further includes: identifying whether there is a voice component from the preferred operator among the at least one voice component.

[0011] According to some embodiments, the operation of identifying whether there is a voice component from the preferred operator in the at least one voice component further includes at least one of the following: (1) in the case that the voice acquisition device includes a directional microphone set for the preferred operator, determining whether the directional microphone has acquired a voice component with a predetermined intensity; (2) in the case that the voice acquisition device includes a plurality of microphones distributed in a distributed manner, determining whether the voice component with the strongest intensity comes from a predetermined direction of the preferred operator; and (3) determining whether the voice acquisition device has acquired a voice component with a voiceprint associated with the preferred operator.

[0012] According to some embodiments, the method further includes: in response to identifying the presence of a voice component from the preferred operator in the at least one voice component, performing wake word recognition on the voice component from the preferred operator based on a selected wake word recognition mode; and in response to successfully recognizing the wake word, performing control associated with the voice component from the preferred operator.

[0013] According to some embodiments, the method further includes: setting a first user as the preferred operator in response to input from a first user; and / or switching the preferred operator from the first user to the second user in response to input from a second user.

[0014] According to some embodiments, the method further includes: entering a silent period in response to a wake word after content playback has begun in response to a prior voice control command.

[0015] According to some embodiments, the method further includes performing at least one of the following during a silent period for the wake word: not performing wake word recognition; not performing control instructions containing the wake word; and performing only pre-specified control instructions.

[0016] According to another aspect of this disclosure, a computer system is provided, comprising: one or more processors, and a memory coupled to the one or more processors, the memory storing computer-readable program instructions that, when executed by the one or more processors, perform the method as described above.

[0017] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores computer-readable program instructions thereon, which, when executed by the processor, perform the method as described above.

[0018] According to another aspect of this disclosure, a device for human-computer voice interaction is provided, including means for implementing the method described above. Attached Figure Description

[0019] The above and other objects, features and advantages of this disclosure will become more apparent from the more detailed description of exemplary embodiments thereof taken in conjunction with the accompanying drawings, wherein like reference numerals generally denote like parts.

[0020] Figure 1 This is a flowchart of a method for human-computer voice interaction according to an embodiment of the present disclosure.

[0021] Figure 2 This is a flowchart of a method for human-computer voice interaction according to an embodiment of the present disclosure.

[0022] Figure 3 A flowchart illustrating a method for setting a priority operator according to an embodiment of this disclosure is shown.

[0023] Figure 4A An exemplary user interface for setting priority operators is shown according to an embodiment of this disclosure.

[0024] Figure 4B An exemplary user interface for setting priority operators is shown according to an embodiment of this disclosure.

[0025] Figure 4C An exemplary user interface for setting priority operators is shown according to an embodiment of this disclosure.

[0026] Figure 5 A flowchart illustrating a method for human-computer voice interaction according to an embodiment of the present disclosure is shown.

[0027] Figure 6 This is a schematic diagram illustrating a general hardware environment in which a device according to an embodiment of the present disclosure can be implemented. Detailed Implementation

[0028] The following description is provided to enable those skilled in the art to implement and use the embodiments, and the description is provided in the context of a particular application and its requirements. Various modifications will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of the embodiments. Therefore, the embodiments are not limited to the embodiments shown, but are to be given the widest scope consistent with the principles and features disclosed herein.

[0029] This disclosure provides improved methods, systems, devices, and media for human-computer voice interaction. Embodiments of this disclosure acquire video of the lip movements of a priority operator simultaneously with voice input, determine the lip movement pattern of the priority operator, and select wake-word recognition modes with different recognition accuracy based on the lip movement pattern. This makes the recognition of wake-words more aligned with the actual needs of the priority operator, providing a better user experience.

[0030] Some embodiments of this disclosure identify / separate a voice component from a priority operator from a voice input containing at least one voice component, and perform wake-up based on a wake-up word identified from the voice component of the priority operator. This distinguishes the priority operator from other operators, enabling priority operator-specific human-computer interaction function wake-up and / or control command execution, providing a better user experience.

[0031] Some embodiments of this disclosure, after initiating content playback in response to a prior voice control command, enter a silent period for the wake word. During this silent period, wake word recognition is not performed, and / or control commands containing the wake word are not executed, and / or only pre-specified control commands are executed. This avoids interrupting playback immediately after content playback begins due to wake word recognition, thereby providing a better user experience.

[0032] Figure 1 This is a flowchart of a method 100 for human-computer voice interaction according to an embodiment of the present disclosure.

[0033] like Figure 1 As shown, method 100 may include operation 101, in which voice input is acquired via a voice acquisition device.

[0034] Voice acquisition devices may include one or more directional microphones and / or one or more omnidirectional microphones.

[0035] In some embodiments, a directional microphone may support capturing only the driver's (typically the vehicle owner's) voice input, without capturing voice input from other directions. In other embodiments, a directional microphone may support capturing voice input from all directions, but voice input from a specific direction is amplified while voice input from other directions is de-amplified.

[0036] For example, a directional microphone positioned near the steering wheel can capture or amplify voice input from the driver's direction while ignoring or weakening voice input from other directions.

[0037] In some embodiments, multiple microphones can be placed at different locations inside the vehicle. Voice input collected by multiple microphones is integrated, processed, and analyzed.

[0038] The collected voice input may include voice components collected by one or more microphones from one or more sound sources (e.g., multiple people talking in the car).

[0039] Method 100 may further include operation 103, in which, while acquiring voice input via a voice acquisition device, video of at least a priority operator is acquired via a video acquisition device, the video of the priority operator including at least a video of the priority operator's lips.

[0040] The video capture device may include one or more cameras. In some embodiments, multiple cameras may be positioned at different locations within the vehicle to capture video of occupants at each location. In some embodiments, only one camera may be positioned near the steering wheel, primarily for capturing video of the driver.

[0041] Priority operators are those whose voice commands are processed first during human-computer voice interaction. Priority operators can be preset through the user interface.

[0042] For example, in a scenario where a father is driving with a child, the father can be designated as the priority operator. During human-computer interaction, if both the father and child issue voice commands containing a wake word, the father's voice command will be executed first, or only the father's voice command will be executed. Similarly, in the same scenario, the child can also be designated as the priority operator. In the same scenario, if both the father and child issue voice commands containing a wake word, the child's voice command will be executed first, or only the child's voice command will be executed.

[0043] In other embodiments, operation priorities can be set for different registered operators (e.g., father, mother, and child, all of whom have obtained operation authorization through pre-registration). For example, the father's operation priority is the highest, the mother's priority is the second highest, and the child's priority is the lowest. Then, when at least two of the father, mother, and child issue voice commands simultaneously, the voice command issued by the operator with the highest priority is executed first.

[0044] The priority operator setting can be switched. Each operator can have a substantially fixed position within the vehicle, and therefore, each authorized operator can have an associated camera and / or directional microphone. Thus, switching the priority operator can also affect the processing of the voice components and / or video from the associated microphone and / or camera. For example, when the priority operator changes from the father in the driver's seat to the child in the rear left seat, the voice component captured by the directional microphone positioned for the rear left seat is the priority operator's voice component, and the video captured by the camera positioned for the rear left seat is the priority operator's video.

[0045] In some embodiments, the voice acquisition device and the video acquisition device may be separate devices. In other embodiments, the voice acquisition device and the video acquisition device may be at least partially integrated.

[0046] Method 100 may also include operation 105, in which the video of the priority operator is analyzed to determine the lip movement pattern of the priority operator during the speech input.

[0047] The lip movement pattern of the priority operator can indicate whether the priority operator is lip-moving.

[0048] No lip movement from the priority operator means that there is no speech component produced by the priority operator in the acquired speech input. In this case, no further processing of the speech input is required.

[0049] The presence of lip movements by the priority operator indicates that the acquired speech input is very likely to include speech input from that operator. In this case, further processing of the speech input is required.

[0050] In some embodiments, the determined lip movement pattern of the priority operator can also indicate whether the priority operator is likely to have said a wake word (e.g., the wake word "hello, car") during the voice input and whether there are continuous lip movements before and after the wake word. A set of lip movement images / video clips previously captured when the priority operator normally output the wake word can be used as a comparison basis. By comparing the currently captured video of the priority operator's lips during the voice input with the previously saved set of lip movement images / video clips of the priority operator normally outputting the wake word, it can be determined whether the priority operator said the wake word during the voice input; that is, whether there is a set of lip movement images / video clips in the currently captured video of the priority operator's lips during the voice input that match (are similar to or identical to) the previously saved set of lip movement images / video clips of the priority operator normally outputting the wake word. When it is determined that the priority operator said the wake word during the voice input, it can be further determined whether there are continuous lip movements of the priority operator in the vicinity of the matching set of lip movement images / video clips (e.g., within a predetermined time period before and after). This can be used to determine whether the priority operator is issuing a normal voice command or engaging in conversation. Generally, when the priority operator issues a voice command, there will be a pause before / after the wake word. However, if the priority operator only mentions the wake word during the chat or if the chat contains other words whose lip movements match the wake word's lip movements, there will generally be no pause before / after the wake word, resulting in continuous lip movements.

[0051] In some embodiments, the pre-saved set of images / video segments associated with the lip movement pattern of the wake word may be specific to a priority operator, for example, acquired by photographing the priority operator. In other embodiments, the pre-saved set of images / video segments associated with the lip movement pattern of the wake word may also be acquired by photographing other operators. In still other embodiments, the pre-saved set of images / video segments associated with the lip movement pattern of the wake word may also be obtained by combining multiple sets of images / video segments associated with the lip movement pattern of the wake word for one or more operators. The pre-saved set of images / video segments associated with the lip movement pattern of the wake word may be a standard set of images / video segments not specific to any operator.

[0052] In some embodiments, multiple image sets / video segments corresponding to different lip-movement patterns can be pre-saved. These lip-movement patterns may include, for example, a no-lip-movement pattern, a command-based lip-movement pattern, and a chat-based lip-movement pattern. By comparing the currently acquired lip-movement pattern of the priority operator with the pre-saved lip-movement patterns, the category of the current priority operator's lip-movement pattern can be determined.

[0053] Method 100 may further include operation 107, in which a corresponding wake word recognition mode is selected from multiple wake word recognition modes based on the lip movement pattern of the preferred operator, wherein different wake word recognition modes correspond to different wake word recognition accuracies.

[0054] If the video of the acquired lip movements of the priority operator contains a video segment that matches the image set / video segment associated with the lip movement pattern of the wake word, and if it is further determined that there are no continuous lip movements near the matching video segment, it means that the acquired speech input is very likely to contain a speech component from the priority operator containing the wake word. In this case, a wake word recognition mode with lower recognition accuracy can be selected from multiple wake word recognition modes. For example, fuzzy recognition can be used to recognize the wake word. Using a wake word recognition mode with lower recognition accuracy means that, for example, for the set wake word "Hello, car", inputs with standard pronunciation and approximate pronunciation will be recognized as the wake word "Hello, car", and / or inputs that strictly match and substantially match the wake word (e.g., "car, hello", "car, hello", "hello, car") will also be recognized as the wake word "Hello, car". Those skilled in the art can select the algorithm and technique for fuzzy recognition as needed.

[0055] If the video footage of the acquired lip movements of the priority operator contains a video segment that matches the image set / video segment associated with the lip movement pattern of the wake word, and if it is further determined that continuous lip movements exist near the matching video segment, it means that even if the acquired speech input contains a speech component from the priority operator, that speech component from the priority operator is likely to be the priority operator's chatter rather than a control command. In this case, a wake word recognition mode with high recognition accuracy can be selected from multiple wake word recognition modes. For example, accurate recognition can be used to identify the wake word. Using a wake word recognition mode with high recognition accuracy means that, for example, for the set wake word "Hello, car," only standard pronunciation input will be recognized as the wake word "Hello, car," and / or only input that strictly matches the wake word will be recognized as the wake word "Hello, car." Those skilled in the art can select the algorithm and technique for accurate recognition as needed.

[0056] Although not shown, method 100 may also include identifying a wake word based on a selected wake word recognition pattern.

[0057] In one embodiment, when the acquired voice input includes voice input from multiple operators, the voice component from the priority operator can be first identified / separated from the acquired voice input. Then, based on the selected wake-word recognition mode, it can be determined whether a wake-word exists in the voice component from the priority operator. If it is determined that a wake-word exists in the voice component from the priority operator, then wake-up and / or execution of control commands associated with that voice component are performed according to the voice component from the priority operator.

[0058] In another embodiment, one or more wake words in the voice input can be identified first based on a selected wake word recognition pattern, and then it can be determined whether there is a voice component containing a wake word associated with a priority operator. If there is a voice component containing a wake word associated with a priority operator, then wake-up is performed based on that voice component and / or control commands associated with that voice component are executed.

[0059] Figure 2 This is a flowchart of a method 200 for human-computer voice interaction according to an embodiment of the present disclosure.

[0060] like Figure 2 As shown, method 200 may include operation 201, in which voice input is acquired via a voice acquisition device. This operation is similar to operation 101 in method 100, and will not be described again here.

[0061] Method 200 may further include operation 203, in which, while acquiring voice input via a voice acquisition device, video of at least a priority operator is acquired via a video acquisition device, the video of the priority operator including at least a video of the priority operator's lips. This operation is similar to operation 103 in method 100 and will not be described further here.

[0062] Here, assuming the driver is set as the priority operator, the priority operator's video will at least include video of the driver's lips.

[0063] Method 200 may further include operation 205, in which the video of the priority operator is analyzed to determine the lip movements of the priority operator during the speech input. This operation is similar to operation 5 in method 100 and will not be described again here.

[0064] Method 200 may also include operation 207, in which it is determined whether the preferred operator has lip movement during the speech input based on the determined lip movement pattern.

[0065] If it is determined at operation 207 that the priority operator has no lip movement during the voice input, then method 200 proceeds to operation 209, in which no further processing of the voice input is performed, and method 200 ends. In other words, in this case, no further processing of the voice input is performed, the wake word is not recognized, and no wake-up and / or control associated with the voice input is performed.

[0066] If it is determined at operation 207 that the priority operator has lip movements during the voice input, then method 200 proceeds to operation 211, in which it is determined whether there is a voice component from the priority operator (i.e., the driver) in the acquired voice input.

[0067] In some embodiments, the voice acquisition device may include a directional microphone positioned for the driver (e.g., positioned near the steering wheel and oriented towards the driver). The driver's voice component can be determined by the directional microphone capturing a voice component exceeding a predetermined intensity. In some embodiments, the directional microphone only captures voice components in the driver's direction; as long as the directional microphone captures a voice component, it means that the acquired voice input includes a voice component from the driver. In this case, the voice component from the directional microphone or the voice component with the highest intensity captured by the directional microphone can be considered the voice component of the preferred operator.

[0068] In other embodiments, the voice acquisition device may include multiple microphones distributed throughout, and the acquired voice input may contain a voice component from the driver by determining that the strongest voice component originates from the driver's direction. In this case, the voice component from the driver's direction can be separated as the voice component of the priority operator.

[0069] In other embodiments, the pre-registered operator may have already recorded a corresponding voiceprint. In this case, the presence of a voice component from the driver in the collected voice input can be determined by identifying a voice component in the collected voice input that has a voiceprint matching a pre-set voiceprint associated with the driver. In this case, the voice component with the matching voiceprint can be isolated as the voice component of the preferred operator.

[0070] The above are merely exemplary methods for identifying / separating speech components from a specific operator. Identifying and / or separating different speech components in speech input can be achieved using techniques well-known in the art or developed in the future.

[0071] If it is determined in operation 211 that there is no speech component from the priority operator (i.e., the driver) in the collected speech input, then method 200 proceeds to operation 221, in which no further processing of the speech input is performed, and method 200 ends.

[0072] If it is determined in operation 211 that there are speech components from the priority operator (i.e., the driver) in the acquired speech input, then method 200 proceeds to operation 213, in which the wake word recognition mode to be used is determined based on the determined lip movement pattern.

[0073] As previously mentioned, the identified lip movement pattern can indicate whether the priority operator said a wake word (e.g., "Hello, car!") during the speech input and whether there were continuous lip movements before and after the wake word.

[0074] If the determined lip movement pattern indicates that the speech component from the priority operator is very likely to contain a wake word and there are continuous lip movements before and after the wake word, then a wake word recognition pattern with high recognition accuracy is selected. That is, in this case, accurate recognition can be used.

[0075] If the determined lip movement pattern indicates that the speech component from the priority operator is likely to contain a wake word and there are no consecutive lip movements before and after the wake word, then a wake word recognition pattern with low recognition accuracy is selected. That is, in this case, fuzzy recognition can be used.

[0076] Next, the method proceeds to operation 215, in which wake word recognition is performed using the determined wake word recognition pattern.

[0077] The determined wake word recognition pattern is used to identify the wake word from the voice component of the priority operator.

[0078] If the wake word is successfully recognized in operation 215, method 200 proceeds to operation 217, in which control associated with the voice component of the priority operator is performed, such as wake-up and / or specific control actions (e.g., playing music, starting a fan, etc.).

[0079] If the wake word is not successfully recognized in operation 215, then method 200 proceeds to operation 219, where no further processing is performed, and method 200 ends.

[0080] In the illustrated method 200, the lip movement pattern of the priority operator is analyzed in operation 205, including a coarse lip movement pattern (e.g., whether there is lip movement) and, if there is lip movement, a more specific lip movement pattern (e.g., whether it is a command-type lip movement pattern or a chat-type lip movement pattern). In some variations, only the coarse lip movement pattern may be analyzed in operation 205, and the specific lip movement pattern may be analyzed in operation 213 at an appropriate time before or immediately preceding the wake-word recognition pattern.

[0081] In the illustrated method 200, the presence of a speech component from a priority operator is first determined, and then a wake-up word is identified. Those skilled in the art will understand that in some variations, the wake-up word can be identified first, followed by determining whether a speech component from a priority operator is present.

[0082] Without departing from the teachings of this disclosure, those skilled in the art can make various modifications to the specific operation of the method and the order of operation.

[0083] Figure 3 A flowchart is shown for a method 300 for setting a priority operator according to an embodiment of the present disclosure.

[0084] like Figure 3 As shown, method 300 may include operation 301, in which a first user is set as the preferred operator in response to first user input.

[0085] Figure 4A An exemplary user interface for setting a priority operator according to an embodiment of this disclosure is shown. In the user interface for setting a priority operator, user interface elements corresponding to multiple authorized users, such as user 1 to user 3, can be presented, as shown in 401, 403, and 405. The corresponding user can be set as the priority operator by selecting one of the user interface elements. For example, in... Figure 4A In the middle, users can select the interface element by clicking on the user interface element 401.

[0086] The system can store information such as a user's voiceprint, the microphone associated with the user, and the camera associated with the user, all associated with each authorized username.

[0087] Method 300 may further include operation 302, in which the preferred operator is switched from the first user to the second user in response to input from the second user.

[0088] Figure 4B An exemplary user interface for setting priority operators is shown according to an embodiment of this disclosure. For example... Figure 4B As shown, by selecting the interface element corresponding to User 2, User 2 can be set as the priority operator, thereby switching the priority operator from User 1 to User 2.

[0089] Figure 4C An exemplary user interface for setting priority operators according to an embodiment of the present disclosure is shown. The user interface for setting priority operators may also provide an interface element 406 for adding new authorized users, whereby the user can enter a new authorized username through interaction with the user interface element 406. Interface elements 407 and 408 may also be provided for entering information about the user's associated seat and / or voiceprint.

[0090] In some embodiments, the priority of different users can be specified by reordering interface elements 401, 403, and 405. For example, the user at the top of the list is the first operator by default. Alternatively, the user at the top of the list has the highest priority by default, and the priority decreases in sequence. When multiple users are assigned different priorities, wake-word recognition can be performed first on the voice component of the user with the highest priority. If recognition fails, wake-word recognition can then be performed on the voice component of the user with the second highest priority.

[0091] Those skilled in the art will understand that the user interface and associated user interface elements for setting priority operators can be designed as needed.

[0092] Figure 5 A flowchart is shown for a method 500 for human-computer voice interaction according to an embodiment of the present disclosure.

[0093] Method 500 may include operation 501, in which, after content playback has commenced in response to a prior voice control instruction, a silence period for a wake-word is entered. During the wake-word silence period, at least one of the following is performed: wake-word recognition is not performed; control instructions containing the wake-word are not performed; and only pre-specified control instructions are performed.

[0094] In this scenario, upon receiving voice input, it can be first determined whether the input occurs during the wake-up word's silence period. If the input occurs during the wake-up word's silence period, it can be left unprocessed, or even if the wake-up word is recognized, the corresponding control (e.g., wake-up) can be omitted. Alternatively, after recognizing the wake-up word and specific control command, wake-up and specific control can only be executed if the specific control command matches at least one of one or more pre-set control commands. If the voice input does not occur during the wake-up word's silence period, then actions such as those described above can be performed. Figure 1-2 The method described.

[0095] Users can set the length of the silent period, such as a few minutes after playback begins. The length of the silent period can also be based on the length of the track being played, such as the silent period lasting until the end of the track. In some embodiments, the silent period can continue indefinitely until the user ends it via a user interface operation. Those skilled in the art will understand that various designs for the silent period can be made as needed without departing from the teachings of this disclosure.

[0096] Figure 6 This is a schematic diagram illustrating a general hardware environment in which a device according to an embodiment of the present disclosure can be implemented.

[0097] Now for reference Figure 6 The diagram illustrates an example of compute node 600. Compute node 600 is merely one example of a suitable compute node and is not intended to imply any limitation on the scope of use or functionality of the embodiments described herein. In any case, compute node 600 is capable of implementing and / or performing any of the functions set forth above.

[0098] Within compute node 600, there is a computer system / server 6012 that can operate with a wide variety of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with computer system / server 6012 include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the aforementioned systems or devices, etc.

[0099] The computer system / server 6012 can be described in the general context of computer system executable instructions (such as program modules) executed by the computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, etc., that perform specific tasks or implement specific abstract data types. The computer system / server 6012 can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules can reside on both local and remote computer system storage media, including memory storage devices.

[0100] like Figure 6 As shown, the computer system / server 6012 in computing node 600 is illustrated in the form of a general-purpose computing device. The components of the computer system / server 6012 may include, but are not limited to: one or more processors or processing units 6016, system memory 6028, and a bus 6018 that couples the various system components, including the system memory 6028, to the processing unit 6016.

[0101] Bus 6018 represents any one or more of several types of bus architectures, including memory buses or memory controllers, peripheral buses, accelerated graphics ports, processors, or local buses using any of the various bus architectures. By way of example and not limitation, these architectures include, but are not limited to, Industry Standard Architecture (ISA) buses, Microchannel Architecture (MAC) buses, Enhanced ISA buses, Video Electronics Standards Association (VESA) local buses, Peripheral Component Interconnect (PCI) buses, Peripheral Component Interconnect High Speed ​​(PCIe) buses, and Advanced Microcontroller Bus Architecture (AMBA).

[0102] Computer system / server 6012 typically includes various computer system readable media. These media can be any available media accessible by computer system / server 6012, including volatile and non-volatile media, removable and non-removable media.

[0103] System memory 6028 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 6032. Computer system / server 6012 may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 6034 may be provided for reading from and writing to a non-removable non-volatile magnetic medium (not shown, and generally referred to as a "hard disk drive"). Although not shown, a disk drive may be provided for reading from and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disc drive may be provided for reading from and writing to a removable non-volatile optical disc (such as a CD-ROM, DVD-ROM, or other optical media). In these cases, each may be connected to bus 6018 via one or more data media interfaces. As will be further described and illustrated below, memory 6028 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of embodiments of this disclosure.

[0104] By way of example and not limitation, a program / utility 6040 having a set (at least one) of program modules 6042, along with an operating system, one or more application programs, other program modules, and program data, may be stored in memory 6028. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, may include an implementation of a network environment. Program modules 6042 generally perform functions and / or methods as described in the embodiments herein.

[0105] The computer system / server 6012 can also communicate with one or more external devices 6014 (such as a keyboard, indicating device, display 6024, etc.), one or more devices that enable a user to interact with the computer system / server 6012, and / or any device that enables the computer system / server 6012 to communicate with one or more other computing devices (e.g., a network card, modem, etc.). This communication can occur via input / output (I / O) interface 22. Furthermore, the computer system / server 6012 can communicate with one or more networks (such as a local area network (LAN), a general area network (WAN), and / or a public network (e.g., the Internet)) via network adapter 20. As depicted, network adapter 20 communicates with other components of the computer system / server 6012 via bus 6018. It should be understood that, although not shown, other hardware and / or software components can be used in conjunction with the computer system / server 6012. Examples include, but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archiving storage systems.

[0106] This disclosure can be implemented as a system, method, and / or computer program product. The computer program product may include one or more computer-readable storage media having computer-readable program instructions thereon for causing a processor to perform aspects of this disclosure.

[0107] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example (but not limited to), electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital universal disc (DVD), memory sticks, floppy disks, mechanical encoding devices (such as punch cards or recessed protrusions storing instructions thereon), and any suitable combination of the foregoing. As used herein, computer-readable storage media is not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0108] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network (e.g., the Internet, local area network, wide area network, and / or wireless network) to an external computer or external storage device. The network may include copper cables, fiber optic cables, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to computer-readable storage media within the respective computing / processing device.

[0109] Computer-readable program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Smalltalk, C++, etc.) and conventional procedural programming languages ​​(such as the "C" programming language or similar programming languages). The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may be personalized by utilizing state information from the computer-readable program instructions to perform aspects of this disclosure.

[0110] This document describes aspects of the present disclosure with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0111] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that, when executed by the processor of the computer or other programmable data processing apparatus, these instructions create means for implementing the functions / behaviors specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that directs a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, thereby including an article of manufacture comprising instructions for implementing aspects of the functions / behaviors specified in one or more blocks of the flowchart and / or block diagram.

[0112] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, thereby causing the instructions to be executed on the computer, other programmable apparatus, or other device to perform the functions / behaviors specified in one or more boxes of a flowchart and / or block diagram.

[0113] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a portion of a module, segment, or instruction containing one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the figures. For example, depending on the functions involved, two consecutive blocks may actually be executed substantially in parallel, or these blocks may sometimes be executed in reverse order. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or behavior or executes a combination of dedicated hardware and computer instructions.

[0114] Those skilled in the art should also understand that the various operations illustrated in sequence in the embodiments of this disclosure do not necessarily have to be performed in the illustrated order. Those skilled in the art can adjust the order of operations as needed. They can also add more operations or omit some operations as needed.

[0115] Various embodiments of this disclosure have been described for illustrative purposes, but are not intended to be exhaustive or limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles of the embodiments, their practical application, or technical improvements to technologies found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for human-computer voice interaction, comprising: Voice input is collected via a voice acquisition device; While acquiring voice input via a voice acquisition device, video of at least the priority operator is acquired via a video acquisition device, wherein the video of the priority operator includes at least a video of the priority operator's lips. Analyze the video of the priority operator to determine the lip movement pattern of the priority operator during the voice input, wherein the determined lip movement pattern of the priority operator indicates whether there is continuous lip movement before and after the wake word; and Selecting a corresponding wake word recognition mode from multiple wake word recognition modes based on the lip movement pattern of the preferred operator further includes: In response to a lip movement pattern indication indicating continuous lip movement before and after a wake-up word, a first wake-up word recognition mode is selected from the plurality of wake-up word recognition modes, and In response to the lip movement pattern indication that there is no continuous lip movement before and after the wake word, a second wake word recognition mode is selected from the plurality of wake word recognition modes. Among them, the wake word recognition accuracy of the first wake word recognition mode is lower than that of the second wake word recognition mode.

2. The method according to claim 1, wherein, Analyzing the video of the priority operator to determine the lip movement pattern of the priority operator during the voice input further includes: Compare the acquired video of the priority operator's lips with previously acquired image sets / video segments associated with lip movement patterns associated with spoken wake words to determine whether the acquired video of the priority operator's lips contains a video segment that matches the previously acquired image set / video segments associated with lip movement patterns associated with spoken wake words; and In response to determining that the video of the captured lip movements of the priority operator contains a video segment that matches a previously captured set of images / video segments associated with lip movement patterns associated with uttering a wake word, it is further determined whether there are continuous lip movements in the vicinity of the matching video segment.

3. The method according to claim 2, wherein, The operation of selecting a corresponding wake word recognition mode from multiple wake word recognition modes based on the lip movement pattern of the preferred operator further includes: In response to determining that there is no continuous lip movement in the vicinity of the matched video segment, a first wake-word recognition mode is selected from the plurality of wake-word recognition modes; and In response to determining that there is continuous lip movement in the vicinity of the matched video segment, a second wake word recognition mode is selected from the plurality of wake word recognition modes.

4. The method according to claim 1 further includes identifying a wake word based on a selected wake word recognition pattern.

5. The method according to claim 1, further comprising: When the lip movement pattern of the priority operator indicates that the priority operator has no lip movement during the voice input, the voice input is not processed.

6. The method according to claim 1, wherein, The voice input includes at least one voice component from at least one operator, and the method further includes: Identify whether there is a speech component from the preferred operator in the at least one speech component.

7. The method according to claim 6, wherein, The operation of identifying whether there is a speech component from the preferred operator in the at least one speech component further includes at least one of the following: (1) In the case that the voice acquisition device includes a directional microphone set for the preferred operator, determine whether the directional microphone has acquired a voice component exceeding a predetermined intensity; (2) In the case where the voice acquisition device includes multiple microphones distributed in a distribution, determine whether the voice component with the strongest intensity comes from the direction of the predetermined priority operator; and (3) Determine whether the voice acquisition device has acquired a voice component with a voiceprint associated with the priority operator.

8. The method according to claim 6, further comprising: In response to identifying that there is a speech component from the priority operator in the at least one speech component, wake word recognition is performed on the speech component from the priority operator based on the selected wake word recognition mode; and In response to successful recognition of the wake word, control associated with the voice component from the priority operator is executed.

9. The method according to claim 1, further comprising: In response to input from the first user, the first user is set as the preferred operator; and / or In response to input from a second user, the priority operator is switched from the first user to the second user.

10. The method according to claim 1, further comprising: After responding to the preceding voice control command and starting content playback, a silent period is entered in response to the wake word.

11. The method of claim 10, further comprising performing at least one of the following during a silence period for the wake word: Do not perform wake word recognition; Do not execute control commands containing wake words; and Only execute pre-specified control commands.

12. A computer system, comprising: One or more processors; and A memory coupled to the one or more processors, the memory storing computer-readable program instructions that, when executed by the one or more processors, perform the operation of the method as described in any one of claims 1-11.

13. A computer-readable storage medium storing computer-readable program instructions thereon, which, when executed by a processor, perform the operation of the method as described in any one of claims 1-11.

14. An apparatus comprising means for implementing the method as described in any one of claims 1-11.

Citation Information

Patent Citations

  • Awakening method and device of vehicle-mounted voice system, vehicle and medium

    CN111833870A

  • Wake-up method and device of intelligent terminal and electronic equipment

    CN112669837A