Voice interaction method, device, equipment, storage medium and computer program product

By switching the voice interaction mode according to the window type, the problem of combining voice interaction technology with large-screen devices is solved, and a voice interaction experience without frequent use of wake-up words on large-screen devices is realized.

CN114038458BActive Publication Date: 2025-05-13BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111314317.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-08
Publication Date
2025-05-13
Estimated Expiration
2041-11-08

AI Technical Summary

Technical Problem

How to combine voice interaction technology originally adapted to small-screen devices with large-screen devices to provide a better user experience.

Method used

By determining the window type of the currently displayed window at the uppermost level of the information display area, switching to the preset mode in response to different types of windows, and collecting voice signals as control instructions. For window types that do not require wake-up words, all voice signals are collected directly; for window types that require wake-up words, subsequent voice signals are collected only after receiving the voice trigger signal containing the wake-up words.

Benefits of technology

There is no need to frequently say wake-up words during information browsing, which improves the human-computer interaction experience and is especially suitable for large-screen smart devices such as smart mirrors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114038458B_ABST
    Figure CN114038458B_ABST
Patent Text Reader

Abstract

The present disclosure provides a voice interaction method, device, electronic device, computer-readable storage medium and computer program product, which relate to the fields of artificial intelligence technology such as smart home, intelligent voice, information interaction, etc. The method includes: determining the window type of the topmost window currently displayed in the information display area; in response to the window type being a first type that does not require a wake-up word as a prerequisite for triggering voice interaction, switching to a preset first mode, and collecting all incoming voice signals as control instructions for execution; in response to the window type being a second type that requires a wake-up word as a prerequisite for triggering voice interaction, switching to a preset second mode, and only collecting subsequent incoming voice signals as control instructions for execution after first receiving a voice trigger signal containing a wake-up word. The present disclosure is aimed at large-screen smart devices such as smart mirrors, and voice interaction can be achieved without frequently saying the wake-up word in the window of browsing streaming information, thereby improving the human-computer interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of human-computer information interaction processing, specifically to the field of artificial intelligence technologies such as smart home, smart voice, information interaction, and more particularly to a voice interaction method, device, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] As a means of human-computer interaction, voice interaction technology has been widely used in various small-screen smart devices, such as smartphones, tablets, smart watches, and various smart wearable devices that are not convenient for touch operation.

[0003] With the increasing popularity of the concept of smart home and people's yearning for an increasingly better life, various large-screen devices in homes or indoors are gradually combined with intelligence to form smart large-screen devices.

[0004] How to combine the voice interaction technology originally adapted to small-screen devices with large-screen devices to provide users with a better user experience is a technical problem that needs to be urgently solved by technical personnel in this field. Summary of the invention

[0005] The embodiments of the present disclosure provide a voice interaction method, device, electronic device, computer-readable storage medium, and computer program product.

[0006] In the first aspect, an embodiment of the present disclosure proposes a voice interaction method, including: determining the window type of the topmost window currently displayed in the information display area; in response to the window type being a first type that does not require a wake-up word as a prerequisite for triggering voice interaction, switching to a preset first mode, and collecting all incoming voice signals as control instructions for execution; in response to the window type being a second type that requires a wake-up word as a prerequisite for triggering voice interaction, switching to a preset second mode, and collecting subsequent incoming voice signals as control instructions for execution only after first receiving a voice trigger signal containing a wake-up word.

[0007] In the second aspect, an embodiment of the present disclosure proposes a voice interaction device, including: a window type determination unit, configured to determine the window type of the topmost window currently displayed in the information display area; a first mode switching and processing unit, configured to switch to a preset first mode in response to the window type being a first type that does not require a wake-up word as a prerequisite for triggering voice interaction, and collect all incoming voice signals as control instructions for execution; a second mode switching and processing unit, configured to switch to a preset second mode in response to the window type being a second type that requires a wake-up word as a prerequisite for triggering voice interaction, and collect subsequent incoming voice signals as control instructions for execution only after first receiving a voice trigger signal containing a wake-up word.

[0008] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can implement the voice interaction method described in any implementation method in the first aspect when executing.

[0009] In a fourth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, which are used to enable a computer to implement the voice interaction method described in any implementation method in the first aspect when executed.

[0010] In a fifth aspect, an embodiment of the present disclosure provides a computer program product comprising a computer program, which, when executed by a processor, can implement the voice interaction method described in any implementation manner in the first aspect.

[0011] Since the biggest difference between large-screen smart devices such as smart mirrors and traditional small-screen smart devices is that they have information display areas with larger aspect ratios, they are more suitable for presenting more and different information in an enumerated, waterfall, or streaming manner. Therefore, the present disclosure targets the characteristics of large-screen smart devices such as smart mirrors. When the window presenting such information content is at the top of the information display area, the first mode without the need for a wake-up word as a prerequisite for triggering voice interaction is used as an adapted voice interaction mode, thereby eliminating the need to frequently say the wake-up word during information browsing, thereby improving the human-computer interaction experience.

[0012] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Other features, objects and advantages of the present disclosure will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0014] Figure 1 is an exemplary system architecture in which the present disclosure may be applied;

[0015] Figure 2 A flow chart of a voice interaction method provided by an embodiment of the present disclosure;

[0016] Figure 3 A flowchart of a method for predetermining window types of different windows provided in an embodiment of the present disclosure;

[0017] Figure 4 A flowchart of another voice interaction method provided by an embodiment of the present disclosure;

[0018] Figure 5 A schematic diagram of voice interaction for a smart mirror provided in an embodiment of the present disclosure;

[0019] Figure 6 A structural block diagram of a voice interaction device provided in an embodiment of the present disclosure;

[0020] Figure 7 A schematic diagram of the structure of an electronic device suitable for executing a voice interaction method provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description. It should be noted that the embodiments in the present disclosure and the features in the embodiments may be combined with each other without conflict.

[0022] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0023] Figure 1 An exemplary system architecture 100 is shown to which embodiments of the voice interaction method, apparatus, electronic device, and computer-readable storage medium of the present disclosure can be applied.

[0024] like Figure 1 As shown, the system architecture 100 may include a smart mirror 101 and a user 102 using the smart mirror 101 .

[0025] The smart mirror 101 can interact with other terminal devices and servers through the network or other means, so as to utilize other terminal devices and servers to provide more functions for the user 102. The user 102 can also interact with the smart mirror 101 in various ways, such as voice interaction, touch interaction, gesture interaction, etc.

[0026] A variety of applications or programs can be installed on the smart mirror 101 to implement the above functions, such as news applications, voice interaction applications, search engine applications, instant messaging applications, online shopping applications, smart clothing applications, etc.

[0027] The smart mirror 101 can provide various services through various built-in applications. Taking the voice interaction application that can provide voice interaction service as an example, the smart mirror 101 can achieve the following effects when running the voice interaction application: first, continuously monitor the window changes that appear in the top layer of the information display area, and determine the window type of each window displayed on the top layer according to the window changes; then, when it is determined that its window type is the first type that does not require a wake-up word as a prerequisite for triggering voice interaction, switch to the preset first mode, and collect all incoming voice signals as control instructions for execution; conversely, when it is determined that its window type is the second type that requires a wake-up word as a prerequisite for triggering voice interaction, switch to the preset second mode, and collect subsequent incoming voice signals as control instructions for execution only after receiving a voice trigger signal containing the wake-up word.

[0028] The voice interaction methods provided in the subsequent embodiments of the present disclosure are generally executed by the smart mirror 101 , and accordingly, the voice interaction device is generally also disposed in the smart mirror 101 .

[0029] It should be understood that Figure 1 The shape, size, number and positional relationship between the smart mirrors and the user are only schematic and can be adaptively adjusted according to the needs of implementation.

[0030] Please refer to Figure 2 , Figure 2 A flow chart of a voice interaction method provided by an embodiment of the present disclosure, wherein process 200 includes the following steps:

[0031] Step 201: Determine the window type of the topmost window currently displayed in the information display area;

[0032] The information display area is the area on the smart mirror used to display information externally, which usually accounts for the vast majority of the overall size of the smart mirror, and its shape is usually consistent with the overall shape of the smart mirror, usually a long vertical strip (that is, the length of the plumb line that serves as the long side of the information display area is significantly greater than the length of the horizontal line that serves as the short side). Of course, in some special scenarios, the overall shape of the smart mirror will also be designed as a horizontal strip or square according to needs.

[0033] Under the control of current smart operating systems, smart devices like smart mirrors can run multiple devices in the background. However, if the split-screen function is not enabled, usually only the window of the application at the top layer can be seen by the user in front of the mirror. In other words, the information display area only displays the content in the window at the top layer, and other windows that are not at the top layer usually run in the background or are temporarily closed. Of course, if the split-screen function is enabled, the top layer can also display multiple windows from different applications or the same application at the same time.

[0034] This step is intended to be performed by the execution subject of the voice interaction method (such as Figure 1 The smart mirror 101 shown in the figure determines the window type of the window currently displayed at the top layer of its information display area. In order to determine the window type of the window currently displayed at the top layer, the window types corresponding to different windows have actually been defined before this step, and the window type of the actually displayed window can be accurately determined based on the definition result.

[0035] For example, all windows under a certain application may be determined to belong to a certain window type according to the function of the application, or different windows under the same application may be further divided into different window types. Alternatively, in order to improve efficiency, only certain special windows may be set to a certain preset window type, while other windows are still defaulted to the default window type, and so on. No specific limitations are made here.

[0036] It should be noted that the above-mentioned execution subject can be in a state of monitoring the window displayed on the top layer on a normal basis, that is, following the instructions of the user of the above-mentioned execution subject, new windows will inevitably be opened, old windows will be closed, and the display levels of different windows will be switched. Therefore, whenever a change is detected in the window displayed on the top layer, this step should be triggered to judge the window type of the window, so as to determine the voice interaction mode to be adopted according to the judgment result.

[0037] Step 202: In response to the window type being the first type that does not require a wake-up word as a prerequisite for triggering voice interaction, switching to a preset first mode, and collecting all incoming voice signals as control instructions for execution;

[0038] For the case where the window type is the first type that does not require a wake-up word as a prerequisite for triggering voice interaction, this step is intended to enable the above-mentioned execution subject to switch the current voice interaction mode to the preset first mode, and collect all incoming voice signals as control instructions for execution.

[0039] That is, in the preset first mode, the above-mentioned execution subject will collect all incoming voice signals and execute all collected voice signals as control instructions. In other words, in the preset first mode, the above-mentioned execution subject is already in a state of waiting to receive voice control instructions, which is equivalent to skipping the link of the conventional voice interaction mechanism first receiving the wake-up word issued by the user and directly entering the subsequent link of receiving subsequent instructions.

[0040] Taking the wake-up word "Xiao X" set for the smart mirror of XX brand as an example, in order to prevent false triggering of voice recognition, the conventional voice interaction mechanism usually needs to first receive a voice trigger signal containing "Xiao X" from the user. After the smart mirror recognizes the voice trigger signal to confirm that the user wants to send a voice command to itself, it puts itself in a state of waiting to receive voice control commands, that is, all subsequent voice signals transmitted by the user are recognized and executed as control commands.

[0041] When this step determines that the window type of the currently displayed window on the top layer is the first type that does not require a wake-up word as a prerequisite for triggering voice interaction, it is directly in a state of waiting to receive voice control instructions by switching to the preset first mode, which means that the process of first receiving the "small X" sent by the user is eliminated. For example, assuming that the current window is a window that continuously presents an information flow downward in a waterfall-like manner, the above-mentioned execution subject determines that the window type it belongs to is the first type, and by switching to the preset first mode, the user can directly say "slide down", "pause", "play this video" to control the smart mirror while keeping the window, without adding a wake-up word such as "small X" before each instruction.

[0042] Step 203: In response to the window type being the second type that requires a wake-up word as a prerequisite for triggering voice interaction, switch to the preset second mode, and collect subsequent incoming voice signals as control instructions for execution only after receiving a voice trigger signal containing the wake-up word.

[0043] For the case where the window type is the second type that requires a wake-up word as a prerequisite for triggering voice interaction, this step is intended to enable the above-mentioned execution subject to switch the current voice interaction mode to the preset second mode, and only collect the subsequent incoming voice signal as a control instruction for execution after receiving the voice trigger signal containing the wake-up word.

[0044] The situation described in this step is the conventional voice interaction mechanism mentioned above. For example, assuming that the current window is the desktop homepage of the smart mirror, the above-mentioned execution subject determines that the window type to which it belongs is the second type, and switches to the preset second mode, so that the user must speak a voice command with the structure of "wake-up word + action command" while keeping the window in order to effectively control the smart mirror, that is, at least a wake-up word such as "small X" must be added before the first action command to allow the smart mirror to determine that the subsequent action command is issued by the user to itself, and the premise that the subsequent action command does not need to be prefixed by the wake-up word is that the smart mirror has not ended the last state of waiting to receive voice control commands.

[0045] According to the comparison between step 202 and step 203, it can be seen that the reason why the present disclosure sets different voice interaction modes by distinguishing window types is to take into account the differences between performing certain operations on a large screen such as a smart mirror and the same operations on a conventional small-screen device. For example, information flow display windows, list-style information display windows, and waterfall-style information display windows can usually better utilize the characteristics of the large information display area of ​​the smart mirror. This type of window browses massive amounts of information, and it is obvious that by improving the conventional voice interaction mechanism, it can bring users a better user experience.

[0046] Since the biggest difference between large-screen smart devices such as smart mirrors and traditional small-screen smart devices is that they have information display areas with larger aspect ratios, they are more suitable for presenting more and different information in an enumerated, waterfall, or streaming manner. Therefore, the voice interaction method provided by the embodiment of the present disclosure for large-screen smart devices such as smart mirrors is based on the characteristics of large-screen smart devices such as smart mirrors. When the window presenting such information content is at the top of the information display area, the first mode without the need for a wake-up word as a prerequisite for triggering voice interaction is used as an adapted voice interaction method, thereby eliminating the need to frequently say the wake-up word during information browsing, thereby improving the human-computer interaction experience.

[0047] In order to deepen the understanding of how to determine the appropriate window type for different windows, the present disclosure also provides Figure 3 An implementation is shown, mainly for a smart mirror with an overall shape of a longitudinally long strip, wherein process 300 includes the following steps:

[0048] Step 301: according to the vertical long strip shape presented by the information display area, determining the window type of the enumerated information display window adapted to the vertical long strip shape as the first type;

[0049] This step is intended to allow the execution subject to determine the window type of the window that is consistent with the shape of the information display area and displays a large amount of information in an enumerated manner as the first type.

[0050] Step 302: Determine the window type of other windows that are not determined to be of the first type as the second type.

[0051] On the basis of step 301, this step aims to determine the window types of other windows that are not determined as the first type as the second type by the above-mentioned execution subject, that is, to quickly complete the determination of the window types of all windows by using the exclusive method.

[0052] The solution provided in this embodiment is obtained from the consideration of the adaptability of the shape of the information display area of ​​the smart mirror and the way the information is presented in the window. Different window type determination mechanisms can also be considered in combination with other aspects to better bring convenience to users when using large-screen smart devices such as smart mirrors.

[0053] Based on any of the above embodiments, the present disclosure further provides Figure 4 Another more specific voice interaction method is provided, wherein the process 400 includes the following steps:

[0054] Step 401: Reading the started application from the cache in the storage information display area;

[0055] This step is intended to enable the execution subject to read the activated application from the cache in the storage information display area, that is, the application data of the activated application is usually stored in the cache space provided by the power-off volatile storage medium, so as to provide a higher data reading and writing speed by taking advantage of the characteristics of the power-off volatile storage medium. On the contrary, the data of the unactivated application is stored in the power-off non-volatile storage medium.

[0056] Step 402: Determine the display level of each started application respectively, and determine the target application displayed at the top layer according to the display level;

[0057] On the basis of step 401, this step aims to determine the display level of each launched application by the above-mentioned execution subject, and determine the target application displayed at the top layer according to the display level. That is, according to the user's selection, usually only one of each launched application can stay at the top layer, and the remaining launched applications will determine their real display level according to the startup sequence and level adjustment operation. The above-mentioned execution subject can clarify the current display level of each launched application according to the user's operation log, and determine the target application displayed at the top layer.

[0058] Step 403: Determine the window type of the current display window of the target application;

[0059] In order to prevent different window types from being set for different windows of the target application in advance, this step will first clarify the current display window of the target application, and then determine the window type of the current display window. Specifically, when determining the window type of the window, it can be determined based on a correspondence table that records the window identifier and the window type.

[0060] Step 404: Determine whether the window type is the first type, if so, execute step 405, otherwise execute step 406;

[0061] Step 405: Switch to the preset first mode and collect all incoming voice signals as control instructions for execution;

[0062] This step is based on the judgment result of step 404 that the window type is the first type, and is intended for the above-mentioned execution subject to switch the current voice interaction mode to the preset first mode, collect all incoming voice signals, and execute all collected voice signals as control instructions.

[0063] Step 406: In response to the monitored voice signal containing the wake-up word, it is determined that the prerequisite for receiving the voice trigger signal is met, and the current preset second mode is switched to the preset first mode.

[0064] This step is based on the judgment result of step 404 that the window type is not the second type. In the scenario of this embodiment, when there are only two window types, the first type and the second type, the window type of the currently displayed window should be the second type. Therefore, the wake-up word will be monitored first, and after the monitored voice signal contains the wake-up word and it is determined that the premise of receiving the voice trigger signal is met, the current preset second mode is switched to the preset first mode, so that the subsequent state is the same as step 405.

[0065] It can also be simply understood as follows: the conventional voice interaction mechanism is to first enter the preset second mode to monitor the wake-up word, and then enter the preset first mode to receive and process the voice control instructions subsequently spoken by the user after monitoring the voice trigger signal containing the wake-up word. That is to say, in the embodiment, the preset second mode is not a complete voice interaction mode, but a prefix component selectively combined with the preset first mode.

[0066] Of course, in other embodiments, the preset first mode and the preset second mode may also exist as a complete voice interaction mode. In this case, if the window type is the first type of window is no longer displayed on the top layer due to the user's subsequent operation (for example, it is closed or adjusted to run in the background), it can also be switched from the current preset first mode to the preset second mode, that is, from the wake-up word-free mode to the wake-up word-required mode.

[0067] On the basis of any of the above embodiments, considering that the incoming voice control command will be directly executed in the preset first mode, in order to avoid misoperation and recognize the voice signal of non-users as control commands, when it is found that the received voice signal contains multiple voiceprints (meaning that the voice signals of multiple users are recognized at the same time), the valid voice signal corresponding to the valid voiceprint is determined according to at least one of the sound source position, sound intensity, and first appearance time of the sound signal corresponding to the different voiceprints, and finally the valid voice signal is executed as the control command. For example, the user who has previously received the voice signal can be identified as a valid user according to the continuity of the voiceprint, and the user who does not come from the sound source in the non-mirror front area can be determined as an invalid user according to the sound source position, and so on.

[0068] It should also be noted that, considering the different installation scenarios and large-screen characteristics of the smart mirror, the voice command library, wake-up words, and keywords can also be supplemented and updated. For example, when the smart mirror is installed in a fitness venue and used as a smart fitness mirror, a batch of fitness-related voice commands and wake-up words can be added to improve the recognition accuracy of voice commands and shorten the recognition time.

[0069] To deepen understanding, the present disclosure also provides a specific implementation solution in combination with a specific application scenario, see Figure 5 The schematic diagram shown:

[0070] 1) User A wakes up the smart mirror from the black screen standby state to the bright screen state by speaking the voice wake-up word "Good morning, Xiao X";

[0071] 2) After the smart mirror is awakened and the screen is on, it is in the desktop homepage by default. The window type recognition component recognizes that the window type of the desktop homepage is the second type that requires a voice wake-up word as a trigger for voice interaction, and switches the current voice interaction mode to the preset second mode;

[0072] 3) User A walks in front of the smart mirror and sends two voice signals, “XiaoX” and “Open News App”, in sequence;

[0073] 4) In the second mode, the smart mirror detects that the first voice signal contains the wake-up word "XiaoX", and then enters the state of waiting to receive voice control instructions, and then executes "open the news application" in the next voice signal as a voice control instruction, which opens the default news application and displays its window at the top of the information display area. After the execution is completed, the smart mirror exits the state of waiting to receive voice control instructions and re-enters the state of listening to the wake-up word;

[0074] 5) Smart mirror in the news application window (such as Figure 5 After the waterfall news display window shown in the figure is opened, the window type recognition component recognizes that the window type of the news application window is the first type that does not require a voice wake-up word as a trigger for voice interaction, switches the current voice interaction mode from the preset second mode to the preset first mode, and maintains a state of waiting to receive a voice control instruction;

[0075] 6) The smart mirror directly executes the voice control commands of "swipe down", "pause", and "play this video" issued by user A while waiting to receive the voice control command;

[0076] 7) After receiving the last "close news application" instruction from user A, the smart mirror closes the news application window and makes the desktop home page the topmost window again. At this time, the window type recognition component recognizes that the window type of the desktop home page is the second type that requires a voice wake-up word as a trigger for voice interaction, and switches the current voice interaction mode from the preset first mode to the preset second mode.

[0077] Further references Figure 6 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a voice interaction device, and the device embodiment is Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0078] like Figure 6 As shown, the voice interaction device 600 of this embodiment may include: a window type determination unit 601, a first mode switching and processing unit 602, and a second mode switching and processing unit 603. Among them, the window type determination unit 601 is configured to determine the window type of the window currently displayed in the top layer of the information display area; the first mode switching and processing unit 602 is configured to switch to the preset first mode in response to the window type being the first type that does not require a wake-up word as a prerequisite for triggering voice interaction, and collect all incoming voice signals as control instructions for execution; the second mode switching and processing unit 603 is configured to switch to the preset second mode in response to the window type being the second type that requires a wake-up word as a prerequisite for triggering voice interaction, and collect subsequent incoming voice signals as control instructions for execution only after receiving a voice trigger signal containing a wake-up word.

[0079] In this embodiment, in the voice interaction device 600, the specific processing of the window type determination unit 601, the first mode switching and processing unit 602, and the second mode switching and processing unit 603 and the technical effects thereof can be referred to in Figure 2 The relevant descriptions of steps 201 - 203 in the corresponding embodiment are not repeated here.

[0080] In some optional implementations of this embodiment, the window type determining unit 601 may be further configured to:

[0081] Read the started applications from the cache storing the information display area;

[0082] Determine the display level of each started application respectively, and determine the target application displayed at the top layer according to the display level;

[0083] Determines the window type of the target application's currently displayed window.

[0084] In some optional implementations of this embodiment, the second mode switching and processing unit 603 may include a second processing subunit configured to collect a subsequent incoming voice signal as a control instruction to execute only after first receiving a voice trigger signal containing a wake-up word, and the second processing subunit may be further configured to:

[0085] In response to the monitored voice signal containing the wake-up word, it is determined that the prerequisite for receiving the voice trigger signal is met, and the current preset second mode is switched to the preset first mode.

[0086] In some optional implementations of this embodiment, the voice interaction device 600 may further include:

[0087] The mode switching unit is configured to switch from the current preset first mode to the preset second mode in response to the window of the first type being no longer displayed on the uppermost layer.

[0088] In some optional implementations of this embodiment, the voice interaction device 600 may further include:

[0089] A first window type predetermining unit is configured to determine a first type of window type of the enumerated information display window adapted to the vertically long strip shape according to the vertically long strip shape presented by the information display area;

[0090] The second window type predetermining unit is configured to determine the window type of other windows that are not determined as the first type as the second type.

[0091] In some optional implementations of this embodiment, the voice interaction device 600 may further include:

[0092] A multi-voiceprint processing unit is configured to determine, in response to a voice signal received in a preset first mode containing multiple voiceprints, a valid voice signal corresponding to a valid voiceprint according to at least one of a sound source position, a sound intensity, and a first appearance time of sound signals corresponding to different voiceprints;

[0093] The execution unit is configured to execute the effective voice signal as a control instruction.

[0094] This embodiment exists as an apparatus embodiment corresponding to the above method embodiment.

[0095] Since the biggest difference between large-screen smart devices such as smart mirrors and traditional small-screen smart devices is that they have information display areas with larger aspect ratios, they are more suitable for presenting more and different information in an enumerated, waterfall, or streaming manner. Therefore, the voice interaction device provided by the embodiment of the present disclosure for a large-screen smart device such as a smart mirror is based on the characteristics of a large-screen smart device such as a smart mirror. When the window presenting such information content is at the top of the information display area, the first mode without the need for a wake-up word as a prerequisite for triggering voice interaction is used as an adapted voice interaction mode, thereby eliminating the need to frequently say the wake-up word during information browsing, thereby improving the human-computer interaction experience.

[0096] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the voice interaction method described in any of the above embodiments can be implemented when the at least one processor executes.

[0097] According to an embodiment of the present disclosure, the present disclosure further provides a readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to implement the voice interaction method described in any of the above embodiments when executed.

[0098] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which, when executed by a processor, can implement the various steps of the voice interaction method described in any of the above embodiments.

[0099] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0100] like Figure 7As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0101] A number of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0102] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as a voice interaction method. For example, in some embodiments, the voice interaction method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the voice interaction method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the voice interaction method in any other appropriate manner (e.g., by means of firmware).

[0103] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0104] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0105] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0106] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0107] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0108] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and virtual private servers (VPS) services.

[0109] Since the biggest difference between large-screen smart devices such as smart mirrors and traditional small-screen smart devices is that they have information display areas with larger aspect ratios, they are more suitable for presenting more and different information in an enumerated, waterfall, or streaming manner. Therefore, the disclosed embodiment targets the characteristics of large-screen smart devices such as smart mirrors. When the window presenting such information content is at the top of the information display area, the first mode without the need for a wake-up word as a prerequisite for triggering voice interaction is used as an adapted voice interaction mode, thereby eliminating the need to frequently say the wake-up word during information browsing, thereby improving the human-computer interaction experience.

[0110] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0111] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A voice interaction method, applied to a smart mirror, comprising: Determine the window type of the topmost window currently displayed in the information display area; In response to the window type being the first type that does not require a wake-up word as a prerequisite for triggering voice interaction, switching to a preset first mode, and collecting all incoming voice signals as control instructions for execution; In response to the window type being the second type that requires a wake-up word as a prerequisite for triggering voice interaction, switching to the preset second mode, and collecting subsequent incoming voice signals as control instructions for execution only after receiving a voice trigger signal containing the wake-up word; According to the longitudinal long strip shape presented by the information display area, determining the first type as the window type of the enumerated information display window adapted to the longitudinal long strip shape; The window types of other windows that are not determined to be of the first type are determined to be of the second type.

2. The method according to claim 1, wherein: The step of determining the window type of the topmost window currently displayed in the information display area includes: Reading the started application from the cache storing the information display area; Determine the display level of each of the started applications respectively, and determine the target application displayed at the top layer according to the display level; The window type of the current display window of the target application is determined.

3. The method according to claim 1, wherein: The method of collecting the subsequent incoming voice signal as a control instruction for execution only after receiving the voice trigger signal containing the wake-up word first includes: In response to the monitored voice signal containing the wake-up word, it is determined that the prerequisite for receiving the voice trigger signal is met, and the current preset second mode is switched to the preset first mode.

4. The method according to claim 1, further comprising: In response to the window of the first type being no longer displayed on the uppermost layer, the current preset first mode is switched to the preset second mode.

5. The method according to any one of claims 1 to 4, further comprising: In response to the voice signal received in the preset first mode including a plurality of voiceprints, determining a valid voice signal corresponding to the valid voiceprint according to at least one of the sound source position, sound intensity, and first appearance time of the sound signals corresponding to the different voiceprints; The effective voice signal is executed as a control instruction.

6. A voice interaction device, applied to a smart mirror, comprising: a window type determination unit configured to determine the window type of the window currently displayed at the top layer of the information display area; A first mode switching and processing unit is configured to switch to a preset first mode in response to the window type being a first type that does not require a wake-up word as a prerequisite for triggering voice interaction, and collect all incoming voice signals as control instructions for execution; A second mode switching and processing unit is configured to switch to a preset second mode in response to the window type being a second type that requires a wake-up word as a prerequisite for triggering voice interaction, and collect subsequent incoming voice signals as control instructions for execution only after receiving a voice trigger signal containing the wake-up word; A first window type predetermining unit is configured to determine the first type of window type of the enumerated information display window adapted to the vertical long strip shape according to the vertical long strip shape presented by the information display area; The second window type predetermining unit is configured to determine the window type of other windows that are not determined as the first type as the second type.

7. The device according to claim 6, wherein: The window type determination unit is further configured to: Reading the started application from the cache storing the information display area; Determine the display level of each of the started applications respectively, and determine the target application displayed at the top layer according to the display level; The window type of the current display window of the target application is determined.

8. The device according to claim 6, wherein: The second mode switching and processing unit includes a second processing subunit configured to collect a subsequent incoming voice signal as a control instruction for execution only after receiving a voice trigger signal containing a wake-up word, and the second processing subunit is further configured to: In response to the monitored voice signal containing the wake-up word, it is determined that the prerequisite for receiving the voice trigger signal is met, and the current preset second mode is switched to the preset first mode.

9. The apparatus according to claim 6, further comprising: The mode switching unit is configured to switch from the current preset first mode to the preset second mode in response to the window of the first type no longer being displayed on the uppermost layer.

10. The device according to any one of claims 6 to 9, further comprising: a multi-voiceprint processing unit configured to, in response to the voice signal received in the preset first mode containing multiple voiceprints, determine a valid voice signal corresponding to the valid voiceprint according to at least one of the sound source position, sound intensity, and first appearance time of the sound signals corresponding to different voiceprints; The execution unit is configured to execute the valid voice signal as a control instruction.

11. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the voice interaction method described in any one of claims 1-5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the voice interaction method according to any one of claims 1 to 5.

13. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the voice interaction method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Voice control method and device, electronic equipment and readable storage medium

    CN112581946A

  • Wake-up response prompting method and display equipment

    CN113066490A