Voice interaction method and washing machine
By obtaining user status information to determine whether to respond to voice wake-up instructions, the problem of voice interaction devices being accidentally awakened is solved, and more accurate voice interaction and improved user experience is achieved.
Patent Information
- Application Number
- CN202010448371.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-05-25
AI Technical Summary
In some scenarios, the user may cause the voice interactive device to be accidentally awakened because the speech or ambient noise is similar to the wake-up words of the device.
Determine whether to respond to the voice wake-up command by collecting voice wake-up commands and obtaining user status information, including the distance between the device and the user's behavior. The user's status information can reflect whether the user has the intention to wake up the voice interactive device.
This method can accurately determine whether to respond to voice wake-up instructions, thereby avoiding false wake-up situations and improving the accuracy and user experience of voice interactive devices.
Smart Images

Figure CN113793616B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of voice interaction technology, and in particular, to a voice interaction method and a washing machine. Background Art
[0002] In the field of voice interaction technology, voice interaction is usually implemented in the following manner: a user speaks a wake-up word to a voice interaction device, and the voice interaction device is awakened after recognizing the wake-up word. Then, the user can communicate with the voice interaction device, and the voice interaction device makes corresponding responses or actions.
[0003] For example, in a mobile phone navigation software, the mobile phone navigation software presets a wake-up word "A". When the user needs the mobile phone navigation software to navigate to location B, the user first speaks the wake-up word "A". After the mobile phone navigation software recognizes the wake-up word "A", it is awakened. At this time, the user then speaks "navigate to location B", and the mobile phone navigation software can plan a navigation route to location B according to the user's request and display it to the user.
[0004] For another example, for a washing machine supporting voice interaction, the washing machine presets a wake-up word "C". When the user needs to control the washing machine to dehydrate by voice, the user first speaks the wake-up word "C". After the washing machine recognizes the wake-up word "C", it is awakened. At this time, the user then speaks "dehydrate", and the washing machine can start to perform the dehydration action according to the user's request.
[0005] However, in some scenarios, when the user is making a call, chatting, or playing music, the voice interaction device may be accidentally awakened because the words or environmental noise are similar to the wake-up word of the device. Summary of the Invention
[0006] The voice interaction method and the washing machine provided by the embodiments of the present invention are used to overcome the problem that the voice interaction device is accidentally awakened.
[0007] The first aspect of the present application provides a voice interaction method, which is applied to a voice interaction device. The method includes:
[0008] Collect a voice wake-up instruction;
[0009] Obtain the status information of the user, where the status information includes at least one of the following: the distance between the voice interaction device and the user, the behavior of the user;
[0010] Determine whether to respond to the voice wake-up instruction according to the status information of the user.
[0011] Optionally, the obtaining the status information of the user specifically includes:
[0012] Obtain the status information of the user according to the collected user image.
[0013] Optionally, determining whether to respond to the voice wake-up instruction according to the status information of the user specifically includes:
[0014] If the distance between the voice interaction device and the user is greater than or equal to a preset distance, or the behavior of the user is a preset behavior, it is determined not to respond to the voice wake-up instruction.
[0015] Optionally, determining whether to respond to the voice wake-up instruction according to the status information of the user specifically includes:
[0016] If the distance between the voice interaction device and the user is less than the preset distance, and the behavior of the user is not the preset behavior, determine the wake-up value of the user, where the wake-up value is related to at least one of the following: the orientation of the user, whether the user pronounces;
[0017] Determine whether to respond to the voice wake-up instruction according to the wake-up value of the user.
[0018] Optionally, there is one user, and determining whether to respond to the voice wake-up instruction according to the wake-up value of the user includes:
[0019] If the wake-up value of the user is greater than or equal to the wake-up threshold, it is determined to respond to the voice wake-up instruction; or,
[0020] If the wake-up value of the user is less than the wake-up threshold, it is determined not to respond to the voice wake-up instruction.
[0021] Optionally, there are multiple users; determining whether to respond to the voice wake-up instruction according to the wake-up value of the user includes:
[0022] According to the wake-up value of each user, obtain the wake-up value corresponding to the voice wake-up instruction;
[0023] If the wake-up value corresponding to the voice wake-up instruction is greater than or equal to the wake-up threshold, it is determined to respond to the voice wake-up instruction; or,
[0024] If the wake-up value corresponding to the voice wake-up instruction is less than the wake-up threshold, it is determined not to respond to the voice wake-up instruction.
[0025] Optionally, obtaining the wake-up value corresponding to the voice wake-up instruction according to the wake-up value of each user includes:
[0026] Taking the distance between each user and the voice interaction device as the weight of the wake-up value of each user, and performing weighted averaging on the wake-up values of the multiple users to obtain the wake-up value corresponding to the voice wake-up instruction.
[0027] Optionally, the method further includes:
[0028] When performing voice playback, obtaining the status information of the user;
[0029] Determining the playback volume during voice playback according to the status information of the user.
[0030] Optionally, the determining the playback volume during voice playback according to the status information of the user includes:
[0031] If the distance between the voice interaction device and the user is less than a preset distance, and the behavior of the user is a preset behavior, then reducing the playback volume during voice playback; or,
[0032] If the distance between the voice interaction device and the user is greater than or equal to the preset distance, or the behavior of the user is not the preset behavior, then using the preset volume as the playback volume during voice playback.
[0033] The second aspect of the present application provides a voice interaction device, which is applied to a voice interaction device, and the device includes:
[0034] An acquisition module, configured to acquire a voice wake-up instruction;
[0035] A processing module, configured to obtain the status information of the user; determining whether to respond to the voice wake-up instruction according to the status information of the user; the status information includes at least one of the following: the distance between the voice interaction device and the user, the behavior of the user.
[0036] Optionally, the processing module is specifically configured to obtain the status information of the user according to the acquired user image.
[0037] Optionally, the processing module is specifically configured to determine not to respond to the voice wake-up instruction when the distance between the voice interaction device and the user is greater than or equal to a preset distance, or the behavior of the user is a preset behavior.
[0038] Optionally, the processing module is specifically configured to determine the wake-up value of the user when the distance between the voice interaction device and the user is less than the preset distance, and the behavior of the user is not the preset behavior; determining whether to respond to the voice wake-up instruction according to the wake-up value of the user; the wake-up value is related to at least one of the following: the orientation of the user, whether the user pronounces.
[0039] Optionally, the user is one;
[0040] The processing module is specifically configured to determine to respond to the voice wake-up instruction when the wake-up value of the user is greater than or equal to the wake-up threshold; or, when the wake-up value of the user is less than the wake-up threshold, determine not to respond to the voice wake-up instruction.
[0041] Optionally, there are multiple users;
[0042] The processing module is specifically configured to obtain the wake-up value corresponding to the voice wake-up instruction according to the wake-up value of each user; when the wake-up value corresponding to the voice wake-up instruction is greater than or equal to the wake-up threshold, determine to respond to the voice wake-up instruction; or, when the wake-up value corresponding to the voice wake-up instruction is less than the wake-up threshold, determine not to respond to the voice wake-up instruction.
[0043] Optionally, the processing module is specifically configured to use the distance between each user and the voice interaction device as the weight value of the wake-up value of each user, and perform weighted averaging on the wake-up values of the multiple users to obtain the wake-up value corresponding to the voice wake-up instruction.
[0044] Optionally, the processing module is further configured to obtain the status information of the user during voice playback; and determine the playback volume during voice playback according to the status information of the user.
[0045] Optionally, the processing module is specifically configured to reduce the playback volume during voice playback when the distance between the voice interaction device and the user is less than a preset distance and the behavior of the user is a preset behavior; or, when the distance between the voice interaction device and the user is greater than or equal to the preset distance, or the behavior of the user is not the preset behavior, use the preset volume as the playback volume during voice playback.
[0046] A third aspect of the present application provides a voice interaction device, including: at least one processor and a memory;
[0047] The memory stores computer execution instructions;
[0048] The at least one processor executes the computer execution instructions stored in the memory, so that the device executes the method according to any one of the first aspects.
[0049] A fourth aspect of the present application provides a computer-readable storage medium, on which computer execution instructions are stored. When the computer execution instructions are executed by a processor, the method according to any one of the first aspects is implemented.
[0050] In a fifth aspect, an embodiment of the present invention provides a washing machine, and the washing machine is used to execute the method according to any one of the first aspects.
[0051] The voice interaction method and washing machine provided by the embodiments of the present invention. The voice interaction device collects a voice wake-up instruction and obtains the status information of the user. Since the status information of the user can reflect whether the user has the intention to wake up the voice interaction device, therefore, according to the status information of the user, this method can accurately determine whether to respond to the voice wake-up instruction, thereby avoiding the situation of false wake-up. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The preferred embodiments of the program update method provided by the present application will be described below with reference to the accompanying drawings and in combination with specific embodiments. The drawings are as follows:
[0053] Figure 1 is a schematic structural diagram of a washing machine provided by the embodiments of the present invention;
[0054] Figure 2 is a schematic flowchart of a voice interaction method provided by the embodiments of the present invention;
[0055] Figure 3 is a schematic flowchart of another voice interaction method provided by the embodiments of the present invention;
[0056] Figure 4 is a schematic flowchart of yet another voice interaction method provided by the embodiments of the present invention;
[0057] Figure 5 is a schematic flowchart of still another voice interaction method provided by the embodiments of the present invention;
[0058] Figure 6 is a schematic flowchart of still another voice interaction method provided by the embodiments of the present invention;
[0059] Figure 7 is a schematic structural diagram of a voice interaction device provided by the embodiments of the present invention;
[0060] Figure 8 is a schematic structural diagram of another voice interaction device provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0062] In recent years, voice interaction technology has been applied in many application scenarios, such as navigation software, smart speakers, household appliances, etc. Among them, household appliances such as washing machines, refrigerators, air conditioners, microwave ovens, televisions, etc. that support voice interaction.
[0063] However, in some scenarios, when the user is making a call, chatting, or playing music, the voice interaction device may be accidentally awakened because the words or environmental noise are similar to the wake-up word of the device.
[0064] Exemplarily, a certain washing machine supports voice interaction. Assume that the wake-up word of the washing machine is "Dabai". When the user says "Chinese cabbage" during a call, the washing machine misidentifies the first two words of the "Chinese cabbage" spoken by the user as the wake-up word, and then wakes up the washing machine, resulting in the washing machine being accidentally awakened.
[0065] Since the user's state can reflect whether the user has the intention to wake up the voice interaction device, therefore, in order to overcome the above-described technical problems, the embodiments of the present invention utilize the user's state information to determine whether it affects the voice wake-up instruction input by the user, so as to avoid accidental wake-up.
[0066] The execution subject of the embodiments of the present invention is a voice interaction device, which is provided with a microphone, a speaker, and a camera. Among them, the microphone is used to collect voice, the speaker is used to play voice, and the camera is used to collect images.
[0067] The voice interaction device can be, for example, a washing machine, a refrigerator, an air conditioner, a smart speaker, a mobile phone, etc.
[0068] Taking the voice interaction device as a washing machine as an example below, Figure 1 is a schematic structural diagram of a washing machine provided by an embodiment of the present invention. As Figure 1 shown, the washing machine includes a main body 11, a microphone 12, a speaker 13, and a camera 14. The microphone 12 and the speaker can be set on the front or side of the main body 11 ( Figure 1 taking the side as an example), and the camera 14 can be set on the front of the main body 11 to collect images within the field of view directly in front of the washing machine.
[0069] Next, in combination with specific embodiments, the technical solutions of the voice interaction method provided by the present invention will be described in detail. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0070] Figure 2 is a schematic flowchart of a voice interaction method provided by an embodiment of the present invention. As Figure 2 shown, the method of the present invention may include:
[0071] S101. Collect a voice wake-up command.
[0072] Exemplarily, when the voice interaction device is a washing machine, referring to Figure 1 , the washing machine can collect voice through the microphone 12, and based on existing voice recognition technology, identify whether the voice is a voice wake-up command. For example, the washing machine can determine whether the voice is a preset wake-up word or includes a preset wake-up word. When it is recognized that the voice is a preset wake-up word or includes a preset wake-up word, it is determined that the voice is a voice wake-up command.
[0073] S102. Obtain the status information of the user.
[0074] The status information includes at least one of the following: the distance between the voice interaction device and the user, the behavior of the user.
[0075] Exemplarily, the status information of the user can be obtained according to the collected user image, or the status information of the user can be obtained according to the collected user voice, or the motion trajectory of the user can be collected through a sensor installed on the user to obtain the status information of the user.
[0076] Taking the collected user image as an example, at least the following two methods can be used to obtain the distance between the voice interaction device and the user:
[0077] In a possible implementation, the voice interaction device can collect a user image through the camera 14. The user image can be a single-frame user image or a continuous multi-frame user image, which can be specifically determined according to the actual situation. Among them, the user image can include one user or multiple users. The voice interaction device can extract the size of the reference object in the user image from the user image. The voice interaction device can determine the distance between the voice interaction device and the user based on the size of the reference object in the image, the actual size of the reference object, and the size of the user in the image. Among them, the reference object is within the field of view of the camera, and the actual size of the reference object can be pre-stored in the voice interaction device. When the voice interaction device is connected to the network, the actual size of the reference object can also be obtained through the network.
[0078] In another possible implementation, the voice interaction device can collect a user image through the camera 14, then calibrate the installation height of the camera and the distance to the nearest end of the field of view, then obtain the distance value between the top and bottom of the user in the user image and the horizontal distance of the user in the image, and then use the perspective geometric relationship between the angle of the ground corresponding to the user image captured by the camera and the central projection of the camera to convert the distance between the voice interaction device and the user into the ground view angle of the user image, so that the distance between the voice interaction device and the user can be determined according to the calibrated installation height of the camera and the distance to the nearest end of the field of view.
[0079] When there is one user, the distance from the voice interaction device to the user can be obtained; when there are multiple users, the distance from the voice interaction device to each user can be obtained.
[0080] Taking the collected user images as an example, at least the following two methods can be used to obtain the user's behavior:
[0081] In a possible implementation, the voice interaction device can collect user images through the camera 14, and then input the user images into the user behavior recognition model to recognize the user's behavior. The user behavior recognition model can be obtained by training by inputting a set of user image samples calibrated with user behaviors into the user behavior recognition model.
[0082] In another possible implementation, the voice interaction device can collect user images through the camera 14, then extract the user's contour in the user image, compare the contour of the user in the user image with the contour model, and each contour model indicates a corresponding behavior. When the similarity between the user's contour and a certain contour model is greater than a certain threshold, the behavior of the user can be recognized as the behavior corresponding to the contour model.
[0083] When there is one user, the behavior of the user can be obtained; when there are multiple users, the behavior of each user can be obtained.
[0084] S103. Determine whether to respond to the voice wake-up instruction according to the user's status information.
[0085] When the distance between the voice interaction device and the user and the user's behavior indicate that the user has the intention to wake up the voice interaction device, it is determined to respond to the voice wake-up instruction; when the distance between the voice interaction device and the user and the user's behavior indicate that the user does not have the intention to wake up the voice interaction device, it is determined not to respond to the voice wake-up instruction to avoid the voice interaction device being woken up by mistake.
[0086] Among them, responding to the voice wake-up instruction, for example, in Figure 1 the washing machine shown, a preset voice message can be played to the user through the speaker 13 to respond to the voice wake-up instruction.
[0087] In the embodiment of the present invention, the voice interaction device collects the voice wake-up instruction and obtains the user's status information. Since the user's status information can reflect whether the user has the intention to wake up the voice interaction device, therefore, this method can accurately judge whether to respond to the voice wake-up instruction according to the user's status information, so as to avoid the situation of mis-wake-up.
[0088] On the basis of the above embodiment, the following embodiment will focus on introducing how to determine whether to respond to the voice wake-up instruction according to the distance between the voice interaction device and the user and the user's behavior when there is one user.
[0089] Figure 3 It is a schematic flowchart of another voice interaction method provided by an embodiment of the present invention. Based on Figure 2 , as Figure 3 shown, the method may further include:
[0090] S201. Determine whether the distance between the voice interaction device and the user is greater than or equal to a preset distance.
[0091] Wherein, the preset distance can be set according to the actual situation.
[0092] If the distance between the voice interaction device and the user is greater than or equal to the preset distance, it indicates that the user is at a relatively far distance from the voice interaction device. In the case of a relatively far distance, the user usually has no intention of waking up the voice interaction device. At this time, step S203 is executed.
[0093] If the distance between the voice interaction device and the user is less than the preset distance, it indicates that the user may have the intention of waking up the voice interaction device. At this time, step S202 is executed.
[0094] S202. Determine whether the user's behavior is a preset behavior.
[0095] Wherein, the preset behavior can be, for example, making a call, having a conversation, writing, etc., and can be specifically set according to the user's needs. Among them, the user's needs can be selected through the menu in the washing machine control panel, or the user can also set it in the mobile phone APP and send it to the washing machine through the mobile phone after setting.
[0096] If the user's behavior is a preset behavior, it indicates that the user is doing other things, such as making a call, having a conversation, writing, etc., and the user has no intention of waking up the voice interaction device. At this time, step S203 is executed.
[0097] If the user's behavior is not a preset behavior, it indicates that the user may have the intention of waking up the voice interaction device. At this time, step S204 is executed.
[0098] S203. Determine not to respond to the voice wake-up instruction.
[0099] At this time, by not responding to the voice wake-up instruction, the voice interaction device can be prevented from being accidentally woken up.
[0100] S204. Determine the user's wake-up value.
[0101] The wake-up value is related to at least one of the following: the user's orientation, whether the user pronounces.
[0102] Among them, for example, the user's orientation can be obtained by recognizing the user's image.
[0103] For example, it is also possible to obtain whether the user is pronouncing by recognizing the lip movement of the user in the user image. Exemplarily, based on the continuously acquired multi-frame user images, it can be determined whether the user's lips have closing and opening actions, and further whether the user is pronouncing can be obtained. Further, it can also be determined according to the collected voice that the user's lip movement is not chewing or the like.
[0104] The wake-up value of the user can be obtained based on any one of the user's orientation and whether the user is pronouncing.
[0105] For example, when determining the wake-up value of the user based on the user's orientation, when the user faces the voice interaction device directly, the wake-up value of the user is determined to be a1; when the user faces the voice interaction device with the back, the wake-up value of the user is determined to be a2; when the user faces the voice interaction device sideways, the wake-up value of the user is a3, where a1 > a3 > a2. For example, multiple angular ranges can also be set for the angle of the user's side body. When the angle of the user's side body is within the corresponding range, the wake-up value corresponding to this range is determined.
[0106] For example, when determining the wake-up value of the user based on whether the user is pronouncing, when the user is pronouncing, the wake-up value of the user is determined to be b1; when the user is not pronouncing, the wake-up value of the user is determined to be b2, where b1 > b2. For example, when the duration of the user's pronunciation is the same as the duration of the user saying the wake-up word, the wake-up value of the user is determined to be c1; when the user is not pronouncing, the wake-up value of the user is determined to be c2; when the duration of the user's pronunciation is different from the duration of the user saying the wake-up word, the wake-up value of the user is determined to be c3, where c1 > c3 > c2.
[0107] The wake-up value of the user can be obtained based on the user's orientation and whether the user is pronouncing. Specifically, the wake-up value of the user determined based on the user's orientation and the wake-up value of the user determined based on whether the user is pronouncing can be added together to obtain the wake-up value of the user. Weights can also be set for the user's orientation and whether the user is pronouncing. After multiplying the wake-up value of the user determined based on the user's orientation and the wake-up value of the user determined based on whether the user is pronouncing by the corresponding weights, they are added together to obtain the wake-up value of the user.
[0108] S205. Determine whether to respond to the voice wake-up instruction according to the wake-up value of the user.
[0109] When the wake-up value of the user is greater than or equal to the wake-up threshold, it indicates that the user has the intention to wake up the voice interaction device. At this time, it can be determined to respond to the voice wake-up instruction.
[0110] When the wake-up value of the user is less than the wake-up threshold, it indicates that the user has no intention to wake up the voice interaction device. At this time, it can be determined not to respond to the voice wake-up instruction.
[0111] Among them, the wake-up threshold can be set according to the actual situation.
[0112] Exemplarily, as Figure 1 shown in the washing machine, the washing machine obtains the user image through the camera 14, and obtains the distance between the washing machine and the user and the user's behavior according to the user image. When the distance between the washing machine and the user is greater than the preset distance of 4 meters, or when the user is on the phone, the washing machine will not be woken up. When the distance between the washing machine and the user is less than the preset distance of 4 meters, and the user is not on the phone, talking or writing, the wake-up value of the user is further determined according to the orientation of the user and whether the user is pronouncing. Assuming that the wake-up threshold is 60, when the wake-up value is 70, the washing machine is woken up, and when the wake-up value is 30, the washing machine is not woken up.
[0113] In the embodiment of the present invention, it is judged whether the user has the intention to wake up the voice interaction device according to the distance between the voice interaction device and the user and the user's behavior. When the user has no intention to wake up the voice interaction device, it is determined not to respond to the voice wake-up instruction; when the user has the intention to wake up the voice interaction device, the wake-up value is further determined according to at least one of the orientation of the user and whether the user is pronouncing, and compared with the wake-up threshold to further determine whether the user has the intention to wake up the voice interaction device. Based on the above method, the present invention can accurately identify whether the user has the intention to wake up the voice interaction device, and thus can effectively avoid the voice interaction device from being woken up by mistake.
[0114] The following embodiments will focus on how to determine whether to respond to the voice wake-up instruction according to the distance between the voice interaction device and the user and the user's behavior when there are multiple users.
[0115] Figure 4 is a schematic flowchart of another voice interaction method provided by the embodiment of the present invention. On the basis of Figure 2 as Figure 4 shown, the method may further include:
[0116] S301. Judge whether the distances between the voice interaction device and multiple users are greater than or equal to the preset distance.
[0117] If the distances between the voice interaction device and each user are greater than or equal to the preset distance, it means that the distance between each user and the voice interaction device is relatively far. In the case of a relatively far distance, multiple users have no intention to wake up the voice interaction device. At this time, step S303 is executed.
[0118] If the distance between the voice interaction device and at least one user is less than the preset distance, it means that at least one user may have the intention to wake up the voice interaction device. At this time, step S302 is executed.
[0119] S302. Judge whether the behaviors of multiple users are preset behaviors.
[0120] If the behavior of each user is a preset behavior, it indicates that the user is doing other things, such as making a call, having a conversation, writing, etc. Multiple users have no intention of waking up the voice interaction device. At this time, step S303 is executed.
[0121] If the behavior of at least one user is not a preset behavior, it indicates that at least one user may have the intention of waking up the voice interaction device. At this time, step S304 is executed.
[0122] S303. Determine not to respond to the voice wake-up instruction.
[0123] At this time, by not responding to the voice wake-up instruction, the voice interaction device can be prevented from being woken up by mistake.
[0124] S304. Determine the wake-up values of multiple users.
[0125] The wake-up value of each user is related to at least one of the following: the orientation of the user, whether the user is pronouncing.
[0126] Among them, the specific way in which the wake-up value of the user can be obtained based on any one of the orientation of the user and whether the user is pronouncing can refer to the description of the foregoing step S204 and will not be elaborated here.
[0127] S305. Determine whether to respond to the voice wake-up instruction according to the wake-up values of multiple users.
[0128] According to the wake-up value of each user, obtain the wake-up value corresponding to the voice wake-up instruction.
[0129] In a possible implementation manner, the average value of the wake-up values of multiple users can be calculated, and this average value is used as the wake-up value corresponding to the voice wake-up instruction.
[0130] For example, when there are 3 users, the wake-up values of the 3 users are respectively denoted as x1, x2, x3, then the wake-up value xn corresponding to the voice wake-up instruction can be denoted as: xn = (x1 + x2 + x3) / 3.
[0131] In another possible implementation manner, the distance between each user and the voice interaction device is used as the weight value of the wake-up value of each user, and the wake-up values of multiple users are weighted and averaged to obtain the wake-up value corresponding to the voice wake-up instruction.
[0132] For example, when there are 3 users, the wake-up values of the 3 users are respectively denoted as x1, x2, x3, and the distances between the 3 users and the voice interaction device are respectively denoted as r1, r2, r3, then the wake-up value xn corresponding to the voice wake-up instruction can be denoted as: xn = (x1*r1 + x2*r2 + x3*r3) / 3.
[0133] When the wake-up value corresponding to the voice wake-up instruction is greater than or equal to the wake-up threshold, it indicates that the user intends to wake up the voice interaction device. At this time, it can be determined to respond to the voice wake-up instruction.
[0134] When the wake-up value corresponding to the voice wake-up instruction is less than the wake-up threshold, it indicates that the user does not intend to wake up the voice interaction device. At this time, it can be determined not to respond to the voice wake-up instruction.
[0135] In the embodiment of the present invention, it is determined whether multiple users have the intention to wake up the voice interaction device according to the distance between the voice interaction device and multiple users and the behaviors of multiple users. When multiple users do not have the intention to wake up the voice interaction device, it is determined not to respond to the voice wake-up instruction; when at least one of multiple users has the intention to wake up the voice interaction device, the wake-up values of multiple users are further determined according to at least one of the orientations of multiple users and whether multiple users pronounce, and based on the comparison between the wake-up values of multiple users and the wake-up threshold, it is further determined whether multiple users have the intention to wake up the voice interaction device. Based on the above method, the present invention can accurately identify whether multiple users have the intention to wake up the voice interaction device, and thus can effectively avoid the voice interaction device from being woken up by mistake.
[0136] On the basis of the above embodiment, when there is one user, the present invention can also, after the voice interaction device is woken up, adjust the volume of voice playback and delay the response to voice answering based on the status information of the user, so as to avoid affecting the user when the user is on the phone or in conversation, and improve the user experience.
[0137] Figure 5 is a schematic flowchart of another voice interaction method provided by the embodiment of the present invention. As Figure 5 shown, the method may further include:
[0138] S401. When performing voice playback, obtain the status information of the user.
[0139] Among them, the status information of the user may include at least one of the following: the distance between the voice interaction device and the user, the behavior of the user, etc., and specifically, reference may be made to step S102 above.
[0140] S402. Determine the playback volume during voice playback according to the status information of the user.
[0141] In a possible implementation manner, if the distance between the voice interaction device and the user is less than a preset distance, and the behavior of the user is a preset behavior, the playback volume during voice playback can be reduced at this time to avoid affecting the user and improve the user experience. If the distance between the voice interaction device and the user is greater than or equal to the preset distance, or the behavior of the user is not a preset behavior, the preset volume can be used as the playback volume during voice playback to facilitate the interaction between the user and the voice interaction device. Among them, the preset volume can be set according to the actual situation.
[0142] In another possible implementation, only the distance between the voice interaction device and the user is considered, and the user's behavior is not considered. For example, if the distance between the voice interaction device and the user is less than a preset distance, the playback volume during voice playback can be reduced at this time to avoid disturbing the user and improve the user experience. If the distance between the voice interaction device and the user is greater than or equal to the preset distance, the preset volume can be used as the playback volume during voice playback to facilitate the interaction between the user and the voice interaction device.
[0143] In yet another possible implementation, only the user's behavior is considered, and the distance between the voice interaction device and the user is not considered. For example, if the user's behavior is a preset behavior, the playback volume during voice playback can be reduced at this time to avoid disturbing the user and improve the user experience. If the user's behavior is not a preset behavior, the preset volume can be used as the playback volume during voice playback to facilitate the interaction between the user and the voice interaction device.
[0144] In addition to using the method in step S402, the voice response can also be delayed according to the user's status information.
[0145] In one possible implementation, if the distance between the voice interaction device and the user is less than the preset distance and the user's behavior is a preset behavior, the voice response can be temporarily not responded to at this time to avoid disturbing the user and improve the user experience. When the distance between the voice interaction device and the user is greater than or equal to the preset distance, or the user's behavior is not a preset behavior, the voice response can be made at this time to facilitate the interaction between the user and the voice interaction device.
[0146] For example, assume that the preset distance is 4 meters and the user is making a call 2 meters away from the voice interaction device. At this time, when the washing machine finishes washing and is about to play the response of "Washing completed", in order to avoid disturbing the user's call, it is not played temporarily. It will be played again after the user stops making the call or the distance between the voice interaction device and the user exceeds 4 meters.
[0147] In another possible implementation, only the distance between the voice interaction device and the user is considered, and the user's behavior is not considered. For example, if the distance between the voice interaction device and the user is less than the preset distance, the voice response can be temporarily not responded to at this time to avoid disturbing the user and improve the user experience. When the distance between the voice interaction device and the user is greater than or equal to the preset distance, the voice response can be made at this time to facilitate the interaction between the user and the voice interaction device.
[0148] In another possible implementation, only the user's behavior is considered, and the distance between the voice interaction device and the user is not considered. For example, if the user's behavior is a preset behavior, the voice response can be temporarily not responded to at this time to avoid affecting the user and improve the user experience. When the user's behavior is not a preset behavior, the voice response can be responded to at this time to facilitate the interaction between the user and the voice interaction device.
[0149] In the embodiment of the present invention, when playing voice, the state information of the user is obtained, and according to the state information of the user, the playing volume when playing voice is determined, or the voice response of the user is delayed, so that the user can be avoided from being affected when making a call or having a conversation, and the user experience is improved.
[0150] On the basis of the above embodiment, when there are multiple users, the present invention can also, after the voice interaction device is awakened, based on the state information of the users, adjust the volume of voice playback and delay the response to the voice response, so as to avoid affecting the users when the users are making a call or having a conversation, and improve the user experience.
[0151] Figure 6 It is a schematic flowchart of another voice interaction method provided by the embodiment of the present invention. As Figure 6 shown, the method may further include:
[0152] S501. When playing voice, obtain the state information of multiple users.
[0153] Among them, the state information of the user may include at least one of the following: the distance between the voice interaction device and the user, the behavior of the user, etc., which can be specifically referred to the above embodiment.
[0154] S502. Determine the playing volume when playing voice according to the state information of multiple users.
[0155] In a possible implementation, if the distance between the voice interaction device and at least one user is less than a preset distance, and the behavior of the at least one user is a preset behavior, the playing volume when playing voice can be reduced at this time to avoid affecting the user and improve the user experience. If the distance between the voice interaction device and each user is greater than or equal to the preset distance, or the behavior of each user is not a preset behavior, the preset volume can be used as the playing volume when playing voice to facilitate the interaction between the user and the voice interaction device. Among them, the preset volume can be set according to the actual situation.
[0156] In another possible implementation, only the distance between the voice interaction device and the user is considered, and the user's behavior is not considered. For example, if the distance between the voice interaction device and at least one user is less than a preset distance, the playback volume during voice playback can be reduced at this time to avoid disturbing the user and improve the user experience. If the distance between the voice interaction device and each user is greater than or equal to the preset distance, the preset volume can be used as the playback volume during voice playback to facilitate the interaction between the user and the voice interaction device.
[0157] In yet another possible implementation, only the user's behavior is considered, and the distance between the voice interaction device and the user is not considered. For example, if the behavior of at least one user is a preset behavior, the playback volume during voice playback can be reduced at this time to avoid disturbing the user and improve the user experience. If the behavior of each user is not a preset behavior, the preset volume can be used as the playback volume during voice playback to facilitate the interaction between the user and the voice interaction device.
[0158] In addition to using the method in step S502, the voice response can also be delayed according to the user's status information. Here, the voice response is an instruction for which the user requests the voice interaction device to give a voice response.
[0159] In a possible implementation, if the distance between the voice interaction device and at least one user is less than a preset distance, and the behavior of the at least one user is a preset behavior, the voice response can be temporarily not responded to at this time to avoid disturbing the user and improve the user experience. When the distance between the voice interaction device and each user is greater than or equal to the preset distance, or the behavior of each user is not a preset behavior, the voice response can be responded to at this time to facilitate the interaction between the user and the voice interaction device.
[0160] In another possible implementation, only the distance between the voice interaction device and the user is considered, and the user's behavior is not considered. For example, if the distance between the voice interaction device and at least one user is less than a preset distance, the voice response can be temporarily not responded to at this time to avoid disturbing the user and improve the user experience. When the distance between the voice interaction device and each user is greater than or equal to the preset distance, the voice response can be responded to at this time to facilitate the interaction between the user and the voice interaction device.
[0161] In yet another possible implementation, only the user's behavior is considered, and the distance between the voice interaction device and the user is not considered. For example, if the behavior of at least one user is a preset behavior, the voice response can be temporarily not responded to at this time to avoid disturbing the user and improve the user experience. When the behavior of each user is not a preset behavior, the voice response can be responded to at this time to facilitate the interaction between the user and the voice interaction device.
[0162] In an embodiment of the present invention, when playing voice, the status information of multiple users is obtained, and based on the status information of the multiple users, the playing volume during voice playing is determined, or the response to the user's voice response is delayed, which can avoid disturbing the user during a call or conversation and improve the user experience.
[0163] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program code.
[0164] Figure 7 is a schematic structural diagram of a voice interaction device provided by an embodiment of the present invention, as Figure 7 shown, this device is applied to a voice interaction device, and this device may include: an acquisition module 21 and a processing module 22. Among them,
[0165] The acquisition module 21 is used to acquire a voice wake-up instruction.
[0166] The processing module 22 is used to obtain the status information of the user; determine whether to respond to the voice wake-up instruction according to the status information of the user; the status information includes at least one of the following: the distance between the voice interaction device and the user, the behavior of the user.
[0167] Optionally, in some possible implementation manners, the processing module 22 is specifically configured to obtain the status information of the user according to the acquired user image.
[0168] Optionally, in some possible implementation manners, the processing module 22 is specifically configured to determine not to respond to the voice wake-up instruction when the distance between the voice interaction device and the user is greater than or equal to a preset distance, or when the behavior of the user is a preset behavior.
[0169] Optionally, in some possible implementation manners, the processing module 22 is specifically configured to determine the wake-up value of the user when the distance between the voice interaction device and the user is less than the preset distance and the behavior of the user is not a preset behavior; determine whether to respond to the voice wake-up instruction according to the wake-up value of the user; the wake-up value is related to at least one of the following: the orientation of the user, whether the user is pronouncing.
[0170] Optionally, in some possible implementation manners, there is one user;
[0171] The processing module 22 is specifically configured to determine to respond to the voice wake-up instruction when the wake-up value of the user is greater than or equal to the wake-up threshold; or determine not to respond to the voice wake-up instruction when the wake-up value of the user is less than the wake-up threshold.
[0172] Optionally, in some possible implementation manners, there are multiple users;
[0173] The processing module 22 is specifically configured to obtain the wake-up value corresponding to the voice wake-up instruction according to the wake-up value of each user; when the wake-up value corresponding to the voice wake-up instruction is greater than or equal to the wake-up threshold, determine to respond to the voice wake-up instruction; or, when the wake-up value corresponding to the voice wake-up instruction is less than the wake-up threshold, determine not to respond to the voice wake-up instruction.
[0174] Optionally, in some possible implementation manners, the processing module 22 is specifically configured to use the distance between each user and the voice interaction device as the weight value of the wake-up value of each user, and perform weighted averaging on the wake-up values of multiple users to obtain the wake-up value corresponding to the voice wake-up instruction.
[0175] Optionally, in some possible implementation manners, the processing module 22 is further configured to obtain the status information of the user during voice playback; and determine the playback volume during voice playback according to the status information of the user.
[0176] Optionally, in some possible implementation manners, the processing module 22 is specifically configured to reduce the playback volume during voice playback when the distance between the voice interaction device and the user is less than a preset distance and the behavior of the user is a preset behavior; or, when the distance between the voice interaction device and the user is greater than or equal to the preset distance or the behavior of the user is not a preset behavior, use the preset volume as the playback volume during voice playback.
[0177] The present invention Figure 7 The voice interaction device provided by the embodiment shown can perform the actions of the voice interaction device in the above method embodiment. For example, the voice interaction device can be the voice interaction device itself or a chip of the voice interaction device.
[0178] Figure 8 is a schematic structural diagram of another voice interaction device provided by an embodiment of the present invention. As Figure 8 shown, the device includes: a memory 91 and at least one processor 92.
[0179] The memory 91 is used to store program instructions.
[0180] The processor 92 is configured to implement the voice interaction method in the embodiment of the present invention when the program instructions are executed. The specific implementation principle can be referred to the above embodiment, and will not be elaborated here in this embodiment.
[0181] The voice interaction device may further include an input / output interface 93.
[0182] The input / output interface 93 may include an independent output interface and an input interface, or may be an integrated interface that integrates input and output. Among them, the output interface is used to output data, and the input interface is used to obtain the input data. The output data mentioned above is a general term for the output in the above method embodiments, and the input data is a general term for the input in the above method embodiments.
[0183] This application also provides a readable storage medium, in which execution instructions are stored. When at least one processor of the voice interaction device executes the execution instructions, when the computer execution instructions are executed by the processor, the voice interaction method in the above embodiments is implemented.
[0184] This application also provides a program product, which includes execution instructions stored in a readable storage medium. At least one processor of the voice interaction device can read the execution instructions from the readable storage medium, and the execution of the execution instructions by at least one processor enables the voice interaction device to implement the voice interaction methods provided by the above various embodiments.
[0185] An embodiment of the present invention provides a voice interaction device, which is used to implement the above method embodiments. The voice interaction device may be, for example: a washing machine, a refrigerator, an air conditioner, etc.
[0186] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be in electrical, mechanical or other forms.
[0187] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0188] In addition, in each embodiment of this application, the various functional modules can be integrated in one processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware, or in the form of hardware plus software functional modules.
[0189] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above-mentioned software functional modules stored in a storage medium include several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (English: Read-Only Memory, abbreviated as: ROM), random access memories (English: Random Access Memory, abbreviated as: RAM), magnetic disks, or optical discs that can store program codes.
[0190] In the above embodiments of the server or the terminal, it should be understood that the processing module can be a central processing unit (English: Central Processing Unit, abbreviated as: CPU), and can also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (English: Application SpecificIntegrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the present application can be directly embodied as being executed and completed by a hardware processor, or can be executed and completed by a combination of hardware and software modules in the processor.
[0191] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present application.
Claims
1. A voice interaction method, characterized in that, The method is applied to a voice interaction device, and the method includes: Collecting a voice wake-up instruction; Obtaining the status information of the user, where the status information includes at least one of the following: the distance between the voice interaction device and the user, the behavior of the user, and the behavior of the user includes at least one of the following: making a call, having a conversation, writing; Determining whether to respond to the voice wake-up instruction according to the status information of the user; When there is one user, the determining whether to respond to the voice wake-up instruction according to the status information of the user specifically includes: If the distance between the voice interaction device and the user is less than a preset distance, and the behavior of the user is not a preset behavior, then determining the wake-up value of the user, where the wake-up value is related to at least one of the following: the orientation of the user, whether the user is pronouncing; Determining whether to respond to the voice wake-up instruction according to the wake-up value of the user; The method further includes: obtaining the status information of the user during voice playback; If the distance between the voice interaction device and the user is less than a preset distance, and the behavior of the user is a preset behavior, then reducing the playback volume during voice playback and delaying the response to the voice response.
2. The method according to claim 1, characterized in that, The obtaining the status information of the user specifically includes: Obtaining the status information of the user according to the collected user image.
3. The method according to claim 1, wherein The determining whether to respond to the voice wake-up instruction according to the status information of the user further includes: If the distance between the voice interaction device and the user is greater than or equal to the preset distance, or the behavior of the user is a preset behavior, then determining not to respond to the voice wake-up instruction.
4. The method according to claim 1, wherein When there is one user, the determining whether to respond to the voice wake-up instruction according to the wake-up value of the user includes: If the wake-up value of the user is greater than or equal to the wake-up threshold, then determining to respond to the voice wake-up instruction; or, If the wake-up value of the user is less than the wake-up threshold, then determining not to respond to the voice wake-up instruction.
5. The method according to claim 4, characterized in that, When there are multiple users; the determining whether to respond to the voice wake-up instruction according to the wake-up value of the user includes: Obtaining the wake-up value corresponding to the voice wake-up instruction according to the wake-up value of each user; If the wake-up value corresponding to the voice wake-up instruction is greater than or equal to the wake-up threshold, then determining to respond to the voice wake-up instruction; or, If the wake-up value corresponding to the voice wake-up instruction is less than the wake-up threshold, then determining not to respond to the voice wake-up instruction.
6. The method according to claim 4, characterized in that The obtaining the wake-up value corresponding to the voice wake-up instruction according to the wake-up value of each user includes: Taking the distance between each user and the voice interaction device as the weight of the wake-up value of each user, and performing weighted averaging on the wake-up values of the multiple users to obtain the wake-up value corresponding to the voice wake-up instruction.
7. The method according to claim 1, characterized in that, It further includes: If the distance between the voice interaction device and the user is greater than or equal to the preset distance, or the behavior of the user is not the preset behavior, then taking the preset volume as the playback volume during voice playback.
8. A washing machine, characterized in that, The washing machine is used to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Voice recognition method, intelligent device and storage medium
CN108711430A
Interactive method and equipment
CN109767774A
Natural interactive voice control method and device
CN110136714A