Method, device and electronic device for waking up voice assistant
By detecting voice data and motion data, combined with the display screen occlusion status, the false wake-up of voice assistant applications is avoided, and the power consumption and user experience problems caused by false wake-up in the prior art are solved, achieving a more efficient wake-up mechanism.
Patent Information
- Application Number
- CN202411596210.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-11-11
AI Technical Summary
In the prior art, voice assistant applications are susceptible to environmental factors and accidentally wake up, resulting in increased power consumption of electronic devices and user experience.
By detecting voice data and motion data and combining whether the display is in an obstructed state, it is determined whether the current scene is a false wake-up scene, thereby avoiding the false wake-up of the voice assistant application.
It effectively avoids the false wake-up of voice assistant applications, reduces the power consumption of electronic devices, and improves the user experience.
Smart Images

Figure CN119299560B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of terminal technology, and in particular to a method, device and electronic device for waking up a voice assistant. Background Art
[0002] With the development of voice recognition technology, voice assistant applications are installed in many electronic devices to help users complete the human-computer interaction process with the electronic devices. Generally speaking, the voice assistant application is in a dormant state, and when the user wants to use the voice assistant, the voice assistant application can be awakened.
[0003] Currently, waking up the voice assistant application can be achieved through breath awakening; in the breath awakening method, the electronic device wakes up the voice assistant by receiving voice data and gesture data related to breath awakening the voice assistant.
[0004] However, in the above-mentioned implementation method of breath awakening, if affected by environmental factors, there is a situation where the voice assistant application is mistakenly awakened. If the voice assistant application is frequently mistakenly awakened, it will increase the power consumption of the electronic device and affect the user experience. Summary of the invention
[0005] The present application provides a method, device and electronic device for waking up a voice assistant, which can ensure that the voice assistant will not be woken up by mistake as much as possible, reduce the power consumption of the electronic device, and ensure user experience.
[0006] In a first aspect, the present application provides a method for waking up a voice assistant, wherein the electronic device includes a voice assistant application, and the method includes:
[0007] During the operation of the electronic device, if first voice data is detected and the first voice data satisfies a first condition, it is determined whether the display screen of the electronic device is in an obstructed state, and the first condition includes that the sound corresponding to the first voice data conforms to the sound regularity of the sound produced by the user's speech; if the display screen is not in an obstructed state, it is determined whether first motion data is detected, and the first motion data is greater than a preset threshold (the preset threshold includes 0); if the first motion data is detected, it is determined whether the first motion data satisfies a second condition, and the second condition includes that the first motion data is motion data generated when the electronic device performs a first action, and the first action is an action of the mobile electronic device related to waking up the voice assistant application; if the first motion data satisfies the second condition, a first interface is displayed, and the first interface is a voice interaction interface of the voice assistant application.
[0008] The preset threshold is usually 0, and the first motion data can be understood as acceleration data.
[0009] In the above method, it is considered that in the scenario of normal awakening by breath, the electronic device should not be in an obstructed state. Based on this, during the operation of the electronic device, if voice data is detected and the sound corresponding to the voice data conforms to the sound law of the sound produced by the user's speech, it can be determined whether the display screen is in an obstructed state. If it is in an obstructed state, it may be a false awakening scene, and the awakening process ends. If it is not in an obstructed state, it can be determined that the current scene may not be a false awakening scene, and it can be further determined whether motion data is detected. If motion data is detected, and the motion data is motion data generated by the first action that the user wants to perform by breath awakening, the voice assistant application can be started. Therefore, by determining whether the display screen is in an obstructed state, the scenario of false awakening of the voice assistant application can be avoided, the situation where the power consumption of the electronic device is increased due to the false awakening of the voice assistant application can be avoided, and the intrusion of the user by the false awakening of the voice assistant application can be avoided, thereby ensuring the user experience.
[0010] In conjunction with the first aspect, in some implementations of the first aspect, determining whether a display screen of an electronic device is in a blocked state includes:
[0011] A first distance value between a display screen and a target body is obtained, wherein the distance between the target body and a first plane where the display screen is located is less than the distance between the target body and a plane where a first component in the electronic device is located, and the first component is any device in the electronic device except the display screen; determining whether the first distance value is greater than or equal to a first threshold value; if the first distance value is greater than or equal to the first threshold value, determining that the display screen is not in an obstructed state; if the first distance value is less than the first threshold value, determining that the display screen is in an obstructed state.
[0012] Among them, the distance between the target body and the first plane where the display screen is located is smaller than the distance between the target body and the plane where the first component in the electronic device is located. The first component is any device in the electronic device except the display screen, that is, the target body is any object located on the same side as the display screen.
[0013] In the above method, the electronic device can determine whether the display screen is in an obstructed state by determining the distance between the display screen and the target object. If the distance is small, it means that it is in an obstructed state; if the distance is large, it means that it is not in an obstructed state.
[0014] In combination with the first aspect, in some implementations of the first aspect, the electronic device includes a distance sensor, and obtaining a first distance value between the display screen and the target object includes: obtaining the first distance value between the display screen and the target object through the distance sensor.
[0015] The distance sensor and the display screen are located on the same side of the electronic device.
[0016] For example, a distance sensor is provided in the display screen.
[0017] In combination with the first aspect, in some implementations of the first aspect, if the first motion data is detected, determining whether the first motion data satisfies the second condition includes:
[0018] If the first motion data is detected, determine whether the electronic device is in a shaking state, and the shaking state is used to indicate that the electronic device moves in accordance with the motion law of the second motor in the electronic device; if the electronic device is not in a shaking state, determine whether the first motion data meets the second condition.
[0019] It should be understood that in the scenario where the electronic device wakes up the voice assistant application by breathing, the user is required to wake up the voice assistant application by picking up the electronic device and speaking through the first action. Taking into account that the electronic device may be in a shaking state due to alarms, incoming calls, etc., the electronic device may recognize the motion data generated by alarms, incoming calls, etc. as the motion data generated by the user picking up the electronic device through the first action, causing the voice assistant application to be woken up by mistake.
[0020] Thus, in the above method, if the electronic device detects the first motion data, the electronic device can further determine whether the current scene is a false awakening scene by determining whether the electronic device is in a shaking state. If the electronic device is in a shaking state, it indicates that the current scene is not a scene for waking up the voice assistant by breath, that is, it is a false awakening scene, and the process of waking up the voice assistant by breath ends. If the electronic device is not in a shaking state, the electronic device can further determine whether the current scene is a scene for waking up the voice assistant by breath by determining whether the first motion data satisfies the second condition. In this way, the breath awakening accuracy of the voice assistant application can be improved, and false awakening of the voice assistant application can be avoided.
[0021] In conjunction with the first aspect, in some implementations of the first aspect, determining whether the electronic device is in a shaking state includes:
[0022] The first motion data is input into the vibration model, and a first similarity value is output. The vibration model is used to predict the similarity between the first motion data and the motion data corresponding to the first motor when the first motor is in a working state. The vibration model is trained by the motion data corresponding to the first motor when the first motor is in a working state. The first motor and the second motor in the electronic device are motors of the same type. If the first similarity value is greater than or equal to the second threshold value, it is determined that the electronic device is in a shaking state.
[0023] In the above method, since the vibration model is trained by the motion data corresponding to the first motor being in a working state, the vibration model can recognize the motion data corresponding to the first action.
[0024] In combination with the first aspect, in certain implementations of the first aspect, the first motion data includes acceleration data collected by an acceleration sensor in an electronic device, the acceleration data includes first acceleration data and second acceleration data within a first time period, the first time period is a time period between a current moment and a first moment, the first moment is a moment earlier than the current moment, the first acceleration data includes data in the acceleration data that is greater than an acceleration threshold, the acceleration threshold includes 0, the first motion data is input into a vibration model, and a first similarity value is output, including: inputting the first acceleration data into the vibration model to obtain a first similarity value.
[0025] For example, the acceleration threshold is 0, the first acceleration data includes data that is not 0 in the acceleration data, and the second acceleration data includes data that is 0 in the acceleration data.
[0026] In the above method, the second acceleration data is data when the electronic device is in a stationary state, and the first acceleration data is data when the electronic device performs a first action. The second acceleration data is irrelevant to determining whether the electronic device is in a shaking state and need not be included in the calculation.
[0027] In combination with the first aspect, in some implementations of the first aspect, the method further includes:
[0028] If the electronic device is in a shaking state, determine whether the second motor in the electronic device is in a working state; if the second motor is not in a working state, determine whether the first motion data meets the second condition.
[0029] In the above method, it is taken into account that the electronic device is in a shaking state, which may be because the user shakes the electronic device when picking up the electronic device and speaking to wake up the voice assistant application. Therefore, when the electronic device is in a shaking state, it can also be determined whether the second motor of the electronic device is in a working state. If the second motor is in a working state, it is a false awakening scenario, and the process of waking up the voice assistant by breath ends. If the second motor is not in a working state, it may not be a false awakening scenario, and the electronic device can continue to determine whether the user is waking up the voice assistant application by breath awakening through the first motion data. Thus, by judging whether the second motor is in a working state, the false awakening of the voice assistant application can be ruled out, and the breath awakening accuracy of the voice assistant application can be further improved, and the false awakening of the voice assistant application can be avoided. At the same time, events such as the alarm clock stopping due to false awakening can be avoided.
[0030] In conjunction with the first aspect, in some implementations of the first aspect, the vibration model generation process includes:
[0031] Obtain N sample motion data and sample similarity values corresponding to the N sample motion data, where each of the N sample motion data is collected when the first motor is in a working state, and N is a positive integer greater than or equal to 2; input the N sample motion data and the sample similarity values corresponding to the N sample motion data into a classification training model, train the classification training model, and obtain a vibration model.
[0032] The vibration model is trained by the motion data corresponding to the first motor in the working state, so the vibration model can identify the motion data corresponding to the first action. The N sample motion data include the motion data when the first motor is located on various objects and is in the working state.
[0033] In the above method, specifically, each of the N sample motion data includes three-axis acceleration data. During training, each positive data unit in the three-axis acceleration data can be marked with a confidence level of [0.9, 1]; and each negative data unit can be marked with a confidence level of [0, 0.3]. The positive data unit can be understood as data that is not 0 in the three-axis acceleration data, and the negative data unit can be understood as data that is 0 in the three-axis acceleration data.
[0034] In conjunction with the first aspect, in some implementations of the first aspect, determining whether the first motion data satisfies the second condition includes:
[0035] The first motion data is input into the gesture model to obtain a second similarity value. The gesture model is used to predict the similarity between the first motion data and the motion data corresponding to the first action. The gesture model is trained with the motion data corresponding to the first action. If the second similarity value is greater than or equal to the third threshold, it is determined that the first motion data meets the second condition.
[0036] The gesture model is trained by the motion data corresponding to the first action, so the gesture model can recognize the motion data corresponding to the first action.
[0037] In the above method, when the first motion data is detected, the first motion data can be input into the gesture model. If the second similarity value output by the gesture model is greater than or equal to the third threshold, it means that the first motion data has a high similarity with the motion data corresponding to the first action, and the first motion data is the motion data generated when the electronic device performs the first action, then the electronic device can determine that the first motion data meets the second condition.
[0038] In combination with the first aspect, in some implementations of the first aspect, the electronic device includes a first microphone and a second microphone, the first voice data includes first sub-voice data collected by the first microphone and second sub-voice data collected by the second microphone, the first microphone and the second microphone are located on different sides of the electronic device, and if the first voice data is detected and the first voice data satisfies a first condition, determining whether a display screen of the electronic device is in an obstructed state includes:
[0039] If the first voice data is detected, the sound pressure amplitude difference between the first sub-voice data and the second sub-voice data is determined; if the sound pressure amplitude difference is greater than or equal to the fourth threshold, the first voice data satisfies the first condition, and it is determined whether the display screen is in an obstructed state.
[0040] The present application does not limit the specific positions of the first microphone and the second microphone in the electronic device. For example, the first microphone is located at the top of the electronic device, and the second microphone is located at the bottom of the electronic device. For another example, the first microphone is located at the side of the electronic device, and the second microphone is located at the bottom of the electronic device.
[0041] In the above method, due to the different positions of the two microphones, if the user speaks into one of the microphones, the sound source has obvious directionality, and the time and intensity of the sound reaching the two microphones are quite different. Then, the relative position relationship between the user's mouth and the microphone will cause the sound signals received by the two microphones to have differences in sound pressure amplitude.
[0042] Usually, the sound of background noise comes from the surroundings of electronic equipment, the sound source does not have obvious directionality, and the difference in time and intensity between the sound reaching the two microphones is small.
[0043] Based on this, the electronic device can determine whether the sound corresponding to the first voice data is the sound of the breath produced by the user's speech through the fourth threshold.
[0044] In combination with the first aspect, in some implementations of the first aspect, if first motion data is detected, determining whether the first motion data satisfies a second condition includes:
[0045] If first motion data is detected, determine whether the first motion data satisfies the second condition, and determine whether the first voice data satisfies the third condition, the third condition including that a third similarity value between a first voiceprint feature corresponding to the first voice data and a voiceprint feature of a target sound pre-stored in the electronic device is greater than or equal to a fifth threshold; if the first motion data satisfies the second condition, display the first interface, including: if the first motion data satisfies the second condition, and the first voice data satisfies the third condition, display the first interface.
[0046] In the above method, the electronic device can also further determine whether the current scene is a scene for waking up the voice assistant application by breath awakening by determining whether the first motion data meets the second condition and whether the first voice data meets the third condition. In this way, the breath awakening accuracy of the voice assistant application can be improved and the voice assistant application can be avoided from being awakened by mistake.
[0047] In combination with the first aspect, in certain implementations of the first aspect, the third condition also includes a first sound source angle being within a first range, the first sound source angle being used to indicate the angle of the direction of the sound source of the first voice data relative to the electronic device, and the first sound source angle being determined based on the sound pressure amplitude difference.
[0048] In a second aspect, the present application provides a device for waking up a voice assistant, wherein the device for waking up a voice assistant includes a module for executing the method in the first aspect and any possible implementation manner of the first aspect.
[0049] In a third aspect, the present application provides an electronic device, comprising: one or more processors, and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method in the first aspect and any possible implementation of the first aspect.
[0050] In a fourth aspect, the present application provides a chip system, which is applied to an electronic device. The chip system includes one or more processors, and the one or more processors are used to call computer instructions to enable the electronic device to execute the method in the first aspect and any possible implementation of the first aspect.
[0051] In a fifth aspect, the present application provides a computer-readable storage medium, which includes instructions. When the instructions are executed on an electronic device, the electronic device executes the method in the first aspect and any possible implementation of the first aspect.
[0052] In a sixth aspect, the present application provides a computer program product. When the computer program product runs on a computer, it enables the computer to execute the method in the first aspect and any possible implementation manner of the first aspect.
[0053] It can be understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 A schematic diagram of an electronic device in three directions provided by an embodiment of the present application;
[0055] Figure 2 A diagram of a human-computer interaction interface provided in one embodiment of the present application;
[0056] Figure 3 A diagram of a human-computer interaction interface provided in one embodiment of the present application;
[0057] Figure 4 A diagram of a human-computer interaction interface provided in one embodiment of the present application;
[0058] Figure 5 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application;
[0059] Figure 6 A schematic diagram of a software architecture of an electronic device provided in one embodiment of the present application;
[0060] Figure 7 A flowchart of a method for waking up a voice assistant provided in one embodiment of the present application;
[0061] Figure 8 A schematic diagram of microphone positions provided in one embodiment of the present application;
[0062] Fig. 9 A flowchart of a method for waking up a voice assistant provided in one embodiment of the present application;
[0063] Fig.10 A schematic diagram of the structure of a classification training model provided in one embodiment of the present application;
[0064] Fig.11 A schematic diagram of acceleration data provided by an embodiment of the present application;
[0065] Fig.12 A flowchart of a method for waking up a voice assistant provided in one embodiment of the present application;
[0066] Fig.13 A flowchart of a method for waking up a voice assistant provided in one embodiment of the present application;
[0067] Fig.14 A schematic diagram of the structure of a device for waking up a voice assistant provided in one embodiment of the present application. DETAILED DESCRIPTION
[0068] In this application, "at least one" means one or more, and "plurality" means two or more. "And / or" describes the first conversion relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c alone can represent: a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c can be single or multiple. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance.
[0069] To facilitate understanding, some of the examples given are provided for reference to the description of concepts related to the embodiments of the present application.
[0070] 1. Acceleration sensor
[0071] The accelerometer is a common inertial detection sensor that can measure the three-axis acceleration signals of electronic devices on three perpendicular axes (x-axis, y-axis and z-axis).
[0072] The three-axis acceleration signal reflects the motion state of the electronic device in three-dimensional space, including static, linear motion, and rotational motion. The three-axis acceleration signal can be converted into digital three-axis acceleration data through the conversion device inside the acceleration sensor, such as the analog-to-digital converter (ADC).
[0073] Among them, Figure 1 As shown, for any frame of acceleration signal collected by the acceleration sensor, the electronic device can obtain three-axis acceleration data, which includes acceleration signal values corresponding to the three directions of the X-axis, Y-axis and Z-axis of the electronic device.
[0074] With the development of voice recognition technology, voice assistant applications are installed in many electronic devices to help users complete the human-computer interaction process with the electronic devices. Generally speaking, the voice assistant application is in a dormant state, and when the user wants to use the voice assistant, the voice assistant application can be awakened.
[0075] At present, the three main technologies for waking up voice assistant applications include button wake-up, keyword wake-up and breath wake-up.
[0076] In the button wake-up technology, the voice assistant is woken up by the user's triggering operation on a button (such as the power button). In the keyword wake-up technology, the voice assistant is woken up by receiving a specific wake-up word (for example, "Hello, YOYO", "Xiaoyi, Xiaoyi", "Hi Siri") input by the user's voice.
[0077] Next, combine Figure 2-Figure 4 , taking the electronic device as a mobile phone as an example, the specific implementation method of breath awakening is explained.
[0078] The phone can display Figure 2 The interface 11 shown in (a) is a setting interface of the mobile phone. The interface 12 may include: a control 101. The control 101 is used to trigger the setting interface of the voice assistant application.
[0079] Upon receiving the user's Figure 2 After the control 101 shown in (a) is triggered, the mobile phone can Figure 2 The interface 11 shown in (a) in FIG. 1 changes to display as shown in FIG. Figure 2 The interface 12 shown in (b) in the figure may include: a control 102. The control 102 is used to trigger the display of the breath awakening setting interface.
[0080] Upon receiving the user's Figure 2 After the control 102 shown in (b) in FIG. 1 is triggered, the mobile phone can be used as follows: Figure 2 The interface 12 shown in (b) in FIG. 1 changes to display as shown in FIG. Figure 3 The interface 13 shown in (a) of FIG. 1 may include: a control 103. The control 103 is used to trigger the breath wake-up function of the mobile phone.
[0081] Upon receiving the user's Figure 3 After the control 103 shown in (a) is triggered, the mobile phone starts the breath wake-up function of the voice assistant application.
[0082] If the mobile phone receives voice data and motion data related to the breath-activated voice assistant, the voice assistant application is activated, so that the mobile phone can display the following information: Figure 3 The interface 14 shown in (b) .
[0083] The interface 14 may include: a control 104. The control 104 is used to remind the user that the voice assistant application has been started, and to trigger the voice assistant application to end obtaining the first voice data of the user.
[0084] After receiving the voice data, the voice assistant application can convert the voice data into text data. At this time, the interface 14 can also include a control 105. The control 105 is used to display the text data corresponding to the voice data.
[0085] For example, when the voice data is "Today's weather", the text data displayed in the control 105 is "Today's weather". It should be understood that when the voice data is "Today's weather", it means that the user wants to know today's weather conditions.
[0086] When the voice assistant application finishes voice recognition based on the voice data, the phone can display the following Figure 4 Interface 15 is shown.
[0087] The interface 15 may further include a display area 106. The display area 106 is used to display the speech recognition result of the voice assistant application.
[0088] For example, when the voice data is "Today's weather", the voice recognition result displayed in the display area 106 may be "City A, mostly cloudy, 10% chance of rainfall, current temperature 14°C, maximum temperature 17°C, minimum temperature 11°C".
[0089] In the above-mentioned implementation method of breath awakening, there is a situation where the voice assistant application is woken up by mistake due to environmental influences. If the voice assistant is woken up by mistake frequently, it will increase the power consumption of the electronic device and affect the user experience.
[0090] Taking a mobile phone as an example, for the breath-activated voice assistant application, there may be the following false-activated scenarios: The voice assistant application may be activated by zippers or friction when the phone is in a moving user's pocket. In a moving user's pocket, when the motion data detected by the phone meets the gesture model, and the voice data corresponding to the zipper sound or friction sound detected by the phone is close to the sound of the breath produced by the user's speech, it will trigger the voice assistant application to be activated by breath. The triggering of the false-activated breath will invoke some behaviors of the voice assistant application, which will increase unnecessary power consumption.
[0091] According to some user feedback, the following scenarios may occur: the phone is placed in the pocket, and the voice assistant application suddenly wakes up; the phone is placed in the pocket, and the voice assistant application suddenly wakes up while the user is talking to another user; the phone is placed in the pocket, and the voice assistant application suddenly wakes up while the user is chatting with another user through headphones.
[0092] In view of the above problems, the present application can provide a method, device, electronic device, chip system, computer-readable storage medium and computer program product for waking up a voice assistant. Considering that in the scenario of normal awakening by breath, the electronic device should not be in an obstructed state, based on this, during the operation of the electronic device, if voice data is detected and the sound corresponding to the voice data conforms to the sound law of the sound generated by the user's speech (the sound of the breath generated by the user's speech), it can be determined whether the display screen is in an obstructed state. If it is in an obstructed state, it may be a false awakening scene, and the awakening process ends. If the display screen is not in an obstructed state, it can be determined that the current scene may not be a false awakening scene, and it can be further determined whether motion data is detected. If motion data is detected, and the motion data is motion data generated by the action that the user wants to perform by breath awakening, the voice assistant application can be started. Therefore, the scenario of false awakening of the voice assistant application can be avoided by whether the display screen is in an obstructed state, the situation where the power consumption of the electronic device is increased due to the false awakening of the voice assistant application can be avoided, and the intrusion of the user by the false awakening of the voice assistant application can be avoided, thereby ensuring the user experience.
[0093] Among them, the above-mentioned electronic device can be an electronic device with a voice assistant application, the voice assistant application has a breath wake-up function, and the electronic device has display hardware and corresponding software support.
[0094] The electronic devices may include mobile phones, tablet computers, vehicle-mounted devices, laptop computers, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), smart cars, smart TVs, robots and other devices.
[0095] It should be noted that, in some possible implementations, the electronic device may also be referred to as a terminal device (station), a user equipment (UE), etc., which is not limited in the embodiments of the present application.
[0096] For ease of explanation, Figure 5 In the description, the electronic device 100 is taken as a mobile phone as an example.
[0097] like Figure 5 As shown, in some embodiments, the electronic device 100 may include a processor 101, a communication module 102, a display screen 103, a camera 104, a sensor 105, an internal memory 106, a USB interface 107, an external memory interface 108, a charging management module 109, a power management module 110, and a battery 111, etc.
[0098] Among them, the processor 101 may include one or more processing units. For example, the processor 101 may include an application processor (AP), a modem processor, a graphics processor, an image signal processor (ISP), a controller, a memory, a video stream codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors 101.
[0099] The controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0100] The digital signal processor may also include an audio digital signal processor (ADSP).
[0101] A memory may also be provided in the processor 101 for storing instructions and data.
[0102] The communication module 102 may include antenna 1 and antenna 2, a mobile communication module, and / or a wireless communication module.
[0103] Among them, the sensor 105 may include a sound acquisition sensor, an acceleration sensor, and a distance sensor.
[0104] Optionally, the electronic device 100 may further include peripheral devices such as a mouse, a button, an indicator light, a keyboard, a speaker, a microphone, etc.
[0105] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100.
[0106] In other embodiments, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0107] Please refer to Figure 6 , which is a schematic diagram of the software architecture of an electronic device provided by an embodiment of this application. The method for waking up the voice assistant provided by the embodiment of this application is applied to Figure 5 When the electronic device 100 shown, the software in the electronic device 100 may be divided as follows Figure 6 The illustrated application layer 201 , application framework layer (framework, FWK) 202 , hardware abstraction layer (hardware abstraction layer, HAL) 203 and driver layer 204 .
[0108] A plurality of applications may be installed in the application layer 201 .
[0109] For example, the application layer 201 includes a voice assistant application.
[0110] The application framework layer 202 provides a set of basic functions and services for the application layer 201 to call and use.
[0111] The hardware abstraction layer 203 is a software located between the operating system kernel and the hardware circuit, and is usually used to abstract the hardware to achieve the interaction between the operating system and the hardware circuit at the logic layer.
[0112] The driver layer 204 may be installed with multiple drivers for driving the hardware.
[0113] It should be noted that the application layer 201, the application framework layer 202, the hardware abstraction layer 203 and the driver layer 204 may also include other contents, which are not specifically limited here.
[0114] Based on the above description, the method for waking up the voice assistant provided in the embodiment of the present application is elaborated in detail below in combination with the accompanying drawings and application scenarios.
[0115] See also Figure 7 , Figure 7 A flow chart of a method for waking up a voice assistant provided in one embodiment of the present application is shown.
[0116] like Figure 7 As shown, the method for waking up the voice assistant provided in this application may include:
[0117] S301. During the operation of the electronic device, if first voice data is detected and the first voice data satisfies a first condition, determine whether the display screen of the electronic device is in an obstructed state. The first condition includes that the sound corresponding to the first voice data conforms to the sound regularity of the sound generated by the user's speech.
[0118] The electronic device may include a sound collection sensor, which can collect voice signals in real time and convert the voice signals into first voice data. In some embodiments, the sound collection sensor is a microphone (MIC).
[0119] If the electronic device detects the user's first voice data, the electronic device can determine whether the first voice data meets the first condition, that is, determine whether the sound corresponding to the first voice data is the sound of the breath produced by the user's speech. The sound of the breath produced by the user's speech can be understood as the sound produced by the airflow of the user's speech hitting the microphone.
[0120] If the sound corresponding to the first voice data is not the sound of the breath produced by the user speaking, then the sound corresponding to the first voice data may be environmental noise, indicating that the user may not have awakened the voice assistant application by breath awakening, and the process of awakening the voice assistant by breath ends.
[0121] If the sound corresponding to the first voice data is the sound of the user's breath when speaking, then it means that the user may be waking up the voice assistant application by breathing, and the electronic device can further determine whether the display screen of the electronic device is in an obstructed state.
[0122] In some embodiments, the electronic device includes a first microphone and a second microphone, the first voice data includes first sub-voice data collected by the first microphone and second sub-voice data collected by the second microphone, the first microphone and the second microphone are located on different sides of the electronic device, and if the first voice data is detected and the first voice data satisfies a first condition, determining whether the display screen of the electronic device is in an obstructed state includes:
[0123] If the first voice data is detected, the sound pressure amplitude difference between the first sub-voice data and the second sub-voice data is determined; if the sound pressure amplitude difference is greater than or equal to the fourth threshold, the first voice data satisfies the first condition, and it is determined whether the display screen is in an obstructed state.
[0124] The present application does not limit the specific positions of the first microphone and the second microphone in the electronic device. Below, two situations of the positions of the first microphone and the second microphone in the electronic device are described.
[0125] Case 1: The first microphone is located at the top of the electronic device, and the second microphone is located at the bottom of the electronic device, that is, the first microphone and the second microphone are respectively located at two short sides of the electronic device.
[0126] Taking a non-folding screen mobile phone as an example, the position of the first microphone can be found in Figure 8 The position of the microphone A in the electronic device 100 is shown in (a), and the position of the second microphone can be seen in Figure 8 FIG. 1 is a diagram showing a position of the microphone B in the electronic device 100 as shown in FIG. 1 (a).
[0127] Case 2: The first microphone is located on the side of the electronic device, and the second microphone is located on the bottom of the electronic device, that is, the first microphone is located on the long side of the electronic device, and the second microphone is located on the short side of the electronic device.
[0128] Taking the electronic device as a small folding screen mobile phone as an example, the position of the first microphone can be seen in Figure 8 The position of the microphone C in the electronic device 100 is shown in (b) of FIG. 1 . The position of the second microphone can be seen in FIG. Figure 8 FIG. 2 shows a position of the microphone D in the electronic device 100 as shown in FIG.
[0129] The sound pressure amplitude difference is used to represent the difference in sound intensity between the first sub-voice data and the second sub-voice data.
[0130] Due to the different positions of the two microphones, if the user speaks into one of the microphones, the sound source has obvious directionality, and the time and intensity of the sound reaching the two microphones are quite different. Then, the relative position relationship between the user's mouth and the microphone will cause the sound signals received by the two microphones to have differences in sound pressure amplitude.
[0131] Usually, the sound of background noise comes from the surroundings of electronic equipment, the sound source does not have obvious directionality, and the difference in time and intensity between the sound reaching the two microphones is small.
[0132] Based on this, the electronic device can determine whether the sound corresponding to the first voice data is the sound of the breath produced by the user's speech through the fourth threshold.
[0133] In some embodiments, the electronic device can determine the sound pressure amplitude difference between the first sub-voice data and the second sub-voice data through a near-talk algorithm; specifically, the electronic device can input the first sub-voice data and the second sub-voice data into the near-talk algorithm, and the near-talk algorithm can output the sound pressure amplitude difference between the first sub-voice data and the second sub-voice data, and determine whether the first voice data meets the first condition by whether the sound pressure amplitude difference is greater than or equal to a fourth threshold.
[0134] Among them, the electronic device can also determine the sound pressure amplitude difference between the first sub-voice data and the second sub-voice data in other ways, and this application does not limit this.
[0135] It should be understood that in the scenario of waking up the voice assistant application through breath, the user needs to pick up the electronic device and speak to wake up the voice assistant application, and the display screen of the electronic device should not be blocked.
[0136] Based on this, if the electronic device detects the first voice data and the first voice data meets the first condition, if the display screen of the electronic device is in an obstructed state (for example, the electronic device is in the user's pocket, the electronic device is placed on the table, and the display screen faces the table), then the current scene does not belong to the scene of waking up the voice assistant application by breath awakening, it is a false awakening scene, and the process of waking up the voice assistant by breath ends.
[0137] If the display screen of the electronic device is not blocked, then it means that the current scene may be a scene of waking up the voice assistant application by breath awakening. The electronic device can perform the next operation, such as S302, to further determine whether the current scene is a scene of waking up the voice assistant application by breath awakening.
[0138] S302: If the display screen is not in an obstructed state, determine whether first motion data is detected, and the first motion data is greater than a preset threshold, and the preset threshold includes 0.
[0139] Among them, the electronic device may include an inertial detection sensor, which can collect inertial signals in real time and convert the inertial signals into first motion data; the inertial signal can also be called an inertial measurement unit (IMU) signal, and the first motion data can also be called IMU data.
[0140] In the scenario of waking up the voice assistant application through breath, the user needs to pick up the electronic device and speak to wake up the voice assistant application. When the user picks up the electronic device, the electronic device can detect the first motion data.
[0141] In this way, if the display screen of the electronic device is not in an obstructed state, the electronic device can continue to determine whether the first motion data is detected to further determine whether the current scene belongs to a scene for waking up the voice assistant application by breath waking.
[0142] If the first motion data is not detected, it means that the current scene may not be a scene for waking up the voice assistant application by breath awakening, and it is a false awakening scene, and the process of waking up the voice assistant by breath ends.
[0143] If the first motion data is detected, it means that the current scene may be a scene of waking up the voice assistant application by breath waking.
[0144] S303. If first motion data is detected, determine whether the first motion data satisfies a second condition, where the second condition includes that the first motion data is motion data generated when the electronic device performs a first action, and the first action is an action of the mobile electronic device related to waking up the voice assistant application.
[0145] In the scenario of waking up the voice assistant application through breath, the user is required to pick up the electronic device, and the action of the user picking up the electronic device needs to be a preset action, that is, the first action, and the first action is usually a wrist raising action.
[0146] Therefore, if the first motion data is detected, the electronic device also needs to determine whether the first motion data is motion data generated when the user moves the electronic device through the first action. If the first motion data is motion data generated when the user moves the electronic device through the first action, the first motion data satisfies the second condition, indicating that the current scene belongs to a scene of waking up the voice assistant application by breath waking.
[0147] Similarly, if the first motion data is not the motion data generated when the user moves the electronic device through the first action, the first motion data does not meet the second condition, indicating that the current scene does not belong to the scene of waking up the voice assistant application through breath awakening.
[0148] In some embodiments, the electronic device may determine whether the first motion data satisfies the second condition through a gesture model, specifically including:
[0149] The first motion data is input into the gesture model to obtain a second similarity value. The gesture model is used to predict the similarity between the first motion data and the motion data corresponding to the first action. The gesture model is trained with the motion data corresponding to the first action. If the second similarity value is greater than or equal to the third threshold, it is determined that the first motion data meets the second condition.
[0150] The gesture model is trained by the motion data corresponding to the first action, so the gesture model can recognize the motion data corresponding to the first action.
[0151] When the first motion data is detected, the first motion data can be input into the gesture model. If the second similarity value output by the gesture model is greater than or equal to the third threshold, it means that the first motion data has a high similarity with the motion data corresponding to the first action, and the first motion data is the motion data generated when the electronic device performs the first action. Then, the electronic device can determine that the first motion data meets the second condition.
[0152] S304: If the first motion data meets the second condition, a first interface is displayed, where the first interface is a voice interaction interface of a voice assistant application.
[0153] If the first motion data satisfies the second condition, it indicates that the current scene belongs to a scene of waking up the voice assistant application by breath waking, and therefore, the electronic device can start the voice assistant application.
[0154] After the voice assistant application is started, the electronic device can display a first interface, which is a voice interaction interface of the voice assistant application.
[0155] The first interface may display the content of the first voice data. This application does not limit the specific implementation of the first interface.
[0156] For example, the specific implementation of the first interface can be found in Figure 3 The interface 14 shown in (b) will not be described in detail here.
[0157] In addition, after displaying the first interface, the electronic device may also display a second interface, and the second interface may include the content of the electronic device's response data to the first voice data. The present application does not limit the specific implementation of the second interface.
[0158] For example, the specific implementation of the second interface can be found in Figure 4 The interface 15 shown is not described in detail here.
[0159] The method for waking up a voice assistant of the present application, during the operation of an electronic device, if first voice data is detected, a first condition can be used to determine whether the sound corresponding to the first voice data is the sound of the breath produced by the user's speech. If the first voice data meets the first condition, the electronic device can determine that the user may be waking up the voice assistant application by waking up with the breath. In this way, the electronic device can further determine whether the current scene is a false awakening scene by determining whether the display screen is in an obstructed state. If the display screen is in an obstructed state, it can be determined that the current scene is not a false awakening scene, and the process of waking up the voice assistant application by breath ends. If the display screen is not in an obstructed state, it can be determined that the current scene is not a false awakening scene. In this way, the electronic device The sub-device can continue to determine whether the current scene belongs to the scene of waking up the voice assistant application by breath through the second condition. If the first motion data meets the second condition, the electronic device can determine that the current scene belongs to the scene of waking up the voice assistant application by breath. Therefore, the electronic device can display the first interface of the voice assistant application. Thus, the electronic device can avoid the occurrence of false awakening of the voice assistant application by breath caused by other non-active operations when the electronic device is in a blocked state, can avoid the situation where the power consumption of the electronic device increases due to false awakening of the voice assistant application, and can avoid the intrusion of the user by the situation where the voice assistant application receives and broadcasts sound due to false awakening, thereby ensuring the user experience.
[0160] Based on the above Figure 7According to the description of the illustrated embodiment, the electronic device can not only determine whether the current scene is a false awakening scene by determining whether the display screen is in an obstructed state, but further, the electronic device can also determine whether the current scene is a false awakening scene by determining whether the electronic device is in a shaking state.
[0161] Next, combine Fig. 9 , which introduces in detail the method of waking up the voice assistant of this application.
[0162] See also Fig. 9 , Fig. 9 A flow chart of a method for waking up a voice assistant provided in one embodiment of the present application is shown.
[0163] like Fig. 9 As shown, the method for waking up the voice assistant provided in this application may include:
[0164] S401. During operation of the electronic device, first voice data and first motion data are detected in real time.
[0165] In the process of the electronic device running, the specific implementation method of real-time detection of the first voice data can be found in Figure 7 The description of S301 in the illustrated embodiment regarding real-time detection of the first voice data during operation of the electronic device will not be repeated here.
[0166] During the operation of the electronic device, the specific implementation method of real-time detection of the first motion data can be found in Figure 7 The description of S302 in the illustrated embodiment regarding real-time detection of the first motion data during the operation of the electronic device will not be repeated here.
[0167] S402: If first voice data is detected, determine whether the first voice data satisfies a first condition.
[0168] The specific implementation method of determining whether the first voice data meets the first condition can be found in Figure 7 The description of whether the first voice data satisfies the first condition in S301 in the illustrated embodiment will not be repeated here.
[0169] When the first condition is met, the electronic device can execute S403; when the first condition is not met, it means that the user may not have awakened the voice assistant application by breath awakening, and the electronic device does not execute any steps, and the process of awakening the voice assistant by breath ends.
[0170] S403, obtaining a first distance value between the display screen and the target body, wherein the distance between the target body and a first plane where the display screen is located is smaller than the distance between the target body and a plane where a first component in the electronic device is located, and the first component is any device in the electronic device except the display screen.
[0171] Among them, the distance between the target body and the first plane where the display screen is located is smaller than the distance between the first component in the electronic device and the plane where the first component is located. The first component is any device in the electronic device except the display screen, that is, the target body is any object located on the same side as the display screen.
[0172] Considering that when the display screen of the electronic device is in an obstructed state, the distance between the display screen and the obstruction is smaller, and when the display screen of the electronic device is not in an obstructed state, the distance between the display screen and the obstruction is larger.
[0173] In this way, the electronic device can obtain the first distance between any object located on the same side of the display screen and the display screen, so as to determine that the display screen is not in an obstructed state when the first distance value is greater than or equal to the first threshold, or determine that the display screen is in an obstructed state when the first distance value is less than the first threshold.
[0174] In some embodiments, the electronic device includes a distance sensor, and the distance sensor and the display screen are located on the same side of the electronic device. Obtaining a first distance value between the display screen and the target includes: obtaining the first distance value between the display screen and the target through the distance sensor.
[0175] Specifically, the distance sensor and the display screen are located on the same side. When the first voice data meets the first condition, the electronic device can control the distance sensor to send an ultrasonic signal. When the ultrasonic signal encounters the target object, it can be reflected back to form a response signal. The electronic device can determine the first distance value between the display screen and the target object through the response signal.
[0176] For example, when the electronic device is placed on a table and the display screen is facing the table, the electronic device controls the distance sensor to send an ultrasonic signal. When the ultrasonic signal encounters the table, it can be reflected back to form a response signal. The electronic device can determine the first distance value between the display screen and the table through the response signal.
[0177] S404: Determine whether the first distance value is greater than or equal to a first threshold.
[0178] The first threshold is preset; the first threshold is used to avoid the situation where the display screen of the electronic device is in an obstructed state.
[0179] For example, the first threshold is 2 cm. If the first distance value is less than the first threshold, the electronic device can determine that the display screen is in an obstructed state. If the first distance value is greater than or equal to 2 cm, the electronic device can determine that the display screen is not in an obstructed state.
[0180] For example, the first threshold is 2cm. When the electronic device is placed on a table and the display screen is facing the table, the electronic device controls the distance sensor to send an ultrasonic signal. When the ultrasonic signal encounters the table, it can be reflected back to form a response signal. The electronic device can determine through the response signal that the first distance value between the display screen and the table is 0.1cm, and 0.1 is less than 2, that is, the first distance value is less than the first threshold value. The electronic device determines that the display screen is in an obstructed state.
[0181] If the first distance value is greater than or equal to the first threshold, the electronic device determines that the display screen is not in an obstructed state, and the electronic device may execute S405.
[0182] If the first distance value is less than the first threshold, the electronic device determines that the display screen is in a blocked state, indicating that the user may not have awakened the voice assistant application by breath awakening. The electronic device does not execute any steps, and the process of awakening the voice assistant by breath ends.
[0183] S405: Determine whether first motion data is detected.
[0184] Based on S401, it can be known that the electronic device can detect the first motion data in real time. If the electronic device detects the first motion data, the electronic device can execute S406; if the electronic device does not detect the first motion data, it means that the user may not have awakened the voice assistant application by breath awakening, and the electronic device does not execute any steps, and the process of awakening the voice assistant by breath ends.
[0185] S406: Determine whether the electronic device is in a shaking state.
[0186] In the scenario where the electronic device is in a shaking state due to alarms, incoming calls, etc., the motion data generated by the electronic device is misidentified by the gesture model, and the current sound is close to the sound of the user's breath when speaking, which will cause the voice assistant application to be woken up by mistake. The triggering of the breath-induced false wake-up will invoke some behaviors of the voice assistant application, increasing unnecessary power consumption.
[0187] In some scenarios, for example, when the alarm rings and vibrates, the voice assistant application may be woken up by mistake, which may cause serious events such as the alarm stopping ringing; in some scenarios, the voice assistant application is woken up by mistake, causing the voice assistant application to mistakenly receive and broadcast the sound, which also does not meet user expectations.
[0188] According to some user feedback, the following scenarios may occur: at 7:50 a.m., the phone is in the pocket, the alarm rings and vibrates, and suddenly the voice assistant application wakes up; the alarm rings and vibrates, the voice assistant application wakes up, and the phone turns off the alarm.
[0189] In the scenario where the electronic device wakes up the voice assistant application by breathing, the user is required to pick up the electronic device and speak to wake up the voice assistant application through the first action. Considering that the electronic device may be in a shaking state due to alarms, incoming calls, etc., the electronic device may recognize the motion data generated by alarms, incoming calls, etc. as the motion data generated by the user picking up the electronic device through the first action, causing the voice assistant application to be woken up by mistake.
[0190] Based on this, if the electronic device detects the first motion data, the electronic device can further determine whether the electronic device is in a shaking state, thereby avoiding the situation where the voice assistant application is mistakenly awakened due to alarms, incoming calls, etc.
[0191] If the electronic device is not in a shaking state, the electronic device can execute S407; if the electronic device is in a shaking state, it means that the user may not have awakened the voice assistant application by breath awakening, and the electronic device does not execute any steps, and the process of awakening the voice assistant by breath ends.
[0192] In some embodiments, determining whether the electronic device is in a shaking state includes:
[0193] The first motion data is input into the vibration model, and a first similarity value is output. The vibration model is used to predict the similarity between the first motion data and the motion data corresponding to the first motor when the first motor is in a working state. The vibration model is trained by the motion data corresponding to the first motor when the first motor is in a working state. The first motor and the second motor in the electronic device are motors of the same type. If the first similarity value is greater than or equal to the second threshold value, it is determined that the electronic device is in a shaking state.
[0194] The second threshold is preset and can be adjusted according to actual conditions. The second threshold is used to avoid the situation where the electronic device is in a shaking state and the voice assistant application is mistakenly awakened. For example, the second threshold can be set to 0.8.
[0195] The vibration model is trained by using the motion data corresponding to the first motor being in a working state, so the vibration model can recognize the motion data corresponding to the first action.
[0196] When the first motion data is detected, the first motion data can be input into the vibration model. If the first similarity value output by the vibration model is greater than or equal to the second threshold, it means that the first motion data is highly similar to the motion data corresponding to the first motor in the working state, and the first motion data is the motion data generated when the electronic device is in a shaking state, then the electronic device can determine that the first motion data meets the second condition.
[0197] In some embodiments, the first motion data includes acceleration data collected by an acceleration sensor in an electronic device, the acceleration data includes first acceleration data and second acceleration data within a first time period, the first time period is a time period between a current moment and a first moment, the first moment is a moment earlier than the current moment, the first acceleration data includes data in the acceleration data that is greater than an acceleration threshold, the acceleration threshold includes 0, the first motion data is input into a vibration model, and a first similarity value is output, including: inputting the first acceleration data in the first motion data into the vibration model to obtain a first similarity value.
[0198] The acceleration data includes first acceleration data and second acceleration data, the first acceleration data includes data in the acceleration data that is greater than an acceleration threshold, and the second acceleration data includes data in the acceleration data that is equal to the acceleration threshold.
[0199] Specifically, the acceleration threshold is 0, the first acceleration data includes data that is not 0 in the acceleration data, and the second acceleration data includes data that is 0 in the acceleration data.
[0200] The first acceleration data may also be referred to as vibration segment data, and the second acceleration data may also be referred to as idle segment data.
[0201] For example, the current time is 10:00:00 and the preset duration is 2s, then the first time is 9:59:58, the first time period is the time period between 9:59:58 and 10:00:00, and the acceleration data includes the acceleration data between 9:59:58 and 10:00:00.
[0202] In some embodiments, the acceleration data includes three-axis acceleration data (a matrix formed by three-axis acceleration data), such as Figure 1 As shown, the three axes include the X-axis, the Y-axis and the Z-axis, and the three-axis acceleration data include the acceleration values corresponding to the three-axis directions, namely the acceleration value in the X-axis direction, the acceleration value in the Y-axis direction and the acceleration value in the Z-axis direction.
[0203] The acceleration data in the first time period includes acceleration values in the X-axis direction, the Y-axis direction, and the Z-axis direction in a time series. The acceleration data in the first time period can reflect the movement state of the electronic device from a stationary state to a first action.
[0204] The second acceleration data is data when the electronic device is in a stationary state, and the second acceleration data includes data that is 0 in the acceleration data. The first acceleration data is data when the electronic device performs a first action, and the first acceleration data includes data that is not 0 in the acceleration data. The second acceleration data is irrelevant to determining whether the electronic device is in a shaking state and may not be involved in the calculation.
[0205] In other embodiments, the three-axis acceleration data includes acceleration values corresponding to the three-axis directions respectively. The three-axis acceleration data is filtered using a low-pass filter to obtain a gravity coefficient, a linear acceleration coefficient, a rotational linear acceleration coefficient, and a normalized value of the linear acceleration, the absolute value, variance, mean, first-order variance, and mean difference of the three-axis signals, for a total of 37 dimensions; there are 40 input parameters in total.
[0206] Specifically, the three-axis acceleration data include the acceleration values corresponding to the three-axis directions, the gravity coefficients corresponding to the three-axis directions, the linear acceleration coefficients corresponding to the three-axis directions, the rotational linear acceleration coefficients corresponding to the three-axis directions, the normalized values of the linear acceleration coefficients corresponding to the three-axis directions, the absolute values of the acceleration values corresponding to the three-axis directions, the variances of the acceleration values corresponding to the three-axis directions, the average values of the acceleration values corresponding to the three-axis directions, the first-order variances of the acceleration values corresponding to the three-axis directions, and the differences of the means of the acceleration values corresponding to the three-axis directions.
[0207] The three-axis acceleration data may also include the sum of the acceleration values corresponding to the three-axis directions, the root mean square of the acceleration values corresponding to the three-axis directions, and any one of the maximum value of the acceleration values corresponding to the three-axis directions, the minimum value of the acceleration values corresponding to the three-axis directions, and the difference between the maximum value and the minimum value of the acceleration values corresponding to the three-axis directions.
[0208] Take the preset duration of 2s, where the acceleration data includes the acceleration values in the X-axis direction, the Y-axis direction, and the Z-axis direction within 2s as an example:
[0209] Fig.11 The acceleration data shown in (a) is the acceleration data collected by the acceleration sensor when the user wakes up the voice assistant application with normal breath and the user lifts the electronic device through the first action.
[0210] If you will Fig.11 The acceleration data shown in (a) is input into the vibration model, the confidence (similarity value) output by the vibration model is 0.21, and the second threshold corresponding to the vibration model is 0.7. The electronic device can determine that the electronic device is not in a shaking state.
[0211] If you will Fig.11 The acceleration data shown in (a) is input into the gesture model, the confidence of the gesture model output is 0.96, and the third threshold corresponding to the gesture model is 0.7. The electronic device can determine that the acceleration data (first motion data) meets the second condition.
[0212] If you will Fig.11In (b), the acceleration data is input into the vibration model, the confidence of the vibration model output is 0.63, and the second threshold corresponding to the vibration model is 0.7. The electronic device can determine that the electronic device is not in a shaking state.
[0213] If you will Fig.11 In (b), the acceleration data is input into the gesture model, the confidence of the gesture model output is 0.71, and the third threshold corresponding to the gesture model is 0.7. The electronic device can determine that the acceleration data (first motion data) meets the second condition.
[0214] Based on this, electronic devices can be based on Fig.11 (a) and Fig.11 The acceleration data (first motion data) in (b) satisfies the second condition, the electronic device is not in a shaking state, and it is determined that the current scene is a scene of waking up the voice assistant with normal breath. The electronic device can start the voice assistant application.
[0215] Fig.11 The acceleration data shown in (c) is the acceleration data collected by the acceleration sensor when the second motor of the electronic device is in a working state (the electronic device is in a vibration mode).
[0216] If you will Fig.11 The acceleration data shown in (c) is input into the vibration model. The confidence of the vibration model output is 0.98. The second threshold corresponding to the vibration model is 0.7. The electronic device can determine that the electronic device is in a shaking state. The current scenario is a scenario in which the voice assistant is mistakenly awakened by breath.
[0217] If there is no vibration model, directly Fig.11 The acceleration data shown in (c) is input into the gesture model, the confidence of the gesture model output is 0.74, and the third threshold corresponding to the gesture model is 0.7. The electronic device can determine that the acceleration data (first motion data) meets the second condition.
[0218] Based on this, through the vibration model, the electronic device can identify whether the current scene is a scene in which the voice assistant application is mistakenly awakened by breath, and can avoid directly inputting the first motion data when the electronic device is in a shaking state into the gesture model, causing the gesture model to recognize that the acceleration data meets the second condition, thereby avoiding the voice assistant application being mistakenly awakened due to the shaking of the electronic device.
[0219] In some embodiments, the training process of the vibration model includes:
[0220] Step 4061: Obtain N sample motion data and sample similarity values corresponding to the N sample motion data, where each of the N sample motion data is collected when the first motor is in a working state, and N is a positive integer greater than or equal to 2.
[0221] Step 4062: Input the N sample motion data and the sample similarity values corresponding to the N sample motion data into the classification training model, train the classification training model, and obtain a vibration model.
[0222] The vibration model is trained by the motion data corresponding to the first motor in the working state, so the vibration model can recognize the motion data corresponding to the first action. The sample similarity values corresponding to the N sample motion data can be the same (1 or close to 1) or different (similar, 1 or close to 1).
[0223] Specifically, the N sample motion data include motion data when the first motor is located on various objects and is in a working state. Each of the N sample motion data includes three-axis acceleration data. During training, each positive data unit in the three-axis acceleration data can be marked with a confidence of [0.9, 1], that is, a sample similarity value of [0.9, 1]; and the negative data unit can be marked with a confidence of [0, 0.3], that is, a sample similarity value of [0, 0.3]. The positive data unit can be understood as data that is not 0 in the three-axis acceleration data, and the negative data unit can be understood as data that is 0 in the three-axis acceleration data.
[0224] like Fig.10 As shown in Figure 1, the classification training model includes a feature layer, a convolutional neural network (CNN) layer, and a fully connected layer.
[0225] The feature layer is usually located at the front end of the network and is used for feature extraction. The data received in each batch in the feature layer is 40 frames. For details, please refer to the 40 types of data mentioned above.
[0226] The CNN layer is used for feature learning. The CNN layer includes 3 layers. The first layer includes 1 1-dimensional convolution layer, the second layer includes 1 1-dimensional convolution layer, and the third layer includes 4 1-dimensional convolution layers.
[0227] The fully connected layer (FCL) is used for feature integration. The fully connected layer consists of 2 layers. The first layer includes a fully connected layer fc1, and the second layer includes a fully connected layer fc2.
[0228] S407: Whether the first motion data satisfies the second condition.
[0229] Among them, S407 and Figure 7 The implementation of S303 in the illustrated embodiment is similar and will not be described in detail here.
[0230] If the electronic device meets the second condition, the electronic device can execute S408; if the electronic device does not meet the second condition, it means that the user may not have awakened the voice assistant application by breath awakening, and the electronic device does not execute any steps, and the process of awakening the voice assistant by breath ends.
[0231] S408. Display a first interface, where the first interface is a voice interaction interface of the voice assistant application.
[0232] Among them, S408 and Figure 7 The implementation of S304 in the illustrated embodiment is similar and will not be described in detail here.
[0233] In the present application, the electronic device can determine whether the display screen is in an obstructed state by determining a first distance value between the display screen and the target object. If the first distance value is less than a first threshold value, it can be determined that the display screen is in an obstructed state, and the current scene is not a scene for waking up the voice assistant by breath, that is, it is a false awakening scene, and the process of waking up the voice assistant by breath ends.
[0234] In addition, if the first distance value is greater than or equal to the first threshold, it can be determined that the display screen is not in an obstructed state. If the electronic device detects the first motion data, the electronic device can further determine whether the electronic device is in a shaking state to determine whether the current scene is a false awakening scene. If the electronic device is in a shaking state, it means that the current scene is not a scene for waking up the voice assistant by breath, that is, it is a false awakening scene, and the process of waking up the voice assistant by breath ends; if the electronic device is not in a shaking state, the electronic device can further determine whether the current scene is a scene for waking up the voice assistant by breath by whether the first motion data meets the second condition. In this way, the breath awakening accuracy of the voice assistant application can be improved, and the false awakening of the voice assistant application can be avoided.
[0235] Based on the above Fig. 9 The description of the illustrated embodiment takes into account that the electronic device is in a shaking state, which may be because the user shakes the electronic device when picking up the electronic device and speaking to wake up the voice assistant application. Therefore, when the electronic device is in a shaking state, it is also possible to determine whether the second motor of the electronic device is in a working state to eliminate the scenario where the electronic device is mistakenly awakened.
[0236] In addition, when determining that the second motor is not in the working state, it is not only necessary to determine whether the first motion data satisfies the second condition, but also necessary to determine whether the first voice data satisfies the third condition.
[0237] Among them, the execution order of determining whether the first motion data meets the second condition and determining whether the first voice data meets the third condition is not particular, and can be executed sequentially or simultaneously.
[0238] When executing sequentially, the electronic device may first determine whether the first motion data satisfies the second condition, and if so, determine whether the first voice data satisfies the third condition; the electronic device may also first determine whether the first voice data satisfies the third condition, and if so, determine whether the first motion data satisfies the second condition.
[0239] When executed simultaneously, when the first motion data does not meet the second condition, or the first voice data does not meet the third condition, the electronic device can determine that the current scene is a false awakening scene; when the first motion data meets the second condition, and the first voice data meets the third condition, the electronic device can determine that the current scene is a normal breath awakening scene, and can start the voice assistant application.
[0240] The present application does not limit the execution order of determining whether the first motion data satisfies the second condition and determining whether the first voice data satisfies the third condition.
[0241] In the following, it is taken as an example to determine whether the first motion data satisfies the second condition, and if so, to determine whether the first voice data satisfies the third condition. Fig.12 , which introduces in detail the method of waking up the voice assistant of this application.
[0242] See also Fig.12 , Fig.12 A flow chart of a method for waking up a voice assistant provided in one embodiment of the present application is shown.
[0243] like Fig.12 As shown, the method for waking up the voice assistant provided in this application may include:
[0244] S501. During operation of the electronic device, first voice data and first motion data are detected in real time.
[0245] S502: If first voice data is detected, determine whether the first voice data satisfies a first condition.
[0246] S503: Acquire a first distance value between the display screen and a target object, where the target object is any object located on the same side as the display screen.
[0247] S504: Determine whether the first distance value is greater than or equal to a first threshold.
[0248] S505: Determine whether first motion data is detected.
[0249] S506: Determine whether the electronic device is in a shaking state.
[0250] It should be noted that when a user picks up an electronic device and speaks to wake up the voice assistant application, the electronic device may shake. In this case, it is not a scenario of accidentally waking up the voice assistant. Therefore, when it is determined that the electronic device is in a jitter state, it is also necessary to further determine whether the second motor of the electronic device is in a working state to exclude the situation of accidentally waking up the voice assistant application.
[0251] Based on this, if the electronic device is in a jitter state, the electronic device can execute S507; if the electronic device is not in a jitter state, the electronic device can execute S508.
[0252] Among them, S501, S502, S503, S504, S505, and S506 are respectively similar to Fig. 9 S401, S402, S403, S404, S405, and S406 in the illustrated embodiment in implementation, and will not be elaborated here.
[0253] S507, determine whether the second motor is in a working state.
[0254] Generally, the second motor in the electronic device may be in a working state due to situations such as an alarm or an incoming call, thus causing the electronic device to be in a jitter state and resulting in an accidental wake-up of the voice assistant application.
[0255] Based on this, if the second motor is not in a working state, the electronic device can execute S508; if the second motor is in a working state, which is a scenario of accidentally waking up the voice assistant application, the electronic device can not execute any steps, and the process of waking up the voice assistant by breath ends.
[0256] S508, whether the first motion data meets the second condition.
[0257] Among them, S508 is similar to Fig. 9 S407 in the illustrated embodiment in implementation, and will not be elaborated here.
[0258] S509, whether the first voice data meets the third condition.
[0259] Among them, the third condition includes two cases.
[0260] In some embodiments, the third condition includes that the third similarity value between the first voiceprint feature corresponding to the first voice data and the voiceprint feature of the target sound pre-stored in the electronic device is greater than or equal to the fifth threshold.
[0261] Among them, the target sound is the sound of the user recorded by the electronic device when the breath wake-up function of the electronic device is turned on.
[0262] The third similarity value between the first voiceprint feature corresponding to the first voice data and the voiceprint feature of the target sound is greater than or equal to the fifth threshold, which indicates that the sound corresponding to the first voice data and the target sound are the voices of the same user, and the first voice data is voice data spoken by the user of the electronic device.
[0263] In other embodiments, the third condition includes that a third similarity value between a first voiceprint feature corresponding to the first voice data and a voiceprint feature of a target sound pre-stored in the electronic device is greater than or equal to a fifth threshold, and that a first sound source angle is within a first range, and the first sound source angle is used to indicate the angle of the direction of the sound source of the first voice data relative to the electronic device, and the first sound source angle is determined based on the sound pressure amplitude difference.
[0264] Among them, the first sound source angle can be understood as the angle between the user's mouth and the microphone. The first range is the range of the angle between the user's mouth and the microphone required when the voice assistant application is awakened by breath. The first sound source angle is within the first range, which can indicate that the angle between the user's mouth and the microphone is within the first range, which meets the angle between the user's mouth and the microphone required by the breath to awaken the voice assistant application.
[0265] S510. Display a first interface, where the first interface is a voice interaction interface of a voice assistant application.
[0266] Among them, S510 and Fig. 9 The implementation of S408 in the illustrated embodiment is similar and will not be described in detail here.
[0267] In the present application, it is taken into account that when the electronic device is in a shaking state, it may be because the user shakes the electronic device when picking up the electronic device and speaking to wake up the voice assistant application. Therefore, when the electronic device is in a shaking state, it can also be determined whether the second motor of the electronic device is in a working state. If the second motor is in a working state, it is a false awakening scenario, and the process of awakening the voice assistant by breath ends. If the second motor is not in a working state, it may not be a false awakening scenario, and the electronic device can continue to determine whether the user is waking up the voice assistant application by breath awakening through the first motion data and the first voice data. Thus, by judging whether the second motor is in a working state, the false awakening of the voice assistant application can be ruled out, and the breath awakening accuracy of the voice assistant application can be further improved, and the false awakening of the voice assistant application can be avoided. At the same time, events such as the alarm clock stopping ringing due to false awakening can be avoided.
[0268] In addition, after determining that the second motor is not in a working state, the electronic device can further determine whether the current scene is a scene for waking up the voice assistant application by breath awakening by determining whether the first motion data meets the second condition and whether the first voice data meets the third condition. In this way, the accuracy of the breath awakening of the voice assistant application can be improved, and the situation of false awakening of the voice assistant application can be avoided.
[0269] Based on the above description, in a specific embodiment, the electronic device includes an audio digital signal processor and an application processor, the audio digital signal processor runs a first module, and the application processor runs a second module.
[0270] The audio digital signal processor obtains the first voice data through the first module, determines whether the first voice data meets the first condition, and determines whether the display screen of the electronic device is in an obstructed state. The application processor obtains the first motion data through the second module, determines whether the electronic device is in a shaking state, and determines whether the first motion data meets the second condition and whether the first voice data meets the third condition.
[0271] Next, combine Fig.13 , which introduces in detail the method of waking up the voice assistant of this application.
[0272] See also Fig.13 , Fig.13 A flow chart of a method for waking up a voice assistant provided in one embodiment of the present application is shown.
[0273] like Fig.13 As shown, the method for waking up the voice assistant provided in this application may include:
[0274] S11. The first module detects first voice data in real time during the operation of the electronic device.
[0275] S12. The second module detects the first motion data in real time during the operation of the electronic device.
[0276] Among them, S11 and S12 are Fig.12 The implementation of S501 in the illustrated embodiment is similar and will not be described in detail here.
[0277] S13: If the first module detects the first voice data, it determines whether the first voice data satisfies the first condition.
[0278] S14. The first module obtains a first distance value between the display screen and a target object, where the target object is any object located on the same side as the display screen.
[0279] S15. The first module determines whether the first distance value is greater than or equal to a first threshold.
[0280] Among them, S13, S14 and S15 are respectively Fig.12 The implementation methods of S502, S503 and S504 in the illustrated embodiment are similar and will not be described in detail here.
[0281] S16. The first module sends the first voice data to the second module.
[0282] If the first distance value is greater than or equal to the first threshold, it means that the electronic device is not in an obstruction state, and the first module can send the first voice data to the second module, so that the second module can verify the first voice data again through the third condition.
[0283] After receiving the first voice data, the second module may perform front-end processing, which is usually filtering processing.
[0284] S17. The second module determines whether the first motion data is detected.
[0285] S18. The second module determines whether the electronic device is in a shaking state.
[0286] S19. The second module determines whether the second motor is in operation.
[0287] S20. The second module determines whether the first motion data satisfies a second condition.
[0288] S21. The second module determines whether the first voice data meets a third condition.
[0289] S22. The second module displays a first interface, which is a voice interaction interface of the voice assistant application.
[0290] Among them, S17, S18, S19, S20, S21 and S22 are respectively Fig.12 The implementation methods of S505, S506, S507, S508, S509 and S510 in the illustrated embodiment are similar and will not be described in detail here.
[0291] Exemplarily, the present application provides a device for waking up a voice assistant.
[0292] See also Fig.14 , Fig.14 A schematic block diagram of a device for waking up a voice assistant provided in an embodiment of the present application is shown.
[0293] like Fig.14As shown, the device 600 for waking up the voice assistant can exist independently or be integrated into other devices to achieve mutual communication with the above-mentioned electronic devices, and is used to implement the operations corresponding to the electronic devices in any of the above-mentioned method embodiments. The device 600 for waking up the voice assistant provided in the embodiment of the present application includes a first module 601 and a second module 602.
[0294] The first module 601 is used to determine whether the display screen of the electronic device is in an obstructed state if first voice data is detected during the operation of the electronic device and the first voice data satisfies a first condition, wherein the first condition includes that the sound corresponding to the first voice data conforms to the sound regularity of the sound generated by the user's speech.
[0295] The second module 602 is used to determine whether first motion data is detected if the display screen is not in an obstruction state, and the first motion data is greater than a preset threshold.
[0296] The second module 602 is also used to determine whether the first motion data satisfies a second condition if the first motion data is detected, the second condition including that the first motion data is motion data generated when the electronic device performs a first action, and the first action is an action of the mobile electronic device related to waking up the voice assistant application.
[0297] The second module 602 is also used to display a first interface if the first motion data meets a second condition, where the first interface is a voice interaction interface of the voice assistant application.
[0298] In some embodiments, the first module 601 is specifically used to:
[0299] A first distance value between a display screen and a target body is obtained, wherein the distance between the target body and a first plane where the display screen is located is less than the distance between the target body and a plane where a first component in the electronic device is located, and the first component is any device in the electronic device except the display screen; determining whether the first distance value is greater than or equal to a first threshold value; if the first distance value is greater than or equal to the first threshold value, determining that the display screen is not in an obstructed state; if the first distance value is less than the first threshold value, determining that the display screen is in an obstructed state.
[0300] In some embodiments, the first module 601 is specifically used to:
[0301] A first distance value between the display screen and the target object is obtained through the distance sensor.
[0302] In some embodiments, the second module 602 is specifically configured to:
[0303] If the first motion data is detected, determining whether the electronic device is in a shaking state, the shaking state is used to indicate that the electronic device moves in accordance with the motion law of the second motor in the electronic device;
[0304] If the electronic device is not in a shaking state, it is determined whether the first motion data meets a second condition.
[0305] In some embodiments, the second module 602 is specifically configured to:
[0306] The first motion data is input into the vibration model, and a first similarity value is output. The vibration model is used to predict the similarity between the first motion data and the motion data corresponding to the first motor when the first motor is in a working state. The vibration model is trained by the motion data corresponding to the first motor when the first motor is in a working state. The first motor and the second motor in the electronic device are motors of the same type. If the first similarity value is greater than or equal to the second threshold value, it is determined that the electronic device is in a shaking state.
[0307] In some embodiments, the first motion data includes acceleration data collected by an acceleration sensor in an electronic device, the acceleration data includes first acceleration data and second acceleration data within a first time period, the first time period is a time period between a current moment and a first moment, the first moment is a moment earlier than the current moment, the first acceleration data includes data in the acceleration data that is greater than an acceleration threshold, the acceleration threshold includes 0, and the second module 602 is specifically used to:
[0308] The first acceleration data is input into the vibration model to obtain a first similarity value.
[0309] In some embodiments, the second module 602 is specifically configured to:
[0310] If the electronic device is in a shaking state, determine whether the second motor in the electronic device is in a working state; if the second motor is not in a working state, determine whether the first motion data meets the second condition.
[0311] In some embodiments, the second module 602 may further include a generation module, and the generation module is specifically used to:
[0312] Obtain N sample motion data and sample similarity values corresponding to the N sample motion data, where each of the N sample motion data is collected when the first motor is in a working state, and N is a positive integer greater than or equal to 2; input the N sample motion data and the sample similarity values corresponding to the N sample motion data into a classification training model, train the classification training model, and obtain a vibration model.
[0313] In some embodiments, the second module 602 is specifically configured to:
[0314] The first motion data is input into the gesture model to obtain a second similarity value. The gesture model is used to predict the similarity between the first motion data and the motion data corresponding to the first action. The gesture model is trained with the motion data corresponding to the first action. If the second similarity value is greater than or equal to the third threshold, it is determined that the first motion data meets the second condition.
[0315] In some embodiments, the electronic device includes a first microphone and a second microphone, the first voice data includes first sub-voice data collected by the first microphone and second sub-voice data collected by the second microphone, the first microphone and the second microphone are located on different sides of the electronic device, if the first voice data is detected and the first voice data meets the first condition, the first module 601 is specifically used to:
[0316] If the first voice data is detected, the sound pressure amplitude difference between the first sub-voice data and the second sub-voice data is determined; if the sound pressure amplitude difference is greater than or equal to the fourth threshold, the first voice data satisfies the first condition, and it is determined whether the display screen is in an obstructed state.
[0317] In some embodiments, the second module 602 is specifically configured to:
[0318] If the first motion data is detected, determine whether the first motion data satisfies the second condition, and determine whether the first voice data satisfies the third condition, the third condition including that the third similarity value between the first voiceprint feature corresponding to the first voice data and the voiceprint feature of the target sound pre-stored in the electronic device is greater than or equal to a fifth threshold; if the first motion data satisfies the second condition and the first voice data satisfies the third condition, display the first interface.
[0319] In some embodiments, the third condition also includes a first sound source angle being within a first range, the first sound source angle being used to represent the angle of the direction of the sound source of the first voice data relative to the electronic device, and the first sound source angle being determined based on the sound pressure amplitude difference.
[0320] Exemplarily, the present application provides an electronic device, comprising a processor; when the processor executes computer code or instructions in the memory, the electronic device executes the method of waking up the voice assistant in the previous embodiment.
[0321] Exemplarily, the present application provides an electronic device, comprising: a memory and a processor; the memory is coupled to the processor, and the memory is used to store program codes or instructions; the processor is used to call the program codes or instructions in the memory so that the electronic device executes the method of waking up the voice assistant in the previous embodiment.
[0322] Exemplarily, the present application provides a chip system, which is applied to an electronic device including a memory, a display screen and a sensor; the chip system includes: one or more interface circuits and one or more processors; the interface circuit and the processor are interconnected through lines; the interface circuit is used to receive signals from the memory and send signals to the processor, and the signals include computer codes or instructions stored in the memory; when the processor executes the computer code or instructions, the electronic device executes the method of waking up the voice assistant in the previous embodiment.
[0323] Exemplarily, the present application provides a computer-readable storage medium, in which codes or instructions are stored. When the codes or instructions are executed on an electronic device, the electronic device implements the method of waking up a voice assistant in the foregoing embodiment when executing.
[0324] Exemplarily, the present application provides a computer program product, which, when executed on a computer, enables an electronic device to implement the method for waking up a voice assistant in the foregoing embodiments.
[0325] In the above embodiments, all or part of the functions can be implemented by software, hardware, or a combination of software and hardware. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer codes or instructions. When the computer program code or instructions are loaded and executed on the computer, the process or function according to the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer code or instructions can be stored in a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be accessed by the computer or a data storage device such as a server or a data center that includes one or more available media integration. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.
[0326] A person skilled in the art can understand that to implement all or part of the processes in the above-mentioned embodiments, the processes can be completed by a computer program to instruct the relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned method embodiments. The aforementioned storage medium includes: a read-only memory (ROM) or a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
Claims
1. A method for waking up a voice assistant, characterized in that: Applied to an electronic device, the electronic device includes a voice assistant application, and the method includes: During the operation of the electronic device, if first voice data is detected and the first voice data satisfies a first condition, determining whether the display screen of the electronic device is in a blocked state, the first condition including that a sound corresponding to the first voice data conforms to a sound regularity of a sound generated by a user speaking; Acquire a first distance value between the display screen and a target object, wherein the distance between the target object and a first plane where the display screen is located is smaller than the distance between the target object and a plane where a first component in the electronic device is located, and the first component is any device in the electronic device except the display screen; If the first distance value is greater than or equal to a first threshold, it is determined that the display screen is not in an obstruction state; determining whether first motion data is detected, the first motion data being greater than a preset threshold; If the first motion data is detected, determining whether the first motion data satisfies a second condition, wherein the second condition includes that the first motion data is motion data generated when the electronic device performs a first action, and the first action is an action of moving the electronic device related to waking up the voice assistant application; If the first motion data meets the second condition, a first interface is displayed, where the first interface is a voice interaction interface of the voice assistant application.
2. The method according to claim 1, characterized in that The method further comprises: If the first distance value is less than the first threshold, it is determined that the display screen is in the blocked state.
3. The method according to claim 2, characterized in that The electronic device includes a distance sensor, and obtaining a first distance value between the display screen and the target object includes: The first distance value between the display screen and the target object is acquired through the distance sensor.
4. The method according to claim 1, characterized in that: If the first motion data is detected, determining whether the first motion data satisfies a second condition includes: If the first motion data is detected, determining whether the electronic device is in a shaking state, the shaking state being used to indicate that the electronic device moves in accordance with a motion law of a second motor in the electronic device; If the electronic device is not in the shaking state, it is determined whether the first motion data meets the second condition.
5. The method according to claim 4, characterized in that The determining whether the electronic device is in a shaking state comprises: inputting the first motion data into a vibration model and outputting a first similarity value, wherein the vibration model is used to predict the similarity between the first motion data and motion data corresponding to when the first motor is in a working state, the vibration model is trained by the motion data corresponding to when the first motor is in a working state, and the first motor and the second motor in the electronic device are motors of the same type; If the first similarity value is greater than or equal to a second threshold, it is determined that the electronic device is in the shaking state.
6. The method according to claim 5, characterized in that The first motion data includes acceleration data collected by an acceleration sensor in the electronic device, the acceleration data includes first acceleration data and second acceleration data within a first time period, the first time period is a time period between a current moment and a first moment, the first moment is a moment earlier than the current moment, the first acceleration data includes data in the acceleration data that is greater than an acceleration threshold, the acceleration threshold includes 0, and the inputting of the first motion data into a vibration model and outputting the first similarity value includes: The first acceleration data is input into the vibration model to obtain the first similarity value.
7. The method according to claim 4, characterized in that The method further comprises: If the electronic device is in the shaking state, determining whether the second motor in the electronic device is in a working state; If the second motor is not in the working state, it is determined whether the first motion data meets the second condition.
8. The method according to claim 5 or 6, characterized in that: The generation process of the vibration model includes: Acquire N sample motion data and sample similarity values corresponding to the N sample motion data, wherein each sample motion data of the N sample motion data is collected when the first motor is in a working state, and N is a positive integer greater than or equal to 2; The N sample motion data and the sample similarity values corresponding to the N sample motion data are input into a classification training model, and the classification training model is trained to obtain the vibration model.
9. The method according to any one of claims 1 to 7, characterized in that Determining whether the first motion data satisfies a second condition includes: inputting the first motion data into a gesture model to obtain a second similarity value, the gesture model being used to predict the similarity between the first motion data and motion data corresponding to the first action, the gesture model being trained using the motion data corresponding to the first action; If the second similarity value is greater than or equal to a third threshold, it is determined that the first motion data meets the second condition.
10. The method according to any one of claims 1 to 7, characterized in that The electronic device includes a first microphone and a second microphone, the first voice data includes first sub-voice data collected by the first microphone and second sub-voice data collected by the second microphone, the first microphone and the second microphone are located at different sides of the electronic device, and if the first voice data is detected and the first voice data satisfies a first condition, determining whether the display screen of the electronic device is in an obstructed state includes: If the first voice data is detected, determining a difference in sound pressure amplitude between the first sub-voice data and the second sub-voice data; If the sound pressure amplitude difference is greater than or equal to a fourth threshold value, the first voice data satisfies the first condition, and it is determined whether the display screen is in the blocked state.
11. The method according to claim 10, characterized in that If the first motion data is detected, determining whether the first motion data satisfies a second condition includes: If the first motion data is detected, determining whether the first motion data satisfies a second condition, and determining whether the first voice data satisfies a third condition, wherein the third condition includes that a third similarity value between a first voiceprint feature corresponding to the first voice data and a voiceprint feature of a target sound pre-stored in the electronic device is greater than or equal to a fifth threshold; If the first motion data satisfies the second condition, displaying the first interface includes: If the first motion data satisfies the second condition and the first voice data satisfies the third condition, the first interface is displayed.
12. The method according to claim 11, characterized in that The third condition also includes a first sound source angle being within a first range, wherein the first sound source angle is used to represent an angle of a direction of a sound source of the first voice data relative to the electronic device, and the first sound source angle is determined based on the sound pressure amplitude difference.
13. A device for waking up a voice assistant, characterized in that: The device for waking up a voice assistant includes a module for executing the method as claimed in any one of claims 1 to 12.
14. An electronic device, characterized in that: The electronic device comprises: one or more processors, and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code comprises computer instructions, and the one or more processors call the computer instructions so that the electronic device executes the method as described in any one of claims 1 to 12.
15. A chip system, characterized in that: The chip system is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions so that the electronic device executes the method as described in any one of claims 1 to 12.
16. A computer-readable storage medium, characterized in that: The computer-readable storage medium comprises instructions, which, when executed on an electronic device, enable the electronic device to perform the method as claimed in any one of claims 1 to 12.
17. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is run on an electronic device, the electronic device is caused to perform the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Voice interaction method and related electronic equipment
CN115881118A