Method for waking up voice assistant and electronic equipment

CN120915864APending Publication Date: 2025-11-07HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410549297.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-11-07

Smart Images

  • Figure CN120915864A_ABST
    Figure CN120915864A_ABST
Patent Text Reader

Abstract

The invention discloses a method for waking up a voice assistant and electronic equipment, and relates to the technical field of voice interaction. The electronic equipment is provided with a folding screen and comprises a folding shaft, a first microphone, a second microphone and a third microphone. The folding shaft is parallel to the width direction of the electronic equipment; when the electronic equipment is in the folded state, the distance between the first microphone and the second microphone is smaller than the first distance, and the distance between the first microphone and the third microphone is larger than the first distance. And when the electronic equipment is in a non-folded state, if the posture change of the electronic equipment and the sound signals acquired by the first microphone and the second microphone meet a first condition, waking up the voice assistant. And when the electronic equipment is in the folded state, if the posture change of the electronic equipment and the sound signals acquired by the first microphone and the third microphone meet a second condition, waking up the voice assistant. More microphones are arranged for the electronic equipment with the folding screen, and when the electronic equipment is in the folding state, the electronic equipment responds to a user breath awakening mode to awaken a voice assistant.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of voice interaction, and in particular to a method for waking up a voice assistant and an electronic device. BACKGROUND

[0002] In the use process of an electronic device, a user can wake up a voice assistant in a foldable screen device through breath awakening, such as YOYO of Honor.

[0003] However, it is found in actual use that for an electronic device with a foldable screen (hereinafter referred to as a foldable screen device), it is generally difficult to wake up a voice assistant in the foldable screen device through breath awakening when the foldable screen device is in a folded state, and the voice assistant cannot be quickly woken up, and the human-computer interaction efficiency is low. SUMMARY

[0004] Embodiments of the present application provide a method for waking up a voice assistant and an electronic device, and more microphones are provided for a foldable screen device, so that the foldable screen device can still respond to the breath awakening of a user when the foldable screen device is in a folded state, thereby conveniently waking up a voice assistant.

[0005] To achieve the above object, embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, a method for waking up a voice assistant is provided, applied to an electronic device with a foldable screen, the electronic device comprising a folding shaft, a first microphone, a second microphone and a third microphone. The folding shaft is parallel to the width direction of the electronic device, that is, the user can fold the electronic device inward or outward based on the width direction of the electronic device, so that the electronic device gradually reaches a folded state. When the electronic device is in a folded state, the distance 2 between the first microphone and the second microphone on the electronic device is less than a first distance, and the distance 1 between the first microphone and the third microphone is greater than the first distance, that is, the positions of the second microphone and the third microphone are different. The first distance can be a distance with any value. For example, the first distance indicates the minimum distance between two microphones constituting breath awakening. For example, one of the conditions for the electronic device to execute the voice assistant is that the minimum distance between the two microphones is 10 cm. When the electronic device is in a folded state, the distance 1 between the first microphone and the second microphone is less than 10 cm, and the electronic device cannot realize the voice assistant based on the first microphone and the second microphone. When the electronic device is in a folded state, the distance between the first microphone and the third microphone is greater than the first distance, so the electronic device can realize the voice assistant based on the first microphone and the third microphone.

[0007] It can be understood that the state of the electronic device also includes a non-folded state, such as an unfolded state and a hovering state. When the electronic device is in the unfolded state or the hovering state, the distance between the first microphone and the second microphone of the electronic device can be greater than the first distance, and then the electronic device can implement the wake-up voice assistant in the unfolded state or the hovering state based on the distance between the first microphone and the second microphone.

[0008] The method for waking up the voice assistant includes: in the case that the electronic device is in the non-folded state, if the posture change of the electronic device and the sound signals collected by the first microphone and the second microphone satisfy the first condition, the voice assistant in the electronic device is woken up. In the case that the electronic device is in the folded state, if the posture change of the electronic device and the sound signals collected by the first microphone and the third microphone satisfy the second condition, the voice assistant is woken up. In the folded state, the distance between the first microphone and the second microphone is less than the first distance, and in this case, if the electronic device is in the non-folded state, the sound signals are collected by the first microphone and the second microphone, the voice assistant can not be woken up. By setting the third microphone, and in the case that the electronic device is in the folded state, the distance between the first microphone and the third microphone is greater than the first distance. In this way, the electronic device can use different microphones to collect sound signals when it is in different states, so as to determine whether to wake up the voice assistant.

[0009] By using the above implementation, the electronic device selects the first microphone and the second microphone to collect sound signals when it is in the non-folded state, so as to determine whether to wake up the voice assistant according to the posture change of the electronic device and the sound signals collected by the first microphone and the second microphone, and selects the first microphone and the third microphone to collect sound signals when it is in the folded state, so as to determine whether to wake up the voice assistant according to the posture change of the electronic device and the sound signals collected by the first microphone and the third microphone, thereby realizing that the electronic device can use microphones at different positions to collect sound signals in different states, so as to determine whether to wake up the voice assistant of the electronic device, and the accuracy of determining whether to wake up the voice assistant of the electronic device can be improved.

[0010] In a possible implementation of the first aspect, the method for waking up the voice assistant further includes that the electronic device uses the first microphone and the third microphone to collect sound signals in the folded state, and uses the first microphone and the second microphone to collect sound signals in the non-folded state. After detecting that the electronic device switches from the folded state to the non-folded state, the electronic device can start the second microphone and close the third microphone, that is, switch to using the first microphone and the second microphone to collect sound signals. After the electronic device switches from the non-folded state to the folded state, the electronic device can start the third microphone and close the second microphone, that is, switch to using the first microphone and the third microphone to collect sound signals.

[0011] With the above implementation, the electronic device can close the microphone corresponding to the pre-switching state and start the microphone corresponding to the post-switching state when detecting the state switching, so as to collect sound signals based on the first microphone and another different microphone in different states, and determine whether the condition of waking up the voice assistant corresponding to the post-switching state is met.

[0012] In a possible implementation of the first aspect, in the case where the electronic device is in the folded state, if the displacement of the electronic device being lifted upward exceeds a first displacement, the posture change of the electronic device includes the displacement of the electronic device being lifted upward; and the energy difference between the first sound signal collected by the first microphone and the second sound signal collected by the third microphone exceeds a first difference value, that is, when the posture change of the electronic device, the first sound signal and the second sound signal meet a second condition, the electronic device wakes up the voice assistant. The first displacement can indicate whether the posture change of the current electronic device meets the condition of waking up the voice assistant, and the first difference value can indicate whether the sound signal collected by the microphone meets the condition of waking up the voice assistant. For example, the first displacement is 5 cm, and the first difference value is 10. In the case where the electronic device is in the folded state, if the electronic device is lifted upward by 10 cm or more, and the energy difference between the first sound signal and the second sound signal exceeds 10, it is determined that the condition of waking up the voice assistant, that is, the second condition, is met, and the electronic device can wake up the voice assistant.

[0013] With the above implementation, the electronic device can wake up the voice assistant only when the electronic device is lifted upward by a certain distance (such as exceeding the first displacement), and the energy difference between the first sound signal collected by the first microphone and the second sound signal collected by the third microphone is large enough (such as exceeding the first difference value). When the above second condition is not met, the electronic device does not wake up the voice assistant, which can reduce the number of false wake-up of the voice assistant and avoid unnecessary energy consumption.

[0014] In a possible implementation manner of the first aspect, the electronic device comprises a screen management service, an audio service and an audio driver. After detecting that the electronic device switches from the unfolded state to the folded state, the screen management service can generate the folded state information and send it to the audio service to instruct the audio service that the state of the electronic device changes to the folded state, and the audio service can perform the microphone switching operation. When the audio service receives the folded state information, the audio service sends the microphone switching instruction to the audio driver to instruct the audio driver to use the first microphone and the third microphone to collect the sound signal in the folded state. The audio driver generates the first instruction and the second instruction based on the microphone switching instruction, and sends the first instruction to the second microphone and the second instruction to the third microphone. In this way, in response to receiving the first instruction, the second microphone is started, and in response to receiving the second instruction, the third microphone is started, thereby realizing the microphone switching operation from the unfolded state to the folded state.

[0015] It can be understood that the screen management service can send the unfolded state information to the audio service after detecting that the electronic device switches from the folded state to the unfolded state. For example, the unfolded state includes the unfolded state and the hovering state, and the unfolded state information includes the unfolded state information and the hovering state information. In this way, the audio service can instruct the audio driver to switch to use the first microphone and the third microphone to collect the sound signal in the unfolded state when the audio service receives the unfolded state information.

[0016] With the above implementation manner, the electronic device can determine the state of the electronic device based on the screen management service, and when detecting the state switching, generate the corresponding state information to instruct the audio service and the audio driver to switch the microphone, so as to realize the use of different microphones to collect the sound signal in different states.

[0017] In a possible implementation manner of the first aspect, the electronic device further comprises an ADSP (analog digital signal processor) and a voice assistant service. After detecting that the state of the electronic device switches from the unfolded state to the folded state, the screen management service can further send the folded state information to the ADSP. The ADSP can obtain the folded state parameter matched with the folded state information and send it to the voice assistant service. In this way, the voice assistant service can identify whether the posture change of the electronic device, the first sound signal and the second sound signal satisfy the second condition in the folded state based on the folded state parameter.

[0018] With the above implementation manner, the electronic device can identify whether the posture change of the electronic device, the sound signal collected by different microphones satisfy the wake-up condition in the corresponding state based on the state information in different states for different states, thereby realizing the voice assistant wake-up operation in different states.

[0019] In a possible implementation manner of the first aspect, the electronic device further comprises an IMU sensor, the IMU sensor is configured to collect IMU data and send the folding state parameter IMU data to the ADSP. The method of waking up the voice assistant further comprises that the first microphone sends the collected first sound signal to the ADSP, and the third microphone sends the collected second sound signal to the ADSP, so that the ADSP can send the folding state parameter, the IMU data, the first sound signal, and the second sound signal to the voice assistant, thereby facilitating the voice assistant to determine the posture change of the electronic device based on the folding state parameter, the IMU data, the first sound signal, and the second sound signal in the folding state, and identify whether the posture change, the first sound signal, and the second sound signal meet the second condition.

[0020] It can be understood that when the electronic device is in the unfolded state, the screen management service can also send unfolded state information to the ADSP. In this way, the ADSP can send the unfolded state information, the IMU data, the first sound signal collected by the first microphone, and the third sound signal collected by the second microphone to the voice assistant service, so as to facilitate the voice assistant to determine the posture change of the electronic device based on the unfolded state parameter, the IMU data, the first sound signal, and the third sound signal in the unfolded state, and identify whether the posture change, the first sound signal, and the third sound signal meet the first condition.

[0021] With the above implementation manner, the electronic device can obtain the folding state information corresponding to the folding state based on the ADSP, and the IMU data, the first sound signal, and the second sound signal of the electronic device in the folding state, thereby identifying whether to wake up the voice assistant in the folding state.

[0022] In a possible implementation manner of the first aspect, when the electronic device is in the folding state, the voice assistant can perform data processing on the first sound signal, the second sound signal, and the IMU data based on the folding state parameter when receiving the first sound signal, the second sound signal, and the IMU data, thereby determining the posture change of the electronic device, and identifying whether the posture change, the first sound signal, and the third sound signal meet the second condition. When the posture change of the electronic device, the first sound signal, and the third sound signal meet the second condition, a wake-up instruction is sent to the voice assistant application, and in response to receiving the wake-up instruction, the voice assistant application is woken up, thereby realizing the operation of waking up the voice assistant in the folding state. The data processing includes at least one of the following: voice / wind noise identification processing, front-end enhancement processing, directional sound pickup, sound source angle processing, recording attack playback processing, large model processing, and text-independent voiceprint processing.

[0023] With the above implementation, the electronic device processes the collected first sound signal, the second sound signal and the IMU data based on the state parameter by the voice assistant, improves the accuracy and reliability of the data, so as to identify the accurate posture change of the electronic device, and timely wakes up when the posture change of the electronic device, the first sound signal and the third sound signal meet the second condition, and realizes the wake-up voice assistant operation in the folded state.

[0024] In a possible implementation of the first aspect, the electronic device further includes a first display screen and a second display screen, and the first display screen of the electronic device is invisible and the second display screen is visible when the electronic device is in the folded state. The first display screen can indicate the display screen used by the electronic device in the unfolded state and the hovering state, and the second display screen can indicate the display screen used by the electronic device in the folded state. When the electronic device is in the folded state, if the posture change of the electronic device and the sound signals collected by the first microphone and the third microphone meet the second condition, the voice assistant is woken up. Then the first interface is displayed on the second display screen, and the first interface includes the response result of the voice instruction by the voice assistant application in the electronic device, wherein the sound signals collected by the first microphone and the third microphone include the voice instruction.

[0025] It can be understood that when the electronic device is in the unfolded state, if the posture change of the electronic device and the sound signals collected by the first microphone and the second microphone meet the first condition, the voice assistant in the electronic device is woken up, and the second interface is displayed on the first display screen, and the second interface includes the response result of the second voice instruction by the voice assistant application in the electronic device, and the sound signals collected by the first microphone and the second microphone include the second voice instruction.

[0026] With the above implementation, the electronic device can display the response result of the voice instruction carried in the sound signal by the voice assistant based on different display screens in different states, which improves the user experience.

[0027] In the second aspect, the embodiments of the present application provide an electronic device, which includes a folding screen, at least three microphones, a memory and one or more processors, the memory is coupled with the processor; the memory stores computer program code, and the computer program code includes computer instructions, when the computer instructions are executed by the processor, the electronic device executes the method in the first aspect and any possible implementation thereof.

[0028] In the third aspect, the embodiments of the present application further provide a computer readable storage medium, which stores computer instructions, when the computer instructions run on the electronic device, the electronic device executes the method in the first aspect and any possible implementation thereof.

[0029] In a fourth aspect, the embodiments of the present application further provide a computer program product, comprising computer instructions, when the computer program product is executed on an electronic device, causing the electronic device to perform the method in the first aspect and any possible implementation manner thereof.

[0030] It can be understood that the beneficial effects that can be achieved by the electronic device, the computer readable storage medium, and the computer program product provided above can refer to the beneficial effects in the first aspect and any possible implementation manner thereof, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0031] Figure 1 A state schematic diagram of a transverse folding screen mobile phone provided by the embodiments of the present application;

[0032] Figure 2 A state schematic diagram of a transverse folding screen mobile phone provided by the embodiments of the present application;

[0033] Figure 3 A scene schematic diagram of breath awakening provided by the embodiments of the present application;

[0034] Figure 4 A state schematic diagram of a transverse folding screen mobile phone provided by the embodiments of the present application;

[0035] Figure 5 A scene schematic diagram of breath awakening provided by the embodiments of the present application;

[0036] Figure 6 A hardware structure schematic diagram of a transverse folding screen mobile phone provided by the embodiments of the present application;

[0037] Figure 7 A hardware structure schematic diagram of a transverse folding screen mobile phone provided by the embodiments of the present application;

[0038] Figure 8 A structure schematic diagram of a transverse folding screen mobile phone provided by the embodiments of the present application;

[0039] Figure 9 An interaction schematic diagram of switching a microphone provided by the embodiments of the present application;

[0040] Figure 10 An interaction schematic diagram of a wake-up voice assistant method provided by the embodiments of the present application. DETAILED DESCRIPTION

[0041] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing the specific embodiments and are not intended to be limiting of the present application. As used in the specification and the appended claims of the present application, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that “at least one” and “one or more” as used in the embodiments herein indicates one or two or more (including two), unless otherwise indicated. The term “and / or” is used to describe the association relationship of associated objects, which means that there can be three relationships; for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character “ / ” generally represents an “or” relationship between the associated objects before and after it.

[0042] In this specification, the reference “one embodiment” or “some embodiments” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. Thus, the appearances of the phrases “in one embodiment”, “in some embodiments”, “in other embodiments”, “in additional embodiments”, and so on, in various places in the specification are not necessarily all referring to the same embodiment, unless otherwise specified. The terms “comprising”, “including”, “having” and their variants mean “including but not limited to”, unless otherwise specified. The term “connected” includes both direct and indirect connections, unless otherwise specified. “First”, “second”, etc. are used only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features.

[0043] In the embodiments of the present application, the words “exemplary” or “for example” are used to mean serving as an example, instance, or illustration. Any embodiment or design presented as “exemplary” or “for example” in the embodiments of the present application is not necessarily to be construed as preferred or advantageous over other embodiments or design solutions. Rather, the use of the words “exemplary” or “for example” is intended to present concepts in a particular manner.

[0044] Currently, common folding screen devices mainly include two product forms of horizontal folding screen devices and vertical folding screen devices. Among them, the horizontal folding screen device has a folding axis parallel to the height direction of the horizontal folding screen device, which can also be called a vertical folding axis. The horizontal folding screen device can be folded inward or outward around the vertical folding axis to achieve a folded state. The vertical folding screen device has a folding axis parallel to the width direction of the vertical folding screen device, which can also be called a horizontal folding axis. The vertical folding screen device can be folded inward or outward around the horizontal folding axis to achieve a folded state.

[0045] Taking the horizontal folding screen device as a horizontal folding screen mobile phone and taking folding inward as an example, referring to FIG. 1, Figure 1 , the horizontal folding screen mobile phone includes a folding axis 1011 parallel to the height direction of the horizontal folding screen device. A user can fold the horizontal folding screen mobile phone inward based on the folding axis 1011 parallel to the height direction of the horizontal folding screen device, so that the horizontal folding screen mobile phone changes from an unfolded state 101 to a hovering state 102 and then to a folded state 103. The horizontal folding screen mobile phone generally includes at least two microphones. Among them, one microphone 1012 is arranged at the bottom of the horizontal folding screen mobile phone, and the other microphone 1013 is arranged at the top of the horizontal folding screen mobile phone. According to Figure 1 It can be seen that there is a certain distance between the microphone 1012 and the microphone 1013 in the unfolded state 101, the hovering state 102 and the folded state 103 of the horizontal folding screen device. In this way, there is always a difference in voice energy between the audio data collected by the microphone 1012 with the authorization of the user and the audio data collected by the microphone 1013 with the authorization of the user.

[0046] Taking the vertical folding screen device as a vertical folding screen mobile phone and taking folding inward as an example, referring to FIG. 2, Figure 2 , the vertical folding screen mobile phone includes a folding axis 2013 parallel to the width direction of the vertical folding screen device. A user can fold the vertical folding screen mobile phone inward based on the folding axis 2013 parallel to the width direction of the vertical folding screen device, so that the vertical folding screen mobile phone changes from an unfolded state 201 to a hovering state 202 and then to a folded state 203. The vertical folding screen mobile phone includes at least two microphones. Among them, one microphone 2011 (hereinafter referred to as a main microphone) is arranged at the bottom of the vertical folding screen mobile phone, and the other microphone 2012 (hereinafter referred to as a second microphone) is arranged at the top of the vertical folding screen mobile phone. According to Figure 2 It can be seen that there is a certain distance between the main microphone 2011 and the second microphone 2012 in the unfolded state 201 and the folded state 202 of the vertical folding screen device. The distance between the microphone 2011 and the second microphone 2012 is smaller in the folded state 103 of the vertical folding screen mobile phone. For ease of description, the main microphone can also be called the first microphone.

[0047] The voice assistant wake-up method provided in this application is mainly applied to vertically foldable screen devices. For example, a vertically foldable screen device can be a vertically foldable screen phone, a vertically foldable screen tablet, etc., and this application does not impose any special limitations on the specific type of such device.

[0048] The following text mainly uses vertically foldable screen devices, specifically vertically foldable screen mobile phones, as an example to explain how to wake up the voice assistant.

[0049] The above-described method for waking up the voice assistant can be applied to scenarios where vertically foldable phones use the breath-activated wake-up function to wake up the voice assistant. The breath-activated voice assistant scenario refers to a situation where the user lifts the vertically foldable phone upwards and, without uttering a wake-up phrase such as "Hello, voice assistant," directly issues a command to the voice assistant. The vertically foldable phone then detects the user's breath sound and wakes up the voice assistant to execute the operation corresponding to the command.

[0050] Taking a vertically foldable screen phone in its unfolded state as an example, see... Figure 3 The scenario for breath-activated voice assistant can be as follows: Initially, the user holds the vertically foldable phone in state 301. After the user lifts the phone upwards, the user's holding state changes to state 302. When the user speaks the voice message "Play TV series" into the microphone of the vertically foldable phone, the phone detects the user's breath and activates the breath-activated function to wake up the voice assistant, launch the video application, and play the TV series (as in state 303).

[0051] Typically, in scenarios like the breath-activated voice assistant described above, because there is a certain distance between the main microphone 2011 and the second microphone 2012 on the vertically foldable screen phone in the unfolded state 201 or the hovered state 202, there will be a certain voice energy difference between the audio data collected by the main microphone 2011 and the audio data collected by the second microphone 2012. Thus, the vertically foldable screen phone can realize the breath-activated function to wake up the voice assistant when the user lifts the vertically foldable screen phone upwards, detects the user's breath, and there is a voice energy difference between the main microphone 2011 and the second microphone 2012.

[0052] However, in the folded state 203, in the scenario of waking up the voice assistant by breath as described above, after the user lifts up the vertically folding screen mobile phone and directs the mouth to the microphone of the vertically folding screen mobile phone to send voice information, the vertically folding screen mobile phone can detect the breath of the user, but the distance between the main microphone 2011 and the second microphone 2012 of the vertically folding screen mobile phone is too close, which causes the vertically folding screen mobile phone to fail to detect the voice energy difference between the audio data collected by the main microphone 2011 and the audio data collected by the second microphone 2012, and fail to implement the breath wake-up function in the folded state 203, so as to fail to wake up the voice assistant.

[0053] To solve the above problem, an embodiment of the present application provides a vertically folding screen mobile phone, which is provided with a third microphone, and the distance between the third microphone and the main microphone is greater than a preset distance in the folded state. The vertically folding screen mobile phone collects audio data by using the main microphone and the third microphone in the folded state. In this way, the vertically folding screen mobile phone can implement the breath wake-up function based on the voice energy difference between the main microphone and the third microphone in the folded state, so as to wake up the voice assistant. For the convenience of description, the preset distance can also be referred to as the first distance.

[0054] Specifically, in the folded state, the distance 1 between the third microphone and the main microphone is greater than the preset distance, and the distance 2 between the second microphone and the main microphone is less than the preset distance. For example, the preset distance is 10 cm, and in the folded state, the distance 1 between the third microphone and the main microphone is greater than 10 cm. The vertically folding screen mobile phone can detect the voice energy difference between the audio data collected by the main microphone and the audio data collected by the third microphone, and can indicate that the audio data collected by the main microphone and the audio data collected by the third microphone constitute the voice energy difference for waking up the voice assistant by breath. In this way, in the folded state, in the scenario of waking up the voice assistant by breath as described above, after the user lifts up the vertically folding screen mobile phone and directs the mouth to the microphone such as the main microphone or the third microphone of the vertically folding screen mobile phone and sends voice information, the vertically folding screen mobile phone can detect the breath of the user and detect the voice energy difference between the audio data collected by the main microphone and the audio data collected by the third microphone, so as to implement the breath wake-up function to wake up the voice assistant, and meet the demand of the user for waking up the voice assistant by breath in the folded screen state. The vertically folding screen mobile phone can display the result of the voice assistant responding to the voice information of the user on the second display screen 2031 when the voice assistant is woken up in the folded state 203. For the convenience of description, the above-mentioned voice energy difference can also be referred to as the energy difference.

[0055] In some embodiments, the user lifts the height of the vertically folding screen mobile phone by more than a distance threshold, and the vertically folding screen mobile phone detects that the voice energy difference between the audio data collected by the main microphone and the audio data collected by the third microphone is greater than a voice energy value, and then the voice assistant can be woken up. For the sake of convenience, the user lifting the height of the vertically folding screen mobile phone can also be referred to as the displacement of the vertically folding screen mobile phone being lifted upward, and the above-mentioned distance threshold can also be referred to as the first displacement. The above-mentioned voice energy value can also be referred to as the first difference value.

[0056] In some embodiments, the third microphone of the vertically folding screen mobile phone can be arranged on the side of the body of the vertically folding screen mobile phone (see Figure 4 the third microphone 4011 in FIG. 4) is arranged on the side of the body of the vertically folding screen mobile phone, which is convenient for collecting the voice information of the user.

[0057] It can be understood that when the vertically folding screen mobile phone is in a non-folded state, such as an unfolded state or a hovering state, the distance 2 between the second microphone and the main microphone is greater than the preset distance, and the vertically folding screen mobile phone can use the main microphone and the second microphone to collect audio data. In this way, when the user lifts the vertically folding screen mobile phone upward and detects the user's breath in the non-folded state of the vertically folding screen mobile phone, the breath wake-up function can be realized based on the voice energy difference between the audio data collected by the main microphone and the audio data collected by the second microphone, so as to wake up the voice assistant. For the sake of convenience, the audio data can also be referred to as the voice signal.

[0058] In some embodiments, the vertically folding screen mobile phone can use the microphone located at the bottom as the second microphone and use the microphone located at the top as the main microphone. In this way, in the folded state, the distance 1 between the third microphone and the main microphone of the vertically folding screen mobile phone is greater than the preset distance, and the distance 2 between the second microphone and the main microphone is less than the preset distance. In the non-folded state, the distance 2 between the main microphone and the second microphone of the vertically folding screen mobile phone is greater than the preset distance.

[0059] Referring to Figure 4 , the vertically folding screen mobile phone uses the main microphone 2011 and the second microphone 2012 to collect audio data in the unfolded state 201 and the hovering state 202. The vertically folding screen mobile phone uses the main microphone 2011 and the third microphone 4011 to collect audio data in the folded state. It should be understood that Figure 4 in FIG. 4, the solid circle represents the microphone in the enabled state, and the hollow circle represents the microphone in the disabled state.

[0060] Referring to Figure 5In the scenario of waking up the voice assistant by breath as described above, after the user lifts the vertically folding screen mobile phone in the folded state 203, aims the mouth at the microphone of the vertically folding screen mobile phone such as the main microphone 2011 or the third microphone 4011, and utters the voice information "play TV series", the vertically folding screen mobile phone detects the breath of the user, and realizes the breath wake-up function to wake up the voice assistant according to the voice energy difference between the audio data collected by the main microphone 2011 and the audio data collected by the third microphone 4011. The voice assistant starts the video application according to the voice information "play TV series", and the second display screen 2031 can display the video played by the video application after the voice assistant responds to the voice information "play TV series" of the user. For the convenience of description, the interface displayed on the second display screen can also be called the first interface.

[0061] Referring to Figure 6 A hardware structure diagram of a vertically folding screen mobile phone is provided for the embodiments of the present application.

[0062] The vertically folding screen mobile phone 100 can include a processor 110, an internal memory 120, a universal serial bus (USB) connector 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a microphone 170B, a sensor module 180, a display screen 190, and the like. The sensor module 180 can include an inertial measurement unit (IMU) 181 (hereinafter referred to as an IMU sensor), a proximity light sensor 184, and the like, and the IMU sensor 181 includes an accelerometer 183, a gyroscope sensor 182, and the like.

[0063] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the vertically folding screen mobile phone 100. In other embodiments of the present application, the vertically folding screen mobile phone 100 can include more or fewer components than the illustration, or combine certain components, or split certain components, or different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0064] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices, or can be integrated in one or more processors.

[0065] In some embodiments, the digital signal processor is used to process digital signals, and can also process other digital signals. For example, when the longitudinal folding screen mobile phone 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc. Also for example, when the longitudinal folding screen mobile phone 100 is in an unfolded state, the digital signal processor can call unfolded state parameters, and when the longitudinal folding screen mobile phone is in a folded state, the digital signal processor can call folded state parameters. Exemplarily, the digital signal processor can include an advanced digital signal processor (ADSP), such as ADSP TigerSharc 101, ADSP TigerSharc 201, ADSP Blackfin series, etc.

[0066] The processor can generate operation control signals according to instruction operation codes and timing signals to complete the control of fetching and executing instructions.

[0067] The memory in the processor 110 can also be provided for storing instructions and data. In some embodiments, the memory in the processor 110 can be a cache memory.

[0068] The USB connector 130 is an interface conforming to the USB standard specification, which can be used to connect the longitudinal folding screen mobile phone 100 and peripheral devices, and can be a Mini USB connector, a Micro USB connector, a USB Type C connector, etc. The USB connector 130 can be used to connect a charger to charge the longitudinal folding screen mobile phone 100, and can also be used to connect other longitudinal folding screen mobile phones to transmit data between the longitudinal folding screen mobile phone 100 and other devices.

[0069] The charging management module 140 is used to receive the charging input of the charger.

[0070] The power management module 141 is configured to connect the battery 142 and the charging management module 140 to the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, the internal memory 120, the display screen 190, the wireless communication module 160, and the like.

[0071] The wireless communication function of the longitudinally folding screen mobile phone 100 can be realized by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, and the like.

[0072] The antenna 1 and the antenna 2 are configured to transmit and receive electromagnetic wave signals. Each of the antennas in the longitudinally folding screen mobile phone 100 can be configured to cover a single or multiple communication frequency bands.

[0073] The mobile communication module 150 can provide a solution for wireless communication including 2G / 3G / 4G / 5G and the like applied to the longitudinally folding screen mobile phone 100. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), and the like.

[0074] The wireless communication module 160 can provide a solution for wireless communication including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) network), Bluetooth (BT), Bluetooth low energy (BLE), ultra wideband (UWB), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, and the like applied to the longitudinally folding screen mobile phone 100. The wireless communication module 160 can be one or more devices integrated with at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive signals to be sent from the processor 110, perform frequency modulation and amplification on the signals, and radiate the signals as electromagnetic waves via the antenna 2.

[0075] The longitudinally folding screen mobile phone 100 can realize the display function by the GPU, the display screen 190, and the application processor, and the like.

[0076] The display screen 190 is used to display images, videos, etc. The display screen 190 includes a display panel. In some embodiments, the vertically foldable screen phone 100 may include one or more display screens 190. For example, the vertically foldable screen phone 100 includes two displays; the vertically foldable phone includes a foldable display screen and a second display screen (such as...). Figure 4 The second display screen 2031 in the middle). The folding display screen includes screen A positioned above the lateral folding axis (see...). Figure 4 Screen A in the middle), and screen B located below the horizontal folding axis (see Figure 4 (Screen B in the text). The folding display screen composed of screens A and B described above can also be referred to as the first display screen.

[0077] The internal memory 120 can be used to store computer executable program code, including instructions. The internal memory 120 may include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of the vertically foldable screen phone 100 (such as audio data, phonebook, etc.).

[0078] The vertically foldable screen phone 100 can achieve audio functions through an audio module 170, a speaker 170A, a microphone 170B, and an application processor. These functions include music playback and recording.

[0079] The audio module 170 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be located in the processor 110, or some functional modules of the audio module 170 can be located in the processor 110. For example, the audio module 170 can perform processing on the acquired authorized audio data, including voice / wind noise recognition, NN front-end enhancement, recording playback attack, directional sound pickup, sound source angle, and text-independent voiceprint processing. Specifically, the voice / wind noise recognition processing can identify human voices and / or noise in the audio data; the NN front-end enhancement processing can enhance human voices in the audio data; the recording playback attack processing can confirm that the audio data is human voices; the directional sound pickup processing can pick up audio data in a specific direction; and the sound source angle processing can determine the angle between the audio data and the vertically foldable screen phone.

[0080] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The vertically folding screen phone 100 can listen to music through the speaker 170A or output audio signals for hands-free calls.

[0081] The microphone 170B, also referred to as a "microphone", "microphone", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message to wake up a voice assistant, a user can speak into the microphone 170B by placing the mouth close to the microphone 170B, and input a sound signal into the microphone 170B. In some embodiments, the vertically folding screen mobile phone 100 can be provided with at least three microphones 170B, and there is a certain distance between each two microphones, which can realize the collection of sound signals, noise reduction, and also can identify the sound source, realize the functions of directional recording and breath wake-up, etc.

[0082] The IMU sensor 181 can be used to determine the state of the vertically folding screen mobile phone 100 (including the unfolded state, the hovering state, and the folded state).

[0083] Specifically, the gyroscope sensor 182 can be used to determine the motion posture of the vertically folding screen mobile phone 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor 182.

[0084] Specifically, the acceleration sensor 183 can detect the magnitude of the acceleration of the vertically folding screen mobile phone 100 in various directions (generally three axes). When the vertically folding screen mobile phone 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the vertically folding screen mobile phone 100, and applied to functions such as landscape / portrait screen switching and breath wake-up voice assistant.

[0085] In some embodiments, the vertically folding screen mobile phone 100 can determine the included angle between screen A and screen B in the folding display screen on the vertically folding screen mobile phone 100 based on the data collected by the gyroscope sensor 182 and the acceleration sensor 183. For example, the processor 110 can determine the included angle between screen A and screen B in the display screen according to the acceleration data and the posture data, and determine whether the vertically folding screen mobile phone 100 is in an unfolded state, a hovering state, or a folded state according to the included angle between screen A and screen B in the display screen. For example, when the included angle between screen A and screen B in the display screen is greater than or equal to 170°, it is determined that the vertically folding screen mobile phone 100 is in an unfolded state, and when the included angle between screen A and screen B in the display screen is less than or equal to 10°, it is determined that the vertically folding screen mobile phone 100 is in a folded state.

[0086] In some embodiments, the longitudinal folding screen mobile phone can also determine the state of the longitudinal folding screen mobile phone 100 based on other manners. For example, the longitudinal folding screen mobile phone 100 can determine the state of the longitudinal folding screen mobile phone 100 based on the sensing data collected by the magnetic sensor, the distance sensor, etc. For example, the processor 110 can determine the distance between the main microphone and the second microphone according to the distance data collected by the distance sensor, and determine whether the longitudinal folding screen mobile phone 100 is in an unfolded state, a hovering state or a folded state. When the distance between the main microphone and the second microphone is less than or equal to 1 cm, it is determined that the longitudinal folding screen mobile phone 100 is in a folded state, when the distance between the main microphone and the second microphone is equal to the height of the longitudinal folding screen mobile phone 100, it is determined that the longitudinal folding screen mobile phone 100 is in an unfolded state, otherwise, it is determined that the longitudinal folding screen mobile phone 100 is in a hovering state.

[0087] The proximity light sensor 184 can include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode can be an infrared light emitting diode. The longitudinal folding screen mobile phone 100 emits infrared light outwardly through the light emitting diode. The longitudinal folding screen mobile phone 100 detects infrared reflected light from nearby objects using the photodiode. When the intensity of the detected reflected light is greater than a threshold value, it can be determined that there is an object near the longitudinal folding screen mobile phone 100. When the intensity of the detected reflected light is less than a threshold value, the longitudinal folding screen mobile phone 100 can determine that there is no object near the longitudinal folding screen mobile phone 100. The longitudinal folding screen mobile phone 100 can use the proximity light sensor 184 to detect that the user holds the longitudinal folding screen mobile phone 100 close to the mouth in order to sense the user's breath, thereby realizing the function of waking up the voice assistant by the breath.

[0088] The software system in the longitudinal folding screen mobile phone described above can adopt a layered architecture, an event-driven architecture or a cloud architecture. The embodiment of the present application takes the Android system of the longitudinal folding screen mobile phone as an example to illustrate the software structure of the longitudinal folding screen mobile phone. TM

[0089] Referring to Figure 7 , a software and hardware architecture diagram of the longitudinal folding screen mobile phone provided by the embodiment of the present application.

[0090] As shown in Figure 7 , the layered architecture can divide the software and hardware of the longitudinal folding screen mobile phone into several layers, each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android TM ​The system is divided into four layers, from top to bottom, an application layer, an application framework layer, a kernel layer, a sensor hub, and a hardware layer.

[0091] The application layer 701 can include various application packages, such as a voice assistant application 7011, a video application 7012, and the like.

[0092] For example, the video application 7012 can implement a data playing service (such as playing a video). The voice assistant application 7011 is used to execute an operation corresponding to a voice instruction of a user in a wake-up state (such as starting the video application 7012 to play a video according to the voice instruction).

[0093] Hereinafter, the voice assistant application 7011 will be mainly taken as an example to illustrate the process of breath wake-up.

[0094] The voice assistant application 7011 can send an instruction to other applications to instruct the other applications to execute an operation corresponding to voice information issued by a user. For example, after the user raises the vertically folding screen mobile phone, breath of the user is detected after the user issues a voice information of “playing a TV series”, and the energy difference between the two microphones is detected, the voice assistant application 7011 is woken up, and the voice assistant application 7011 sends an instruction to the video application 7012 based on the voice information of “playing a TV series” to instruct the video application 7012 to start, so that the video application 7012 plays a TV series.

[0095] The voice assistant application 7011 can also send a request to the application framework layer 702 to instruct the microphone 7051 to collect audio data. For example, in a wake-up state, the voice assistant application 7011 sends a request to collect audio data to the application framework layer 702 to instruct the microphone 7051 to collect audio data, so as to respond to the needs of the user.

[0096] The application framework layer (Framework) 702 provides an application programming interface (API) and a programming framework for the application layer 701. The application framework layer 702 includes some pre-defined functions. For example, Figure 7 As shown, the application framework layer 702 can provide a screen management service 7021, an audio service (Audio service) 7022, a voice assistant service 7023, and the like.

[0097] The screen management service 7021 is also used to receive and process data transmitted by the hardware layer 705.

[0098] The screen management service 7021 is specifically configured to determine the state of the vertically folding screen mobile phone according to the IMU data. When it is determined that the vertically folding screen mobile phone is in the folded state, the screen management service 7021 is configured to send the state information to the audio service 7022 and the sensor hub (Sensorhub) 704 (such as the digital signal processor ADSP 7041).

[0099] For example, the screen management service 7021 is specifically configured to determine that the vertically folding screen mobile phone is in the folded state according to the IMU data when the included angle between the screen A and the screen B in the folding display screen of the vertically folding screen mobile phone is less than a preset angle threshold, and send the state information to the audio service 7022 and the ADSP 7041. The audio service 7022 sends the start instruction and the close instruction to the audio driver in the kernel layer. The audio driver controls the third microphone to start based on the start instruction and controls the second microphone to close based on the close instruction. The preset angle threshold is used to determine whether the folding display screen reaches the folded state, and can be set according to actual needs. For example, if the preset angle threshold is set to 10°, the screen management service 7021 determines that the vertically folding screen mobile phone is in the folded state according to the IMU data when the included angle between the screen A and the screen B in the folding display screen of the vertically folding screen mobile phone is 5°.

[0100] The audio service 7022 can receive a request from the application program layer 701 (such as the voice assistant application 7011), such as a request for collecting audio data, and send it to the kernel layer 703.

[0101] The audio service 7022 can also receive state information from the application program framework layer (such as the screen management service), and the state information includes the unfolded state, the hovering state or the folded state. The audio service generates the start instruction and the close instruction when the state information is the folded state. The start instruction is used to instruct the third microphone to start, and the close instruction is used to instruct the second microphone to close.

[0102] The audio service 7022 can also receive data transmitted from the hardware layer 705 (such as the microphone). For example, the audio service 7022 receives the audio data collected by each start microphone and transmits it to the voice assistant service 7023.

[0103] The voice assistant service 7023 is configured to determine the voice energy of the audio data collected by each activated microphone, and determine the voice energy difference according to the audio data collected by the two activated microphones. In the unfolded state and the hovering state, the voice energy of the audio data collected by the main microphone and the voice energy of the audio data collected by the second microphone are determined, and the voice energy difference between the audio data collected by the main microphone and the second microphone is determined. In the folded state, the voice energy of the audio data collected by the main microphone and the voice energy of the audio data collected by the third microphone are determined, and the voice assistant service 7023 determines the voice energy difference between the audio data collected by the main microphone and the third microphone.

[0104] The voice assistant service 7023 is configured to determine whether to perform the breath wake-up function according to the IMU data and the voice energy difference when the longitudinal folding screen mobile phone detects the user's breath, and send a wake-up instruction to the voice assistant application 7011 to wake up the voice assistant application 7011 when it is determined to perform the breath wake-up function.

[0105] The kernel layer (Kernel) 703 includes an audio driver 7031.

[0106] The audio driver 7031 can receive a notification (such as a notification indicating the activation or deactivation of the microphone) from the application framework layer 702, and send an activation or deactivation instruction to the microphone according to the notification.

[0107] The sensor hub (Sensorhub) 704 includes a digital signal processor (ADSP) 7041.

[0108] The ADSP 7041 is configured to receive data (including IMU data and audio data collected by each activated microphone) transmitted by the hardware layer 705, and send the data to the audio service 7022.

[0109] The hardware layer (hardware layer) 705 includes a microphone 7051, an IMU sensor 7052, a proximity light sensor 7055, etc. The IMU sensor 7052 includes a gyroscope sensor 7053 and an acceleration sensor 7054, etc.

[0110] The microphone 7051 is configured to collect audio data and send the audio data to the ADSP 7041.

[0111] The IMU sensor 7052 is configured to collect IMU data and send the IMU data to the ADSP 7041, wherein the IMU data includes acceleration data and motion posture data.

[0112] The gyroscope sensor 7053 is configured to collect motion posture data and send the motion posture data to the ADSP 7041, and the acceleration sensor 7054 is configured to collect acceleration data and send the acceleration data to the ADSP 7041.

[0113] The proximity light sensor 7055 is used to collect proximity light data and transmit to the ADSP 7041.

[0114] The following will be described in combination with Figure 8 The process of waking up the voice assistant 7011 by breath in the folded state is exemplarily described as follows:

[0115] The ADSP includes an IMU module, a voice module and a proximity light module.

[0116] The IMU sensor sends the monitored IMU data to the IMU module in the ADSP, the third microphone and the main microphone collect audio data, and send the collected audio data to the voice module in the ADSP. The proximity light sensor sends the collected proximity light data to the proximity light module in the ADSP.

[0117] The ADSP sends the IMU data, audio data and proximity light data to the AP, specifically to the voice assistant service in the AP.

[0118] The voice assistant service in the AP pre-processes the IMU data, proximity light data and voice data, which includes but is not limited to voice / wind noise identification, NN front-end enhancement, recording playback attack, directional sound pickup, sound source angle, etc. Then the pre-processed IMU data, proximity light data and voice data are input to the large model module and the text-independent voiceprint module.

[0119] When the longitudinal folding screen mobile phone detects the breath of the user, the large model module and the text-independent voiceprint module in the voice assistant service determine whether to execute the breath wake-up function according to the pre-processed IMU data and the audio data collected by the two activated microphones. When the determination result is to execute the breath wake-up function, the voice assistant service sends a wake-up instruction and voice information to the voice assistant application to wake up the voice assistant application. The voice assistant application can respond to the voice information to perform corresponding operations. Specifically, the large model module and the text-independent voiceprint module determine that the longitudinal folding screen mobile phone is lifted according to the pre-processed IMU data, determine the voice energy difference according to the pre-processed audio data collected by the main microphone and the third microphone, and determine that the voiceprint information in the audio data collected by the main microphone and the third microphone matches the voiceprint information of the user, so as to comprehensively determine to execute the breath wake-up operation. For ease of description, the above-mentioned pre-processing and the processing of the large model module and the text-independent voiceprint module can be collectively referred to as data processing.

[0120] The method for waking up the voice assistant provided by the embodiments of the present application can be implemented in the longitudinal folding screen mobile phone with the above-mentioned hardware structure and software structure.

[0121] In the embodiment of the present application, the longitudinal folding screen mobile phone is in a non-folded state (such as an unfolded state or a hovering state), the main microphone and the second microphone are started to collect audio data. Since the main microphone and the second microphone are at a certain distance, the distance 1 between the main microphone and the user's mouth is not the same as the distance 2 between the third microphone and the user's mouth, so there is a voice energy difference between the audio data 1 collected by the main microphone and the audio data 2 collected by the second microphone. In this way, when the user lifts the longitudinal folding screen mobile phone and the longitudinal folding screen mobile phone detects the user's breath, the longitudinal folding screen mobile phone can realize the breath wake-up function according to the voice energy difference between the audio data 1 collected by the main microphone and the audio data 2 collected by the second microphone, to wake up the voice assistant.

[0122] When the longitudinal folding screen mobile phone detects the switching from the non-folded state to the folded state, the second microphone is closed and the third microphone is started to collect audio data with the main microphone. Since the main microphone and the third microphone are at a certain distance, the distance 1 between the main microphone and the user's mouth is not the same as the distance 3 between the third microphone and the user's mouth, so there is a voice energy difference between the audio data 1 collected by the main microphone and the audio data 3 collected by the third microphone. In this way, when the user lifts the longitudinal folding screen mobile phone and the longitudinal folding screen mobile phone detects the user's breath, the longitudinal folding screen mobile phone can realize the breath wake-up function according to the voice energy difference between the audio data 1 collected by the main microphone and the audio data 3 collected by the third microphone, to wake up the voice assistant.

[0123] In some embodiments, in combination with the hardware and software structure of the mobile phone, the embodiment can realize the process of switching the microphone of the longitudinal folding screen mobile phone from the non-folded state to the folded state through the steps as shown in Figure 9

[0124] S900, the IMU sensor collects IMU data.

[0125] Specifically, the longitudinal folding screen mobile phone starts the main microphone and the second microphone in the unfolded state or the hovering state to collect audio data with the main microphone and the second microphone. In the breath wake-up scene, after the user lifts the longitudinal folding screen mobile phone and uses the mouth to aim at the microphone of the longitudinal folding screen mobile phone to issue voice information, the longitudinal folding screen mobile phone can realize the breath wake-up operation based on the voice energy difference between the audio data 1 collected by the main microphone and the audio data 2 collected by the second microphone.

[0126] Specifically, the IMU sensor is in a started state and can continuously collect IMU data.

[0127] ​It should be understood that when the user folds the longitudinal folding screen mobile phone from the unfolded state to the folded state, the IMU data collected by the IMU sensor can indicate the above-mentioned operation of the user. Among them, the IMU sensor includes an acceleration sensor and a gyroscope sensor, and the IMU data includes but is not limited to acceleration data and motion posture data.

[0128] S901, the IMU sensor sends the IMU data to the screen management service.

[0129] Specifically, the IMU sensor sends the IMU data collected in real time to the screen management service.

[0130] S902, the screen management service generates folding state information when determining that the state of the longitudinal folding screen mobile phone is switched from the unfolded state to the folded state according to the IMU data.

[0131] Specifically, when the screen management service receives the IMU data, it can determine the state of the longitudinal folding screen mobile phone according to the IMU data.

[0132] For example, the screen management service can determine the included angle between screen A and screen B in the folding display screen on the longitudinal folding screen mobile phone according to the motion posture data in the IMU data. When the included angle between screen A and screen B is less than a preset angle threshold, the screen management service determines that the state of the longitudinal folding screen mobile phone is the folded state.

[0133] Specifically, the state information of the longitudinal folding screen mobile phone includes unfolded state information, hovering state information, and folded state information. For example, the screen management service can use a state parameter to represent the state information. For example, the state parameter includes "0", "1", and "2". When the state parameter is "0", it indicates that the longitudinal folding screen mobile phone is in the unfolded state, when the state parameter is "1", it indicates that the longitudinal folding screen mobile phone is in the folded state, and when the state parameter is "2", it indicates that the longitudinal folding screen mobile phone is in the hovering state.

[0134] S903, the screen management service sends the folding state information to the digital signal processor ADSP.

[0135] Specifically, when the screen management service detects a change in the state of the longitudinal folding screen mobile phone, the screen management service sends the changed state information to the digital signal processor ADSP, so that the ADSP can call the corresponding state parameter based on the changed state information. For example, when the longitudinal folding screen mobile phone changes from the unfolded state to the folded state, the screen management service sends the folding state information to the digital signal processor ADSP, and the ADSP can call the folding state parameter and send it to the voice assistant service in the AP, so that the voice assistant service determines whether to start the voice assistant application based on the folding state parameter.

[0136] S904, the screen management service sends folding state information to the audio service.

[0137] Specifically, the screen management service sends folding state information to the audio service when detecting that the vertically folding screen mobile phone changes from the unfolded state to the folded state, the folding state information being used to indicate that the state of the vertically folding screen mobile phone is the folded state, so that the audio service drives the audio driver to start the microphone used in the folded state.

[0138] For example, the vertically folding screen mobile phone uses the main microphone and the third microphone to collect audio data in the folded state, and the audio service needs to drive the audio driver to start the main microphone and the third microphone in the folded state.

[0139] S905, the audio service generates switching instructions in response to receiving the folding state information.

[0140] Specifically, the vertically folding screen mobile phone can perform the switching microphone operation when the state of the vertically folding screen mobile phone is the folded state. The audio service generates switching instructions in response to receiving the folding state information. For ease of description, the above switching instructions can also be referred to as microphone switching instructions.

[0141] S906, the audio service sends the switching instructions to the audio driver.

[0142] Specifically, the audio service sends the switching instructions to the audio driver, the switching instructions being used to instruct the audio driver to switch the second microphone to the third microphone.

[0143] S907, the audio driver generates first driving instructions and second driving instructions in response to receiving the switching instructions.

[0144] Specifically, the audio driver switches the second microphone to the third microphone in response to receiving the switching instructions, and generates the first driving instructions and the second driving instructions accordingly. The first driving instructions are used to instruct the second microphone to close, and the second driving instructions are used to instruct the third microphone to start. For ease of description, the above first driving instructions can also be referred to as first instructions, and the above second driving instructions can also be referred to as second instructions.

[0145] S908, the audio driver sends the first driving instructions to the second microphone, the first driving instructions being used to instruct the second microphone to close.

[0146] S909, the second microphone closes in response to receiving the first driving instructions.

[0147] Specifically, when the second microphone is closed, the vertically folding screen mobile phone cannot use the second microphone to collect audio data.

[0148] S910, the audio driver sends a second driving instruction to the third microphone, the second driving instruction being used to instruct the third microphone to start.

[0149] It should be understood that during the switching microphone process in which the second microphone is closed and the third microphone is started, a pop noise may occur. Since the process of detecting user identification is machine identification, the process of switching microphones does not affect the determination result of whether the breath wake-up function is performed by the vertically folding screen mobile phone.

[0150] S911, in response to receiving the second driving instruction, the third microphone starts.

[0151] Specifically, when the third microphone starts, the vertically folding screen mobile phone can use the main microphone to collect audio data 1 and the third microphone to collect audio data 3. In this way, the vertically folding screen mobile phone can realize breath wake-up operation based on the voice energy difference between the audio data 1 collected by the main microphone and the audio data 3 collected by the third microphone in the breath wake-up scene, so as to wake up the voice assistant application.

[0152] It should be understood that in the folded state and the unfolded state, the vertically folding screen mobile phone will use the main microphone to collect audio data, that is, the main microphone is always in the started state, so after the vertically folding screen mobile phone is switched from the unfolded state to the folded state, the audio service does not need to generate an instruction for driving the main microphone.

[0153] It can be understood that the above steps S908-S9011 are described by taking the example of first closing the second microphone and then starting the third microphone, which is not limited in actual application. Exemplarily, steps S908-S909 can be executed before steps S910-S911, that is, the third microphone is started before the second microphone is closed. Exemplarily, steps S910-S911 and steps S908-S909 can be executed simultaneously, that is, the third microphone is started and the second microphone is closed at the same time.

[0154] S912, the main microphone collects audio data 1.

[0155] Specifically, the main microphone is always in the started state, so the main microphone can always collect audio data 1.

[0156] It should be understood that when the user lifts the vertically folding screen mobile phone and directs the mouth to the microphone (which can be the main microphone or the third microphone) of the vertically folding screen mobile phone to issue voice information, the audio data 1 collected by the main microphone includes the voice information of the user.

[0157] S913, the third microphone collects audio data 3.

[0158] Specifically, the third microphone can continuously collect audio data 3 after being activated. It should be understood that, in the operation of the user lifting the vertically folding screen mobile phone and directing the mouth to the microphone (which can be the main microphone or the third microphone) of the vertically folding screen mobile phone to send voice information, the audio data 3 collected by the third microphone includes the voice information of the user. For the convenience of description, the above voice information can also be referred to as voice instruction.

[0159] It should be noted that, during the process of step S901 to step S913, the main microphone is normally working (such as continuously collecting audio data 1 and transmitting to the digital signal processor). And, before step S909, the vertically folding screen mobile phone is in a non-folded state (such as an unfolded state, a hovering state), and the second microphone is normally working (such as continuously collecting audio data 2 and transmitting to the digital signal processor).

[0160] It can be understood that, when the vertically folding screen mobile phone changes from the folded state to the non-folded state, it needs to switch the working of the main microphone and the second microphone, that is, the vertically folding screen mobile phone drives the third microphone to be closed and drives the second microphone to be activated. The specific implementation process can be referred to in the description of step S909. Figure 9 Here, no longer described.

[0161] In some embodiments, the vertically folding screen mobile phone can record the above process of switching the microphone in the offline log. For example, the offline log can be used to record the process of switching the microphone caused by each state change of the vertically folding screen mobile phone. For example, the vertically folding screen mobile phone can obtain the offline log shown in Table 1 as follows:

[0162] Table 1

[0163]

[0164]

[0165] In some embodiments, each microphone has a unique identification, such as a unique microphone ID. For example, in Table 1 above, microphone ID: 1 represents the primary microphone, microphone ID: 2 represents the second microphone, and microphone ID: 3 represents the third microphone. The state is used to represent the state of the vertically folding screen mobile phone; wherein state: 1 represents that the vertically folding screen mobile phone is in an unfolded state, state: 2 represents that the vertically folding screen mobile phone is in a hovering state, and state: 3 represents that the vertically folding screen mobile phone is in a folded state. The status is used to represent the state of the microphone, and status: 0 represents that the microphone is started, and status: 0 represents that the microphone is closed. Based on this, in combination with the three parameters of microphone ID, state and status, the microphone switching process of the vertically folding screen mobile phone in the case of state change can be determined. For example, microphone ID: 3, state: 3 and status: 0 represent that the third microphone is in a started state when the vertically folding screen mobile phone is in a folded state.

[0166] In combination with the above method of switching the microphone of the vertically folding screen mobile phone in the folded state, that is, switching to the primary microphone and the third microphone to work (such as Figure 9 S912 and S913 in the above method), further in combination with Figure 10 the process of realizing breath wake-up voice assistant of the vertically folding screen mobile phone in the folded state is explained:

[0167] S1001, the primary microphone sends audio data 1 to the digital signal processor ADSP.

[0168] S1002, the third microphone sends audio data 3 to the digital signal processor ADSP.

[0169] Specifically, in response to receiving the folded state information, the digital signal processor can call the folded state parameter and send the folded state parameter to the voice assistant service, so that the voice assistant service determines whether to start the voice assistant application based on the audio data 1, the audio data 3 and the IMU data in the folded state according to the folded state parameter.

[0170] It can be understood that the above steps S1001, S1002 and S1003 are steps of sending sensing data (such as audio data 1, audio data 3 and IMU data) to the digital signal processor, and the execution order is not limited to Figure 10The exemplary, the step S1003 can be executed first, and then the step S1001 is executed, and then the step S1002 is executed, that is, the digital signal processor receives the audio data 3 sent by the third microphone first, then receives the IMU data sent by the IMU sensor, and then receives the audio data 1 collected by the main microphone. Exemplarily, the steps S1002, S1003 and S1001 can be executed simultaneously, that is, the digital signal processor receives the audio data 1 sent by the main microphone, the audio data 3 sent by the third microphone and the IMU data sent by the IMU sensor at the same time.

[0171] It should be understood that when the user speaks voice information with the mouth aligned with the main microphone or the third microphone, because there is a certain distance between the main microphone and the third microphone, the distance 1 between the main microphone and the mouth of the user is not the same as the distance 2 between the third microphone and the mouth of the user, so there is a certain voice energy difference between the audio data 1 collected by the main microphone and the audio data 2 collected by the third microphone.

[0172] S1003, the IMU sensor collects IMU data.

[0173] It can be understood that the IMU sensor can continuously collect IMU data, and the IMU sensor collects IMU data in S1000 is not essentially different from the IMU sensor collecting IMU data in S901, but only uses different labels to represent that the IMU sensor collects IMU data at different time points. Among them, S1000 represents that the IMU sensor collects IMU data after switching to the main microphone and the third microphone.

[0174] When the user lifts the vertically folding screen mobile phone, the IMU data collected by the IMU sensor can indicate the above operation of the user.

[0175] S1004, the IMU sensor sends the IMU data to the digital signal processor ADSP.

[0176] After S1000, the IMU sensor can also send the IMU data to the screen management service to the screen management service. When the screen management service determines that the state of the vertically folding screen mobile phone changes based on the IMU data, the corresponding state information is generated and sent to the ADSP. For example, when the screen management service determines that the state of the vertically folding screen mobile phone changes from the folded state information to the unfolded state based on the IMU data, such as the unfolded state or the hovering state, the unfolded state information or the hovering state information is sent to the ADSP. For details, please refer to the description in the foregoing Figure 9 , which will not be described here again. In the flowchart shown in Figure 10 , the specific implementation of waking up the voice assistant is mainly concerned.

[0177] S1005, The digital signal processor ADSP sends the audio data 1, the audio data 3 and the IMU data to the voice assistant service.

[0178] Specifically, the digital signal processor ADSP sends the received audio data 1, the audio data 3 and the IMU data to the AP.

[0179] It can be understood that when the digital signal processor ADSP receives the state information, the digital signal processor ADSP can call the corresponding state parameter according to the state information, and send the state parameter to the voice assistant service in the AP, so that the voice assistant service processes the audio data 1, the audio data 3 and the IMU data based on the folding state parameter to determine whether to start the voice assistant application. For example, when the digital signal processor ADSP receives the folding state information sent by the screen management service, the ADSP can call the folding state parameter and send it to the voice assistant service in the AP, so that the voice assistant service determines whether to start the voice assistant application based on the folding state parameter.

[0180] S1006, In response to receiving the audio data 1, the audio data 3 and the IMU data, the AP determines whether to start the voice assistant application according to the audio data 1, the audio data 3 and the IMU data.

[0181] Specifically, in response to receiving the audio data 1, the audio data 3 and the IMU data, the AP can determine whether the voiceprint of the voice information in the audio data 1 matches the voiceprint of the user, determine whether the voiceprint of the voice information in the audio data 3 matches the voiceprint of the user, determine whether there is a voice energy difference between the audio data 1 and the audio data 3, determine whether the vertically folding screen mobile phone is lifted according to the IMU data, etc., to comprehensively determine whether the breath wake-up operation can be performed to wake up the voice assistant application.

[0182] S1007, When the discrimination result is to perform the breath wake-up operation, the AP generates a wake-up instruction.

[0183] Specifically, when the discrimination result is to perform the breath wake-up operation, the AP determines that the voice assistant application can be awakened by using the breath wake-up function, and generates a wake-up instruction.

[0184] S1008, The AP sends the wake-up instruction to the voice assistant application, and the wake-up instruction is used to instruct the voice assistant application to wake up.

[0185] Specifically, the AP sends the wake-up instruction to the voice assistant application, and the wake-up instruction is used to instruct the voice assistant application to wake up, thereby realizing the function of waking up the voice assistant by breath.

[0186] S1009, In response to receiving the wake-up instruction, the voice assistant application wakes up.

[0187] Specifically, in response to receiving the wake-up instruction, the voice assistant application is woken up so as to respond to the user operation b.

[0188] For ease of illustration, the audio data 1 collected by the main microphone can also be referred to as a first sound signal, the audio data 3 collected by the third microphone can also be referred to as a second sound signal, and the audio data 2 collected by the second microphone can also be referred to as a third sound signal.

[0189] In some embodiments, the wake-up instruction carries the audio data 1 and / or the audio data 3. The voice information corresponding to the user operation b is included in the audio data 1 and / or the audio data 3, so that the voice assistant application can respond to the user operation b based on the voice information of the user. For example, the voice information carried in the audio data 1 and / or the audio data 3 is "play a TV series", and the voice assistant application can start the video application to play the TV series in response to "play a TV series".

[0190] It should be understood that during the process of the longitudinal folding screen mobile phone from the folded state to the unfolded state, the longitudinal folding screen mobile phone can switch the third microphone to the second microphone, that is, close the third microphone and start the second microphone. The longitudinal folding screen mobile phone takes the audio data 1 collected by the main microphone and the audio data 2 collected by the second microphone as the audio data to be processed to realize various functions (including the breath wake-up function). The normal working process of the longitudinal folding screen mobile phone for realizing the breath wake-up function by using the audio data 1 and the audio data 2 can be referred to the process of the foregoing steps S1001-S1009, which will not be described here again.

[0191] When the longitudinal folding screen mobile phone detects that the state is the folded state, the audio data 1 collected by the main microphone and the audio data 3 collected by the third microphone are used, so that in the breath wake-up scene, the breath wake-up operation can be realized based on the voice energy difference between the audio data 1 and the audio data 3 to wake up the voice assistant to respond to the voice information of the user, thereby meeting the demand of the user for waking up the voice assistant by using the breath in the folded state.

[0192] Some other embodiments of the present application provide an electronic device, which can include a memory and one or more processors. The memory and the processor are coupled. The memory is configured to store computer program code including computer instructions. When the processor executes the computer instructions, the electronic device can perform each step in the above method embodiments. The structure of the electronic device can be referred to the structure shown in Figure 6 .

[0193] The embodiments of the present application also provide a computer readable storage medium, which includes computer instructions. When the computer instructions run on the above-mentioned electronic device, the electronic device executes each step in the above-mentioned method embodiments.

[0194] The embodiment of the present application further provides a computer program product, which, when running on an electronic device, causes the above-mentioned electronic device to perform each step in the above-mentioned method embodiment.

[0195] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above-described functions.

[0196] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0197] The units described as separate components can or can not be physically separated, and the components displayed as units can be one physical unit or multiple physical units, that is, can be located in one place, or can be distributed to multiple different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0198] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0199] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The software product is stored in a storage medium, including a plurality of instructions to make a device (which can be a single-chip microcomputer, a chip, etc.) or a processor execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0200] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any change or replacement within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method of waking up a voice assistant, the method comprising: The application is applied to an electronic device with a folding screen, the electronic device comprising a folding shaft, a first microphone, a second microphone and a third microphone; the folding shaft is parallel to the width direction of the electronic device; when the electronic device is in a folded state, the distance between the first microphone and the second microphone is less than a first distance, and the distance between the first microphone and the third microphone is greater than the first distance; The method comprises: In the case where the electronic device is in an unfolded state, if the posture change of the electronic device and the sound signals collected by the first microphone and the second microphone meet a first condition, a voice assistant in the electronic device is woken up; In the case where the electronic device is in a folded state, if the posture change of the electronic device and the sound signals collected by the first microphone and the third microphone meet a second condition, the voice assistant is woken up.

2. The method of claim 1, wherein, The method further comprises: After detecting that the electronic device switches from the folded state to the unfolded state, the second microphone is started and the third microphone is closed; After detecting that the electronic device switches from the unfolded state to the folded state, the third microphone is started and the second microphone is closed.

3. The method according to claim 1 or 2, characterized in that, The second condition comprises: The displacement of the electronic device being lifted upward exceeds a first displacement, and the posture change of the electronic device comprises the displacement of the electronic device being lifted upward; and The energy difference between the first sound signal collected by the first microphone and the second sound signal collected by the third microphone exceeds a first difference value.

4. The method of claim 2, wherein, The electronic device comprises a screen management service, an audio service and an audio driver; The starting of the third microphone and the closing of the second microphone after detecting that the electronic device switches from the unfolded state to the folded state comprises: After detecting that the electronic device switches from the unfolded state to the folded state, the screen management service sends folding state information to the audio service; The audio service sends a microphone switching instruction to the audio driver; The audio driver generates a first instruction and a second instruction based on the microphone switching instruction; The audio driver sends the first instruction to the second microphone and sends the second instruction to the third microphone; In response to the first instruction, the second microphone is closed; In response to the second instruction, the third microphone is started.

5. The method of claim 4, wherein, The electronic device further comprises a digital signal processor (ADSP) and a voice assistant service, and the method further comprises: After detecting that the electronic device switches from the unfolded state to the folded state, the screen management service sends folding state information to the ADSP; The ADSP acquires folding state parameters matched with the folding state information and sends them to the voice assistant service; The voice assistant service identifies whether the posture change of the electronic device and the sound signals collected by the first microphone and the third microphone meet the second condition based on the folding state parameters.

6. The method of claim 5, wherein, The electronic device further comprises an IMU sensor configured to collect IMU data and send the IMU data to the ADSP; The method further comprises: The first microphone sends the collected first sound signal to the ADSP; The third microphone sends the collected second sound signal to the ADSP; The ADSP sends the folding state parameter, the IMU data, the first sound signal, and the second sound signal to the voice assistant.

7. The method of claim 5, wherein, The voice assistant service identifies, based on the folding state parameter, whether the posture change of the electronic device and the sound signals collected by the first microphone and the third microphone satisfy a second condition, including: The voice assistant service performs data processing on the first sound signal collected by the first microphone, the second sound signal collected by the third microphone, and the IMU data based on the folding state parameter to determine the posture change of the electronic device and identify whether the posture change of the electronic device, the first sound signal, and the second sound signal satisfy a second condition; The data processing includes at least one of the following: voice / wind noise identification processing, front-end enhancement processing, directional sound pickup, sound source angle processing, recording attack playback processing, large model processing, and text-independent voiceprint processing; When the posture change of the electronic device, the first sound signal, and the third sound signal satisfy the second condition, the voice assistant service sends a wake-up instruction to the voice assistant application; In response to receiving the wake-up instruction, the voice assistant application is woken up.

8. The method according to any one of claims 1-6, characterized in that, The electronic device further comprises a first display screen and a second display screen; when the electronic device is in the folding state, the first display screen is invisible, and the second display screen is visible; In the case where the electronic device is in the folding state, if the posture change of the electronic device and the sound signals collected by the first microphone and the third microphone satisfy the second condition, after the voice assistant is woken up, the method further comprises: The second display screen displays a first interface, and the first interface includes a response result of a voice assistant application in the electronic device to a voice instruction, and the sound signals collected by the first microphone and the third microphone include the voice instruction.

9. An electronic device, comprising: The electronic device comprises a folding screen, at least three microphones, a memory, and one or more processors; the memory and the processors are coupled; the memory is configured to store computer program code, the computer program code comprising computer instructions, when the computer instructions are executed by the processors, causing the electronic device to perform the method of any one of claims 1-8.

10. A computer-readable storage medium having stored computer instructions therein, characterized in that, When the computer instructions run on the electronic device, the electronic device performs the method of any one of claims 1-8.

11. A computer program product comprising computer instructions, characterized in that, When the computer program product runs on the electronic device, the electronic device performs the method of any one of claims 1-8.