An audio acquisition method, an electronic device, and a storage medium

The audio capture method on electronic devices uses a microphone array to detect proximity and speech, automatically recording and saving audio, addressing the inefficiencies of traditional methods by reducing user interaction and improving user experience.

CN119182851BActive Publication Date: 2025-07-15HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311870062.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-30
Publication Date
2025-07-15
Estimated Expiration
2043-12-30

AI Technical Summary

Technical Problem

In the prior art, users have complicated operations when collecting audio on electronic devices, and the voice wake-up method is prone to mistouching, resulting in low collection efficiency and poor user experience.

Method used

A microphone array is used to recognize near-field voice, collect audio signals through the microphone array, automatically determine the distance from the target sounding object, and automatically collect and save audio signals when the distance is less than the preset distance, or convert them into text information to reduce user operations.

Benefits of technology

It realizes automatic collection and saving of audio after near-field voice is recognized, improving audio acquisition efficiency, reducing user operations, and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119182851B_ABST
    Figure CN119182851B_ABST
Patent Text Reader

Abstract

The present application provides an audio acquisition method, an electronic device, and a storage medium. The first electronic device includes a microphone array, and the microphone array includes at least two microphones. The method includes: the first electronic device acquires a first audio signal output by a target sound-emitting object through the microphone array; the first electronic device determines a first distance between the first electronic device and the target sound-emitting object based on the first audio signal; when the first distance is less than a first preset distance, the first electronic device acquires a second audio signal output by the target sound-emitting object through the microphone array; the first electronic device saves the second audio signal in a first application. It realizes that the electronic device automatically picks up audio, reduces user operations, improves audio acquisition efficiency, and enhances the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technologies, and particularly to an audio acquisition method, an electronic device, and a storage medium. Background Art

[0002] With the development of terminal technologies, electronic devices can support functions of audio output and audio acquisition. For example, an electronic device can play audio through a speaker, and the electronic device can also acquire audio through a microphone.

[0003] The electronic device can save the audio acquired by the microphone locally, or the electronic device can also send the audio acquired by the microphone to other devices. However, the user needs to find the audio acquisition entry in an application that supports the voice recording function and control the electronic device to acquire audio, and the user operation is relatively cumbersome. Or the user can wake up the application that supports the voice recording function on the electronic device through voice to acquire audio, and false touch is likely to occur during voice wake-up. How to provide a method for conveniently, quickly, and accurately acquiring audio remains to be further studied. Summary of the Invention

[0004] This application provides an audio acquisition method, an electronic device, and a storage medium, which improve the audio acquisition efficiency, reduce user operations, and enhance the user experience.

[0005] In a first aspect, this application provides an audio acquisition method. The first electronic device includes a microphone array, and the microphone array includes at least two microphones. The method includes: the first electronic device acquires a first audio signal output by a target sound-emitting object through the microphone array; the first electronic device determines a first distance between the first electronic device and the target sound-emitting object based on the first audio signal; when the first distance is less than a first preset distance, the first electronic device acquires a second audio signal output by the target sound-emitting object through the microphone array; the first electronic device saves the second audio signal in a first application.

[0006] Optionally, the electronic device may also save the first audio signal, or may not save the first audio signal.

[0007] Optionally, the first electronic device saves the second audio signal in the first application. It may be that the first electronic device directly saves the second audio signal in the first application, or it may be that the second audio signal is converted into text information and then the text information is saved in the first application.

[0008] The electronic device can determine a first distance from a target sound - emitting object based on the collected audio signal. When the first distance between the electronic device and the target sound - emitting object is greater than a first preset distance, or when the first distance between the electronic device and the target sound - emitting object is greater than the first preset distance for a certain period of time, the electronic device can determine that near - field voice is recognized. The electronic device can automatically start collecting audio and automatically save the audio in the first application. This realizes the automatic picking up of audio by the electronic device, reduces user operations, improves the audio collection efficiency, and enhances the user experience.

[0009] In combination with the first aspect, in a possible implementation manner, the first electronic device saves the second audio signal in the first application, specifically including:

[0010] In response to the first distance being less than the first preset distance, the first electronic device starts the first application and saves the second audio signal in the first application, or the first electronic device converts the second audio signal into first text information and saves the first text information in the first application.

[0011] In this way, after recognizing near - field voice, the electronic device can automatically save the collected audio signal or the text information corresponding to the collected audio signal in the first application.

[0012] In combination with the first aspect, in a possible implementation manner, the first electronic device saves the second audio signal in the first application, specifically including: In response to the first distance being less than the first preset distance, the first electronic device starts the first application and sends the second audio signal to the second electronic device through the first application, or the first electronic device converts the second audio signal into first text information and sends the first text information to the second electronic device through the first application.

[0013] In this way, after recognizing near - field voice, the electronic device can automatically send the collected audio signal to the second electronic device with which a communication connection is established, or send the text information corresponding to the collected audio signal to the second electronic device with which a communication connection is established.

[0014] In combination with the first aspect, in a possible implementation manner, the first electronic device is connected to a third electronic device via Bluetooth; the first electronic device saves the second audio signal in the first application, specifically including: The first electronic device sends the second audio signal to the third electronic device through the Bluetooth connection; the third electronic device saves the second audio signal in the first application.

[0015] Optionally, the first electronic device can be a Bluetooth headset. After the Bluetooth headset recognizes near - field voice, the Bluetooth headset can automatically send the collected audio signal to the third electronic device with which a Bluetooth connection is established, or send the text information corresponding to the collected audio signal to the third electronic device with which a Bluetooth connection is established.

[0016] In combination with the first aspect, in a possible implementation, while the first electronic device acquires the second audio signal through the microphone array, the first electronic device also acquires the third audio signal output by other sound - producing objects through the microphone array; the electronic device saves the second audio signal in the first application, specifically including: when the first electronic device determines that the distance between the first electronic device and other sound - producing objects exceeds the first preset distance, the electronic device saves the second audio signal in the first application.

[0017] Optionally, the second audio signal and the third audio signal can be audio signals in the same audio file. The electronic device 100 can extract the second audio signal or the third audio signal from the same audio file.

[0018] Optionally, the second audio signal and the third audio signal can also be audio signals in two different audio files respectively.

[0019] Near - field voice can refer to the audio output by a target sound - producing object within the first preset distance from the electronic device. Far - field voice can refer to the audio output by a target sound - producing object beyond the first preset distance from the electronic device.

[0020] In this way, after recognizing near - field voice and starting to collect audio signals, the electronic device can determine whether the collected audio signal is a near - field audio signal or a far - field audio signal, and only save the near - field audio signal, without saving the far - field audio signal, which can avoid the interference of the far - field audio signal.

[0021] In combination with the first aspect, in a possible implementation, when the first distance is less than the first preset distance, the first electronic device acquires the second audio signal output by the target sound - producing object through the microphone array, specifically including: the first electronic device determines the probability of the voice signal contained in the first audio signal; when the first distance is less than the first preset distance and the probability of the voice signal contained in the first audio signal is greater than the first threshold, the first electronic device acquires the second audio signal output by the target sound - producing object through the microphone array.

[0022] In this way, before starting to collect audio, the electronic device can determine the probability of the voice signal contained in the audio signal. When the probability of the voice signal contained in the audio signal is greater than the first threshold, that is, the probability of recognizing that the user is speaking is relatively large, the electronic device can collect and save the second audio signal.

[0023] In combination with the first aspect, in a possible implementation, the first electronic device is connected to a third electronic device via Bluetooth; the first electronic device obtains a second audio signal output by a target sound-emitting object through a microphone array, specifically including: when the first electronic device determines that the first electronic device is in a handheld state, the first distance is less than a first preset distance, and the probability of the voice signal included in the first audio signal is greater than a first threshold, the first electronic device obtains the second audio signal output by the target sound-emitting object through the microphone array.

[0024] Optionally, a sensor is pre-installed on the first electronic device, and whether it is in a handheld state can be confirmed through the sensor signal collected by the sensor.

[0025] In this way, when the first electronic device is a Bluetooth headset, when the Bluetooth headset is in a handheld state, it can be considered that the current user has the intention to speak to the Bluetooth headset. When the Bluetooth headset is in a handheld state, the distance between the Bluetooth headset and the target sound-emitting object is within the first preset distance, and the probability of the voice signal included in the first audio signal is greater than the first threshold, the Bluetooth headset can enter the near-field voice mode and start automatically collecting and saving audio signals, which can improve the accuracy of the Bluetooth headset entering the near-field voice mode.

[0026] In combination with the first aspect, in a possible implementation, the first electronic device includes a speaker; before the first electronic device obtains the second audio signal output by the target sound-emitting object through the microphone array, the method further includes: the first electronic device emits a first ultrasonic signal through the speaker; the first electronic device receives the reflected second ultrasonic signal; the first electronic device determines a second distance between the first electronic device and the target sound-emitting object and the vibration frequency of the target sound-emitting part of the target sound-emitting object based on the first ultrasonic signal and the second ultrasonic signal; when the second distance is less than a second preset distance, and the vibration frequency of the target sound-emitting part of the target sound-emitting object is within a first range, the first electronic device obtains the second audio signal output by the target sound-emitting object through the microphone array.

[0027] Exemplarily, the first range can be between 20hz - 40hz.

[0028] Optionally, the second preset distance can be less than the first preset distance. Optionally, the second preset distance can also be equal to the first preset distance.

[0029] Optionally, when the target sound-emitting object of the target sound-emitting object emits sound, it can emit sound at a certain vibration frequency. The first electronic device can obtain the vibration frequency of the target sound-emitting part of the target sound-emitting object through the first ultrasonic signal and the second ultrasonic signal, and determine whether the target sound-emitting object is emitting sound based on the obtained vibration frequency of the target sound-emitting object. When the vibration frequency of the target sound-emitting object is within the first range, it can be considered that the target sound-emitting object is emitting sound.

[0030] In this way, when the first electronic device enters the near-field voice mode, it can confirm whether the first electronic device and the target sound-emitting object are close enough, and whether the target sound-emitting part of the target sound-emitting object is making a sound. This can improve the accuracy of the first electronic device entering the near-field voice mode.

[0031] Combined with the first aspect, in a possible implementation, after the first electronic device saves the second audio signal in the first application, the method further includes: when it is determined based on the first ultrasonic signal and the second ultrasonic signal that the second distance is greater than the second preset distance and / or the vibration frequency of the target sound-emitting part of the target sound-emitting object is not within the first range for a continuous first preset duration, the first electronic device stops acquiring the audio signal through the microphone array.

[0032] After the first electronic device enters the near-field voice mode, the first electronic device also needs to continuously monitor whether the first electronic device and the target sound-emitting object are close enough, and whether the target sound-emitting part of the target sound-emitting object is making a sound. When it is monitored that the distance between the first electronic device and the target sound-emitting object exceeds the second preset distance, and / or the target sound-emitting part of the target sound-emitting object does not make a sound, the first electronic device can automatically exit the near-field voice mode and stop collecting the audio.

[0033] Combined with the first aspect, in a possible implementation, the microphone array includes a first microphone and a second microphone, and the positions of the first microphone and the second microphone on the first electronic device are different; the first electronic device acquires the first audio signal output by the target sound-emitting object through the microphone array, specifically including: the first electronic device acquires the first audio signal output by the target sound-emitting object through the first microphone and the second microphone respectively; the first electronic device determines the first distance between the first electronic device and the target sound-emitting object based on the first audio signal, specifically including: the first electronic device obtains the first energy value of the first audio signal acquired by the first microphone and the second energy value of the first audio signal acquired by the second microphone; the first electronic device determines the first distance between the first electronic device and the target sound-emitting object based on the difference between the first energy value and the second energy value; and / or, the first electronic device obtains the first time of the first audio signal acquired by the first microphone and the second time of the first audio signal acquired by the second microphone; the first electronic device determines the first distance between the first electronic device and the target sound-emitting object based on the difference between the first time and the second time.

[0034] In this way, the electronic device can determine the distance between the first electronic device and the target sound-emitting object based on the difference in the energy values of the same audio signal acquired by multiple microphones in the microphone array and / or the difference in the times of receiving the same audio signal.

[0035] In combination with the first aspect, in a possible implementation manner, the microphone array includes a first microphone and a second microphone, and the positions of the first microphone and the second microphone on the first electronic device are different. The speaker includes a first speaker and a second speaker. The first speaker is located near the first microphone, and the second speaker is located near the second microphone. The first electronic device emits a first ultrasonic signal through the speaker, specifically including: when the first electronic device determines that the distance between the first microphone and the target sound-emitting object is less than the distance between the second microphone and the target sound-emitting object, the first electronic device emits a first ultrasonic signal through the first speaker, where the first speaker is the speaker closest to the target sound-emitting object.

[0036] In combination with the first aspect, in a possible implementation manner, the first electronic device receives the reflected second ultrasonic signal, specifically including: the first electronic device receives the reflected second ultrasonic signal through the first microphone.

[0037] In this way, the first electronic device can emit an ultrasonic signal through the speaker closest to the target sound-emitting object and receive the reflected ultrasonic signal through the microphone closest to the target sound-emitting object, which can improve the accuracy of identifying the target sound-emitting object.

[0038] In combination with the first aspect, in a possible implementation manner, the first electronic device determines that the distance between the first microphone and the target sound-emitting object is less than the distance between the second microphone and the target sound-emitting object, specifically including: the first electronic device obtains a first energy value of a first audio signal collected by the first microphone and a second energy value of the first audio signal collected by the second microphone; when the first energy value is greater than the second energy value, the first electronic device determines that the distance between the first microphone and the target sound-emitting object is less than the distance between the second microphone and the target sound-emitting object; and / or, the first electronic device obtains a first moment of the first audio signal collected by the first microphone and a second moment of the first audio signal collected by the second microphone; when the first moment is less than the second moment, the first electronic device determines that the distance between the first microphone and the target sound-emitting object is less than the distance between the second microphone and the target sound-emitting object.

[0039] In a second aspect, the present application provides an electronic device, which includes a microphone array, a memory, and a processor. Among them, the microphone array, the memory, and the processor are coupled. The memory is used to store a computer program. When the processor executes and calls the computer program, the electronic device executes an audio acquisition method provided in any possible implementation manner in the first aspect.

[0040] In a third aspect, the present application provides a computer-readable storage medium, including instructions. When the instructions run on an electronic device, the electronic device executes an audio acquisition method provided in any possible implementation manner in the first aspect.

[0041] In a fourth aspect, the present application provides a computer program product containing instructions. When the computer program product runs on an electronic device, the electronic device is caused to execute an audio acquisition method provided in any possible implementation manner of the first aspect above.

[0042] In a fifth aspect, the present application provides a chip system. The chip system includes one or more processors, and the processors are used to call computer instructions to cause the electronic device to execute an audio acquisition method provided in any possible implementation manner of the first aspect above.

[0043] For the description of the beneficial effects of the second aspect to the fifth aspect, reference may be made to the description of the beneficial effects in the first aspect, and the present application will not elaborate herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figures 1A - 1H Shows a UI diagram of an application in the electronic device 100 for acquiring audio;

[0045] Figure 2 Shows a schematic structural diagram of the electronic device 100;

[0046] Figure 3 Is a software structural block diagram of the electronic device 100 according to an embodiment of the present invention;

[0047] Figures 4A - 4N Shows a schematic diagram of a group of electronic devices 100 automatically acquiring audio when near-field voice is recognized;

[0048] Figure 5 Is a schematic diagram of functional modules of an audio acquisition method provided by the present application;

[0049] Figure 6 Is a schematic diagram of the method flow of an audio acquisition method provided by the present application;

[0050] Figures 7A - 7C A schematic diagram of the setting positions of microphones and speakers on a group of electronic devices 100;

[0051] Figures 7D - 7F Shows a schematic diagram of another group of electronic devices 100 automatically acquiring audio when near-field voice is recognized;

[0052] Figure 8 Is a schematic diagram of functional modules of an audio acquisition method provided by the present application;

[0053] Figure 9 Is a schematic diagram of the method flow of another audio acquisition method provided by the present application;

[0054] Figure 10Schematic flowchart of another audio acquisition method provided by this application. Detailed implementation manners

[0055] Hereinafter, the technical solutions in the embodiments of this application will be described clearly and in detail with reference to the accompanying drawings. Among them, in the description of the embodiments of this application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B. The "and / or" in the text is only an association relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "a plurality of" means two or more than two.

[0056] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of this application, unless otherwise specified, the meaning of "a plurality of" is two or more than two.

[0057] The term "user interface (UI)" in the following embodiments of this application is a media interface for interaction and information exchange between an application program or an operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The common manifestation form of the user interface is the graphical user interface (GUI), which refers to the user interface related to computer operations displayed in a graphical manner. It may be visible interface elements such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, Widgets, etc. displayed on the display screen of the wearable device.

[0058] The user can operate within an application program that supports the audio recording function on the electronic device to control the application program to acquire audio. In some embodiments, the application program can acquire audio and save the audio or convert the audio into text. In other embodiments, the application program can acquire audio and send the audio to other devices.

[0059] Figures 1A - 1D Shows a UI diagram of an application program in an electronic device 100 for acquiring audio.

[0060] Exemplarily, the application program can be a memo application.

[0061] Figure 1AThe desktop of the electronic device 100 is shown. Application icons of multiple applications are displayed on the desktop of the electronic device 100. For example, weather application icon, stock application icon, calculator application icon, settings application icon, mail application icon, music application icon, video application icon, browser application icon, map application icon, gallery application icon, memo application icon, voice assistant application icon, beauty application icon, etc. A page indicator is also displayed below the multiple application icons to indicate the total number of pages on the desktop and the positional relationship between the currently displayed page and other pages. For example, the desktop may include three pages, and the white dot in the page indicator may indicate that the currently displayed page is the rightmost one among the three pages. Further optionally, there are multiple tray icons (such as dial application icon, message application icon, contacts application icon, camera application icon) below the page indicator, and the tray icons remain displayed during page switching. Optionally, a status bar is displayed in the upper part of the desktop. The status bar may include: one or more signal strength indicators of mobile communication signals (also known as cellular signals), battery status indicator, time indicator, etc.

[0062] Exemplarily, as Figure 1A shown, the electronic device 100 may receive an input operation (such as a click) from the user on the memo application icon on the desktop. In response to the user's input operation, the electronic device 100 may display Figure 1B the user interface 1100 as shown. The user interface 1100 is the main interface of the memo application provided in the embodiments of the present application.

[0063] As Figure 1B shown, the user interface 1100 may include historical note information, and the current user has not created a note. The user interface 1100 also includes a to-do item option and a add note option. Among them, the user can view one or more items that the user needs to process within a period of time through the to-do item option. The user can also create a new note through the add note option.

[0064] Exemplarily, as Figure 1B shown, the electronic device 100 may receive an input operation (such as a click) from the user on the add note option. In response to the user's input operation, the electronic device 100 may display Figure 1C the user interface 1200 as shown.

[0065] As Figure 1CAs shown, the user can edit text in the user interface 1200, which includes multiple editing options, such as a view list option, a set style option, an insert picture option, a recording option, a handwriting option, etc. The user can also use the recording option to enable the memo application to collect the user's audio. In one possible implementation, the memo application can directly save the user's audio. In other possible implementations, the memo application can convert the user's audio into text and save the text information.

[0066] Exemplarily, as Figure 1C shown, the electronic device 100 can receive an input operation (such as a click) from the user for the recording option in the user interface 1200. In response to the user's input operation, the memo application can display Figure 1D the user interface 1300 shown and collect the audio output by the user.

[0067] However, when the user needs to use the recording function in the memo application, the user needs to follow the steps in Figures 1A - 1C sequence to turn on the recording function in the memo application and collect the audio output by the user through the memo application, which is cumbersome for the user. If the user does not often use the memo application, the user may not be able to find the location of the recording option in time, and the user experience is also not good.

[0068] Figures 1E - 1H shows a UI diagram of an application in another electronic device 100 for collecting audio.

[0069] Exemplarily, the application can be an instant messaging application.

[0070] Exemplarily, as Figure 1E shown, the electronic device 100 can receive an input operation (such as a click) from the user for the icon of the instant messaging application on the desktop. In response to the user's input operation, the electronic device 100 can display Figure 1F the user interface 1400 shown. The user interface 1400 is the main interface of the instant messaging application provided in the embodiment of the present application. The user interface 1400 includes dialog boxes of multiple contacts. For example, the dialog box of contact "Lisa", the dialog box of contact "Alan", the dialog box of contact "Henry", the dialog box of contact "Lucy", etc. The electronic device 100 can receive the user's operation to display the real-time chat interface of a certain contact.

[0071] Exemplarily, as Figure 1F shown, the electronic device 100 can receive an input operation (such as a click) from the user for the dialog box of contact "Lisa" in the user interface 1400. In response to the user's input operation, the electronic device 100 can display Figure 1GThe user interface 1500 shown. The user interface 1500 is a real-time chat interface for the contact "Lisa".

[0072] As Figure 1G shown, the user interface 1500 includes multiple chat records. The user interface 1500 also includes a voice input option 1501, an emoji input option 1502, a multi-functional option 1503, and a keyboard switching option 1504. Among them, the user can input the audio content entered by the user through the voice input option 1501 and send the audio content entered by the user to the contact "Lisa". The user can select one or more emojis through the emoji input option 1502 and send them to the contact "Lisa". The user can send pictures, send location information, initiate a voice / video call, send files, etc. to the contact "Lisa" through the multi-functional option 1503. The user can display a keyboard input box on the user interface 1500 through the keyboard switching option 1504.

[0073] Exemplarily, as Figure 1G shown, the electronic device 100 can receive an input operation (such as a long press) of the user for the voice input option 1501 in the user interface 1500. In response to the user's input operation, the electronic device 100 can display Figure 1H the options 4205, 4206, the prompt message 4207, and the voice collection area 4208 shown. The prompt message 4207 includes the text "Release to send" to prompt the user on how to send audio. Among them, the user needs to long press the voice collection area 4208 for the electronic device 100 to collect voice. When the user does not long press the voice collection area 4208, the electronic device 100 stops collecting voice and sends the previously collected voice to the electronic device 200 (not shown in the figure). The user can move the finger towards the option 4206 so that the electronic device 100 converts the previously collected voice into text and then sends it to the electronic device 200. The user can move the finger towards the option 4205 so that the electronic device 100 stops collecting and stops sending voice messages to the electronic device 200.

[0074] Similarly, if the user needs to use the audio collection function within the instant messaging application, the user needs to follow the steps in Figures 1E - 1H sequence to turn on the audio collection function within the instant messaging application and collect the audio output by the user through the instant messaging application, and the user operation is cumbersome. If the user does not often use the instant messaging application, the user may not be able to find the location of the voice input option 1501 in time, and the user experience is also not good.

[0075] In other embodiments, the user can also wake up the memo application to collect audio or wake up the instant messaging application to collect audio through the operation method of voice commands. However, the situation of accidental touch is relatively common in the voice wake-up method.

[0076] In other embodiments, the user can also use a proximity sensor to detect whether the user is approaching. When the user approaches the device, the electronic device 100 automatically controls the memo application to collect audio or controls the instant messaging application to collect audio. However, distance detection requires an additional proximity device and can only detect whether the user is approaching. It cannot determine whether the user is speaking, and there are also many cases of accidental touches.

[0077] Based on the above analysis, the present application provides an audio collection method, which includes the following steps;

[0078] Step 1: The electronic device 100 confirms whether there is near-field voice currently. In the case of the existence of near-field voice, the electronic device 100 can confirm that the user needs to use the audio collection function of the electronic device 100.

[0079] Among them, the near-field voice can refer to an audio signal emitted by a target sound-generating object within a first preset distance (such as 1 meter) from the electronic device 100.

[0080] Optionally, a plurality of microphones are pre-installed on the electronic device 100. There are obvious differences in the arrival time and energy value of the same audio signal emitted by the same target sound-generating object at the plurality of microphones. It is possible to determine whether the audio signal is near-field voice based on the difference in the arrival time and energy value of the audio signal at the plurality of microphones. For the introduction of near-field voice, reference can be made to Figure 6 the detailed description in the embodiments. The present application will not elaborate here.

[0081] In some embodiments, the electronic device 100 can further confirm whether the target sound-generating part is recognized. After recognizing the target sound-generating part, the electronic device 100 then determines that the user needs to use the audio collection function of the electronic device 100. This can improve the accuracy of confirming that the user uses the audio collection function of the electronic device 100.

[0082] In a possible implementation manner, the electronic device 100 can send an ultrasonic signal and confirm whether the target sound-generating part is recognized through the reflected ultrasonic signal. Because if the user is speaking, when the ultrasonic signal hits the user's mouth, the movement of the mouth will cause the frequency and / or amplitude of the ultrasonic signal to change. The electronic device 100 can recognize the target sound-generating part based on the characteristics of the reflected ultrasonic signal.

[0083] Not limited to ultrasonic signals, the electronic device 100 can also determine the target sound-generating part in other ways, and the present application does not limit this.

[0084] Step 2: The electronic device 100 collects audio.

[0085] The electronic device 100 can save the collected audio locally, or the electronic device 100 can convert the collected audio into text and save the text, or the electronic device 100 can send the collected audio to other electronic devices, or the electronic device 100 can convert the collected audio into text and then the electronic device 100 can send the text to other electronic devices.

[0086] Through this method, when the electronic device 100 recognizes near-field audio, it can automatically collect the audio and process the audio. There is no need for the user to manually turn on the audio collection function of the application to control the application to collect audio, which improves the audio collection efficiency, reduces user operations, and enhances the user experience.

[0087] The following introduces the hardware structure of an electronic device 100 provided in the embodiments of the present application.

[0088] Figure 2 The schematic diagram of the structure of the electronic device 100 is shown.

[0089] The following takes the electronic device 100 as an example to specifically illustrate the embodiments. It should be understood that Figure 2 The shown electronic device 100 is only an example, and the electronic device 100 can have more or fewer components than those shown Figure 2 in the figure, two or more components can be combined, or different component configurations can be had. Figure 2 The various components shown in the figure can be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application specific integrated circuits.

[0090] The electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0091] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0092] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0093] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.

[0094] A memory can also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0095] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0096] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple groups of I2C buses. The processor 110 can be respectively coupled to the touch sensor 180K, the charger, the flash, the camera 193, etc. through different I2C bus interfaces. For example: The processor 110 can be coupled to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface to implement the touch function of the electronic device 100.

[0097] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple groups of I2S buses. The processor 110 can be coupled to the audio module 170 through the I2S bus to implement communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit an audio signal to the wireless communication module 160 through the I2S interface to implement the function of answering a call through a Bluetooth headset.

[0098] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled through the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 through the PCM interface to implement the function of answering calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0099] The UART interface is a general-purpose serial data bus for asynchronous communication. This bus can be a two-way communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 through the UART interface to implement the Bluetooth function. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 through the UART interface to implement the function of playing music through a Bluetooth headset.

[0100] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 110 and the camera 193 communicate through the CSI interface to implement the shooting function of the electronic device 100. The processor 110 and the display screen 194 communicate through the DSI interface to implement the display function of the electronic device 100.

[0101] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0102] The USB interface 130 is an interface that conforms to the USB standard specification, and can specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used for data transmission between the electronic device 100 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as AR devices, etc.

[0103] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of the present invention are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0104] The charging management module 140 is configured to receive a charging input from a charger. The charger may be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 may receive the charging input from the wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 may receive the wireless charging input through the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 may also supply power to the electronic device through the power management module 141.

[0105] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives the inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, the wireless communication module 160, etc. The power management module 141 may also be used to monitor parameters such as the battery capacity, the number of battery charge cycles, and the battery health status (leakage, impedance). In some other embodiments, the power management module 141 may also be disposed in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 may also be disposed in the same device.

[0106] The wireless communication function of the electronic device 100 may be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.

[0107] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 may be used to cover a single or multiple communication frequency bands. Different antennas may also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 may be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna may be used in combination with a tuning switch.

[0108] The mobile communication module 150 may provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 may receive electromagnetic waves through the antenna 1, filter, amplify, and perform other processing on the received electromagnetic waves, and then transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 may also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 150 may be provided in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be provided in the same device.

[0109] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, receiver 170B, etc.), or displays an image or video through the display screen 194. In some embodiments, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 110 and be provided in the same device as the mobile communication module 150 or other functional modules.

[0110] The wireless communication module 160 may provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency-modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 may also receive signals to be sent from the processor 110, frequency-modulate them, amplify them, and convert them into electromagnetic waves through the antenna 2 for radiation.

[0111] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, such that electronic device 100 can communicate with a network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).

[0112] Electronic device 100 implements a display function through a GPU, display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, and is connected to display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0113] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD). The display screen panel can also adopt an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniled, a microled, a micro-oled, a quantum dot light-emitting diode (QLED), etc. to manufacture. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0114] The electronic device 100 can implement the shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, an application processor, etc.

[0115] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera photosensitive element. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise and brightness of the image. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0116] The camera 193 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0117] The digital signal processor is used to process digital signals. In addition to being able to process digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0118] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0119] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as: image recognition, face recognition, voice recognition, text understanding, etc.

[0120] The external memory interface 120 can be used to connect to an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.

[0121] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.). The data storage area can store data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0122] The electronic device 100 can implement audio functions through the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and the application processor, etc. For example, music playback, recording, etc.

[0123] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.

[0124] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or a hands-free call through the speaker 170A.

[0125] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the voice can be listened to by placing the receiver 170B close to the human ear.

[0126] The microphone 170C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by placing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In some other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and implement a directional recording function, etc.

[0127] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0128] The pressure sensor 180A is used to sense the pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be set on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can be a parallel plate including at least two conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 194, the electronic device 100 detects the touch operation intensity according to the pressure sensor 180A. The electronic device 100 can also calculate the touch position according to the detection signal of the pressure sensor 180A. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities can correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, an instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, an instruction to create a new short message is executed.

[0129] The gyro sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyro sensor 180B. The gyro sensor 180B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyro sensor 180B detects the angle of the electronic device 100 shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the electronic device 100 through reverse movement to achieve anti-shake. The gyro sensor 180B can also be used for navigation and somatosensory game scenes.

[0130] The air pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates the altitude through the air pressure value measured by the air pressure sensor 180C to assist positioning and navigation.

[0131] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip leather case. In some embodiments, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip cover according to the magnetic sensor 180D. Then, according to the detected opening and closing state of the leather case or the opening and closing state of the flip cover, the flip cover can be automatically unlocked.

[0132] The acceleration sensor 180E can detect the magnitude of the acceleration of the electronic device 100 in all directions (generally three axes). When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.

[0133] A distance sensor 180F is used to measure distance. The electronic device 100 can measure distance through infrared or laser. In some embodiments, when shooting a scene, the electronic device 100 can use the distance sensor 180F to measure distance to achieve fast focusing.

[0134] The proximity light sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The light-emitting diode may be an infrared light-emitting diode. The electronic device 100 emits infrared light outward through the light-emitting diode. The electronic device 100 uses the photodiode to detect the infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 can use the proximity light sensor 180G to detect that the user holds the electronic device 100 close to the ear for a call, so as to automatically turn off the screen to achieve the purpose of power saving. The proximity light sensor 180G can also be used for the holster mode and automatic unlocking and locking of the pocket mode.

[0135] The ambient light sensor 180L is used to sense the ambient light brightness. The electronic device 100 can adaptively adjust the brightness of the display screen 194 according to the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also cooperate with the proximity light sensor 180G to detect whether the electronic device 100 is in the pocket to prevent accidental touch.

[0136] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application locks, fingerprint taking pictures, fingerprint answering calls, etc.

[0137] The temperature sensor 180J is used to detect temperature. In some embodiments, the electronic device 100 uses the temperature detected by the temperature sensor 180J to execute a temperature processing strategy. For example, when the temperature reported by the temperature sensor 180J exceeds a threshold, the electronic device 100 reduces the performance of the processor near the temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the electronic device 100 heats the battery 142 to avoid abnormal shutdown of the electronic device 100 caused by low temperature. In other some embodiments, when the temperature is lower than yet another threshold, the electronic device 100 boosts the output voltage of the battery 142 to avoid abnormal shutdown caused by low temperature.

[0138] The touch sensor 180K, also known as the "touch panel". The touch sensor 180K can be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as the "touch control screen". The touch sensor 180K is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a different position from that of the display screen 194.

[0139] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire vibration signals of the vibrating bone mass of the human vocal part. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure pulsation signals. In some embodiments, the bone conduction sensor 180M can also be disposed in the earphone to form a bone conduction earphone. The audio module 170 can parse out voice signals based on the vibration signals of the vibrating bone mass of the vocal part acquired by the bone conduction sensor 180M to implement the voice function. The application processor can parse out heart rate information based on the blood pressure pulsation signals acquired by the bone conduction sensor 180M to implement the heart rate detection function.

[0140] The keys 190 include a power-on key, volume keys, etc. The keys 190 can be mechanical keys. They can also be touch keys. The electronic device 100 can receive key inputs to generate key signal inputs related to the user settings and function controls of the electronic device 100.

[0141] The motor 191 can generate vibration prompts. The motor 191 can be used for incoming call vibration prompts and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, playing audio, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 194, the motor 191 can also correspond to different vibration feedback effects. Different application scenarios (such as time reminder, receiving information, alarm clock, game, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0142] The indicator 192 can be an indicator light and can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc.

[0143] The SIM card interface 195 is used to connect the SIM card.

[0144] The software system of the electronic device 100 may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present invention, taking the Android system with a layered architecture as an example, the software structure of the electronic device 100 is exemplarily described.

[0145] Figure 3 It is a block diagram of the software structure of the electronic device 100 in the embodiments of the present invention.

[0146] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely the application layer, the application framework layer, Android runtime and system libraries, and the kernel layer.

[0147] The application layer may include a series of application packages.

[0148] As Figure 3 shown, the application packages may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.

[0149] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.

[0150] As Figure 3 shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc.

[0151] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.

[0152] The content provider is used to store and obtain data, and make this data accessible to applications. The data may include video, image, audio, dialed and answered calls, browsing history and bookmarks, phone book, etc.

[0153] The view system includes visible controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a short message notification icon may include a view for displaying text and a view for displaying pictures.

[0154] The phone manager is used to provide the communication function of the electronic device 100. For example, the management of call status (including connection, hanging up, etc.).

[0155] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, and so on.

[0156] The notification manager enables applications to display notification information in the status bar. It can be used to convey informational messages, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that a download is complete, a message reminder, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as a notification of a background-running application, or a notification that appears on the screen in the form of a dialogue window. For example, it can prompt text information in the status bar, emit a prompt sound, vibrate the electronic device, blink the indicator light, etc.

[0157] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for the scheduling and management of the Android system.

[0158] The core libraries contain two parts: one is the functional functions that the Java language needs to call, and the other is the core libraries of Android.

[0159] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.

[0160] The system libraries can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing libraries (such as: OpenGL ES), 2D graphics engines (such as: SGL), etc.

[0161] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.

[0162] The media libraries support the playback and recording of various common audio and video formats, as well as static image files, etc. The media libraries can support multiple audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0163] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.

[0164] The 2D graphics engine is the drawing engine for 2D drawing.

[0165] The kernel layer is the layer between hardware and software. The kernel layer includes at least a display driver, a camera driver, an audio driver, and a sensor driver.

[0166] Next, the application scenario of an audio acquisition method provided by this application will be introduced in combination with the UI.

[0167] The electronic device 100 can automatically determine whether there is near-field voice currently. After determining that there is near-field voice, the electronic device 100 can automatically start the audio acquisition function of the first application and acquire audio. In a possible implementation, the electronic device 100 can save the audio acquired by the first application within the first application, or convert the audio acquired by the first application into text and save the text within the first application. In other possible implementations, the electronic device 100 can send the audio acquired by the first application to the electronic device 200 that has established a communication connection with it.

[0168] In some embodiments, when the electronic device 100 is not connected to a Bluetooth headset, the electronic device 100 can acquire audio through the microphone on the electronic device 100. The electronic device 100 can also play audio through the speaker on the electronic device 100.

[0169] Figures 4A - 4N Shows a schematic diagram of a group of electronic devices 100 recognizing near-field voice and automatically acquiring audio.

[0170] Among them, Figures 4A - 4J Shows a schematic diagram of an electronic device 100 recognizing near-field voice, automatically acquiring audio, and saving the audio.

[0171] In some embodiments, the electronic device 100 can save the audio acquired by the first application within the first application, or convert the audio acquired by the first application into text and save the text within the first application.

[0172] Exemplarily, the first application can be a memo application.

[0173] Exemplarily, as Figure 4A shown, the user can pick up the electronic device 100 and actively bring it close to the electronic device 100. The user can output voice. When the electronic device 100 recognizes that the current audio is near-field voice, the electronic device 100 can automatically open the memo application and control the memo application to start acquiring audio.

[0174] Optionally, before the electronic device 100 recognizes that the current audio is near-field voice, the electronic device 100 can be displayed on any user interface. For example, as Figure 4B shown, the electronic device 100 can display the desktop.

[0175] In response to recognizing near-field voice, the electronic device 100 can automatically launch the memo application and display Figure 4C the user interface 4100 shown.

[0176] As Figure 4B shown, a prompt message 4101 is displayed on the user interface 4100. The prompt message 4101 includes the text "Near-field voice mode", and this prompt message 4101 is used to prompt the user that the near-field voice mode has been started currently.

[0177] In some embodiments, after the electronic device 100 recognizes that the current audio is near-field voice and before the electronic device 100 automatically launches the memo application, the electronic device 100 can further confirm whether the target sound source is recognized. After recognizing the target sound source, the electronic device 100 then launches the memo application. This can improve the accuracy of confirming that the user uses the audio acquisition function of the electronic device 100.

[0178] In a possible implementation, the electronic device 100 can send an ultrasonic signal and confirm whether the target sound source is recognized through the reflected ultrasonic signal. Because if the user is speaking, when the ultrasonic signal hits the user's mouth, the movement of the mouth will cause the frequency and / or amplitude of the ultrasonic signal to change, and the electronic device 100 can recognize the target sound source based on the characteristics of the reflected ultrasonic signal.

[0179] Not limited to ultrasonic signals, the electronic device 100 can also determine the target sound source in other ways, and this application does not limit this.

[0180] In some embodiments, after the electronic device 100 recognizes that the current audio is near-field voice, the electronic device 100 may also not display Figure 4C the user interface 4100 shown, directly control the memo application to start collecting audio, and display Figure 4D the user interface 4200 shown.

[0181] As Figure 4D shown, the user interface 4200 can display the process of audio collection, such as a waveform diagram, audio duration, and text information corresponding to the audio, etc. In this way, the user can view whether the audio collected by the memo application is the same as the audio output by the user. When the audio collected by the memo application is different from the audio output by the user, the user can choose to re-enter the audio or modify the part that is different from the audio output by the user, avoiding the occurrence of errors in the audio collection by the memo application.

[0182] In some embodiments, after the electronic device 100 does not recognize near-field voice, for example, after the user stops outputting audio, the electronic device 100 can control the memo application to stop collecting audio.

[0183] After the user stops outputting audio, the memo application can save the collected audio within the memo application.

[0184] In a possible implementation, the memo application can directly save the collected audio within the memo application and automatically generate a voice note.

[0185] After the memo application generates a voice note, the electronic device 100 can display Figure 4E the user interface 4300 as shown. The user interface 4300 includes the name of the voice note, the generation date of the voice note, the duration of the voice note, etc. For example, the name of the voice note can be "Thoughts Memo". Optionally, the name of the voice note can be automatically extracted and generated by the memo application based on the audio content. The generation date of the voice note can be "December 18, 2023". The duration of the voice note can be 1 minute and 26 seconds.

[0186] In some embodiments, after the memo application generates a voice note, the electronic device 100 can display Figure 4F the user interface 4400 as shown. The user interface 4400 not only includes the name of the voice note, the generation date of the voice note, the duration of the voice note, but also includes a note type identifier 4401. The note type identifier 4401 is used to indicate the note type, and the note type can include but is not limited to text notes, voice notes, etc. Optionally, the note type identifier 4401 can also be automatically extracted and generated by the memo application based on the audio content.

[0187] In other possible implementations, the memo application can convert the collected audio into text, automatically generate a text note, and save the text note within the memo application.

[0188] In a possible implementation, after the memo application generates a text note, the electronic device 100 can display Figure 4E the user interface 4300 as shown.

[0189] In other possible implementations, after the memo application generates a text note, the electronic device 100 can display Figure 4G the user interface 4500 as shown. The user interface 4500 not only includes the name of the voice note, the generation date of the voice note, the duration of the voice note, but also includes a note type identifier 4501. The note type identifier 4501 is used to indicate the note type, and the note type can include but is not limited to text notes, voice notes, etc. Optionally, the note type identifier 4501 can also be automatically extracted and generated by the memo application based on the audio content.

[0190] After the electronic device 100 fails to recognize near-field voice, for example, after the user stops outputting audio, the electronic device 100 can control the memo application to stop collecting audio. After the memo application stops collecting audio, the memo application can continue to display Figure 4E or Figure 4F or Figure 4G the user interface shown. Alternatively, after the memo application stops collecting audio, the memo application can continue to display the user interface displayed before collecting audio, such as displaying Figure 4B the desktop shown. Alternatively, after the memo application stops collecting audio, the memo application can also display the main interface of the memo application, such as Figure 4C the user interface 4100 shown. After the memo application stops collecting audio, the memo application can also display other user interfaces, which are not limited in this application.

[0191] In some embodiments, after the memo application generates a voice note or a text note, the electronic device 100 can also receive a user operation to view the voice note or the text note.

[0192] Exemplarily, as Figure 4H shown, the electronic device 100 can receive a user's input operation (such as a click) on the "Mind Note" option in the user interface 4300. In response to the user's input operation, the electronic device 100 can display Figure 4I the user interface 4600 as mentioned. The user interface 4600 can include the name of the note, the creation time of the note, the content of the note, etc. Among them, the name of the note can be "Mind Memo", the creation time of the note can be "December 18, 2023", and the content of the note can be "Today, the spring breeze is gentle on the face and the flowers are blooming all over the garden. Time flies by, and the years pass in a hurry. However, there is something on my mind and I make a memo here."

[0193] Exemplarily, as Figure 4H shown, the electronic device 100 can receive a user's input operation (such as a click) on the "Mind Note" option in the user interface 4300. In response to the user's input operation, the electronic device 100 can also display Figure 4J the user interface 4700 as mentioned. The user interface 4700 can include the name of the note, the creation time of the note, and the duration of the note, etc. Among them, the name of the note can be "Mind Memo", the creation time of the note can be "December 18, 2023", and the duration of the note can be 1 minute and 26 seconds. The user interface 4700 also includes a voice note play option, a voice note delete option, and more options, etc. The user can select to share the voice note to other devices or other contact friends in the more options.

[0194] Through the above method, when the electronic device 100 recognizes near-field voice, it can automatically activate the memo application and control the memo application to collect audio, and save the audio. There is no need for the user to follow the Figures 1A - 1C steps to activate the recording function in the memo application and collect the audio output by the user through the memo application. This saves the user's operation and improves the user experience.

[0195] Figures 4K - 4N FIG. shows a schematic diagram of an electronic device 100 recognizing near-field voice, automatically collecting audio, and sending the audio to an electronic device 200.

[0196] In some embodiments, the electronic device 100 may send the audio collected by the first application to the electronic device 200, or convert the audio collected by the first application into text and then send it to the electronic device 200.

[0197] Exemplarily, the first application may be an instant messaging application.

[0198] Reference may be made to the description in the Figure 4A embodiment. The user can pick up the electronic device 100 and actively bring it close. The user can output voice. When the electronic device 100 recognizes that the current audio is near-field voice, the electronic device 100 can control the instant messaging application to start collecting audio.

[0199] Optionally, before the electronic device 100 recognizes that the current audio is near-field voice, the electronic device 100 can display the chat interface of the contact. For example, as Figure 4K shown, the electronic device 100 can display the chat interface of the contact "Lisa".

[0200] In response to recognizing near-field voice, the electronic device 100 can display Figure 4L the prompt bar 4800 shown in FIG.

[0201] As Figure 4L shown, the prompt bar 4800 displays a prompt message 4801. The prompt message 4801 includes the text "Near-field voice mode", and this prompt message 4801 is used to prompt the user that the near-field voice mode has started.

[0202] In some embodiments, after the electronic device 100 recognizes that the current audio is near-field voice, the electronic device 100 can further confirm whether the target sound-emitting part is recognized. After recognizing the target sound-emitting part, the electronic device 100 then activates the memo application. This can improve the accuracy of confirming that the user uses the audio collection function of the electronic device 100.

[0203] In a possible implementation, the electronic device 100 can send an ultrasonic signal and confirm whether the target sound - emitting part is recognized through the reflected ultrasonic signal. Because if the user is speaking, when the ultrasonic signal hits the user's mouth, the movement of the mouth will cause changes in the frequency and / or amplitude of the ultrasonic signal, and the electronic device 100 can identify the target sound - emitting part based on the characteristics of the reflected ultrasonic signal.

[0204] Not limited to ultrasonic signals, the electronic device 100 can also determine the target sound - emitting part in other ways, and this application does not limit it.

[0205] In some embodiments, after the electronic device 100 recognizes that the current audio is near - field voice, the electronic device 100 may not display the prompt message 4801.

[0206] As Figure 4L shown, the prompt bar 4800 can display the process of audio acquisition, such as waveform diagrams, audio duration, and text information corresponding to the audio, etc. In this way, the user can check whether the audio collected by the instant messaging application is the same as the audio output by the user. When the audio collected by the instant messaging application is different from the audio output by the user, the user can choose to re - record the audio or modify the part that is different from the audio output by the user, to avoid the occurrence of errors in the audio collected by the instant messaging application.

[0207] In some embodiments, after the electronic device 100 fails to recognize near - field voice, for example, after the user stops outputting audio, the electronic device 100 can control the instant messaging application to stop collecting audio.

[0208] After the user stops outputting audio, in a possible implementation, the instant messaging application can send the collected audio to the contact "Lisa" and display Figure 4M the user interface 4900 as shown. The user interface 4900 includes an identifier 4901, and the identifier 4901 is used to indicate the audio message sent by the current user to the contact "Lisa".

[0209] After the user stops outputting audio, in other possible implementations, the instant messaging application can convert the collected audio into text, then send the text to the contact "Lisa", and display Figure 4N the user interface 4910 as shown. The user interface 4910 includes an identifier 4911, and the identifier 4911 is used to indicate the text message sent by the current user to the contact "Lisa".

[0210] After the electronic device 100 fails to recognize near - field voice, for example, after the user stops outputting audio, the electronic device 100 can control the instant messaging application to stop collecting audio. And display Figure 4M the user interface 4900 as shown orFigure 4N The user interface 4910 shown.

[0211] Figure 5 It is a schematic diagram of functional modules of an audio acquisition method provided by this application.

[0212] As Figure 5 shown, the functional modules on the electronic device 100 may include but are not limited to: a voice pickup unit, a central control unit, a near-field mode recognition unit, a near-field voice separation unit, an interaction unit, etc.

[0213] Among them, the voice pickup unit is used to periodically / irregularly pick up audio signals and send the picked-up audio signals to the central control unit.

[0214] The central control unit is used to receive the audio signals sent by the voice pickup unit, process the audio signals, and determine whether there are voice signals.

[0215] The central control unit is also used to send the audio signals to the near-field mode recognition unit after determining that there are voice signals.

[0216] In some embodiments, the central control unit is specifically used to determine the probability of the voice signals existing in the audio signals. When the probability of the voice signals existing in the audio signals is greater than a preset value, the central control unit then sends the audio signals to the near-field mode recognition unit. In this way, noise interference can be avoided.

[0217] The near-field mode recognition unit is used to receive the audio signals sent by the central control unit, process the audio signals, and determine whether there is near-field voice.

[0218] The near-field mode recognition unit is also used to send the audio signals to the near-field voice separation unit when it determines that there is near-field voice in the audio signals.

[0219] The central control unit is also used to control the speaker to emit ultrasonic signals after the near-field mode recognition unit determines that there is near-field voice in the audio signals, so as to determine whether the target sound-emitting part is recognized, and improve the accuracy of the electronic device 100 entering the near-field voice mode. This application can emit ultrasonic signals through the speaker module without adding other devices.

[0220] The voice pickup unit is also used to receive the reflected ultrasonic signals when the central control unit controls the speaker to emit ultrasonic signals, and send the reflected ultrasonic signals to the central control unit.

[0221] The central control unit is further configured to confirm whether a target sound-emitting object is recognized based on the characteristics of the emitted ultrasonic signal and the reflected ultrasonic signal. For how the central control unit recognizes the target sound-emitting object, reference can be made to Figure 6 the description of S603 in the embodiment, which will not be elaborated in this application.

[0222] The central control unit is further configured to, after recognizing the target sound-emitting object, send a message indicating that the target sound-emitting object has been recognized to the near-field mode recognition unit.

[0223] Optionally, Figure 5 Steps 6 to 10 shown in the embodiment may not be executed either, and this application does not make any limitation thereto.

[0224] The near-field mode recognition unit is further configured to, in response to the message indicating that the target sound-emitting object has been recognized sent by the central control unit, send an audio signal to the near-field voice separation unit.

[0225] The near-field voice separation unit is configured to separate near-field voice from the audio signal after receiving the audio signal sent by the near-field mode recognition unit.

[0226] The central control unit is further configured to, after recognizing the target sound-emitting object, send a first instruction to the interaction unit.

[0227] The interaction unit is configured to display an audio acquisition interaction interface after receiving the first instruction sent by the central control unit, so as to prompt the user that the current electronic device 100 is in the near-field audio acquisition mode.

[0228] Optionally, the interaction unit may be a functional module in the first application. The near-field voice separation unit may also be a functional module in the first application.

[0229] Figure 6 It is a schematic flowchart of a method for audio acquisition provided by this application.

[0230] S601. The electronic device 100 periodically / irregularly acquires audio signals.

[0231] Optionally, one or more microphones are preset on the electronic device 100, and the electronic device 100 can periodically / irregularly acquire audio signals through the one or more microphones.

[0232] S602. The electronic device 100 needs to determine whether the audio signal includes near-field voice.

[0233] After the electronic device 100 acquires the audio signal, the electronic device 100 needs to confirm whether the audio signal includes near-field voice.

[0234] In some embodiments, before the electronic device 100 confirms whether the audio signal is near-field voice, the electronic device 100 may first determine that there is a voice signal in the audio signal, and then confirm whether the audio signal includes near-field voice.

[0235] Exemplarily, the electronic device 100 may determine the probability of the voice signal existing in the audio signal. When the probability of the voice signal existing in the audio signal is greater than a first threshold (e.g., 80%), the electronic device 100 then confirms whether the audio signal includes near-field voice. In this way, noise interference can be avoided.

[0236] Among them, the near-field voice may refer to an audio signal emitted by a target sound-emitting object within a first preset distance from the electronic device 100.

[0237] The electronic device 100 includes a microphone array. The microphone array includes at least two microphones, such as a first microphone and a second microphone. The electronic device 100 may determine a first distance between the electronic device 100 and the target sound-emitting object based on the audio signal picked up by the microphone array. When the first distance is less than the first preset distance, or when the first distance is less than the first preset distance for a continuous first duration, the electronic device 100 may determine that near-field voice is recognized.

[0238] Optionally, the microphone array includes at least two microphones. The arrival time and energy of the same audio signal emitted by the same target sound-emitting object at the multiple microphones are different. The first distance between the electronic device 100 and the target sound-emitting object may be determined based on the time difference and energy difference of the same audio signal arriving at at least two microphones, and then it is further determined whether it is near-field voice.

[0239] For example, the microphone array may at least include a first microphone and a second microphone, and the positions of the first microphone and the second microphone on the first electronic device are different. The electronic device 100 may collect a first audio signal through the first microphone and the second microphone. The electronic device 100 may obtain a first energy value of the first audio signal collected by the first microphone and a second energy value of the first audio signal collected by the second microphone. The first electronic device may determine the first distance between the first electronic device and the target sound-emitting object based on the difference between the first energy value and the second energy value, and / or the first electronic device may obtain a first time of the first audio signal collected by the first microphone and a second time of the first audio signal collected by the second microphone. The first electronic device may also determine the first distance between the first electronic device and the target sound-emitting object based on the difference between the first time and the second time.

[0240] Exemplarily, the electronic device 100 can also determine the distance between each microphone and the target sound - emitting object. For example, when the first energy value is greater than the second energy value or the first energy value is continuously less than the second energy value for a certain period of time, the first electronic device can determine that the distance between the first microphone and the target sound - emitting object is less than the distance between the second microphone and the target sound - emitting object, and / or the first electronic device can also obtain the first moment of the first audio signal collected by the first microphone and the second moment of the first audio signal collected by the second microphone. When the first moment is less than the second moment or the first moment is continuously less than the second moment for a certain period of time, the first electronic device determines that the distance between the first microphone and the target sound - emitting object is less than the distance between the second microphone and the target sound - emitting object.

[0241] Exemplarily, multiple microphones pre - installed on the electronic device 100 can be located at different positions. For example, the electronic device 100 is pre - installed with microphone 1, microphone 2, and microphone 3. Microphone 1 is located at the bottom of the electronic device 100, microphone 2 is located at the top of the electronic device 100, and microphone 3 is located on the back of the electronic device 100, for example, at the position where the camera module is located on the back of the electronic device 100.

[0242] In some embodiments, microphone 1 can be referred to as the first microphone, and microphone 2 can be referred to as the first microphone.

[0243] Figure 7A An exemplary schematic diagram of the positions of multiple microphones pre - installed on the electronic device 100 is shown.

[0244] As Figure 7A shown, microphone 1 is located at the bottom of the electronic device 100, microphone 2 is located at the top of the electronic device 100, and microphone 3 is located on the back of the electronic device 100, for example, at the position where the camera module is located on the back of the electronic device 100. In some embodiments, one or more microphones are also used to receive ultrasonic signals to detect the target sound - emitting part.

[0245] In some embodiments, one or more speakers can also be pre - installed on the electronic device 100, and the one or more speakers are used to play audio. In some embodiments, the one or more speakers are also used to emit ultrasonic signals to detect the target sound - emitting part.

[0246] As Figure 7A shown, speaker 1 is located at the bottom of the electronic device 100, and speaker 2 is located at the top of the electronic device 100. In some embodiments, the back of the electronic device 100 can also include speaker 3. For example, speaker 3 can be located at the position where the camera module is located on the back of the electronic device 100. In some embodiments, the back of the electronic device 100 may not include speaker 3.

[0247] It should be noted that it is not limited to 3 microphones and 3 speakers. Figure 7A Only the positions of multiple microphones and multiple speakers on the electronic device 100 are exemplarily illustrated, and this application does not limit this either.

[0248] The multiple microphones are located at different positions on the electronic device 100. The closer the target sound - emitting object is to the microphone, the more energy and the shorter the time of the audio collected by this microphone. The farther the target sound - emitting object is from the microphone, the less energy and the longer the time of the audio collected by this microphone. Then, within the first preset distance, for the same audio signal emitted by the same target sound - emitting object, the arrival times and energy values at the multiple microphones are significantly different. Based on the differences in the arrival times and energy values of the same audio signal at the multiple microphones, it can be determined whether the audio signal is a near - field voice.

[0249] In other embodiments, when the distance between the target sound - emitting object and the electronic device 100 exceeds the first preset distance, that is, when the target sound - emitting object is far from the electronic device 100, the arrival times and energy values of the same audio signal emitted by the same target sound - emitting object at the multiple microphones on the electronic device 100 are not significantly different, and it can be judged that the current audio is a non - near - field voice based on this.

[0250] Exemplarily, when the distance between the target sound - emitting object and the electronic device 100 is within the first preset distance, the target sound - emitting object outputs a first audio. The energy of the first audio received by microphone 1 is a first energy value, the energy of the first audio received by microphone 2 is a second energy value, and the energy of the first audio received by microphone 3 is a third energy value. When the first energy value, the second energy value, and the third energy value are different and meet the preset conditions, the first audio can be considered a near - field voice. Optionally, the preset condition can be that the difference between any two of the first energy value, the second energy value, and the third energy value meets a preset energy difference, for example, is greater than the preset energy difference.

[0251] Exemplarily, when the distance between the target sound - emitting object and the electronic device 100 is within the first preset distance, the target sound - emitting object outputs a first audio. The arrival time of the first audio received by microphone 1 is a first time, the arrival time of the first audio received by microphone 2 is a second time, and the arrival time of the first audio received by microphone 3 is a third time. When the first time, the second time, and the third time are different and meet the preset conditions, the first audio can be considered a near - field voice. Optionally, the preset condition can be that the difference between any two of the first time, the second time, and the third time meets a preset time difference, for example, is greater than the preset time difference.

[0252] It should be noted that the above - mentioned preset conditions can also be other judgment bases, and this application does not limit this.

[0253] Exemplarily, such as Figure 7B As shown, the user can pick up the electronic device 100 and actively bring it close. When the user brings it close to the bottom of the electronic device 100 and outputs audio, the distance relationship between the target sound - emitting part (such as the mouth) and each microphone on the electronic device 100 can be obtained. For example, the distance between the target sound - emitting part and microphone 1 is distance 1, the distance between the target sound - emitting part and microphone 2 is distance 2, and the distance between the target sound - emitting part and microphone 3 is distance 3. Since the user outputs audio close to the bottom of the electronic device 100, distance 1 is less than distance 3 is less than distance 2.

[0254] When the user outputs the first audio, the energy of the first audio received by microphone 1 is the first energy value, the energy of the first audio received by microphone 2 is the second energy value, and the energy of the first audio received by microphone 3 is the third energy value. Since distance 1 is less than distance 3 is less than distance 2, therefore, the first energy value, the second energy value, and the third energy value are different, and the first energy value is greater than the third energy value, and the third energy value is greater than the second energy value. Optionally, when the difference between the first energy value and the third energy value meets the preset energy difference, and the difference between the third energy value and the second energy value meets the preset energy difference, for example, when the difference between the first energy value and the third energy value is greater than the preset energy difference, and the difference between the third energy value and the second energy value is greater than the preset energy difference, then the first audio can be considered as near - field speech.

[0255] When the user outputs the first audio, the moment when microphone 1 receives the first audio is the first moment, the moment when microphone 2 receives the first audio is the second moment, and the moment when microphone 3 receives the first audio is the third moment. Since distance 1 is less than distance 3 is less than distance 2, therefore, the first moment, the second moment, and the third moment are different, and the first moment is less than the third moment, and the third moment is less than the second moment. Optionally, when the difference between the first moment and the third moment meets the preset time difference, and the difference between the third moment and the second moment meets the preset time difference, for example, when the difference between the first moment and the third moment is greater than the preset time difference, and the difference between the third moment and the second moment is greater than the preset time difference, then the first audio can be considered as near - field speech.

[0256] Exemplarily, such as Figure 7CAs shown, the user can pick up the electronic device 100 and actively approach it. When the user approaches the top of the electronic device 100 and outputs audio, the distance relationship between the target sound - emitting part (such as the mouth) and each microphone on the electronic device 100 can be obtained. For example, the distance between the target sound - emitting part and microphone 1 is distance 1, the distance between the target sound - emitting part and microphone 2 is distance 2, and the distance between the target sound - emitting part and microphone 3 is distance 3. Since the user approaches the top of the electronic device 100 to output audio, distance 2 is less than distance 3 is less than distance 1.

[0257] When the user outputs the first audio, the energy of the first audio received by microphone 1 is the first energy value, the energy of the first audio received by microphone 2 is the second energy value, and the energy of the first audio received by microphone 3 is the third energy value. Since distance 2 is less than distance 3 is less than distance 1, therefore, the first energy value, the second energy value, and the third energy value are different, and the second energy value is greater than the third energy value, and the third energy value is greater than the first energy value. Optionally, when the difference between the second energy value and the third energy value meets the preset energy difference, and the difference between the third energy value and the first energy value meets the preset energy difference. For example, when the difference between the second energy value and the third energy value is greater than the preset energy difference, and the difference between the third energy value and the first energy value is greater than the preset energy difference, then the first audio can be considered as near - field speech.

[0258] When the user outputs the first audio, the moment when microphone 1 receives the first audio is the first moment, the moment when microphone 2 receives the first audio is the second moment, and the moment when microphone 3 receives the first audio is the third moment. Since distance 2 is less than distance 3 is less than distance 1, therefore, the first moment, the second moment, and the third moment are different, and the second moment is less than the third moment, and the third moment is less than the first moment. Optionally, when the difference between the second moment and the third moment meets the preset time difference, and the difference between the third moment and the first moment meets the preset time difference. For example, when the difference between the second moment and the third moment is greater than the preset time difference, and the difference between the third moment and the first moment is greater than the preset time difference, then the first audio can be considered as near - field speech.

[0259] Not limited to determining whether the audio is near - field speech based on the above - mentioned method, it is also possible to determine whether the audio is near - field speech based on other methods, and this application does not make any limitations in this regard.

[0260] After determining that the audio signal is near - field speech, the electronic device 100 can determine that the current is the near - field speech mode. The electronic device 100 can control the first application to collect the audio signal and execute S603.

[0261] Optionally, after determining that the audio signal of the continuous first duration is near-field speech, the electronic device 100 may determine that the current is the near-field speech mode, and then the electronic device 100 controls the first application to collect the audio signal and execute S603.

[0262] After determining that the audio signal is not near-field speech, the electronic device 100 may determine that the current is not the near-field speech mode, and the electronic device continues to monitor whether there is near-field speech and execute S601.

[0263] S603. The electronic device 100 needs to determine whether the target sound-producing part is recognized.

[0264] After determining that the audio signal includes near-field speech, before collecting the audio signal, the electronic device 100 needs to determine whether the target sound-producing part is recognized. After recognizing the target sound-producing part, the electronic device 100 then collects the audio signal. In this way, the accuracy of the electronic device 100 entering the near-field speech mode can be improved, and the situation of accidental touch can be further prevented.

[0265] In some embodiments, after determining that the audio signal is near-field speech, the electronic device 100 also needs to confirm whether the target sound-producing part is recognized. Recognizing the target sound-producing part may mean that the distance between the electronic device 100 and the target sound-producing object is within a second preset distance (for example, 10 cm), and the vibration frequency of the target sound-producing part is within a first range. Optionally, the second preset distance may be less than the first preset distance. Optionally, the second preset distance may also be equal to the first preset distance. After determining the target sound-producing object, it is necessary to further determine whether the distance between the target sound-producing part of the target sound-producing object and the electronic device 100 is close enough and whether the target sound-producing part is moving. Only after the distance is close enough and the target sound-producing part is moving, it can be determined that the target sound-producing part is making a sound, and then the electronic device 100 starts to collect the audio, which can further prevent the situation of accidental touch from occurring.

[0266] In some embodiments, after determining that the audio signal is near-field speech, the electronic device 100 may emit a first ultrasonic signal and receive the reflected second ultrasonic signal. The electronic device 100 may determine the second distance between the first electronic device and the target sound-producing object and the vibration frequency of the target sound-producing part of the target sound-producing object based on the first ultrasonic signal and the second ultrasonic signal. When the second distance is less than the second preset distance and the vibration frequency of the target sound-producing part of the target sound-producing object is within the first range, the electronic device 100 may determine that the target sound-producing part is making a sound, and then the electronic device 100 starts to collect the audio, which can further prevent the situation of accidental touch from occurring.

[0267] The first range may be between the first frequency value and the second frequency value. Exemplarily, the first range may be 20 hz - 40 hz.

[0268] Optionally, when it is determined based on the first ultrasonic signal and the second ultrasonic signal that the second distance is greater than the second preset distance and / or the vibration frequency of the target sound - emitting part of the target sound - emitting object is not within the first range during the continuous first preset duration, the electronic device 100 exits the near - field voice mode and stops collecting audio.

[0269] In some embodiments, after determining that the audio signal is near - field voice, the electronic device 100 can emit an ultrasonic signal with preset characteristics, and based on the characteristics of the reflected ultrasonic signal, determine whether the target sound - emitting part is recognized and the distance between the electronic device 100 and the target sound - emitting object. Optionally, the preset characteristics of the emitted ultrasonic signal can include but are not limited to preset frequency, preset amplitude, emission time, etc. The characteristics of the reflected ultrasonic signal can include but are not limited to the frequency and amplitude of the reflected ultrasonic signal, reflection time, etc.

[0270] Because when the target sound - emitting part (such as the mouth) outputs audio, the target sound - emitting part will move with a certain frequency and amplitude. The ultrasonic signal is reflected when hitting the moving target sound - emitting part, resulting in changes in the frequency and amplitude of the ultrasonic signal. For example, the frequency and amplitude of the ultrasonic signal will change.

[0271] The electronic device 100 can determine the distance between the electronic device 100 and the target sound - emitting object and the vibration frequency of the target sound - emitting part of the target sound - emitting object based on the emitted ultrasonic signal and the reflected ultrasonic signal.

[0272] Optionally, the electronic device 100 can determine the distance between the electronic device 100 and the target sound - emitting object based on the difference between the time when the ultrasonic signal is emitted and the time when the reflected ultrasonic signal is received.

[0273] Optionally, the electronic device 100 can determine the vibration frequency of the target sound - emitting part of the target sound - emitting object based on the difference between the frequency and / or amplitude of the emitted ultrasonic signal and the frequency and / or amplitude of the reflected ultrasonic signal.

[0274] Exemplarily, the electronic device 100 can compare the frequency and amplitude of the reflected ultrasonic signal with the frequency and amplitude of the ultrasonic signal emitted by the electronic device 100, determine the vibration frequency of the target sound - emitting part of the target sound - emitting object, and determine whether the target sound - emitting part is recognized based on the vibration frequency of the target sound - emitting part. Exemplarily, when the vibration frequency of the target sound - emitting part is within the first range, it can be determined that the target sound - emitting part is recognized.

[0275] Exemplarily, when the difference between the frequency and amplitude of the reflected ultrasonic signal and the frequency and amplitude of the ultrasonic signal emitted by the electronic device 100 meets a preset value, for example, is greater than the preset value, the electronic device 100 can determine that the target sound-emitting part is recognized.

[0276] If the ultrasonic signal hits a non-moving object, the frequency and amplitude of the ultrasonic signal basically do not change, and the difference between the frequency and amplitude of the reflected ultrasonic signal and the frequency and amplitude of the ultrasonic signal emitted by the electronic device 100 does not meet the preset value, for example, is less than the preset value, and the electronic device 100 can determine that the target sound-emitting part is not recognized.

[0277] In some embodiments, as introduced in S602, a plurality of speakers are pre-installed on the electronic device 100, and the electronic device 100 can emit ultrasonic signals through the speakers to determine whether the target sound-emitting part is recognized.

[0278] Based on the introduction in S602, the electronic device 100 can determine the distance between the target sound-emitting object and each microphone based on the energy value and time difference of the audio signals received by the plurality of microphones.

[0279] After determining the distance between the target sound-emitting object and each microphone, the electronic device 100 can determine the microphone closest to the target sound-emitting object. After determining the microphone closest to the target sound-emitting object, the electronic device 100 can emit ultrasonic signals through the speaker near the microphone, and the microphone receives the reflected ultrasonic signals to determine whether the target sound-emitting part is recognized, where the speaker near the microphone can refer to the speaker closest to the microphone. In this way, the closer the target sound-emitting part is to the speaker and the microphone, the less the signal decays, and the more accurate the detection result is.

[0280] Exemplarily, referring to Figure 7B the description in the embodiment, when the user is close to the bottom of the electronic device 100 and outputs audio, the target sound-emitting part is closest to the microphone 1 and the speaker 1, then the electronic device 100 can emit ultrasonic signals through the speaker 1 and receive the reflected ultrasonic signals through the microphone 1 to confirm whether the target sound-emitting part is recognized.

[0281] Exemplarily, referring to Figure 7C the description in the embodiment, when the user is close to the bottom of the electronic device 100 and outputs audio, the target sound-emitting part is closest to the microphone 2 and the speaker 2, then the electronic device 100 can emit ultrasonic signals through the speaker 2 and receive the reflected ultrasonic signals through the microphone 2 to confirm whether the target sound-emitting part is recognized.

[0282] It is not limited to emitting and receiving ultrasonic signals through the speaker and microphone closest to the target sound - emitting object. It is also possible to simultaneously emit and receive ultrasonic signals based on other multiple sets of speakers and microphones, and confirm whether the target sound - emitting part is recognized through multiple sets of detection results. In this way, the accuracy of recognizing the target sound - emitting part can also be improved.

[0283] It is not limited to confirming whether the target sound - emitting part is recognized through ultrasonic signals. It is also possible to confirm whether the target sound - emitting part is recognized through other means, and this application does not make any limitations in this regard.

[0284] In some embodiments, the electronic device 100 may also not execute S603. After determining the near - field voice in S602, the electronic device 100 may directly execute S604.

[0285] S604: The electronic device 100 collects an audio signal and saves the audio signal in the first application.

[0286] After determining the near - field voice, or after determining the near - field voice and recognizing the target sound - emitting part, the electronic device 100 may collect an audio signal and save the collected audio signal in the first application.

[0287] The electronic device 100 saving the collected audio signal in the first application may include: The electronic device 100 directly saves the audio signal in the first application. Or, the electronic device 100 converts the audio signal into text information and saves the text information in the first application. Exemplarily, the first application may be a memo application. Specifically, reference may be made to Figures 4A - 4J the description in the embodiments.

[0288] The electronic device 100 saving the collected audio signal in the first application may also include: The electronic device 100 sends the audio signal to the electronic device 200 through the first application. Or, the electronic device 100 converts the audio signal into text information and sends the text information to the electronic device 200 through the first application. Exemplarily, the first application may be an instant messaging application. Specifically, reference may be made to Figures 4K - 4N the description in the embodiments.

[0289] In some embodiments, after collecting the audio signal, the electronic device 100 also needs to process the audio signal to filter out the far - field voice signal and obtain the near - field voice signal, so as to avoid the interference of the far - field voice signal.

[0290] Based on the description in S602, it can be known that the electronic device 100 can determine the distance between the electronic device 100 and the target sound - emitting object based on the audio signal collected by the microphone array. For example, the electronic device 100 simultaneously receives the second audio signal sent by the target sound - emitting object and the third audio signal output by other sound - emitting objects. When the electronic device 100 determines that the distance between the electronic device 100 and the target sound - emitting object is less than the first preset distance, and the distance between the electronic device 100 and other sound - emitting objects is greater than the first preset distance, the electronic device 100 can only save the second audio signal and not save the third audio signal. How the electronic device 100 determines the distance from the sound - emitting object can refer to the description in S602, and this application will not elaborate here.

[0291] Exemplarily, based on the description in S602, it can be known that the difference between near - field speech and far - field speech is that when the distance between the target sound - emitting object and the electronic device 100 is within the first preset distance, there are obvious differences in the energy values and time differences of the same audio signal emitted by the same sound - emitting object reaching multiple microphones on the electronic device 100, and based on this, it can be determined that the current audio is near - field speech. When the distance between the target sound - emitting object and the electronic device 100 exceeds the first preset distance, that is, the target sound - emitting object is far from the electronic device 100, there are no obvious differences in the time and energy of the same audio signal emitted by the same target sound - emitting object reaching multiple microphones on the electronic device 100, and based on this, it can be determined that the current audio is far - field speech.

[0292] Exemplarily, obtain the energy values of the same audio signal received by multiple microphones on the electronic device 100, and the difference between any two of the multiple energy values satisfies a preset energy difference, for example, is greater than the preset energy difference, and / or, obtain the time of the same audio signal received by multiple microphones on the electronic device 100, and the difference between any two of the multiple times satisfies a preset time difference, for example, is greater than the preset time difference, then it can be determined that the audio is near - field speech.

[0293] Exemplarily, obtain the energy values of the same audio signal received by multiple microphones on the electronic device 100, and the difference between any two of the multiple energy values does not satisfy the preset energy difference, for example, is less than the preset energy difference, and / or, obtain the time of the same audio signal received by multiple microphones on the electronic device 100, and the difference between any two of the multiple times does not satisfy the preset time difference, for example, is less than the preset time difference, then it can be determined that the audio is far - field speech.

[0294] Based on the above features, near - field speech can be separated from the audio signals collected by the electronic device 100.

[0295] If in the audio signal collected by the electronic device 100, the near-field voice can be expressed as s1(t), the transmission path of the near-field voice is f1(t,n), the far-field voice can be expressed as s2(t), and the transmission path of the far-field voice is f2(t,n). Wherein, t represents time, f1(t,n) represents the transmission path of the near-field voice reaching the nth microphone, and f2(t,n) represents the transmission path of the far-field voice reaching the nth microphone. Then the audio signal collected by the electronic device 100 can be expressed by formula (1).

[0296] y(t,n)= s1(t)*f1(t,n) + s2(t)*f2(t,n) Formula (1)

[0297] As shown in formula (1), y(t,n) represents the audio signal collected by the nth microphone, s1(t) represents the near-field voice collected by the electronic device 100, f1(t,n) represents the transmission path of the near-field voice, s2(t) represents the far-field voice collected by the electronic device 100, and f2(t,n) represents the transmission path of the far-field voice.

[0298] The audio signals received by n microphones on the electronic device 100 can be expressed by formula (2).

[0299] Y(t)=[y(t,1),y(t,2),…,y(t,N)] Formula (2)

[0300] As shown in formula (2), Y(t) represents the audio signals received by n microphones on the electronic device 100, y(t,1) represents the audio signal received by the first microphone on the electronic device 100, y(t,2) represents the audio signal received by the second microphone on the electronic device 100, and y(t,N) represents the audio signal received by the nth microphone on the electronic device 100.

[0301] Performing a frequency-domain transformation on Y(t) can obtain Y(t,f) shown in formula (3).

[0302] Y(t,f)=F(Y(t)) Formula (3)

[0303] As shown in formula (3), Y(t,f) represents the audio signals received by n microphones on the electronic device 100 in the frequency domain, f represents frequency. Y(t) represents the audio signals received by n microphones on the electronic device 100 in the time domain.

[0304] The electronic device 100 can separate the near-field voice from the audio signal collected by the electronic device 100 through the following formula (4).

[0305]

[0306] As shown in formula (4), s1(t) represents the near-field speech separated from the audio signal collected by the electronic device 100, Y(t, f) represents the audio signals received by n microphones on the electronic device 100 in the frequency domain, and W(t, f) represents a solution matrix, which is used to separate the near-field speech from the audio signal collected by the electronic device 100. The solution matrix can represent the difference relationship between the energy values of the same audio emitted by the same sound-emitting object when the audio is near-field speech and reaches multiple microphones on the electronic device 100, and / or the difference relationship between the arrival times of the same audio emitted by the same sound-emitting object and reaching multiple microphones on the electronic device 100. Exemplarily, the difference relationship between the energy values of the same audio emitted by the same sound-emitting object and reaching multiple microphones on the electronic device 100 can be: the energy values of the same audio signal received by multiple microphones on the electronic device 100, and the difference between any two of the multiple energy values satisfies a preset energy difference, for example, is greater than the preset energy difference. The difference relationship between the arrival times of the same audio emitted by the same sound-emitting object and reaching multiple microphones on the electronic device 100 can be: obtaining the arrival times of the same audio signal received by multiple microphones on the electronic device 100, and the difference between any two of the multiple times satisfies a preset time difference, for example, is greater than the preset time difference.

[0307] Optionally, the solution matrix W(t, f) can be obtained by the electronic device 100 through deep learning, and the solution matrix W(t, f) can also be updated periodically / irregularly.

[0308] It should be noted that the above formulas (1) to (4) are only used to explain how to separate the near-field speech from the audio signal in the present application. The near-field speech can also be separated from the audio signal by other means, and the present application does not make any limitations in this regard.

[0309] S605. The electronic device 100 needs to continue to monitor whether the target sound-emitting part is recognized.

[0310] After determining the near-field speech, or after determining the near-field speech and recognizing the target sound-emitting part, the electronic device 100 can collect the audio signal and save the collected audio signal in the first application.

[0311] The electronic device 100 also needs to continue to monitor whether the target sound-emitting part is recognized. In the case where the target sound-emitting part is not recognized, the electronic device 100 exits the near-field speech mode, stops saving the collected audio signal in the first application, and executes S601.

[0312] In the case where the target sound-emitting part is continuously recognized, the electronic device 100 continues to save the collected audio signal in the first application and executes S604.

[0313] In other embodiments, when the electronic device 100 is connected to a Bluetooth headset, the Bluetooth headset can collect audio through the microphone on the Bluetooth headset and send the audio to the electronic device 100 via a Bluetooth connection. The electronic device 100 can also send the audio to the Bluetooth headset via the Bluetooth connection, and the Bluetooth headset then plays the audio sent by the electronic device 100 through the speaker on the Bluetooth headset.

[0314] Figures 7D - 7F FIG. shows a schematic diagram of another group of electronic devices 100 recognizing near-field voice and automatically collecting audio.

[0315] In some embodiments, the electronic device 100 can save the audio collected by the first application within the first application, or convert the audio collected by the first application into text and save the text within the first application.

[0316] Exemplarily, the first application can be a memo application.

[0317] Exemplarily, as Figure 7D shown, the user can pick up the Bluetooth headset, bring the Bluetooth headset close to the user's mouth, and detect a voice signal. When it is detected that the Bluetooth headset is in a hand-held state and a voice signal is detected, the Bluetooth headset can determine that it is in the handset mode. In response to the Bluetooth headset being in the handset mode, the Bluetooth headset can collect an audio signal and send the audio signal to the electronic device 100 with which it has established a Bluetooth connection, and the electronic device 100 saves the audio signal collected by the Bluetooth headset within the first application.

[0318] In some embodiments, the Bluetooth headset can also include multiple microphones and multiple speakers, and the Bluetooth headset can also determine whether there is near-field voice in a manner similar to the above-mentioned electronic device 100. After determining that there is near-field voice, the Bluetooth headset can collect an audio signal and send the audio signal to the electronic device 100 with which it has established a Bluetooth connection, and the electronic device 100 saves the audio signal collected by the Bluetooth headset within the first application.

[0319] Optionally, before the Bluetooth headset sends the collected audio signal to the electronic device 100, the electronic device 100 can be displayed on any user interface. Exemplarily, referring to Figure 4B shown, the electronic device 100 can display the desktop.

[0320] In response to the audio signal sent by the Bluetooth headset, the electronic device 100 can automatically launch the memo application and display the Figure 7E user interface 7100 shown.

[0321] As Figure 7EAs shown, a prompt message 7101 is displayed on the user interface 4100. The prompt message 7101 includes the text "Near-field voice mode of Bluetooth headset", and this prompt message 7101 is used to prompt the user that the current is to identify the near-field voice mode through the Bluetooth headset.

[0322] In some embodiments, before the Bluetooth headset sends the collected audio signal to the electronic device 100, the Bluetooth headset can further confirm whether the target sound source is recognized. After recognizing the target sound source, the Bluetooth headset then sends the collected audio signal to the electronic device 100. This can improve the accuracy of confirming that the user uses the audio collection function of the Bluetooth headset.

[0323] In a possible implementation manner, the Bluetooth headset can send an ultrasonic signal, and confirm whether the target sound source is recognized through the reflected ultrasonic signal and the transmitted ultrasonic signal. Exemplarily, because if the user is speaking, the ultrasonic signal hits the user's mouth, and the movement of the mouth will cause the frequency and / or amplitude of the ultrasonic signal to change, and the Bluetooth headset can identify the target sound source based on the characteristics of the reflected ultrasonic signal.

[0324] Not limited to ultrasonic signals, the Bluetooth headset can also determine the target sound source in other ways, and this application does not make any limitations in this regard.

[0325] In some embodiments, after the electronic device 100 receives the audio signal sent by the Bluetooth headset, the electronic device 100 may also not display Figure 7E the shown user interface 7100, directly control the memo application to start saving the audio, and display Figure 4D the shown user interface 4200.

[0326] After that, the memo application can display the real-time received audio signal, and after the user stops outputting audio, save the received audio signal in the memo application, or convert the received audio signal into text information and save the text information in the memo application. The user can also view the saved audio signal or the text information corresponding to the audio signal in the memo application. Specifically, reference can be made to Figures 4E - 4J the description in the embodiments, and this application will not elaborate here.

[0327] In some embodiments, the electronic device 100 can send the audio signal sent by the Bluetooth headset to the electronic device 200 through the first application, or convert the audio signal sent by the Bluetooth headset into text and then send it to the electronic device 200 through the first application.

[0328] Exemplarily, the first application can be an instant messaging application.

[0329] Reference can be made to Figure 7DAs described in the embodiments, the user can pick up the Bluetooth headset, bring the Bluetooth headset close to the user's mouth, and detect the voice signal. When it is detected that the Bluetooth headset is close to the user's mouth and the voice signal is detected, the Bluetooth headset can determine that it is in the handset mode. In response to the Bluetooth headset being in the handset mode, the Bluetooth headset can collect the audio signal and send the audio signal to the electronic device 100 that has established a Bluetooth connection with it, and save the audio signal collected by the Bluetooth headset in the first application through the electronic device 100.

[0330] Optionally, before the electronic device 100 receives the audio signal sent by the Bluetooth headset, the electronic device 100 can display the chat interface of the contact. For example, referring to Figure 4K As shown, the electronic device 100 can display the chat interface of the contact "Lisa".

[0331] In response to the audio signal sent by the Bluetooth headset, the electronic device 100 can display Figure 7F The prompt bar 7200 shown.

[0332] As Figure 7F shown, the prompt information 7201 is displayed on the prompt bar 7200. The prompt information 7201 includes the text "Bluetooth headset near-field voice mode", and this prompt information 7201 is used to prompt the user that the current is to identify the near-field voice mode through the Bluetooth headset.

[0333] In some embodiments, before the Bluetooth headset sends the collected audio signal to the electronic device 100, the Bluetooth headset can further confirm whether the target sound-producing part is recognized. After recognizing the target sound-producing part, the Bluetooth headset then sends the collected audio signal to the electronic device 100. This can improve the accuracy of confirming that the user uses the audio collection function of the Bluetooth headset.

[0334] In a possible implementation manner, the Bluetooth headset can send an ultrasonic signal and confirm whether the target sound-producing part is recognized through the reflected ultrasonic signal. Because if the user is speaking, the ultrasonic signal hits the user's mouth, and the movement of the mouth will cause the frequency and / or amplitude of the ultrasonic signal to change. The Bluetooth headset can recognize the target sound-producing part based on the characteristics of the reflected ultrasonic signal.

[0335] Not limited to ultrasonic signals, the Bluetooth headset can also determine the target sound-producing part in other ways, and this application does not make a limitation on this.

[0336] In some embodiments, after the electronic device 100 receives the audio signal sent by the Bluetooth headset, the electronic device 100 may also not display Figure 7F the prompt information 7201 shown.

[0337] After receiving the audio signal sent by the Bluetooth headset, the instant messaging application can display the real-time received audio signal, and after the user stops outputting audio, send the received audio signal to the electronic device 200 through the first application, or convert the audio signal sent by the Bluetooth headset into text and then send it to the electronic device 200 through the first application. Specifically, reference can be made to Figures 4K - 4N the description in the embodiments, which will not be elaborated herein in this application.

[0338] Figure 8 It is a schematic diagram of functional modules of another audio acquisition method provided by this application.

[0339] As Figure 8 shown, the functional modules on the Bluetooth headset may include but are not limited to: a voice pickup unit, a central control unit, a near-field mode recognition unit, a near-field voice separation unit, a Bluetooth communication unit, etc. The functional modules on the electronic device 100 may include but are not limited to: an interaction unit, a Bluetooth communication unit.

[0340] Among them, the voice pickup unit is used to periodically / irregularly pick up the audio signal and send the picked-up audio signal to the central control unit.

[0341] The central control unit is used to receive the audio signal sent by the voice pickup unit, process the audio signal, and determine whether the Bluetooth headset is in the earpiece mode.

[0342] The Bluetooth headset being in the earpiece mode may mean that the Bluetooth headset recognizes a voice signal and the Bluetooth headset is in a handheld state.

[0343] In some embodiments, sensors are pre-installed on the Bluetooth headset, and it is possible to determine whether the Bluetooth headset is in a handheld state based on the sensor data collected by the sensors.

[0344] In other embodiments, it is also possible to determine whether the Bluetooth headset is in a handheld state based on the movement trajectory of the Bluetooth headset.

[0345] It is also possible to determine whether the Bluetooth headset is in a handheld state based on other methods, and this application does not make any limitations in this regard.

[0346] In some embodiments, the central control unit is specifically used to determine the probability of the voice signal existing in the audio signal. When the probability of the voice signal existing in the audio signal is greater than a preset value, the central control unit determines that a voice signal is recognized. In this way, noise interference can be avoided.

[0347] The central control unit is further used to send the audio signal to the near-field mode recognition unit after determining that the Bluetooth headset is in the earpiece mode.

[0348] A near-field mode recognition unit, configured to receive an audio signal sent by a central control unit, process the audio signal, and determine whether there is near-field speech.

[0349] The central control unit is further configured to, when the near-field mode recognition unit determines that there is near-field speech in the audio signal, control a speaker to emit an ultrasonic signal to determine whether a target sound-emitting part is recognized, so as to improve the accuracy of the Bluetooth headset entering the near-field speech mode. In this application, an ultrasonic signal can be emitted through a speaker module, and no other devices need to be added.

[0350] A voice pickup unit is further configured to, when the central control unit controls the speaker to emit an ultrasonic signal, receive the reflected ultrasonic signal and send the reflected ultrasonic signal to the central control unit.

[0351] The central control unit is further configured to confirm whether a target sound-emitting object is recognized based on the characteristics of the emitted ultrasonic signal and the reflected ultrasonic signal. For how the central control unit recognizes the target sound-emitting object, reference can be made to Figure 6 the description of S603 in the embodiment, and this application will not elaborate here.

[0352] The central control unit is further configured to send a message indicating that the target sound-emitting object is recognized to the near-field mode recognition unit after recognizing the target sound-emitting object.

[0353] Optionally, Figure 5 Steps 6 to 10 shown in the embodiment may not be executed either, and this application does not make any limitation in this regard.

[0354] The near-field mode recognition unit is further configured to send the audio signal to the near-field voice separation unit in response to the message indicating that the target sound-emitting object is recognized sent by the central control unit.

[0355] A near-field voice separation unit is configured to separate near-field speech from the audio signal after receiving the audio signal sent by the near-field mode recognition unit.

[0356] The near-field voice separation unit is further configured to send the near-field speech to a Bluetooth communication unit on the Bluetooth headset.

[0357] The Bluetooth communication unit on the Bluetooth headset is configured to send the near-field speech to the Bluetooth communication unit on the electronic device 100.

[0358] The Bluetooth communication unit on the Bluetooth headset is further configured to send a first instruction to an interaction unit on the electronic device 100 after receiving the near-field speech sent by the Bluetooth communication unit on the electronic device 100.

[0359] An interaction unit on the electronic device 100 is configured to display an audio acquisition interaction interface after receiving a first instruction sent by the Bluetooth communication unit on the electronic device 100, so as to prompt the user that the current Bluetooth headset is in the near-field audio acquisition mode.

[0360] Optionally, the interaction unit may be a functional module in a first application on the electronic device 100.

[0361] Optionally, the near-field voice separation unit may also be a functional module on the electronic device 100, and the present application does not make any limitation in this regard.

[0362] Figure 9 It is a schematic flowchart of another audio acquisition method provided by the present application.

[0363] S901: The electronic device 100 and the Bluetooth headset establish a Bluetooth communication connection.

[0364] After the electronic device 100 and the Bluetooth headset establish a Bluetooth communication connection, the Bluetooth headset can collect audio through the microphone on the Bluetooth headset and send the audio to the electronic device 100 through the Bluetooth connection. The electronic device 100 can also send the audio to the Bluetooth headset through the Bluetooth connection, and the Bluetooth headset then plays the audio sent by the electronic device 100 through the speaker on the Bluetooth headset.

[0365] S902: The Bluetooth headset periodically / irregularly collects audio signals.

[0366] Optionally, one or more microphones are pre-installed on the Bluetooth headset, and the Bluetooth headset can collect audio signals periodically / irregularly through one or more microphones.

[0367] S903: The Bluetooth headset needs to determine whether the Bluetooth headset is in the earpiece mode.

[0368] The Bluetooth headset being in the earpiece mode may mean that the Bluetooth headset recognizes a voice signal and / or the Bluetooth headset is in a hand-held state.

[0369] In some embodiments, sensors are pre-installed on the Bluetooth headset, and it is possible to determine whether the Bluetooth headset is in a hand-held state based on the sensor data collected by the sensors.

[0370] In other embodiments, it is also possible to determine whether the Bluetooth headset is in a hand-held state based on the movement trajectory of the Bluetooth headset.

[0371] It is also possible to determine whether the Bluetooth headset is in a hand-held state based on other methods, and the present application does not make any limitation in this regard.

[0372] In some embodiments, the Bluetooth headset recognizing a voice signal may refer to the probability that the Bluetooth headset determines the presence of a voice signal in an audio signal. For example, when the probability of determining the presence of a voice signal in the audio signal is greater than a first threshold (e.g., 80%), the Bluetooth headset may determine that it has recognized the voice signal, thus avoiding noise interference.

[0373] In some embodiments, the Bluetooth headset recognizing a voice signal may also refer to the Bluetooth headset recognizing that the energy value of the collected audio signal is greater than a preset value.

[0374] When the Bluetooth headset is in the earpiece mode, S904 is executed.

[0375] When the Bluetooth headset is not in the earpiece mode, S902 is executed, and the Bluetooth headset continues to collect audio and monitors whether it is in the earpiece mode to determine whether the Bluetooth headset enters the near-field audio collection mode.

[0376] Optionally, S903 can also be executed by the electronic device 100. After the electronic device 100 determines that the Bluetooth headset is in the earpiece mode, the electronic device 100 sends a message indicating that the Bluetooth headset is in the earpiece mode to the Bluetooth headset, so as to reduce the computational load of the Bluetooth headset.

[0377] Optionally, the Bluetooth headset may not execute S903.

[0378] S904. The Bluetooth headset needs to determine whether the audio signal includes near-field voice.

[0379] After the Bluetooth headset collects the audio signal and determines that the Bluetooth headset is in the earpiece mode, the Bluetooth headset further needs to determine whether the audio signal includes near-field voice to improve the accuracy of the Bluetooth headset entering the near-field audio collection mode.

[0380] In some embodiments, when the Bluetooth headset recognizes that the energy value of the collected audio signal is greater than a preset value, the Bluetooth headset then determines whether the audio signal includes near-field voice.

[0381] Among them, near-field voice may refer to an audio signal emitted by a target sound-emitting object within a first preset distance from the Bluetooth headset. Optionally, multiple microphones are pre-installed on the Bluetooth headset. For the same audio signal emitted by the same target sound-emitting object, the arrival time and energy value at these multiple microphones are significantly different. It is possible to determine whether the audio signal includes near-field voice based on the differences in the arrival time and energy value of the same audio signal at multiple microphones.

[0382] The specific implementation of how the Bluetooth headset recognizes near-field voice is similar to the specific implementation of how the electronic device 100 recognizes near-field voice. Specifically, reference can be made to Figure 6 the description of S602 in the embodiment, and details are not elaborated herein in this application.

[0383] If it is determined that the audio signal includes near-field speech, S905 is executed.

[0384] If it is determined that the audio signal does not include near-field speech, S902 or S903 is executed, and the Bluetooth headset continues to collect audio and monitor whether it is in the receiver mode and whether near-field speech is recognized to determine whether the Bluetooth headset enters the near-field audio collection mode.

[0385] Optionally, the Bluetooth headset may also not execute S904.

[0386] Optionally, S904 may also be executed by the electronic device 100. After the electronic device 100 recognizes that the audio signal includes near-field speech, the electronic device 100 sends a message indicating that near-field speech is recognized to the Bluetooth headset to reduce the computational load of the Bluetooth headset.

[0387] S905: The Bluetooth headset needs to determine whether the target sound source is recognized.

[0388] After it is determined that the audio signal includes near-field speech, before collecting the audio signal, the Bluetooth headset needs to determine whether the target sound source is recognized. After the target sound source is recognized, the Bluetooth headset then collects the audio signal. In this way, the accuracy of the Bluetooth headset entering the near-field speech mode can be improved, and the occurrence of accidental touch can be further prevented.

[0389] In some embodiments, when the Bluetooth headset recognizes that the energy value of the collected audio signal is greater than a preset value, the Bluetooth headset then determines whether the target sound source is recognized.

[0390] In some embodiments, after it is determined that the audio signal is near-field speech, the Bluetooth headset can emit an ultrasonic signal with a preset characteristic, and based on the characteristic of the reflected ultrasonic signal, determine whether the target sound source is recognized.

[0391] Recognizing the target sound source may mean that the distance between the Bluetooth headset and the target sound object is within a second preset distance (e.g., 10 cm), and the vibration frequency of the target sound source is within a first range. Optionally, the second preset distance may be less than the first preset distance. Optionally, the second preset distance may also be equal to the first preset distance. After the target sound object is determined, it is necessary to further determine whether the distance between the target sound source of the target sound object and the Bluetooth headset is close enough, and whether the target sound source is moving. Only after the distance is close enough and the target sound source is moving can it be determined that the target sound source is making a sound, and the Bluetooth headset then starts to collect audio, which can further prevent the occurrence of accidental touch.

[0392] In some embodiments, after determining that the audio signal is near-field voice, the Bluetooth headset can emit a first ultrasonic signal and receive the reflected second ultrasonic signal. The Bluetooth headset can determine the second distance between the first electronic device and the target sound-emitting object and the vibration frequency of the target sound-emitting part of the target sound-emitting object based on the first ultrasonic signal and the second ultrasonic signal. When the second distance is less than the second preset distance and the vibration frequency of the target sound-emitting part of the target sound-emitting object is within the first range, the Bluetooth headset can determine that the target sound-emitting part is emitting sound, and then the Bluetooth headset starts to collect audio, which can further prevent the occurrence of accidental touch.

[0393] The first range can be between the first frequency value and the second frequency value. Exemplarily, the first range can be 20hz - 40hz.

[0394] Optionally, when it is determined based on the first ultrasonic signal and the second ultrasonic signal that the second distance is greater than the second preset distance and / or the vibration frequency of the target sound-emitting part of the target sound-emitting object is not within the first range for a continuous first preset duration, the electronic device 100 exits the near-field voice mode and stops collecting audio.

[0395] In some embodiments, after determining that the audio signal is near-field voice, the Bluetooth headset can emit an ultrasonic signal with preset characteristics and determine whether the target sound-emitting part is recognized and the distance between the Bluetooth headset and the target sound-emitting object based on the characteristics of the reflected ultrasonic signal. Optionally, the preset characteristics of the emitted ultrasonic signal can include but are not limited to preset frequency, preset amplitude, emission time, etc. The characteristics of the reflected ultrasonic signal can include but are not limited to the frequency and amplitude of the reflected ultrasonic signal, reflection time, etc.

[0396] Because when the target sound-emitting part (such as the mouth) outputs audio, the target sound-emitting part will move with a certain frequency and amplitude. The ultrasonic signal is reflected when hitting the moving target sound-emitting part, resulting in changes in the frequency and amplitude of the ultrasonic signal. For example, the frequency and amplitude of the ultrasonic signal will change.

[0397] The Bluetooth headset can determine the distance between the Bluetooth headset and the target sound-emitting object and the vibration frequency of the target sound-emitting part of the target sound-emitting object based on the emitted ultrasonic signal and the reflected ultrasonic signal.

[0398] Optionally, the Bluetooth headset can determine the distance between the Bluetooth headset and the target sound-emitting object based on the difference between the time when the ultrasonic signal is emitted and the time when the reflected ultrasonic signal is received.

[0399] Optionally, the Bluetooth headset can determine the vibration frequency of the target sound-emitting part of the target sound-emitting object based on the difference between the frequency and / or amplitude of the emitted ultrasonic signal and the frequency and / or amplitude of the reflected ultrasonic signal.

[0400] Exemplarily, the Bluetooth headset can determine the vibration frequency of the target sound - generating part of the target sound - generating object based on the frequency and amplitude of the ultrasonic signal emitted and the frequency and amplitude of the reflected ultrasonic signal, compare it with the frequency and amplitude of the ultrasonic signal emitted by the Bluetooth headset, and determine whether the target sound - generating part is recognized based on the vibration frequency of the target sound - generating part. Exemplarily, when the vibration frequency of the target sound - generating part is within the first range, it can be determined that the target sound - generating part is recognized.

[0401] Exemplarily, when the difference between the frequency and amplitude of the reflected ultrasonic signal and the frequency and amplitude of the ultrasonic signal emitted by the Bluetooth headset meets a preset value, for example, is greater than the preset value, the Bluetooth headset can determine that the target sound - generating part is recognized.

[0402] If the ultrasonic signal hits a non - moving object, the frequency and amplitude of the ultrasonic signal will basically not change, and the difference between the frequency and amplitude of the reflected ultrasonic signal and the frequency and amplitude of the ultrasonic signal emitted by the Bluetooth headset does not meet the preset value, for example, is less than the preset value, and the Bluetooth headset can determine that the target sound - generating part is not recognized.

[0403] The specific implementation of how the Bluetooth headset recognizes the target sound - generating part is similar to the specific implementation of how the electronic device 100 recognizes the target sound - generating part. Specifically, reference can be made to Figure 6 the description of S603 in the embodiment, and the present application will not elaborate here.

[0404] In the case where the target sound - generating part is recognized, S906 is executed.

[0405] In the case where the target sound - generating part is not recognized, S902 or S903 or S904 is executed, and the Bluetooth headset continues to collect audio and monitors whether it is in the earpiece mode, whether near - field voice is recognized, and whether the target sound - generating part is recognized to determine whether the Bluetooth headset enters the near - field audio collection mode.

[0406] Optionally, the Bluetooth headset may also not execute S905.

[0407] Optionally, the Bluetooth headset can execute any one or any two or all of S905, S906, and S907, and the present application does not limit this.

[0408] S906. The Bluetooth headset sends the collected audio signal to the electronic device 100.

[0409] After determining that the Bluetooth headset is in the earpiece mode, or after determining near - field voice, or after recognizing the target sound - generating part, the electronic device 100 can collect the audio signal and send the collected audio signal to the electronic device 100 through the Bluetooth connection.

[0410] In some embodiments, after collecting the audio signal, the Bluetooth headset also needs to process the audio signal to filter out the near-field voice signal and filter off the far-field voice signal, which can avoid the interference of the far-field voice signal.

[0411] Based on the description in S904, the difference between near-field voice and far-field voice is that when the distance between the target sound-emitting object and the Bluetooth headset is within the first preset distance, the energy values and arrival times of the same audio emitted by the same sound-emitting object at multiple microphones on the Bluetooth headset are significantly different, and based on this, it can be determined that the current audio is near-field voice. When the distance between the target sound-emitting object and the Bluetooth headset exceeds the first preset distance, that is, when the target sound-emitting object is far from the Bluetooth headset, the arrival times and energy values of the same audio emitted by the same target sound-emitting object at multiple microphones on the Bluetooth headset are not significantly different, and based on this, it can be determined that the current audio is far-field voice.

[0412] Exemplarily, obtain the energy values of the same audio signal received by multiple microphones on the Bluetooth headset, and the difference between any two of the multiple energy values satisfies a preset energy difference, for example, is greater than the preset energy difference, and / or obtain the arrival times of the same audio signal received by multiple microphones on the Bluetooth headset, and the difference between any two of the multiple arrival times satisfies a preset time difference, for example, is greater than the preset time difference, then it can be determined that the audio is near-field voice.

[0413] Exemplarily, obtain the energy values of the same audio signal received by multiple microphones on the Bluetooth headset, and the difference between any two of the multiple energy values does not satisfy the preset energy difference, for example, is less than the preset energy difference, and / or obtain the arrival times of the same audio signal received by multiple microphones on the Bluetooth headset, and the difference between any two of the multiple arrival times does not satisfy the preset time difference, for example, is less than the preset time difference, then it can be determined that the audio is far-field voice.

[0414] Based on the above characteristics, the near-field voice can be separated from the audio signal collected by the Bluetooth headset.

[0415] Optionally, the step of filtering out the near-field voice signal from the audio signal collected by the Bluetooth headset can also be executed by the electronic device 100, and this application does not make any limitations in this regard.

[0416] The specific implementation of how the Bluetooth headset separates the near-field voice from the collected audio signal is similar to the specific implementation of how the electronic device 100 separates the near-field voice from the collected audio signal. Specifically, reference can be made to Figure 6 the description of S604 in the embodiment, and this application will not elaborate here.

[0417] S907. The electronic device 100 stores the audio signal in the first application.

[0418] In response to the audio signal sent by the Bluetooth headset, the electronic device 100 stores the audio signal in the first application.

[0419] The electronic device 100 storing the collected audio signal in the first application may include: The electronic device 100 directly stores the audio signal in the first application. Or, the electronic device 100 converts the audio signal into text information and stores the text information in the first application. Exemplarily, the first application may be a memo application. Specifically, reference may be made to Figure 7B , Figures 4D - 4J the description in the embodiment.

[0420] The electronic device 100 storing the collected audio signal in the first application may further include: The electronic device 100 sends the audio signal to the electronic device 200 through the first application. Or, the electronic device 100 converts the audio signal into text information and sends the text information to the electronic device 200 through the first application. Exemplarily, the first application may be an instant messaging application. Specifically, reference may be made to Figure 7C , Figures 4M - 4N the description in the embodiment.

[0421] S908. The Bluetooth headset continues to monitor whether the Bluetooth headset recognizes the target sound - producing part.

[0422] After the Bluetooth headset sends the collected audio signal to the electronic device 100, the Bluetooth headset needs to continue to monitor whether the Bluetooth headset recognizes the target sound - producing part.

[0423] In the case where the target sound - producing part is not recognized, the Bluetooth headset exits the near - field voice mode, stops sending the audio signal to the electronic device 100, and executes S902.

[0424] In the case where the target sound - producing part is continuously recognized, the Bluetooth headset continues to send the collected audio signal to the electronic device 100 and executes S906.

[0425] Figure 10 It is a schematic flowchart of another audio acquisition method provided by this application.

[0426] S1001. The first electronic device collects the first audio signal output by the target sound - producing object through the microphone array.

[0427] S1002. The first electronic device determines the first distance between the first electronic device and the target sound - producing object based on the first audio signal.

[0428] S1003. When the first distance is less than the first preset distance, the first electronic device acquires a second audio signal output by the target sound - emitting object through a microphone array.

[0429] S1004. The first electronic device saves the second audio signal in the first application.

[0430] Optionally, the electronic device may or may not save the first audio signal.

[0431] Optionally, when the first electronic device saves the second audio signal in the first application, it may be that the first electronic device directly saves the second audio signal in the first application, or it may be that the second audio signal is converted into text information and then the text information is saved in the first application.

[0432] The electronic device can determine the first distance from the target sound - emitting object based on the collected audio signal. When the first distance between the electronic device and the target sound - emitting object is greater than the first preset distance, or when the first distance between the electronic device and the target sound - emitting object is greater than the first preset distance for a certain period of time, the electronic device can determine that near - field voice is recognized. The electronic device can automatically start collecting audio and automatically save the audio in the first application. This realizes the automatic picking up of audio by the electronic device, reduces user operations, improves the audio collection efficiency, and enhances the user experience.

[0433] In a possible implementation, when the first electronic device saves the second audio signal in the first application, it specifically includes:

[0434] In response to the first distance being less than the first preset distance, the first electronic device starts the first application and saves the second audio signal in the first application, or the first electronic device converts the second audio signal into first text information and saves the first text information in the first application.

[0435] In this way, after recognizing near - field voice, the electronic device can automatically save the collected audio signal or the text information corresponding to the collected audio signal in the first application.

[0436] Exemplarily, the first application may be a memo application.

[0437] In a possible implementation, when the first electronic device saves the second audio signal in the first application, it specifically includes: In response to the first distance being less than the first preset distance, the first electronic device starts the first application and sends the second audio signal to the second electronic device through the first application, or the first electronic device converts the second audio signal into first text information and sends the first text information to the second electronic device through the first application.

[0438] In this way, after near-field voice is recognized, the electronic device can automatically send the collected audio signal to a second electronic device with which a communication connection is established, or send the text information corresponding to the collected audio signal to the second electronic device with which the communication connection is established.

[0439] Exemplarily, the first application can be an instant messaging application.

[0440] In a possible implementation, the first electronic device is connected to a third electronic device via Bluetooth; the first electronic device saves the second audio signal in the first application, specifically including: the first electronic device sends the second audio signal to the third electronic device via the Bluetooth connection; the third electronic device saves the second audio signal in the first application.

[0441] Optionally, the first electronic device can be a Bluetooth headset. After the Bluetooth headset recognizes near-field voice, the Bluetooth headset can automatically send the collected audio signal to a third electronic device with which a Bluetooth connection is established, or send the text information corresponding to the collected audio signal to the third electronic device with which the Bluetooth connection is established.

[0442] In a possible implementation, while the first electronic device obtains the second audio signal through the microphone array, the first electronic device also obtains a third audio signal output by other sounding objects through the microphone array; the electronic device saves the second audio signal in the first application, specifically including: when the first electronic device determines that the distance between the first electronic device and other sounding objects exceeds a first preset distance, the electronic device saves the second audio signal in the first application.

[0443] Optionally, the second audio signal and the third audio signal can be audio signals in the same audio file. The electronic device 100 can extract the second audio signal or the third audio signal from the same audio file.

[0444] Optionally, the second audio signal and the third audio signal can also be audio signals in two different audio files respectively.

[0445] Near-field voice can refer to the audio output by a target sounding object within a first preset distance from the electronic device. Far-field voice can refer to the audio output by a target sounding object that is more than the first preset distance from the electronic device.

[0446] In this way, after near-field voice is recognized and audio signal collection starts, the electronic device can determine whether the collected audio signal is a near-field audio signal or a far-field audio signal, and only save the near-field audio signal, without saving the far-field audio signal, which can avoid the interference of the far-field audio signal.

[0447] For how an electronic device distinguishes between near-field speech and far-field speech, reference may be made to the detailed descriptions in Formulas (1) to (4).

[0448] In a possible implementation, when the first distance is less than the first preset distance, the first electronic device obtains a second audio signal output by a target sound-emitting object through a microphone array, which specifically includes: the first electronic device determines the probability of a speech signal included in the first audio signal; when the first distance is less than the first preset distance and the probability of the speech signal included in the first audio signal is greater than the first threshold, the first electronic device obtains the second audio signal output by the target sound-emitting object through the microphone array.

[0449] In this way, before starting to collect audio, the electronic device can determine the probability of a speech signal included in the audio signal. When the probability of the speech signal included in the audio signal is greater than the first threshold, that is, the probability of recognizing that the user is speaking is relatively high, the electronic device can collect and save the second audio signal.

[0450] In a possible implementation, the first electronic device is connected to a third electronic device via Bluetooth; the first electronic device obtains a second audio signal output by a target sound-emitting object through a microphone array, which specifically includes: when the first electronic device determines that the first electronic device is in a hand-held state, the first distance is less than the first preset distance, and the probability of the speech signal included in the first audio signal is greater than the first threshold, the first electronic device obtains the second audio signal output by the target sound-emitting object through the microphone array.

[0451] Optionally, a sensor is pre-installed on the first electronic device, and whether it is in a hand-held state can be confirmed through the sensor signal collected by the sensor.

[0452] In this way, when the first electronic device is a Bluetooth headset, when the Bluetooth headset is in a hand-held state, it can be considered that the current user has the intention of speaking to the Bluetooth headset. When the Bluetooth headset is in a hand-held state, the distance between the Bluetooth headset and the target sound-emitting object is within the first preset distance, and the probability of the speech signal included in the first audio signal is greater than the first threshold, the Bluetooth headset can enter the near-field speech mode and start automatically collecting and saving audio signals, which can improve the accuracy of the Bluetooth headset entering the near-field speech mode.

[0453] In a possible implementation, the first electronic device includes a speaker; before the first electronic device acquires a second audio signal output by a target sound - emitting object through a microphone array, the method further includes: the first electronic device emits a first ultrasonic signal through the speaker; the first electronic device receives a reflected second ultrasonic signal; the first electronic device determines a second distance between the first electronic device and the target sound - emitting object and the vibration frequency of the target sound - emitting part of the target sound - emitting object based on the first ultrasonic signal and the second ultrasonic signal; when the second distance is less than a second preset distance and the vibration frequency of the target sound - emitting part of the target sound - emitting object is within a first range, the first electronic device acquires the second audio signal output by the target sound - emitting object through the microphone array.

[0454] Exemplarily, the first range can be between 20 hz and 40 hz.

[0455] Optionally, the second preset distance can be less than the first preset distance. Optionally, the second preset distance can also be equal to the first preset distance.

[0456] Optionally, when the target sound - emitting object emits sound, it can emit sound at a certain vibration frequency. The first electronic device can acquire the vibration frequency of the target sound - emitting part of the target sound - emitting object through the first ultrasonic signal and the second ultrasonic signal, and determine whether the target sound - emitting object is emitting sound based on the acquired vibration frequency of the target sound - emitting object. When the vibration frequency of the target sound - emitting object is within the first range, it can be considered that the target sound - emitting object is emitting sound.

[0457] In this way, when the first electronic device enters the near - field voice mode, it can confirm whether the first electronic device and the target sound - emitting object are close enough, and whether the target sound - emitting part of the target sound - emitting object is emitting sound. The accuracy of the first electronic device entering the near - field voice mode can be improved.

[0458] In a possible implementation, after the first electronic device saves the second audio signal in the first application, the method further includes: when it is determined based on the first ultrasonic signal and the second ultrasonic signal that the second distance is greater than the second preset distance and / or the vibration frequency of the target sound - emitting part of the target sound - emitting object is not within the first range for a continuous first preset duration, the first electronic device stops acquiring audio signals through the microphone array.

[0459] After the first electronic device enters the near - field voice mode, the first electronic device also needs to continuously monitor whether the first electronic device and the target sound - emitting object are close enough, and whether the target sound - emitting part of the target sound - emitting object is emitting sound. When it is monitored that the distance between the first electronic device and the target sound - emitting object exceeds the second preset distance, and / or the target sound - emitting part of the target sound - emitting object does not emit sound, the first electronic device can automatically exit the near - field voice mode and stop collecting audio.

[0460] In a possible implementation, the microphone array includes a first microphone and a second microphone, and the positions of the first microphone and the second microphone on the first electronic device are different; the first electronic device collects a first audio signal output by a target sound-emitting object through the microphone array, specifically including: the first electronic device collects the first audio signal output by the target sound-emitting object through the first microphone and the second microphone respectively; the first electronic device determines a first distance between the first electronic device and the target sound-emitting object based on the first audio signal, specifically including: the first electronic device obtains a first energy value of the first audio signal collected by the first microphone and a second energy value of the first audio signal collected by the second microphone; the first electronic device determines the first distance between the first electronic device and the target sound-emitting object based on the difference between the first energy value and the second energy value; and / or, the first electronic device obtains a first time of the first audio signal collected by the first microphone and a second time of the first audio signal collected by the second microphone; the first electronic device determines the first distance between the first electronic device and the target sound-emitting object based on the difference between the first time and the second time.

[0461] In this way, the electronic device can determine the distance between the first electronic device and the target sound-emitting object based on the difference in the energy values of the same audio signal collected by multiple microphones in the microphone array and / or the difference in the times of receiving the same audio signal.

[0462] In a possible implementation, the microphone array includes a first microphone and a second microphone, and the positions of the first microphone and the second microphone on the first electronic device are different. The speaker includes a first speaker and a second speaker, the first speaker is near the first microphone, and the second speaker is near the second microphone; the first electronic device emits a first ultrasonic signal through the speaker, specifically including: when the first electronic device determines that the distance between the first microphone and the target sound-emitting object is less than the distance between the second microphone and the target sound-emitting object, the first electronic device emits the first ultrasonic signal through the first speaker, where the first speaker is the speaker closest to the target sound-emitting object.

[0463] Combined with the first aspect, in a possible implementation, the first electronic device receives a reflected second ultrasonic signal, specifically including: the first electronic device receives the reflected second ultrasonic signal through the first microphone.

[0464] In this way, the first electronic device can emit an ultrasonic signal through the speaker closest to the target sound-emitting object and receive the reflected ultrasonic signal through the microphone closest to the target sound-emitting object, which can improve the accuracy of identifying the target sound-emitting object.

[0465] In a possible implementation, the first electronic device determines that the distance between the first microphone and the target sound-emitting object is less than the distance between the second microphone and the target sound-emitting object, which specifically includes: the first electronic device obtains a first energy value of a first audio signal collected by the first microphone and a second energy value of the first audio signal collected by the second microphone; when the first energy value is greater than the second energy value, the first electronic device determines that the distance between the first microphone and the target sound-emitting object is less than the distance between the second microphone and the target sound-emitting object; and / or, the first electronic device obtains a first time of the first audio signal collected by the first microphone and a second time of the first audio signal collected by the second microphone; when the first time is less than the second time, the first electronic device determines that the distance between the first microphone and the target sound-emitting object is less than the distance between the second microphone and the target sound-emitting object.

[0466] The above are only partial embodiments and implementation manners of the present application, and the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

[0467] It can be understood that the various user interfaces described in the embodiments of the present application are only example interfaces and do not limit the solution of the present application. In other embodiments, the user interface can adopt different interface layouts, can include more or fewer controls, and can add or reduce other function options. As long as it is based on the same inventive concept provided by the present application, it is within the protection scope of the present application.

[0468] It should be noted that, without contradiction or conflict, any feature in any embodiment of the present application, or any part of any feature, can be combined, and the combined technical solution is also within the scope of the embodiments of the present application.

[0469] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. An audio acquisition method, characterized in that The first electronic device includes a microphone array, and the microphone array includes at least two microphones. The method includes: The first electronic device collects a first audio signal output by a target sound - emitting object through the microphone array; The first electronic device determines a first distance between the first electronic device and the target sound - emitting object based on the first audio signal; When the first distance is less than a first preset distance, the first electronic device obtains a second audio signal output by the target sound - emitting object through the microphone array; The first electronic device saves the second audio signal in a first application.

2. The method according to claim 1, wherein The first electronic device saves the second audio signal in a first application, specifically including: In response to the first distance being less than the first preset distance, the first electronic device activates the first application and saves the second audio signal in the first application, or the first electronic device converts the second audio signal into first text information and saves the first text information in the first application.

3. The method according to claim 1, wherein The first electronic device saves the second audio signal in a first application, specifically including: In response to the first distance being less than the first preset distance, the first electronic device activates the first application and sends the second audio signal to a second electronic device through the first application, or the first electronic device converts the second audio signal into first text information and sends the first text information to the second electronic device through the first application.

4. The method according to claim 1, characterized in that, The first electronic device is connected to a third electronic device via Bluetooth; the first electronic device saves the second audio signal in a first application, specifically including: The first electronic device sends the second audio signal to the third electronic device through the Bluetooth connection; The third electronic device saves the second audio signal in the first application.

5. The method according to any one of claims 1 to 4, characterized in that While the first electronic device obtains the second audio signal through the microphone array, the first electronic device also obtains a third audio signal output by other sound - emitting objects through the microphone array; the electronic device saves the second audio signal in a first application, specifically including: When the first electronic device determines that the distance between the first electronic device and the other sound - emitting objects exceeds the first preset distance, the electronic device saves the second audio signal in the first application.

6. The method according to any one of claims 1-5, characterized in that, When the first distance is less than the first preset distance, the first electronic device obtains the second audio signal output by the target sound - emitting object through the microphone array, specifically including: The first electronic device determines the probability of the voice signal included in the first audio signal; When the first distance is less than the first preset distance and the probability of the voice signal included in the first audio signal is greater than a first threshold, the first electronic device obtains the second audio signal output by the target sound - emitting object through the microphone array.

7. The method according to claim 6, characterized in that, The first electronic device is connected to a third electronic device via Bluetooth; the first electronic device obtains a second audio signal output by the target sound - emitting object through the microphone array, specifically including: When the first electronic device determines that the first electronic device is in a handheld state, the first distance is less than a first preset distance, and the probability of the voice signal included in the first audio signal is greater than a first threshold, the first electronic device obtains the second audio signal output by the target sound - emitting object through the microphone array.

8. The method according to any one of claims 1-7, characterized in that The first electronic device includes a speaker; before the first electronic device obtains the second audio signal output by the target sound - emitting object through the microphone array, the method further includes: The first electronic device emits a first ultrasonic signal through the speaker; The first electronic device receives a reflected second ultrasonic signal; The first electronic device determines a second distance between the first electronic device and the target sound - emitting object and the vibration frequency of the target sound - emitting part of the target sound - emitting object based on the first ultrasonic signal and the second ultrasonic signal; When the second distance is less than a second preset distance and the vibration frequency of the target sound - emitting part of the target sound - emitting object is within a first range, the first electronic device obtains the second audio signal output by the target sound - emitting object through the microphone array.

9. The method according to claim 8, wherein After the first electronic device saves the second audio signal in a first application, the method further includes: When it is determined that the second distance is greater than the second preset distance and / or the vibration frequency of the target sound - emitting part of the target sound - emitting object is not within the first range based on the first ultrasonic signal and the second ultrasonic signal for a continuous first preset duration, the first electronic device stops obtaining audio signals through the microphone array.

10. The method according to any one of claims 1-9, characterized in that, The microphone array includes a first microphone and a second microphone, and the positions of the first microphone and the second microphone on the first electronic device are different; The first electronic device collects a first audio signal output by the target sound - emitting object through the microphone array, specifically including: The first electronic device respectively collects the first audio signal output by the target sound - emitting object through the first microphone and the second microphone; The first electronic device determines the first distance between the first electronic device and the target sound - emitting object based on the first audio signal, specifically including: The first electronic device obtains a first energy value of the first audio signal collected by the first microphone and a second energy value of the first audio signal collected by the second microphone; The first electronic device determines the first distance between the first electronic device and the target sound - emitting object based on the difference between the first energy value and the second energy value; and / or The first electronic device obtains a first time of the first audio signal collected by the first microphone and a second time of the first audio signal collected by the second microphone; The first electronic device determines the first distance between the first electronic device and the target sound - emitting object based on the difference between the first moment and the second moment.

11. The method according to claim 8 or 9, characterized in that The microphone array includes a first microphone and a second microphone, and the positions of the first microphone and the second microphone on the first electronic device are different. The speaker includes a first speaker and a second speaker. The first speaker is located near the first microphone, and the second speaker is located near the second microphone. The first electronic device emits a first ultrasonic signal through the speaker, specifically including: When the first electronic device determines that the distance between the first microphone and the target sound - emitting object is less than the distance between the second microphone and the target sound - emitting object, the first electronic device emits the first ultrasonic signal through the first speaker, where the first speaker is the speaker closest to the target sound - emitting object.

12. The method according to claim 11, characterized in that, The first electronic device receives the reflected second ultrasonic signal, specifically including: The first electronic device receives the reflected second ultrasonic signal through the first microphone.

13. The method according to claim 11 or 12, characterized in that, The first electronic device determines that the distance between the first microphone and the target sound - emitting object is less than the distance between the second microphone and the target sound - emitting object, specifically including: The first electronic device obtains a first energy value of the first audio signal collected by the first microphone and a second energy value of the first audio signal collected by the second microphone. When the first energy value is greater than the second energy value, the first electronic device determines that the distance between the first microphone and the target sound - emitting object is less than the distance between the second microphone and the target sound - emitting object. And / or The first electronic device obtains a first moment of the first audio signal collected by the first microphone and a second moment of the first audio signal collected by the second microphone. When the first moment is less than the second moment, the first electronic device determines that the distance between the first microphone and the target sound - emitting object is less than the distance between the second microphone and the target sound - emitting object.

14. An electronic device, characterized in that, The electronic device includes a microphone array, a memory, and a processor. Among them, the microphone array, the memory, and the processor are coupled. The memory is used to store a computer program. When the processor executes and calls the computer program, the electronic device executes the method according to any one of claims 1 - 13.

15. A computer-readable storage medium, comprising instructions, characterized in that, When the instruction runs on the electronic device, the electronic device executes the method according to any one of claims 1 - 13.

Citation Information

Patent Citations

  • Multi-device voice wake-up implementation method and device, electronic device and medium

    CN111812588A

  • Intelligent voice recognition method and system for self-adaptive environment perception

    CN117198295A