Audio collection method, electronic device and storage medium

Through microphone array and distance detection technology, electronic devices automatically recognize near-field voice and collect audio, solving the problems of cumbersome audio acquisition operations and mistouching in the existing technology, and improving the acquisition efficiency and user experience.

WO2025140676A1PCT designated stage expired Publication Date: 2025-07-03HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2024/143535
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-30
Filing Date
2024-12-28
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the prior art, the audio acquisition operation of electronic devices is complicated, and users need to manually open the application and collect audio. The voice wake-up method is prone to mist touching, resulting in poor user experience.

Method used

The microphone array is used to collect audio signals. By determining the distance between the device and the target sounding object and the probability of the voice signal, the near-field voice is automatically recognized and the audio acquisition is initiated, and the user's operations are reduced.

Benefits of technology

It realizes that electronic devices automatically collect and save audio after recognizing near-field voice, improves audio acquisition efficiency, reduces user operation steps, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024143535_03072025_PF_FP_ABST
    Figure CN2024143535_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are an audio collection method, an electronic device and a storage medium. A first electronic device comprises a microphone array, the microphone array comprising at least two microphones. The method comprises: by means of a microphone array, a first electronic device collects a first audio signal output by a target sound-producing object; on the basis of the first audio signal, the first electronic device determines a first distance between the first electronic device and the target sound-producing object; when the first distance is less than a first preset distance, by means of the microphone array, the first electronic device acquires a second audio signal output by the target sound-producing object; and the first electronic device saves the second audio signal in a first application. The present application allows for the automatic audio collection of electronic devices, reduces user operations, improves audio collection efficiency and improves user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Audio acquisition method, electronic device and storage medium

[0001] This application claims priority to the Chinese patent application with application number 202311870062.3 filed with the State Intellectual Property Office of China on December 30, 2023, and priority to the Chinese patent application with the invention name “A method for audio acquisition, an electronic device and a storage medium”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of terminal technology, and in particular to an audio acquisition method, electronic device, and storage medium. Background Art

[0003] With the development of terminal technology, electronic devices can support the functions of audio output and audio collection. For example, electronic devices can play audio through speakers (speakers) and can also collect audio through microphones (mics).

[0004] Electronic devices can save the audio collected by the microphone locally, and they can also send the audio collected by the microphone to other devices. However, users need to find the audio collection entrance in the application that supports the voice recording function and control the electronic device to collect audio, which is cumbersome for users to operate. Alternatively, users can use voice to wake up the application that supports the voice recording function on the electronic device to collect audio, but voice wake-up is prone to accidental touches. How to provide a convenient, fast and accurate method for collecting audio needs further research. Summary of the Invention

[0005] The present application provides an audio acquisition method, electronic device, and storage medium, which improve audio acquisition efficiency, reduce user operations, and enhance user experience.

[0006] In a first aspect, the present application provides an audio acquisition method, wherein a first electronic device includes a microphone array, the microphone array includes at least two microphones, and the method includes: the first electronic device acquires a first audio signal output by a target sound-emitting object through the microphone array; the first electronic device determines a first distance between the first electronic device and the target sound-emitting object based on the first audio signal; when the first distance is less than a first preset distance, the first electronic device acquires a second audio signal output by the target sound-emitting object through the microphone array; the first electronic device saves the second audio signal in a first application.

[0007] Optionally, the electronic device may store the first audio signal or may not store the first audio signal.

[0008] Optionally, the first electronic device stores the second audio signal in the first application. The first electronic device may store the second audio signal in the first application directly, or convert the second audio signal into text information and then store the text information in the first application.

[0009] The electronic device can determine a first distance from a target sound-emitting object based on the collected audio signal. When the first distance between the electronic device and the target sound-emitting object is greater than a first preset distance, or when the first distance between the electronic device and the target sound-emitting object is greater than the first preset distance for a certain period of time, the electronic device can determine that near-field speech has been recognized. The electronic device can automatically begin collecting audio and automatically save the audio within the first application. This enables the electronic device to automatically pick up audio, reduces user operations, improves audio collection efficiency, and enhances the user experience.

[0010] In conjunction with the first aspect, in one possible implementation, the first electronic device stores the second audio signal in the first application, specifically including:

[0011] In response to the first distance being less than the first preset distance, the first electronic device opens the first application and saves the second audio signal in the first application, or converts the second audio signal into a first text message and saves the first text message in the first application.

[0012] In this way, after recognizing the near-field voice, the electronic device can automatically save the collected audio signal or the text information corresponding to the collected audio signal in the first application.

[0013] In combination with the first aspect, in one possible implementation, the first electronic device saves the second audio signal in the first application, specifically including: in response to the first distance being less than the first preset distance, the first electronic device opens the first application and sends the second audio signal to the second electronic device through the first application, or the first electronic device converts the second audio signal into a first text message and sends the first text message to the second electronic device through the first application.

[0014] In this way, after recognizing the near-field voice, the electronic device can automatically send the collected audio signal to the second electronic device with which the communication connection is established, or send the text information corresponding to the collected audio signal to the second electronic device with which the communication connection is established.

[0015] In combination with the first aspect, in one possible implementation, the first electronic device is connected to a third electronic device via Bluetooth; the first electronic device saves the second audio signal in the first application, specifically including: the first electronic device sends the second audio signal to the third electronic device via Bluetooth connection; the third electronic device saves the second audio signal in the first application.

[0016] Optionally, the first electronic device may be a Bluetooth headset. After recognizing the near-field voice, the Bluetooth headset may automatically send the collected audio signal to a third electronic device with which a Bluetooth connection is established, or send a text message corresponding to the collected audio signal to the third electronic device with which a Bluetooth connection is established.

[0017] In combination with the first aspect, in one possible implementation, while the first electronic device obtains the second audio signal through the microphone array, the first electronic device also obtains the third audio signal output by other sound-emitting objects through the microphone array; the electronic device saves the second audio signal in the first application, specifically including: when the first electronic device determines that the distance between the first electronic device and the other sound-emitting objects exceeds a first preset distance, the electronic device saves the second audio signal in the first application.

[0018] Optionally, the second audio signal and the third audio signal may be audio signals in the same audio file. The electronic device 100 may extract the second audio signal or the third audio signal from the same audio file.

[0019] Optionally, the second audio signal and the third audio signal may also be audio signals in two audio files respectively.

[0020] Near-field speech may refer to audio output from a target sound-emitting object within a first preset distance from the electronic device. Far-field speech may refer to audio output from a target sound-emitting object beyond the first preset distance from the electronic device.

[0021] In this way, after recognizing near-field speech and starting to collect audio signals, the electronic device can determine whether the collected audio signal is a near-field audio signal or a far-field audio signal, and only save the near-field audio signal without saving the far-field audio signal, thereby avoiding interference from the far-field audio signal.

[0022] In combination with the first aspect, in a possible implementation method, when the first distance is less than a first preset distance, the first electronic device obtains a second audio signal output by the target sound-emitting object through a microphone array, specifically including: the first electronic device determines the probability of a voice signal contained in the first audio signal; when the first distance is less than the first preset distance, and when the probability of a voice signal contained in the first audio signal is greater than a first threshold, the first electronic device obtains the second audio signal output by the target sound-emitting object through the microphone array.

[0023] In this way, before starting to collect audio, the electronic device can determine the probability of a voice signal contained in the audio signal. When the probability of a voice signal contained in the audio signal is greater than the first threshold, that is, the probability of recognizing that the user is speaking is high, the electronic device can collect and save the second audio signal.

[0024] In combination with the first aspect, in a possible implementation, the first electronic device is connected to a third electronic device via Bluetooth; the first electronic device obtains the second audio signal output by the target sound-emitting object through a microphone array, specifically including: when the first electronic device determines that the first electronic device is in a handheld state, the first distance is less than the first preset distance, and the probability of the voice signal contained in the first audio signal is greater than the first threshold, the first electronic device obtains the second audio signal output by the target sound-emitting object through the microphone array.

[0025] Optionally, a sensor is pre-installed on the first electronic device, and whether the device is in a handheld state can be confirmed through a sensor signal collected by the sensor.

[0026] In this way, when the first electronic device is a Bluetooth headset, and the Bluetooth headset is in a handheld state, it can be assumed that the current user has the intention to speak into the Bluetooth headset. When the Bluetooth headset is in a handheld state, the distance between the Bluetooth headset and the target sound-emitting object is within a first preset distance, and the probability of a voice signal contained in the first audio signal is greater than a first threshold, the Bluetooth headset can enter near-field voice mode and begin automatically collecting and saving audio signals, thereby improving the accuracy of the Bluetooth headset entering near-field voice mode.

[0027] In combination with the first aspect, in a possible implementation, the first electronic device includes a speaker; before the first electronic device obtains the second audio signal output by the target sound-emitting object through the microphone array, the method also includes: the first electronic device emits a first ultrasonic signal through the speaker; the first electronic device receives the reflected second ultrasonic signal; the first electronic device determines the second distance between the first electronic device and the target sound-emitting object and the vibration frequency of the target sound-emitting part of the target sound-emitting object based on the first ultrasonic signal and the second ultrasonic signal; when the second distance is less than the second preset distance and the vibration frequency of the target sound-emitting part of the target sound-emitting object is within the first range, the first electronic device obtains the second audio signal output by the target sound-emitting object through the microphone array.

[0028] Exemplarily, the first range may be between 20 Hz and 40 Hz.

[0029] Optionally, the second preset distance may be smaller than the first preset distance. Optionally, the second preset distance may be equal to the first preset distance.

[0030] Optionally, the target sound-emitting object may generate sound at a certain vibration frequency. The first electronic device may obtain the vibration frequency of the target sound-emitting part of the target sound-emitting object through the first ultrasonic signal and the second ultrasonic signal, and determine whether the target sound-emitting object is generating sound based on the obtained vibration frequency of the target sound-emitting object. If the vibration frequency of the target sound-emitting object is within a first range, it can be determined that the target sound-emitting object is generating sound.

[0031] In this way, when the first electronic device enters the near-field voice mode, it can be confirmed whether the first electronic device and the target sound-emitting object are close enough and whether the target sound-emitting part of the target sound-emitting object is emitting sound, thereby improving the accuracy of the first electronic device entering the near-field voice mode.

[0032] In combination with the first aspect, in a possible implementation, after the first electronic device saves the second audio signal in the first application, the method also includes: when it is determined based on the first ultrasonic signal and the second ultrasonic signal for a continuous first preset time period that the second distance is greater than the second preset distance and / or the vibration frequency of the target sound-emitting part of the target sound-emitting object is not within the first range, the first electronic device stops acquiring the audio signal through the microphone array.

[0033] After the first electronic device enters near-field voice mode, it must continue to monitor whether the first electronic device and the target sound-emitting object are sufficiently close, and whether the target sound-emitting part of the target sound-emitting object is emitting sound. If it is detected that the distance between the first electronic device and the target sound-emitting object exceeds a second preset distance, and / or the target sound-emitting part of the target sound-emitting object is not emitting sound, the first electronic device can automatically exit near-field voice mode and stop collecting audio.

[0034] In combination with the first aspect, in a possible implementation, the microphone array includes a first microphone and a second microphone, and the positions of the first microphone and the second microphone are different on the first electronic device; the first electronic device collects the first audio signal output by the target sound-emitting object through the microphone array, specifically including: the first electronic device collects the first audio signal output by the target sound-emitting object through the first microphone and the second microphone respectively; the first electronic device determines the first distance between the first electronic device and the target sound-emitting object based on the first audio signal, specifically including: the first electronic device obtains a first energy value of the first audio signal collected by the first microphone and a second energy value of the first audio signal collected by the second microphone; the first electronic device determines the first distance between the first electronic device and the target sound-emitting object based on the difference between the first energy value and the second energy value; and / or, the first electronic device obtains a first moment of the first audio signal collected by the first microphone and a second moment of the first audio signal collected by the second microphone; the first electronic device determines the first distance between the first electronic device and the target sound-emitting object based on the difference between the first moment and the second moment.

[0035] In this way, the electronic device can determine the distance between the first electronic device and the target sound-emitting object by the difference in energy values ​​of the same audio signal collected by multiple microphones in the microphone array and / or the difference between the times of receiving the same audio signal.

[0036] In combination with the first aspect, in a possible implementation, the microphone array includes a first microphone and a second microphone, and the positions of the first microphone and the second microphone are different on the first electronic device. The speaker includes a first speaker and a second speaker, and the first speaker is located near the first microphone and the second speaker is located near the second microphone; the first electronic device sends a first ultrasonic signal through the speaker, specifically including: when the first electronic device determines that the distance between the first microphone and the target sound-emitting object is less than the distance between the second microphone and the target sound-emitting object, the first electronic device sends the first ultrasonic signal through the first speaker, wherein the first speaker is the speaker closest to the target sound-emitting object.

[0037] In combination with the first aspect, in a possible implementation, the first electronic device receives the reflected second ultrasonic signal, which specifically includes: the first electronic device receives the reflected second ultrasonic signal through a first microphone.

[0038] In this way, the first electronic device can emit an ultrasonic signal through a speaker closest to the target sound-emitting object and receive a reflected ultrasonic signal through a microphone closest to the target sound-emitting object, thereby improving the accuracy of identifying the target sound-emitting object.

[0039] In combination with the first aspect, in a possible implementation, the first electronic device determines that the distance between the first microphone and the target sound-emitting object is smaller than the distance between the second microphone and the target sound-emitting object, specifically including: the first electronic device obtains a first energy value of the first audio signal collected by the first microphone and a second energy value of the first audio signal collected by the second microphone; when the first energy value is greater than the second energy value, the first electronic device determines that the distance between the first microphone and the target sound-emitting object is smaller than the distance between the second microphone and the target sound-emitting object; and / or, the first electronic device obtains a first moment of the first audio signal collected by the first microphone and a second moment of the first audio signal collected by the second microphone; when the first moment is less than the second moment, the first electronic device determines that the distance between the first microphone and the target sound-emitting object is smaller than the distance between the second microphone and the target sound-emitting object.

[0040] In a second aspect, the present application provides an electronic device, which includes a microphone array, a memory, and a processor; wherein the microphone array, the memory, and the processor are coupled, and the memory is used to store a computer program. When the processor executes and calls the computer program, the electronic device executes an audio acquisition method provided in any possible implementation method in the first aspect.

[0041] In a third aspect, the present application provides a computer-readable storage medium comprising instructions. When the instructions are executed on an electronic device, the electronic device executes an audio acquisition method provided in any possible implementation of the first aspect.

[0042] In a fourth aspect, the present application provides a computer program product comprising instructions. When the computer program product is run on an electronic device, the electronic device executes an audio acquisition method provided in any possible implementation of the first aspect.

[0043] In a fifth aspect, the present application provides a chip system, which includes one or more processors, and the processors are used to call computer instructions to enable an electronic device to execute an audio acquisition method provided in any possible implementation of the first aspect above.

[0044] For the description of the beneficial effects of the second to fifth aspects, reference may be made to the description of the beneficial effects in the first aspect, and this application will not repeat them here. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] 1A-1H show UI diagrams of an application in an electronic device 100 collecting audio;

[0046] FIG2 shows a schematic structural diagram of the electronic device 100;

[0047] FIG3 is a software structure block diagram of the electronic device 100 according to an embodiment of the present invention;

[0048] 4A-4N are schematic diagrams showing a group of electronic devices 100 automatically collecting audio upon recognizing near-field speech;

[0049] FIG5 is a schematic diagram of the functional modules of an audio acquisition method provided by the present application;

[0050] FIG6 is a schematic diagram of a method flow of an audio acquisition method provided by the present application;

[0051] 7A-7C are schematic diagrams showing the arrangement positions of microphones and speakers on a group of electronic devices 100;

[0052] 7D-7F show another set of schematic diagrams of the electronic device 100 automatically collecting audio upon recognizing near-field speech;

[0053] FIG8 is a schematic diagram of the functional modules of an audio acquisition method provided by the present application;

[0054] FIG9 is a schematic diagram of a method flow of another audio acquisition method provided by the present application;

[0055] FIG10 is a flow chart of another audio acquisition method provided in this application. DETAILED DESCRIPTION

[0056] The following is a clear and detailed description of the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in the text is only a description of the association relationship between related objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0057] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.

[0058] The term "user interface (UI)" in the following embodiments of this application refers to the media interface for interaction and information exchange between an application or operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The commonly used form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operations displayed in a graphical manner. It can be a visual interface element such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, widgets, etc. displayed on the display screen of a wearable device.

[0059] A user can operate within an application on an electronic device that supports audio recording to control the application to capture audio. In some embodiments, the application can capture audio and save the audio or convert the audio into text. In other embodiments, the application can capture audio and send the audio to another device.

[0060] FIG. 1A to FIG. 1D show a UI diagram of an application in an electronic device 100 collecting audio.

[0061] Exemplarily, the application program may be a memo application.

[0062] Figure 1A shows the desktop of electronic device 100. The desktop of electronic device 100 displays application icons for multiple applications, such as a weather application icon, a stock application icon, a calculator application icon, a settings application icon, an email application icon, a music application icon, a video application icon, a browser application icon, a map application icon, a gallery application icon, a memo application icon, a voice assistant application icon, and a photo application icon. A page indicator is also displayed below the multiple application icons to indicate the total number of pages on the desktop and the position of the currently displayed page in relation to other pages. For example, the desktop may include three pages, and a white dot in the page indicator may indicate that the currently displayed page is the rightmost of the three pages. Optionally, multiple tray icons (such as a dialer application icon, a messaging application icon, a contacts application icon, and a camera application icon) are located below the page indicator. The tray icons remain displayed when switching between pages. Optionally, a status bar is displayed in a portion of the upper area of ​​the desktop. The status bar may include: one or more signal strength indicators for a mobile communication signal (also known as a cellular signal), a battery status indicator, a time indicator, and the like.

[0063] For example, as shown in FIG1A , electronic device 100 may receive a user input operation (e.g., a single click) on a memo application icon on the desktop. In response to the user input operation, electronic device 100 may display user interface 1100 shown in FIG1B . User interface 1100 is the main interface of the memo application provided in an embodiment of the present application.

[0064] As shown in FIG1B , user interface 1100 may include historical note information, and the user currently has no notes created. User interface 1100 also includes a To-Do List option and an Add Note option. The To-Do List option allows the user to view one or more items that the user needs to address within a certain period of time. The user may also create a new note using the Add Note option.

[0065] For example, as shown in FIG1B , the electronic device 100 may receive a user input operation (eg, a single click) for adding a note option, and in response to the user input operation, the electronic device 100 may display the user interface 1200 shown in FIG1C .

[0066] As shown in FIG1C , a user can edit text in user interface 1200, which includes multiple editing options, such as a view list option, a style option, an insert image option, a recording option, and a handwriting option. The user can also use the recording option to have the memo application capture the user's audio. In one possible implementation, the memo application can directly save the user's audio. In other possible implementations, the memo application can convert the user's audio into text and save the text information.

[0067] For example, as shown in FIG1C , the electronic device 100 may receive a user input operation (e.g., a single click) for the recording option in the user interface 1200 . In response to the user input operation, the memo application may display the user interface 1300 shown in FIG1D and collect the audio output by the user.

[0068] However, if a user needs to use the recording function within the Memo app, the user must follow the steps in Figures 1A to 1C to enable the recording function within the Memo app and capture the user's audio output through the Memo app, which is cumbersome. If the user does not frequently use the Memo app, the user may not be able to find the recording option in time, which is a poor user experience.

[0069] 1E-1H show another UI diagram of an application in an electronic device 100 collecting audio.

[0070] Exemplarily, the application may be an instant messaging application.

[0071] Exemplarily, as shown in FIG1E , the electronic device 100 may receive a user input operation (e.g., a single click) on the instant messaging application icon on the desktop. In response to the user input operation, the electronic device 100 may display the user interface 1400 shown in FIG1F . The user interface 1400 is the main interface of the instant messaging application provided in an embodiment of the present application. The user interface 1400 includes dialog boxes for multiple contacts. For example, a dialog box for the contact "Lisa", a dialog box for the contact "Alan", a dialog box for the contact "Henry", a dialog box for the contact "Lucy", and the like. The electronic device 100 may receive a user operation to display a real-time chat interface for a certain contact.

[0072] For example, as shown in FIG1F , electronic device 100 may receive a user input operation (e.g., a click) for a dialog box of contact “Lisa” in user interface 1400. In response to the user input operation, electronic device 100 may display user interface 1500 shown in FIG1G . User interface 1500 is a real-time chat interface for contact “Lisa”.

[0073] As shown in Figure 1G, user interface 1500 includes multiple chat records. User interface 1500 also includes a voice input option 1501, an emoticon input option 1502, a multi-function option 1503, and a keyboard switch option 1504. Among them, the user can use the voice input option 1501 to record the user's input audio content and send the user's input audio content to the contact "Lisa". The user can use the emoticon input option 1502 to select one or more emoticons and send them to the contact "Lisa". The user can use the multi-function option 1503 to send pictures, location information, initiate voice / video calls, send files, etc. to the contact "Lisa". The user can use the keyboard switch option 1504 to display the keyboard input box on the user interface 1500.

[0074] For example, as shown in FIG1G , the electronic device 100 may receive a user input operation (e.g., a long press) on the voice input option 1501 in the user interface 1500. In response to the user's input operation, the electronic device 100 may display options 4205, 4206, prompt information 4207, and voice collection area 4208 shown in FIG1H . Prompt information 4207 includes the text "Release to Send" to prompt the user how to send audio. The user needs to long press the voice collection area 4208 before the electronic device 100 collects the voice. When the user does not long press the voice collection area 4208, the electronic device 100 stops collecting the voice and sends the previously collected voice to the electronic device 200 (not shown in the figure). The user can move their finger toward option 4206 to cause the electronic device 100 to convert the previously collected voice into text and send it to the electronic device 200. The user can move their finger toward option 4205 to cause the electronic device 100 to stop collecting and stop sending voice messages to the electronic device 200.

[0075] Similarly, if the user needs to use the audio collection function within the instant messaging application, the user needs to follow the steps of Figures 1E to 1H to enable the audio collection function within the instant messaging application and collect the user's output audio through the instant messaging application, which is cumbersome for the user. If the user does not frequently use the instant messaging application, the user may not be able to find the location of the voice input option 1501 in time, and the user experience is not good.

[0076] In other embodiments, the user can also use voice commands to wake up the memo application to collect audio or wake up the instant messaging application to collect audio. However, the voice wake-up method often causes accidental touches.

[0077] In other embodiments, a user can also use a proximity sensor to detect whether the user is approaching. When the user approaches the device, the electronic device 100 automatically controls the memo application to collect audio or controls the instant messaging application to collect audio. However, distance detection requires an additional proximity device and can only detect whether the user is approaching, but cannot determine whether the user is speaking, and accidental touches are more likely to occur.

[0078] Based on the above analysis, this application provides an audio acquisition method, which includes the following steps:

[0079] Step 1: The electronic device 100 determines whether there is near-field speech. If there is near-field speech, the electronic device 100 can determine that the user needs to use the audio collection function of the electronic device 100.

[0080] The near-field voice may refer to an audio signal emitted by a target sound-emitting object within a first preset distance (eg, 1 meter) from the electronic device 100 .

[0081] Optionally, the electronic device 100 is pre-installed with multiple microphones. If the same audio signal emitted by the same target sound-emitting object arrives at the multiple microphones at significantly different times and energy levels, the audio signal can be determined to be near-field speech based on the differences in the times and energy levels of the audio signal arriving at the multiple microphones. For an introduction to near-field speech, please refer to the detailed description of the embodiment in FIG6 , which is not further elaborated in this application.

[0082] In some embodiments, the electronic device 100 may further confirm whether the target sound source is identified. After identifying the target sound source, the electronic device 100 may determine that the user needs to use the audio collection function of the electronic device 100. This may improve the accuracy of confirming that the user uses the audio collection function of the electronic device 100.

[0083] In one possible implementation, the electronic device 100 can transmit an ultrasonic signal and use the reflected ultrasonic signal to confirm whether the target sound source has been identified. This is because when a user is speaking, the ultrasonic signal hits the user's mouth, and mouth movement causes the frequency and / or amplitude of the ultrasonic signal to change. The electronic device 100 can identify the target sound source based on the characteristics of the reflected ultrasonic signal.

[0084] Not limited to ultrasonic signals, the electronic device 100 can also determine the target sound location through other methods, which is not limited in this application.

[0085] Step 2: The electronic device 100 collects audio.

[0086] The electronic device 100 may save the collected audio locally, or the electronic device 100 may convert the collected audio into text and save the text, or the electronic device 100 may send the collected audio to other electronic devices, or the electronic device 100 may convert the collected audio into text and then send the text to other electronic devices.

[0087] By using this method, the electronic device 100 can automatically capture and process audio when near-field audio is recognized. This eliminates the need for the user to manually enable the application's audio capture function, thereby improving audio capture efficiency, reducing user operations, and enhancing the user experience.

[0088] The following describes the hardware structure of an electronic device 100 provided in an embodiment of the present application.

[0089] FIG2 shows a schematic structural diagram of the electronic device 100 .

[0090] The following embodiments are described in detail using electronic device 100 as an example. It should be understood that the electronic device 100 shown in FIG2 is merely an example, and that electronic device 100 may have more or fewer components than shown in FIG2 , may combine two or more components, or may have a different component configuration. The various components shown in FIG2 may be implemented in hardware, including one or more signal processing and / or application-specific integrated circuits, software, or a combination of hardware and software.

[0091] The electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0092] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0093] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0094] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0095] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0096] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0097] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C bus lines. The processor 110 may be coupled to the touch sensor 180K, the charger, the flash, the camera 193, and the like via different I2C bus interfaces. For example, the processor 110 may be coupled to the touch sensor 180K via the I2C interface, enabling communication between the processor 110 and the touch sensor 180K via the I2C bus interface, thereby enabling the touch function of the electronic device 100.

[0098] The I2S interface can be used for audio communication. In some embodiments, the processor 110 can include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface, enabling the function of answering calls through a Bluetooth headset.

[0099] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via a PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering calls via a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0100] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface, enabling the function of playing music through Bluetooth headphones.

[0101] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display 194 and the camera 193. MIPI interfaces include the camera serial interface (CSI) and the display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to implement the camera function of the electronic device 100. The processor 110 and the display 194 communicate via the DSI interface to implement the display function of the electronic device 100.

[0102] The GPIO interface can be configured via software. The GPIO interface can be configured as either a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, display 194, wireless communication module 160, audio module 170, sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0103] The USB interface 130 is an interface that complies with USB standards and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. The USB interface 130 can be used to connect a charger to charge the electronic device 100, or to transfer data between the electronic device 100 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as augmented reality devices.

[0104] It is understood that the interface connection relationship between the modules illustrated in the embodiment of the present invention is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0105] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also provide power to the electronic device via the power management module 141.

[0106] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.

[0107] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0108] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0109] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.

[0110] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.

[0111] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the electronic device 100. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0112] In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with a network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).

[0113] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0114] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD). The display screen panel can also be made of an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniLED, a microLED, a micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than one.

[0115] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.

[0116] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise and brightness. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.

[0117] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.

[0118] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0119] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. This allows electronic device 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.

[0120] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU can enable intelligent cognitive applications in electronic device 100, such as image recognition, face recognition, speech recognition, and text comprehension.

[0121] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0122] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0123] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0124] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.

[0125] The speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or listen to hands-free calls through the speaker 170A.

[0126] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or a voice message, the user can place the receiver 170B close to the ear to hear the voice.

[0127] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the source of sound, realize directional recording function, etc.

[0128] The headphone jack 170D is used to connect a wired headphone and can be the USB interface 130 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0129] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be located on display screen 194. There are many types of pressure sensors 180A, such as resistive, inductive, and capacitive. A capacitive pressure sensor can include at least two parallel plates made of conductive material. When force acts on pressure sensor 180A, the capacitance between the electrodes changes. Electronic device 100 determines the intensity of the pressure based on this change in capacitance. When a touch operation is applied to display screen 194, electronic device 100 detects the touch intensity based on pressure sensor 180A. Electronic device 100 can also calculate the touch location based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch location but with different touch intensities can correspond to different operation instructions. For example, when a touch operation with an intensity less than a first pressure threshold is applied to a short message application icon, a command to view short messages is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to a short message application icon, a command to create a new short message is executed.

[0130] The gyroscope sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the electronic device 100 shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the electronic device 100 through reverse movement to achieve anti-shake. The gyroscope sensor 180B can also be used for navigation and somatosensory game scenes.

[0131] The air pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates the altitude using the air pressure value measured by the air pressure sensor 180C to assist in positioning and navigation.

[0132] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip case. In some embodiments, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip cover based on the magnetic sensor 180D. Based on the detected opening and closing status of the case or flip cover, features such as automatic unlocking of the flip cover can be configured.

[0133] Accelerometer 180E can detect the magnitude of acceleration of electronic device 100 in all directions (generally three axes). It can also detect the magnitude and direction of gravity when electronic device 100 is stationary. It can also be used to identify the electronic device's posture, enabling applications such as switching between landscape and portrait modes and pedometers.

[0134] The distance sensor 180F is used to measure distance. The electronic device 100 can measure distance using infrared or laser. In some embodiments, when shooting a scene, the electronic device 100 can use the distance sensor 180F to measure distance to achieve fast focusing.

[0135] The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode. The electronic device 100 emits infrared light outward through the light emitting diode. The electronic device 100 uses a photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 can use the proximity light sensor 180G to detect that the user is holding the electronic device 100 close to the ear to talk, so as to automatically turn off the screen to save power. The proximity light sensor 180G can also be used in leather case mode and pocket mode to automatically unlock and lock the screen.

[0136] Ambient light sensor 180L is used to sense ambient light brightness. Electronic device 100 can adaptively adjust the brightness of display screen 194 based on the perceived ambient light. Ambient light sensor 180L can also be used to automatically adjust white balance when taking photos. Ambient light sensor 180L can also work with proximity light sensor 180G to detect whether electronic device 100 is in a pocket to prevent accidental touches.

[0137] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to implement fingerprint unlocking, access application locks, fingerprint photography, fingerprint call answering, etc.

[0138] The temperature sensor 180J is used to detect temperature. In some embodiments, the electronic device 100 uses the temperature detected by the temperature sensor 180J to execute a temperature processing strategy. For example, when the temperature reported by the temperature sensor 180J exceeds a threshold, the electronic device 100 reduces the performance of the processor located near the temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the electronic device 100 heats the battery 142 to prevent the electronic device 100 from shutting down abnormally due to low temperature. In other embodiments, when the temperature is lower than another threshold, the electronic device 100 boosts the output voltage of the battery 142 to prevent abnormal shutdown due to low temperature.

[0139] The touch sensor 180K is also called a "touch panel." The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen." The touch sensor 180K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations can be provided via the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, in a location different from that of the display screen 194.

[0140] The bone conduction sensor 180M can obtain vibration signals. In some embodiments, the bone conduction sensor 180M can obtain vibration signals from the vibrating bones of the human body. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure pulse signals. In some embodiments, the bone conduction sensor 180M can also be set in headphones to form bone conduction headphones. The audio module 170 can parse out voice signals based on the vibration signals of the vibrating bones of the human body obtained by the bone conduction sensor 180M to implement voice functions. The application processor can parse heart rate information based on the blood pressure pulse signals obtained by the bone conduction sensor 180M to implement heart rate detection functions.

[0141] The buttons 190 include a power button, a volume button, and the like. The buttons 190 may be mechanical buttons or touch buttons. The electronic device 100 may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device 100.

[0142] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 194, motor 191 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0143] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.

[0144] The SIM card interface 195 is used to connect a SIM card.

[0145] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present invention, the Android system with a layered architecture is used as an example to illustrate the software structure of the electronic device 100.

[0146] FIG3 is a block diagram of the software structure of the electronic device 100 according to an embodiment of the present invention.

[0147] A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0148] The application layer can include a series of application packages.

[0149] As shown in FIG3 , the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and short message.

[0150] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0151] As shown in FIG3 , the application framework layer may include a window manager, a content provider, a view system, a telephony manager, a resource manager, a notification manager, and the like.

[0152] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.

[0153] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.

[0154] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.

[0155] The phone manager is used to provide communication functions of the electronic device 100, such as management of call status (including answering, hanging up, etc.).

[0156] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.

[0157] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically without user interaction. For example, the Notification Manager is used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the top status bar of the system as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include text messages in the status bar, beeps, vibrations on electronic devices, and flashing indicator lights.

[0158] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for scheduling and management of the Android system.

[0159] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.

[0160] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.

[0161] The system library can include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.

[0162] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.

[0163] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0164] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0165] A 2D graphics engine is a drawing engine for 2D drawings.

[0166] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.

[0167] Next, we will introduce an application scenario of an audio acquisition method provided by this application in conjunction with UI.

[0168] The electronic device 100 can automatically determine whether near-field speech is currently present. After determining that near-field speech is present, the electronic device 100 can automatically start the audio collection function of the first application and collect audio. In one possible implementation, the electronic device 100 can save the audio collected by the first application in the first application, or convert the audio collected by the first application into text and save the text in the first application. In other possible implementations, the electronic device 100 can send the audio collected by the first application to the electronic device 200 with which a communication connection is established.

[0169] In some embodiments, when the electronic device 100 is not connected to a Bluetooth headset, the electronic device 100 can collect audio through a microphone on the electronic device 100. The electronic device 100 can also play audio through a speaker on the electronic device 100.

[0170] 4A-4N are schematic diagrams showing a group of electronic devices 100 automatically collecting audio when they recognize near-field speech.

[0171] 4A to 4J show schematic diagrams of an electronic device 100 recognizing near-field speech, automatically collecting audio, and saving the audio.

[0172] In some embodiments, the electronic device 100 may save the audio collected by the first application within the first application, or convert the audio collected by the first application into text and save the text within the first application.

[0173] Exemplarily, the first application may be a memo application.

[0174] For example, as shown in FIG4A , the user can pick up the electronic device 100 and actively move it close to the electronic device 100. The user can output a voice, and when the electronic device 100 recognizes that the current audio is near-field voice, the electronic device 100 can automatically open the memo application and control the memo application to start collecting audio.

[0175] Optionally, before the electronic device 100 recognizes that the current audio is near-field speech, the electronic device 100 may display any user interface. For example, as shown in FIG4B , the electronic device 100 may display a desktop.

[0176] In response to recognizing the near-field voice, the electronic device 100 may automatically open the memo application and display the user interface 4100 shown in FIG. 4C .

[0177] As shown in FIG4B , a prompt message 4101 is displayed on the user interface 4100 . The prompt message 4101 includes the text “near-field voice mode”. The prompt message 4101 is used to prompt the user that the near-field voice mode has been started.

[0178] In some embodiments, after the electronic device 100 identifies the current audio as near-field speech, before the electronic device 100 automatically opens the memo application, the electronic device 100 may further confirm whether the target sound source location has been identified. After the target sound source location has been identified, the electronic device 100 then opens the memo application. This can improve the accuracy of confirming that the user is using the audio collection function of the electronic device 100.

[0179] In one possible implementation, the electronic device 100 can transmit an ultrasonic signal and use the reflected ultrasonic signal to confirm whether the target sound source has been identified. This is because when a user is speaking, the ultrasonic signal hits the user's mouth, and mouth movement causes the frequency and / or amplitude of the ultrasonic signal to change. The electronic device 100 can identify the target sound source based on the characteristics of the reflected ultrasonic signal.

[0180] Not limited to ultrasonic signals, the electronic device 100 can also determine the target sound location through other methods, which is not limited in this application.

[0181] In some embodiments, after the electronic device 100 recognizes that the current audio is near-field speech, the electronic device 100 may not display the user interface 4100 shown in Figure 4C, but directly control the memo application to start collecting audio and display the user interface 4200 shown in Figure 4D.

[0182] As shown in Figure 4D, user interface 4200 can display the audio capture process, such as a waveform, audio duration, and text information corresponding to the audio. This allows the user to check whether the audio captured by the Memo app is the same as the audio output by the user. If the audio captured by the Memo app is different from the audio output by the user, the user can choose to re-record the audio or modify the different parts of the audio output to avoid errors in the audio captured by the Memo app.

[0183] In some embodiments, after the electronic device 100 fails to recognize near-field voice, for example, after the user stops outputting audio, the electronic device 100 may control the memo application to stop collecting audio.

[0184] After the user stops outputting audio, the memo application can save the collected audio in the memo application.

[0185] In one possible implementation, the memo application can directly save the collected audio in the memo application and automatically generate a voice note.

[0186] After the memo app generates a voice note, the electronic device 100 may display user interface 4300 shown in FIG4E , which includes the name of the voice note, the date the voice note was generated, the duration of the voice note, etc. For example, the name of the voice note may be "Xin Shi Bei Lu". Alternatively, the name of the voice note may be automatically extracted and generated by the memo app based on the audio content. The date the voice note was generated may be "December 18, 2023". The duration of the voice note may be 1 minute and 26 seconds.

[0187] In some embodiments, after the memo application generates a voice note, electronic device 100 may display user interface 4400 shown in FIG4F . User interface 4400 includes not only the name of the voice note, the date the voice note was generated, and the duration of the voice note, but also a note type identifier 4401. Note type identifier 4401 is used to indicate the note type, which may include but is not limited to text notes, voice notes, etc. Optionally, note type identifier 4401 may also be automatically extracted and generated by the memo application based on the audio content.

[0188] In other possible implementations, the memo application may convert the collected audio into text, automatically generate a text note, and save the text note in the memo application.

[0189] In one possible implementation, after the memo application generates a text note, the electronic device 100 may display the user interface 4300 shown in FIG4E .

[0190] In another possible implementation, after the memo application generates a text note, electronic device 100 may display user interface 4500 shown in FIG4G . User interface 4500 includes not only the name of the voice note, the date the voice note was generated, and the duration of the voice note, but also a note type identifier 4501. Note type identifier 4501 is used to indicate the note type, which may include, but is not limited to, text notes, voice notes, etc. Optionally, note type identifier 4501 may be automatically extracted and generated by the memo application based on the audio content.

[0191] After the electronic device 100 does not recognize near-field speech, for example, after the user stops outputting audio, the electronic device 100 can control the memo application to stop collecting audio. After the memo application stops collecting audio, the memo application can continue to display the user interface shown in Figure 4E or Figure 4F or Figure 4G. Alternatively, after the memo application stops collecting audio, the memo application can continue to display the user interface displayed before collecting audio, such as displaying the desktop shown in Figure 4B. Alternatively, after the memo application stops collecting audio, the memo application can also display the main interface of the memo application, such as the user interface 4100 shown in Figure 4C. After the memo application stops collecting audio, the memo application can also display other user interfaces, which is not limited in this application.

[0192] In some embodiments, after the memo application generates a voice note or a text note, the electronic device 100 may also receive a user operation to view the voice note or the text note.

[0193] For example, as shown in FIG4H , electronic device 100 may receive a user input operation (e.g., a single click) for the "Heartfelt Notes" option in user interface 4300. In response to the user input operation, electronic device 100 may display user interface 4600 shown in FIG4I . User interface 4600 may include the name of the note, the creation time of the note, and the content of the note. The name of the note may be "Heartfelt Notes," the creation time of the note may be "December 18, 2023," and the content of the note may be "Today, the spring breeze is blowing, and the garden is full of blooming flowers. Time flies, and the years are fleeting. However, I have something on my mind, and I am writing it down here."

[0194] Exemplarily, as shown in Figure 4H, the electronic device 100 can receive a user input operation (such as a single click) for the "Heart Notes" option in the user interface 4300. In response to the user's input operation, the electronic device 100 can also display the user interface 4700 described in Figure 4J. The user interface 4700 may include the name of the note, the creation time of the note, and the duration of the note. Among them, the name of the note may be "Heart Notes", the creation time of the note may be "December 18, 2023", and the duration of the note may be 1 minute and 26 seconds. The user interface 4700 also includes a voice note playback option, a voice note deletion option, and more options. The user can choose to share the voice note to other devices or other contacts and friends in more options.

[0195] Through the above method, the electronic device 100 can automatically open the memo application and control the memo application to collect audio and save the audio when near-field speech is recognized. The user no longer needs to follow the steps of Figures 1A-1C to open the recording function in the memo application and collect the user's audio output through the memo application. This saves user operations and improves the user experience.

[0196] 4K-4N show schematic diagrams of an electronic device 100 recognizing near-field speech, automatically collecting audio, and sending the audio to an electronic device 200 .

[0197] In some embodiments, the electronic device 100 may send the audio collected by the first application to the electronic device 200 , or convert the audio collected by the first application into text and then send it to the electronic device 200 .

[0198] Exemplarily, the first application may be an instant messaging application.

[0199] 4A , the user can pick up the electronic device 100 and actively move it close to the electronic device 100. The user can output a voice, and when the electronic device 100 recognizes that the current audio is near-field voice, the electronic device 100 can control the instant messaging application to start collecting audio.

[0200] Optionally, before the electronic device 100 recognizes that the current audio is near-field voice, the electronic device 100 may display a chat interface of the contact. For example, as shown in FIG4K , the electronic device 100 may display a chat interface of the contact “Lisa”.

[0201] In response to recognizing near-field voice, the electronic device 100 may display the prompt bar 4800 shown in FIG. 4L .

[0202] As shown in FIG4L , prompt information 4801 is displayed on prompt bar 4800 , and prompt information 4801 includes the text “near-field voice mode”. Prompt information 4801 is used to prompt the user that the near-field voice mode has been started.

[0203] In some embodiments, after electronic device 100 identifies the current audio as near-field speech, it can further confirm whether the target sound source is identified. After identifying the target sound source, electronic device 100 then opens the memo application. This can improve the accuracy of confirming that the user is using the audio collection function of electronic device 100.

[0204] In one possible implementation, the electronic device 100 can transmit an ultrasonic signal and use the reflected ultrasonic signal to confirm whether the target sound source has been identified. This is because when a user is speaking, the ultrasonic signal hits the user's mouth, and mouth movement causes the frequency and / or amplitude of the ultrasonic signal to change. The electronic device 100 can identify the target sound source based on the characteristics of the reflected ultrasonic signal.

[0205] Not limited to ultrasonic signals, the electronic device 100 can also determine the target sound location through other methods, which is not limited in this application.

[0206] In some embodiments, after the electronic device 100 recognizes that the current audio is near-field speech, the electronic device 100 may not display the prompt information 4801.

[0207] As shown in FIG4L , prompt bar 4800 can display the audio collection process, such as a waveform, audio duration, and text information corresponding to the audio. This allows the user to check whether the audio collected by the instant messaging application is the same as the audio output by the user. If the audio collected by the instant messaging application is different from the audio output by the user, the user can choose to re-record the audio or modify the different parts of the audio output to avoid errors in the audio collected by the instant messaging application.

[0208] In some embodiments, after the electronic device 100 fails to recognize near-field voice, for example, after the user stops outputting audio, the electronic device 100 can control the instant messaging application to stop collecting audio.

[0209] After the user stops outputting audio, in one possible implementation, the instant messaging application can send the collected audio to the contact "Lisa" and display the user interface 4900 shown in Figure 4M, where the user interface 4900 includes an identifier 4901, which is used to indicate the audio message sent by the current user to the contact "Lisa".

[0210] After the user stops outputting audio, in other possible implementations, the instant messaging application can convert the collected audio into text, then send the text to the contact "Lisa", and display the user interface 4910 shown in Figure 4N. The user interface 4910 includes an identifier 4911, which is used to indicate that the current user is sending a text message to the contact "Lisa".

[0211] After the electronic device 100 fails to recognize near-field speech, for example, after the user stops outputting audio, the electronic device 100 may control the instant messaging application to stop collecting audio and display the user interface 4900 shown in FIG. 4M or the user interface 4910 shown in FIG. 4N .

[0212] FIG5 is a schematic diagram of the functional modules of an audio acquisition method provided in this application.

[0213] As shown in FIG5 , the functional modules on the electronic device 100 may include but are not limited to: a voice pickup unit, a central control unit, a near-field pattern recognition unit, a near-field voice separation unit, and an interaction unit.

[0214] The voice pickup unit is used to periodically / irregularly pick up audio signals and send the picked-up audio signals to the central control unit.

[0215] The central control unit is used to receive the audio signal sent by the voice pickup unit, process the audio signal, and determine whether there is a voice signal.

[0216] The central control unit is further configured to send the audio signal to the near-field pattern recognition unit after determining that a voice signal exists.

[0217] In some embodiments, the central control unit is specifically configured to determine the probability of a voice signal being present in the audio signal. If the probability of a voice signal being present in the audio signal is greater than a preset value, the central control unit then sends the audio signal to the near-field pattern recognition unit. This can avoid noise interference.

[0218] The near-field pattern recognition unit is used to receive the audio signal sent by the central control unit, process the audio signal, and determine whether there is near-field speech.

[0219] The near-field pattern recognition unit is further configured to send the audio signal to the near-field speech separation unit when it is determined that near-field speech exists in the audio signal.

[0220] The central control unit is also used to control the speaker to emit an ultrasonic signal after the near-field mode recognition unit determines that near-field speech is present in the audio signal to determine whether the target sound source is recognized, thereby improving the accuracy of the electronic device 100 entering the near-field speech mode. This application can emit ultrasonic signals through the speaker module without adding other components.

[0221] The voice pickup unit is also used to control the loudspeaker in the central control unit to emit ultrasonic signals, receive reflected ultrasonic signals, and send the reflected ultrasonic signals to the central control unit.

[0222] The central control unit is also used to confirm whether the target sound-emitting object is identified based on the characteristics of the emitted ultrasonic signal and the reflected ultrasonic signal. For how the central control unit identifies the target sound-emitting object, please refer to the description of S603 in the embodiment of Figure 6, and this application will not repeat it here.

[0223] The central control unit is further configured to send a message indicating that the target sound emitting object has been identified to the near-field pattern recognition unit after the target sound emitting object has been identified.

[0224] Optionally, steps 6 to 10 shown in the embodiment of FIG. 5 may not be performed, and this application does not limit this.

[0225] The near-field pattern recognition unit is further configured to send an audio signal to the near-field speech separation unit in response to a message sent by the central control unit indicating that a target sound-emitting object has been recognized.

[0226] The near-field speech separation unit is used to separate the near-field speech from the audio signal after receiving the audio signal sent by the near-field pattern recognition unit.

[0227] The central control unit is further configured to send a first instruction to the interaction unit after identifying the target sound-emitting object.

[0228] The interaction unit is configured to display an audio collection interaction interface after receiving a first instruction sent by the central control unit, so as to prompt the user that the electronic device 100 is currently in the near-field audio collection mode.

[0229] Optionally, the interaction unit may be a functional module in the first application. The near-field speech separation unit may also be a functional module in the first application.

[0230] FIG6 is a flowchart of an audio acquisition method provided by the present application.

[0231] S601: The electronic device 100 collects audio signals periodically or irregularly.

[0232] Optionally, the electronic device 100 is pre-installed with one or more microphones, and the electronic device 100 can collect audio signals periodically / irregularly through the one or more microphones.

[0233] S602: The electronic device 100 needs to determine whether the audio signal includes near-field speech.

[0234] After the electronic device 100 acquires the audio signal, the electronic device 100 needs to confirm whether the audio signal includes near-field speech.

[0235] In some embodiments, before the electronic device 100 confirms whether the audio signal is near-field speech, the electronic device 100 may first determine whether a speech signal exists in the audio signal, and then confirm whether the audio signal includes near-field speech.

[0236] For example, the electronic device 100 may determine the probability of a speech signal being present in the audio signal. If the probability of a speech signal being present in the audio signal is greater than a first threshold (e.g., 80%), the electronic device 100 may then determine whether the audio signal includes near-field speech. This may avoid noise interference.

[0237] The near-field voice may refer to an audio signal emitted by a target sound-emitting object within a first preset distance from the electronic device 100 .

[0238] The electronic device 100 includes a microphone array, which includes at least two microphones, for example, a first microphone and a second microphone. The electronic device 100 can determine a first distance between the electronic device 100 and a target sound-emitting object based on the audio signal picked up by the microphone array. When the first distance is less than a first preset distance, or the first distance is less than the first preset distance for a first continuous period of time, the electronic device 100 can determine that near-field speech has been recognized.

[0239] Optionally, the microphone array includes at least two microphones, and the time and energy of the same audio signal emitted by the same target sound-emitting object arriving at the multiple microphones are different. Based on the time difference and energy difference of the same audio signal arriving at the at least two microphones, the first distance between the electronic device 100 and the target sound-emitting object can be determined, and then it can be further determined whether it is near-field speech.

[0240] For example, the microphone array may include at least a first microphone and a second microphone, and the first microphone and the second microphone are located at different positions on the first electronic device. The electronic device 100 may collect the first audio signal through the first microphone and the second microphone. The electronic device 100 may obtain a first energy value of the first audio signal collected by the first microphone and a second energy value of the first audio signal collected by the second microphone, and the first electronic device may determine the first distance between the first electronic device and the target sound-emitting object based on the difference between the first energy value and the second energy value, and / or, the first electronic device may obtain a first moment of the first audio signal collected by the first microphone and a second moment of the first audio signal collected by the second microphone, and the first electronic device may also determine the first distance between the first electronic device and the target sound-emitting object based on the difference between the first moment and the second moment.

[0241] Exemplarily, the electronic device 100 may also determine the distance between each microphone and the target sound-emitting object. For example, when the first energy value is greater than the second energy value or the first energy value is less than the second energy value for a certain period of time, the first electronic device may determine that the distance between the first microphone and the target sound-emitting object is less than the distance between the second microphone and the target sound-emitting object, and / or, the first electronic device may also obtain the first moment of the first audio signal collected by the first microphone and the second moment of the first audio signal collected by the second microphone. When the first moment is less than the second moment or the first moment is less than the second moment for a certain period of time, the first electronic device determines that the distance between the first microphone and the target sound-emitting object is less than the distance between the second microphone and the target sound-emitting object.

[0242] Exemplarily, the multiple microphones pre-installed on the electronic device 100 can be located in different positions. For example, the electronic device 100 is pre-installed with microphone 1, microphone 2, and microphone 3. Microphone 1 is located at the bottom of the electronic device 100, microphone 2 is located at the top of the electronic device 100, and microphone 3 is located on the back of the electronic device 100, for example, at the location where the camera module is located on the back of the electronic device 100.

[0243] In some embodiments, microphone 1 may be referred to as a first microphone, and microphone 2 may be referred to as a second microphone.

[0244] FIG. 7A exemplarily shows a schematic diagram of the positions of multiple microphones pre-installed on the electronic device 100 .

[0245] As shown in Figure 7A, microphone 1 is located at the bottom of electronic device 100, microphone 2 is located at the top of electronic device 100, and microphone 3 is located at the back of electronic device 100, for example, it can be located at the location of the camera module on the back of electronic device 100. In some embodiments, one or more microphones are also used to receive ultrasonic signals to detect the target sound source.

[0246] In some embodiments, the electronic device 100 may also be pre-installed with one or more speakers for playing audio. In some embodiments, the one or more speakers are also used to emit ultrasonic signals to detect the target sound-producing part.

[0247] As shown in FIG7A , speaker 1 is located at the bottom of electronic device 100, and speaker 2 is located at the top of electronic device 100. In some embodiments, speaker 3 may also be included on the back of electronic device 100. For example, speaker 3 may be located at the location of the camera module on the back of electronic device 100. In some embodiments, speaker 3 may not be included on the back of electronic device 100.

[0248] It should be noted that the arrangement is not limited to three microphones and three speakers. FIG. 7A merely illustrates the positions of the multiple microphones and multiple speakers on the electronic device 100 , and this application does not limit this.

[0249] The multiple microphones are located at different positions on the electronic device 100. The closer the target sound-emitting object is to the microphone, the more energy the audio collected by the microphone is and the shorter the time. The farther the target sound-emitting object is from the microphone, the less energy the audio collected by the microphone is and the longer the time. Then, within the first preset distance, the time and energy values ​​of the same audio signal emitted by the same target sound-emitting object reaching the multiple microphones are significantly different. Based on the difference in the time and energy values ​​of the same audio signal reaching the multiple microphones, it can be determined whether the audio signal is near-field speech.

[0250] In other embodiments, when the distance between the target sound-emitting object and the electronic device 100 exceeds a first preset distance, that is, the target sound-emitting object is far away from the electronic device 100, there is no obvious difference in the time and energy value of the same audio signal emitted by the same target sound-emitting object arriving at multiple microphones on the electronic device 100, and the current audio can be judged as non-near-field speech.

[0251] Exemplarily, when the distance between the target sound-emitting object and the electronic device 100 is within a first preset distance, the target sound-emitting object outputs a first audio. The energy of the first audio received by microphone 1 is a first energy value, the energy of the first audio received by microphone 2 is a second energy value, and the energy of the first audio received by microphone 3 is a third energy value. When the first energy value, the second energy value, and the third energy value are different and meet a preset condition, the first audio can be considered to be near-field speech. Optionally, the preset condition can be that the difference between any two energy values ​​of the first energy value, the second energy value, and the third energy value meets a preset energy difference value, for example, is greater than the preset energy difference value.

[0252] Exemplarily, when the distance between the target sound-emitting object and the electronic device 100 is within a first preset distance, the target sound-emitting object outputs a first audio signal. The first audio signal is received by microphone 1 at the first moment, by microphone 2 at the second moment, and by microphone 3 at the third moment. If the first, second, and third moments are different and meet a preset condition, the first audio signal can be considered to be near-field speech. Optionally, the preset condition can be that the difference between any two of the first, second, and third moments meets a preset time difference, for example, is greater than the preset time difference.

[0253] It should be noted that the above-mentioned preset conditions may also be other judgment bases, and this application does not limit this.

[0254] For example, as shown in FIG7B , the user can pick up the electronic device 100 and actively move it close to the electronic device 100. When the user moves close to the bottom of the electronic device 100 and outputs audio, the distance relationship between the target sound-emitting part (e.g., mouth) and each microphone on the electronic device 100 can be obtained. For example, the distance between the target sound-emitting part and microphone 1 is distance 1, the distance between the target sound-emitting part and microphone 2 is distance 2, and the distance between the target sound-emitting part and microphone 3 is distance 3. Since the user moves close to the bottom of the electronic device 100 to output audio, distance 1 is smaller than distance 3, which is smaller than distance 2.

[0255] When the user outputs the first audio, the energy of the first audio received by microphone 1 is the first energy value, the energy of the first audio received by microphone 2 is the second energy value, and the energy of the first audio received by microphone 3 is the third energy value. Since distance 1 is smaller than distance 3 and smaller than distance 2, the first energy value, the second energy value, and the third energy value are different, and the first energy value is greater than the third energy value, and the third energy value is greater than the second energy value. Optionally, when the difference between the first energy value and the third energy value meets the preset energy difference value, and the difference between the third energy value and the second energy value meets the preset energy difference value, for example, the difference between the first energy value and the third energy value is greater than the preset energy difference value, and the difference between the third energy value and the second energy value is greater than the preset energy difference value, then the first audio can be considered to be near-field speech.

[0256] When the user outputs the first audio, the moment when microphone 1 receives the first audio is the first moment, the moment when microphone 2 receives the first audio is the second moment, and the moment when microphone 3 receives the first audio is the third moment. Since distance 1 is less than distance 3 and less than distance 2, the first moment, the second moment, and the third moment are different, and the first moment is less than the third moment, and the third moment is less than the second moment. Optionally, when the difference between the first moment and the third moment meets the preset time difference, and the difference between the third moment and the second moment meets the preset time difference, for example, the difference between the first moment and the third moment is greater than the preset time difference, and the difference between the third moment and the second moment is greater than the preset time difference, then the first audio can be considered to be near-field speech.

[0257] For example, as shown in FIG7C , the user can pick up the electronic device 100 and actively move it close to the electronic device 100. When the user moves close to the top of the electronic device 100 and outputs audio, the distance relationship between the target sound-emitting part (e.g., mouth) and each microphone on the electronic device 100 can be obtained. For example, the distance between the target sound-emitting part and microphone 1 is distance 1, the distance between the target sound-emitting part and microphone 2 is distance 2, and the distance between the target sound-emitting part and microphone 3 is distance 3. Since the user moves close to the top of the electronic device 100 to output audio, distance 2 is smaller than distance 3, which is smaller than distance 1.

[0258] When the user outputs the first audio, the energy of the first audio received by microphone 1 is the first energy value, the energy of the first audio received by microphone 2 is the second energy value, and the energy of the first audio received by microphone 3 is the third energy value. Since distance 2 is smaller than distance 3 and smaller than distance 1, the first energy value, the second energy value, and the third energy value are different, and the second energy value is greater than the third energy value, and the third energy value is greater than the first energy value. Optionally, when the difference between the second energy value and the third energy value meets the preset energy difference value, and the difference between the third energy value and the first energy value meets the preset energy difference value, for example, the difference between the second energy value and the third energy value is greater than the preset energy difference value, and the difference between the third energy value and the first energy value is greater than the preset energy difference value, then the first audio can be considered to be near-field speech.

[0259] When the user outputs the first audio, the moment when microphone 1 receives the first audio is the first moment, the moment when microphone 2 receives the first audio is the second moment, and the moment when microphone 3 receives the first audio is the third moment. Since distance 2 is smaller than distance 3 and smaller than distance 1, the first moment, the second moment, and the third moment are different, and the second moment is smaller than the third moment, and the third moment is smaller than the first moment. Optionally, when the difference between the second moment and the third moment meets the preset time difference, and the difference between the third moment and the first moment meets the preset time difference, for example, the difference between the second moment and the third moment is greater than the preset time difference, and the difference between the third moment and the first moment is greater than the preset time difference, then the first audio can be considered to be near-field speech.

[0260] Whether the audio is near-field speech is not limited to the above-mentioned method, and whether the audio is near-field speech can also be determined based on other methods. This application does not limit this.

[0261] After determining that the audio signal is near-field speech, the electronic device 100 may determine that the current mode is near-field speech, and the electronic device 100 may control the first application to collect the audio signal and execute S603.

[0262] Optionally, after determining that the audio signal of the first continuous duration is near-field speech, the electronic device 100 may determine that the current mode is near-field speech, and then the electronic device 100 controls the first application to collect the audio signal and executes S603.

[0263] After determining that the audio signal is not near-field speech, the electronic device 100 may determine that the current mode is not near-field speech mode, and the electronic device continues to monitor whether there is near-field speech and executes S601.

[0264] S603: The electronic device 100 needs to determine whether the target sound source is recognized.

[0265] After determining that the audio signal includes near-field speech, the electronic device 100 must determine whether the target sound source has been identified before collecting the audio signal. After identifying the target sound source, the electronic device 100 collects the audio signal. This improves the accuracy of the electronic device 100 entering near-field speech mode and further prevents accidental touches.

[0266] In some embodiments, after determining that the audio signal is near-field speech, the electronic device 100 also needs to confirm whether the target sound-emitting part is recognized. Recognizing the target sound-emitting part may mean that the distance between the electronic device 100 and the target sound-emitting object is within a second preset distance (for example, 10 cm), and the vibration frequency of the target sound-emitting part is within a first range. Optionally, the second preset distance may be less than the first preset distance. Optionally, the second preset distance may also be equal to the first preset distance. After determining the target sound-emitting object, it is necessary to further determine whether the distance between the target sound-emitting part of the target sound-emitting object and the electronic device 100 is close enough, and whether the target sound-emitting part is active. Only after the distance is close enough and the target sound-emitting part is active, can it be determined that the target sound-emitting part is making a sound, and the electronic device 100 starts collecting audio again, which can further prevent the sound from being accidentally touched.

[0267] In some embodiments, after determining that the audio signal is near-field speech, the electronic device 100 can emit a first ultrasonic signal and receive a reflected second ultrasonic signal. The electronic device 100 can determine the second distance between the first electronic device and the target sound-emitting object and the vibration frequency of the target sound-emitting part of the target sound-emitting object based on the first ultrasonic signal and the second ultrasonic signal. When the second distance is less than the second preset distance and the vibration frequency of the target sound-emitting part of the target sound-emitting object is within the first range, the electronic device 100 can determine that the target sound-emitting part is making a sound, and the electronic device 100 starts collecting audio again, which can further prevent the sound from being accidentally touched.

[0268] The first range may be between the first frequency value and the second frequency value. For example, the first range may be 20 Hz-40 Hz.

[0269] Optionally, when it is determined based on the first ultrasonic signal and the second ultrasonic signal for a continuous first preset time period that the second distance is greater than the second preset distance and / or the vibration frequency of the target sound-emitting part of the target sound-emitting object is not within the first range, the electronic device 100 exits the near-field voice mode and stops collecting audio.

[0270] In some embodiments, after determining that the audio signal is near-field speech, the electronic device 100 may emit an ultrasonic signal with preset characteristics and, based on the characteristics of the reflected ultrasonic signal, determine whether the target sound source has been identified and the distance between the electronic device 100 and the target sound-emitting object. Optionally, the preset characteristics of the emitted ultrasonic signal may include, but are not limited to, a preset frequency, preset amplitude, and emission time. The characteristics of the reflected ultrasonic signal may include, but are not limited to, the frequency and amplitude of the reflected ultrasonic signal, and the reflection time.

[0271] When the target sound-producing part (e.g., the mouth) outputs audio, it moves at a certain frequency and amplitude. The ultrasonic signal is reflected by the moving target sound-producing part, causing the frequency and amplitude of the ultrasonic signal to change.

[0272] The electronic device 100 may determine the distance between the electronic device 100 and the target sound-emitting object and the vibration frequency of the target sound-emitting part of the target sound-emitting object based on the emitted ultrasonic signal and the reflected ultrasonic signal.

[0273] Optionally, the electronic device 100 may determine the distance between the electronic device 100 and the target sound-emitting object based on the difference between the time when the ultrasonic signal is emitted and the time when the ultrasonic signal is reflected.

[0274] Optionally, the electronic device 100 may determine the vibration frequency of the target sound-emitting part of the target sound-emitting object based on the difference between the frequency and / or amplitude of the emitted ultrasonic signal and the frequency and / or amplitude of the reflected ultrasonic signal.

[0275] For example, the electronic device 100 may determine the vibration frequency of a target sound-emitting portion of the target sound-emitting object based on the frequency and amplitude of the reflected ultrasonic signal and the frequency and amplitude of the ultrasonic signal emitted by the electronic device 100, and determine whether the target sound-emitting portion has been identified based on the vibration frequency of the target sound-emitting portion. For example, when the vibration frequency of the target sound-emitting portion is within a first range, it may be determined that the target sound-emitting portion has been identified.

[0276] Exemplarily, when the difference between the frequency and amplitude of the reflected ultrasonic signal and the frequency and amplitude of the ultrasonic signal emitted by the electronic device 100 meets a preset value, for example, is greater than a preset value, the electronic device 100 can determine that the target sound source is identified.

[0277] If the ultrasonic signal hits an inactive object, the frequency and amplitude of the ultrasonic signal will basically not change. The difference between the frequency and amplitude of the reflected ultrasonic signal and the frequency and amplitude of the ultrasonic signal emitted by the electronic device 100 does not meet the preset value, for example, is less than the preset value. The electronic device 100 can determine that the target sound-emitting part is not identified.

[0278] In some embodiments, as described in S602 , the electronic device 100 is pre-installed with multiple speakers, and the electronic device 100 can emit ultrasonic signals through the speakers to determine whether the target sound source is recognized.

[0279] Based on the introduction of S602 , it can be known that the electronic device 100 can determine the distance between the target sound object and each microphone based on the energy values ​​and time differences of the audio signals received by the multiple microphones.

[0280] After determining the distance between the target sound-emitting object and each microphone, the electronic device 100 can identify the microphone closest to the target sound-emitting object. After determining the microphone closest to the target sound-emitting object, the electronic device 100 can emit an ultrasonic signal through a speaker near the microphone. The microphone receives the reflected ultrasonic signal to determine whether the target sound-emitting part has been identified. The speaker near a microphone can refer to the speaker closest to the microphone. In this way, the closer the target sound-emitting part is to the speaker and microphone, the less signal attenuation there is, and the more accurate the detection result.

[0281] For example, referring to the description in the embodiment of Figure 7B, when the user is close to the bottom of the electronic device 100 and outputs audio, the target sound-emitting part is closest to the microphone 1 and the speaker 1, then the electronic device 100 can emit an ultrasonic signal through the speaker 1 and receive the reflected ultrasonic signal through the microphone 1 to confirm whether the target sound-emitting part is recognized.

[0282] For example, referring to the description in the embodiment of Figure 7C, when the user is close to the bottom of the electronic device 100 and outputs audio, the target sound-emitting part is closest to the microphone 2 and the speaker 2, then the electronic device 100 can emit an ultrasonic signal through the speaker 2 and receive the reflected ultrasonic signal through the microphone 2 to confirm whether the target sound-emitting part is recognized.

[0283] Rather than just using the speaker and microphone closest to the target sound source to send and receive ultrasonic signals, multiple sets of speakers and microphones can also be used to simultaneously send and receive ultrasonic signals, using multiple sets of detection results to confirm whether the target sound source has been identified. This can also improve the accuracy of identifying the target sound source.

[0284] Whether the target sound source is identified is not limited to being confirmed by ultrasonic signals, but can also be confirmed by other methods, which is not limited in this application.

[0285] In some embodiments, the electronic device 100 may not execute S603. After determining the near-field voice in S602, the electronic device 100 may directly execute S604.

[0286] S604: The electronic device 100 collects an audio signal and saves the audio signal in the first application.

[0287] After determining the near-field voice, or after determining the near-field voice and identifying the target sounding position, the electronic device 100 can collect audio signals and save the collected audio signals in the first application.

[0288] The electronic device 100 storing the collected audio signal in the first application may include: the electronic device 100 directly storing the audio signal in the first application. Alternatively, the electronic device 100 converts the audio signal into text information and stores the text information in the first application. Exemplarily, the first application may be a memo application. For details, please refer to the description of the embodiments in Figures 4A-4J.

[0289] The electronic device 100 stores the collected audio signal in the first application, and may also include: the electronic device 100 sends the audio signal to the electronic device 200 via the first application. Alternatively, the electronic device 100 converts the audio signal into text information, and sends the text information to the electronic device 200 via the first application. Exemplarily, the first application may be an instant messaging application. For details, please refer to the description of the embodiments in Figures 4K-4N.

[0290] In some embodiments, after collecting the audio signal, the electronic device 100 further processes the audio signal, filters it to obtain a near-field voice signal, and filters out the far-field voice signal, thereby avoiding interference from the far-field voice signal.

[0291] Based on the description in S602, it can be known that the electronic device 100 can determine the distance between the electronic device 100 and the target sound-emitting object based on the audio signal collected by the microphone array. For example, the electronic device 100 simultaneously receives a second audio signal sent by the target sound-emitting object and a third audio signal output by other sound-emitting objects. When the electronic device 100 determines that the distance between the electronic device 100 and the target sound-emitting object is less than the first preset distance based on the second audio signal and the third audio signal, and the distance between the electronic device 100 and the other sound-emitting objects is greater than the first preset distance, the electronic device 100 can only save the second audio signal and not save the third audio signal. How the electronic device 100 determines the distance to the sound-emitting object can be referred to the description in S602, and this application will not go into details here.

[0292] For example, based on the description in S602, the difference between near-field speech and far-field speech lies in that, when the distance between the target sound-emitting object and the electronic device 100 is within a first preset distance, there are significant differences in the energy values ​​and time differences of the same audio emitted by the same sound-emitting object reaching multiple microphones on the electronic device 100, and the current audio can be judged as near-field speech based on this. When the distance between the target sound-emitting object and the electronic device 100 exceeds the first preset distance, that is, the target sound-emitting object is far away from the electronic device 100, there is no significant difference in the time and energy of the same audio emitted by the same target sound-emitting object reaching multiple microphones on the electronic device 100, and the current audio can be judged as far-field speech based on this.

[0293] Exemplarily, the energy values ​​of the same audio signal received by multiple microphones on the electronic device 100 are obtained, and the difference between any two of the multiple energy values ​​satisfies a preset energy difference, for example, is greater than the preset energy difference, and / or the moments of the same audio signal received by multiple microphones on the electronic device 100 are obtained, and the difference between any two of the multiple moments satisfies a preset time difference, for example, is greater than the preset time difference, then it can be determined that the audio is near-field speech.

[0294] Exemplarily, the energy values ​​of the same audio signal received by multiple microphones on the electronic device 100 are obtained, and the difference between any two energy values ​​among the multiple energy values ​​does not meet the preset energy difference value, for example, is less than the preset energy difference value, and / or the moments of the same audio signal received by multiple microphones on the electronic device 100 are obtained, and the difference between any two moments among the multiple moments does not meet the preset time difference value, for example, is less than the preset time difference, then it can be determined that the audio is far-field speech.

[0295] Based on the above features, near-field speech can be separated from the audio signal collected by the electronic device 100.

[0296] If the near-field speech in the audio signal collected by the electronic device 100 can be represented as s1(t), the transmission path of the near-field speech is f1(t,n), and the far-field speech can be represented as s2(t), and the transmission path of the far-field speech is f2(t,n). Where t represents time, f1(t,n) represents the transmission path of the near-field speech to the nth microphone, and f2(t,n) represents the transmission path of the far-field speech to the nth microphone. Then the audio signal collected by the electronic device 100 can be expressed by formula (1). y(t,n)=s1(t)*f1(t,n)+s2(t)*f2(t,n) Formula (1)

[0297] As shown in formula (1), y(t,n) represents the audio signal collected by the nth microphone, s1(t) represents the near-field speech collected by the electronic device 100, f1(t,n) represents the transmission path of the near-field speech, s2(t) represents the far-field speech collected by the electronic device 100, and f2(t,n) represents the transmission path of the far-field speech.

[0298] The audio signals received by the n microphones on the electronic device 100 can be expressed by formula (2). Y(t) = [y(t, 1), y(t, 2), ..., y(t, N)] Formula (2)

[0299] As shown in formula (2), Y(t) represents the audio signal received by n microphones on the electronic device 100, y(t,1) represents the audio signal received by the first microphone on the electronic device 100, y(t,2) represents the audio signal received by the second microphone on the electronic device 100, and y(t,N) represents the audio signal received by the nth microphone on the electronic device 100.

[0300] By performing a frequency domain transformation on Y(t), we can obtain Y(t,f) as shown in formula (3). Y(t,f)=F(Y(t)) Formula (3)

[0301] As shown in formula (3), Y(t,f) represents the audio signal received by n microphones on the electronic device 100 in the frequency domain, and f represents the frequency. Y(t) represents the audio signal received by n microphones on the electronic device 100 in the time domain.

[0302] The electronic device 100 can separate the near-field speech from the audio signal collected by the electronic device 100 using the following formula (4).

[0303] As shown in formula (4), s1(t) represents the near-field speech separated from the audio signal collected by the electronic device 100, Y(t, f) represents the audio signal received by n microphones on the electronic device 100 in the frequency domain, and W(t, f) represents the solution matrix, which is used to separate the near-field speech from the audio signal collected by the electronic device 100. The solution matrix can represent the difference relationship between the energy values ​​of the same audio signal emitted by the same sound-emitting object when it reaches multiple microphones on the electronic device 100, and / or the difference relationship between the moments when the same audio signal emitted by the same sound-emitting object reaches multiple microphones on the electronic device 100 when the audio signal is near-field speech. Exemplarily, the difference relationship between the energy values ​​of the same audio signal emitted by the same sound-emitting object when it reaches multiple microphones on the electronic device 100 can be: the energy values ​​of the same audio signal received by the multiple microphones on the electronic device 100, the difference between any two of the multiple energy values ​​meeting a preset energy difference value, for example, greater than the preset energy difference value. The difference relationship between the moments when the same audio emitted by the same sound-emitting object reaches multiple microphones on the electronic device 100 can be: the moments when the same audio signal is received by multiple microphones on the electronic device 100 are obtained, and the difference between any two moments among the multiple moments satisfies the preset time difference, for example, is greater than the preset time difference.

[0304] Optionally, the solution matrix W(t, f) may be obtained by the electronic device 100 through deep learning, and the solution matrix W(t, f) may also be updated periodically / irregularly.

[0305] It should be noted that the above formulas (1) to (4) are only used to explain how the present application separates near-field speech from audio signals. Near-field speech can also be separated from audio signals by other methods, and the present application does not limit this.

[0306] S605: The electronic device 100 needs to continue monitoring whether the target sound source is recognized.

[0307] After determining the near-field voice, or after determining the near-field voice and identifying the target sounding position, the electronic device 100 can collect audio signals and save the collected audio signals in the first application.

[0308] The electronic device 100 needs to continue monitoring whether the target sound source is identified. If the target sound source is not identified, the electronic device 100 exits the near-field voice mode, stops saving the collected audio signal in the first application, and executes S601.

[0309] When the target sound location continues to be recognized, the electronic device 100 continues to save the collected audio signal in the first application and executes S604.

[0310] In other embodiments, when the electronic device 100 is connected to a Bluetooth headset, the Bluetooth headset can collect audio through the microphone on the Bluetooth headset and send the audio to the electronic device 100 through the Bluetooth connection. The electronic device 100 can also send audio to the Bluetooth headset through the Bluetooth connection, and the Bluetooth headset then plays the audio sent by the electronic device 100 through the speaker on the Bluetooth headset.

[0311] 7D-7F show another group of schematic diagrams of the electronic device 100 automatically collecting audio when recognizing near-field speech.

[0312] In some embodiments, the electronic device 100 may save the audio collected by the first application within the first application, or convert the audio collected by the first application into text and save the text within the first application.

[0313] Exemplarily, the first application may be a memo application.

[0314] For example, as shown in FIG7D , a user may pick up a Bluetooth headset and place the Bluetooth headset close to the user's mouth, and a voice signal may be detected. When the Bluetooth headset is detected to be in a handheld state and a voice signal is detected, the Bluetooth headset may be determined to be in receiver mode. In response to the Bluetooth headset being in receiver mode, the Bluetooth headset may collect audio signals and send the audio signals to the electronic device 100 with which the Bluetooth connection is established, and the audio signals collected by the Bluetooth headset may be stored in the first application via the electronic device 100.

[0315] In some embodiments, the Bluetooth headset may also include multiple microphones and multiple speakers, and the Bluetooth headset may also determine whether there is near-field speech in a manner similar to the electronic device 100 described above. After determining the presence of near-field speech, the Bluetooth headset may collect audio signals and send the audio signals to the electronic device 100 with which the Bluetooth connection is established, and the electronic device 100 may store the audio signals collected by the Bluetooth headset in the first application.

[0316] Optionally, before the Bluetooth headset sends the collected audio signal to the electronic device 100, the electronic device 100 may be displayed on any user interface. For example, referring to FIG4B , the electronic device 100 may display a desktop.

[0317] In response to the audio signal sent by the Bluetooth headset, the electronic device 100 can automatically open the memo application and display the user interface 7100 shown in Figure 7E.

[0318] As shown in FIG7E , a prompt message 7101 is displayed on the user interface 4100 . The prompt message 7101 includes the text “Bluetooth headset near-field voice mode”. The prompt message 7101 is used to prompt the user that the near-field voice mode is currently being identified through the Bluetooth headset.

[0319] In some embodiments, before the Bluetooth headset transmits the collected audio signal to the electronic device 100, the Bluetooth headset may further confirm whether the target sound source is identified. After the target sound source is identified, the Bluetooth headset transmits the collected audio signal to the electronic device 100. This can improve the accuracy of confirming that the user is using the audio collection function of the Bluetooth headset.

[0320] In one possible implementation, a Bluetooth headset can transmit an ultrasonic signal and confirm whether the target sound source has been identified based on the reflected ultrasonic signal and the transmitted ultrasonic signal. For example, if a user is speaking, the ultrasonic signal hits the user's mouth. The mouth movement causes the frequency and / or amplitude of the ultrasonic signal to change. The Bluetooth headset can then identify the target sound source based on the characteristics of the reflected ultrasonic signal.

[0321] Not limited to ultrasonic signals, Bluetooth headsets can also determine the target sound location through other methods, which is not limited in this application.

[0322] In some embodiments, after the electronic device 100 receives the audio signal sent by the Bluetooth headset, the electronic device 100 may not display the user interface 7100 shown in Figure 7E, but directly control the memo application to start saving audio and display the user interface 4200 shown in Figure 4D.

[0323] Afterwards, the memo app can display the real-time received audio signal and, after the user stops outputting audio, save the received audio signal in the memo app, or convert the received audio signal into a text message and save the text message in the memo app. The user can also view the saved audio signal or the text message corresponding to the audio signal in the memo app. For details, please refer to the description of the embodiments in Figures 4E-4J, and this application will not repeat them here.

[0324] In some embodiments, the electronic device 100 may send the audio signal sent by the Bluetooth headset to the electronic device 200 through the first application, or convert the audio signal sent by the Bluetooth headset into text and then send it to the electronic device 200 through the first application.

[0325] Exemplarily, the first application may be an instant messaging application.

[0326] Referring to the description in the embodiment of FIG7D , the user can pick up the Bluetooth headset, place the Bluetooth headset close to the user's mouth, and detect a voice signal. When the Bluetooth headset is detected to be close to the user's mouth and a voice signal is detected, the Bluetooth headset can be determined to be in receiver mode. In response to the Bluetooth headset being in receiver mode, the Bluetooth headset can collect audio signals and send the audio signals to the electronic device 100 with which the Bluetooth connection is established, and the audio signals collected by the Bluetooth headset are saved in the first application through the electronic device 100.

[0327] Optionally, before the electronic device 100 receives the audio signal sent by the Bluetooth headset, the electronic device 100 may display a chat interface of the contact. For example, as shown in FIG4K , the electronic device 100 may display a chat interface of the contact “Lisa”.

[0328] In response to the audio signal sent by the Bluetooth headset, the electronic device 100 may display the prompt bar 7200 shown in FIG. 7F .

[0329] As shown in FIG7F , prompt information 7201 is displayed on prompt bar 7200 , and prompt information 7201 includes the text “Bluetooth headset near-field voice mode”. Prompt information 7201 is used to prompt the user that the near-field voice mode is currently being recognized through the Bluetooth headset.

[0330] In some embodiments, before the Bluetooth headset transmits the collected audio signal to the electronic device 100, the Bluetooth headset may further confirm whether the target sound source is identified. After the target sound source is identified, the Bluetooth headset transmits the collected audio signal to the electronic device 100. This can improve the accuracy of confirming that the user is using the audio collection function of the Bluetooth headset.

[0331] In one possible implementation, a Bluetooth headset can transmit an ultrasonic signal and use the reflected ultrasonic signal to confirm whether the target sound source has been identified. This is because when a user is speaking, the ultrasonic signal hits the user's mouth, and mouth movement causes the frequency and / or amplitude of the ultrasonic signal to change. The Bluetooth headset can then identify the target sound source based on the characteristics of the reflected ultrasonic signal.

[0332] Not limited to ultrasonic signals, Bluetooth headsets can also determine the target sound location through other methods, which is not limited in this application.

[0333] In some embodiments, after the electronic device 100 receives the audio signal sent by the Bluetooth headset, the electronic device 100 may not display the prompt information 7201 shown in FIG. 7F .

[0334] After receiving the audio signal sent by the Bluetooth headset, the instant messaging application can display the real-time received audio signal, and after the user stops outputting the audio, send the received audio signal to the electronic device 200 through the first application, or convert the audio signal sent by the Bluetooth headset into text, and then send it to the electronic device 200 through the first application. For details, please refer to the description of the embodiments in Figures 4K to 4N, and this application will not repeat them here.

[0335] FIG8 is a schematic diagram of the functional modules of another audio acquisition method provided by the present application.

[0336] As shown in Figure 8, the functional modules of the Bluetooth headset may include but are not limited to: a voice pickup unit, a central control unit, a near-field pattern recognition unit, a near-field voice separation unit, a Bluetooth communication unit, etc. The functional modules of the electronic device 100 may include but are not limited to: an interaction unit and a Bluetooth communication unit.

[0337] The voice pickup unit is used to periodically / irregularly pick up audio signals and send the picked-up audio signals to the central control unit.

[0338] The central control unit is used to receive the audio signal sent by the voice pickup unit, process the audio signal, and determine whether the Bluetooth headset is in the handset mode.

[0339] The Bluetooth headset being in the handset mode may mean that the Bluetooth headset recognizes a voice signal and is in a handheld state.

[0340] In some embodiments, a sensor is pre-installed on the Bluetooth headset, and whether the Bluetooth headset is in a handheld state can be determined based on sensor data collected by the sensor.

[0341] In other embodiments, whether the Bluetooth headset is in the handheld state may also be determined based on the motion trajectory of the Bluetooth headset.

[0342] Whether the Bluetooth headset is in the handheld state can also be determined based on other methods, which are not limited in this application.

[0343] In some embodiments, the central control unit is specifically configured to determine the probability of a voice signal being present in the audio signal. If the probability of a voice signal being present in the audio signal is greater than a preset value, the central control unit determines that a voice signal has been recognized. This can avoid noise interference.

[0344] The central control unit is further configured to send the audio signal to the near-field pattern recognition unit after determining that the Bluetooth headset is in the handset mode.

[0345] The near-field pattern recognition unit is used to receive the audio signal sent by the central control unit, process the audio signal, and determine whether there is near-field speech.

[0346] The central control unit is also used to control the speaker to emit an ultrasonic signal when the near-field pattern recognition unit determines that near-field speech is present in the audio signal to determine whether the target sound source is recognized, thereby improving the accuracy of the Bluetooth headset entering near-field speech mode. This application can emit ultrasonic signals through the speaker module without adding other components.

[0347] The voice pickup unit is also used to control the loudspeaker in the central control unit to emit ultrasonic signals, receive reflected ultrasonic signals, and send the reflected ultrasonic signals to the central control unit.

[0348] The central control unit is also used to confirm whether the target sound-emitting object is identified based on the characteristics of the emitted ultrasonic signal and the reflected ultrasonic signal. For how the central control unit identifies the target sound-emitting object, please refer to the description of S603 in the embodiment of Figure 6, and this application will not repeat it here.

[0349] The central control unit is further configured to send a message indicating that the target sound emitting object has been identified to the near-field pattern recognition unit after the target sound emitting object has been identified.

[0350] Optionally, steps 6 to 10 shown in the embodiment of FIG. 5 may not be performed, and this application does not limit this.

[0351] The near-field pattern recognition unit is further configured to send an audio signal to the near-field speech separation unit in response to a message sent by the central control unit indicating that a target sound-emitting object has been recognized.

[0352] The near-field speech separation unit is used to separate the near-field speech from the audio signal after receiving the audio signal sent by the near-field pattern recognition unit.

[0353] The near-field voice separation unit is also used to send the near-field voice to the Bluetooth communication unit on the Bluetooth headset.

[0354] The Bluetooth communication unit on the Bluetooth headset is used to send near-field voice to the Bluetooth communication unit on the electronic device 100 .

[0355] The Bluetooth communication unit on the Bluetooth headset is further used to send a first instruction to the interaction unit on the electronic device 100 after receiving the near-field voice sent by the Bluetooth communication unit on the electronic device 100.

[0356] The interaction unit on the electronic device 100 is configured to display an audio collection interaction interface after receiving a first instruction sent by the Bluetooth communication unit on the electronic device 100, so as to prompt the user that the current Bluetooth headset is in the near-field audio collection mode.

[0357] Optionally, the interaction unit may be a functional module in the first application on the electronic device 100 .

[0358] Optionally, the near-field speech separation unit may also be a functional module on the electronic device 100, which is not limited in this application.

[0359] FIG9 is a flowchart of another audio acquisition method provided by the present application.

[0360] S901: The electronic device 100 establishes a Bluetooth communication connection with the Bluetooth headset.

[0361] After the electronic device 100 and the Bluetooth headset establish a Bluetooth communication connection, the Bluetooth headset can collect audio through the microphone on the Bluetooth headset and send the audio through the Bluetooth connection to the electronic device 100. The electronic device 100 can also send audio to the Bluetooth headset through the Bluetooth connection, and the Bluetooth headset then plays the audio sent by the electronic device 100 through the speaker on the Bluetooth headset.

[0362] S902: The Bluetooth headset collects audio signals periodically or irregularly.

[0363] Optionally, the Bluetooth headset is pre-installed with one or more microphones, and the Bluetooth headset can collect audio signals periodically / irregularly through the one or more microphones.

[0364] S903: The Bluetooth headset needs to determine whether the Bluetooth headset is in receiver mode.

[0365] The Bluetooth headset being in the handset mode may mean that: the Bluetooth headset recognizes a voice signal and / or the Bluetooth headset is in a handheld state.

[0366] In some embodiments, a sensor is pre-installed on the Bluetooth headset, and whether the Bluetooth headset is in a handheld state can be determined based on sensor data collected by the sensor.

[0367] In other embodiments, whether the Bluetooth headset is in the handheld state may also be determined based on the motion trajectory of the Bluetooth headset.

[0368] Whether the Bluetooth headset is in the handheld state can also be determined based on other methods, which are not limited in this application.

[0369] In some embodiments, the Bluetooth headset recognizes a voice signal, which may refer to the Bluetooth headset determining the probability of the voice signal existing in the audio signal. For example, when the probability of determining that the voice signal exists in the audio signal is greater than a first threshold (for example, 80%), the Bluetooth headset can determine that the voice signal is recognized, thereby avoiding noise interference.

[0370] In some embodiments, the Bluetooth headset recognizing a voice signal may also mean that the Bluetooth headset recognizes that the energy value of the collected audio signal is greater than a preset value.

[0371] When the Bluetooth headset is in the handset mode, execute S904.

[0372] If the Bluetooth headset is not in the handset mode, S902 is executed, and the Bluetooth headset continues to collect audio and monitors whether it is in the handset mode to determine whether the Bluetooth headset enters the near-field audio collection mode.

[0373] Optionally, S903 may also be executed by the electronic device 100. After the electronic device 100 determines that the Bluetooth headset is in the handset mode, the electronic device 100 sends a message to the Bluetooth headset indicating that the Bluetooth headset is in the handset mode, so as to reduce the computational complexity of the Bluetooth headset.

[0374] Optionally, the Bluetooth headset may not execute S903.

[0375] S904: The Bluetooth headset needs to determine whether the audio signal includes near-field voice.

[0376] After the Bluetooth headset collects the audio signal and determines that the Bluetooth headset is in the handset mode, the Bluetooth headset needs to further determine whether the audio signal includes near-field speech to improve the accuracy of the Bluetooth headset entering the near-field audio collection mode.

[0377] In some embodiments, when the Bluetooth headset recognizes that the energy value of the collected audio signal is greater than a preset value, the Bluetooth headset further determines whether the audio signal includes near-field speech.

[0378] Near-field speech may refer to an audio signal emitted by a target sound-emitting object within a first preset distance from the Bluetooth headset. Optionally, the Bluetooth headset is pre-installed with multiple microphones, and the time and energy values ​​of the same audio signal emitted by the same target sound-emitting object reaching the multiple microphones vary significantly. Whether the audio signal includes near-field speech can be determined based on the difference in the time and energy values ​​of the same audio signal reaching the multiple microphones.

[0379] The specific implementation of how the Bluetooth headset recognizes near-field voice is similar to the specific implementation of how the electronic device 100 recognizes near-field voice. For details, please refer to the description of S602 in the embodiment of Figure 6, and this application will not go into details here.

[0380] In a case where it is determined that the audio signal includes near-field speech, S905 is performed.

[0381] If it is determined that the audio signal does not include near-field speech, execute S902 or S903, the Bluetooth headset continues to collect audio and monitors whether it is in earpiece mode and whether near-field speech is recognized to determine whether the Bluetooth headset enters near-field audio collection mode.

[0382] Optionally, the Bluetooth headset may not execute S904.

[0383] Optionally, S904 may also be executed by the electronic device 100. After the electronic device 100 recognizes that the audio signal includes near-field speech, the electronic device 100 sends a message to the Bluetooth headset indicating that the near-field speech has been recognized, so as to reduce the computational complexity of the Bluetooth headset.

[0384] S905: The Bluetooth headset needs to determine whether it has recognized the target sound source.

[0385] After determining that the audio signal includes near-field speech, the Bluetooth headset must determine whether it has identified the target sound source before collecting the audio signal. After identifying the target sound source, the Bluetooth headset collects the audio signal. This improves the accuracy of the Bluetooth headset entering near-field speech mode and further prevents accidental touches.

[0386] In some embodiments, when the Bluetooth headset recognizes that the energy value of the collected audio signal is greater than a preset value, the Bluetooth headset determines whether the target sound source is recognized.

[0387] In some embodiments, after determining that the audio signal is near-field speech, the Bluetooth headset can emit an ultrasonic signal with preset characteristics, and determine whether the target sound source is recognized based on the characteristics of the reflected ultrasonic signal.

[0388] Identifying the target sound-emitting part may mean that: the distance between the Bluetooth headset and the target sound-emitting object is within a second preset distance (for example, 10 cm), and the vibration frequency of the target sound-emitting part is within a first range. Optionally, the second preset distance may be smaller than the first preset distance. Optionally, the second preset distance may also be equal to the first preset distance. After determining the target sound-emitting object, it is necessary to further determine whether the distance between the target sound-emitting part of the target sound-emitting object and the Bluetooth headset is close enough, and whether the target sound-emitting part is active. Only after the distance is close enough and the target sound-emitting part is active, can it be determined that the target sound-emitting part is making a sound, and the Bluetooth headset starts collecting audio again, which can further prevent the sound from being accidentally touched.

[0389] In some embodiments, after determining that the audio signal is near-field speech, the Bluetooth headset can emit a first ultrasonic signal and receive a reflected second ultrasonic signal. The Bluetooth headset can determine the second distance between the first electronic device and the target sound-emitting object and the vibration frequency of the target sound-emitting part of the target sound-emitting object based on the first ultrasonic signal and the second ultrasonic signal. When the second distance is less than the second preset distance and the vibration frequency of the target sound-emitting part of the target sound-emitting object is within the first range, the Bluetooth headset can determine that the target sound-emitting part is making a sound, and the Bluetooth headset will start collecting audio again, which can further prevent accidental touches from making sounds.

[0390] The first range may be between the first frequency value and the second frequency value. For example, the first range may be 20 Hz-40 Hz.

[0391] Optionally, when it is determined based on the first ultrasonic signal and the second ultrasonic signal for a continuous first preset time period that the second distance is greater than the second preset distance and / or the vibration frequency of the target sound-emitting part of the target sound-emitting object is not within the first range, the electronic device 100 exits the near-field voice mode and stops collecting audio.

[0392] In some embodiments, after determining that the audio signal is near-field speech, the Bluetooth headset may emit an ultrasonic signal with preset characteristics. Based on the characteristics of the reflected ultrasonic signal, the Bluetooth headset may determine whether the target sound source has been identified and the distance between the Bluetooth headset and the target sound source. Optionally, the preset characteristics of the emitted ultrasonic signal may include, but are not limited to, a preset frequency, preset amplitude, and emission time. The characteristics of the reflected ultrasonic signal may include, but are not limited to, the frequency and amplitude of the reflected ultrasonic signal, as well as the reflection time.

[0393] When the target sound-producing part (e.g., the mouth) outputs audio, it moves at a certain frequency and amplitude. The ultrasonic signal is reflected by the moving target sound-producing part, causing the frequency and amplitude of the ultrasonic signal to change.

[0394] The Bluetooth headset can determine the distance between the Bluetooth headset and the target sound-emitting object and the vibration frequency of the target sound-emitting part of the target sound-emitting object based on the emitted ultrasonic signal and the reflected ultrasonic signal.

[0395] Optionally, the Bluetooth headset may determine the distance between the Bluetooth headset and the target sound-emitting object based on the difference between the time when the ultrasonic signal is emitted and the time when the ultrasonic signal is reflected.

[0396] Optionally, the Bluetooth headset may determine the vibration frequency of the target sound-emitting part of the target sound-emitting object based on the difference between the frequency and / or amplitude of the emitted ultrasonic signal and the frequency and / or amplitude of the reflected ultrasonic signal.

[0397] For example, the Bluetooth headset can determine the vibration frequency of a target sound-emitting part of a target sound-emitting object based on the frequency and amplitude of the reflected ultrasonic signal of the emitted ultrasonic signal, and compare it with the frequency and amplitude of the ultrasonic signal emitted by the Bluetooth headset. The Bluetooth headset can then determine whether the target sound-emitting part has been identified based on the vibration frequency of the target sound-emitting part. For example, if the vibration frequency of the target sound-emitting part is within a first range, it can be determined that the target sound-emitting part has been identified.

[0398] Exemplarily, when the difference between the frequency and amplitude of the reflected ultrasonic signal and the frequency and amplitude of the ultrasonic signal emitted by the Bluetooth headset meets a preset value, for example, is greater than a preset value, the Bluetooth headset can determine that the target sound source is identified.

[0399] If the ultrasonic signal hits an inactive object, the frequency and amplitude of the ultrasonic signal will basically not change. The difference between the frequency and amplitude of the reflected ultrasonic signal and the frequency and amplitude of the ultrasonic signal emitted by the Bluetooth headset does not meet the preset value, for example, is less than the preset value. The Bluetooth headset can determine that the target sound source is not recognized.

[0400] The specific implementation of how the Bluetooth headset identifies the target sound source is similar to the specific implementation of how the electronic device 100 identifies the target sound source. For details, please refer to the description of S603 in the embodiment of Figure 6, and this application will not go into details here.

[0401] When the target vocalization site is identified, S906 is executed.

[0402] If the target sound source is not identified, execute S902 or S903 or S904, the Bluetooth headset continues to collect audio and monitors whether it is in earpiece mode and whether near-field voice is recognized and whether the target sound source is recognized to determine whether the Bluetooth headset enters the near-field audio collection mode.

[0403] Optionally, the Bluetooth headset may not execute S905.

[0404] Optionally, the Bluetooth headset can execute any one, any two, or all of S905, S906, and S907, which is not limited in this application.

[0405] S906 : The Bluetooth headset sends the collected audio signal to the electronic device 100 .

[0406] After determining that the Bluetooth headset is in receiver mode, or after determining near-field speech, or identifying the target sound source, the electronic device 100 can collect audio signals and send the collected audio signals to the electronic device 100 via the Bluetooth connection.

[0407] In some embodiments, after collecting the audio signal, the Bluetooth headset needs to process the audio signal, filter it to obtain the near-field voice signal, and filter out the far-field voice signal to avoid interference from the far-field voice signal.

[0408] Based on the description in S904, the difference between near-field speech and far-field speech lies in that when the distance between the target sound-emitting object and the Bluetooth headset is within a first preset distance, the energy values ​​and times at which the same audio emitted by the same sound-emitting object reaches multiple microphones on the Bluetooth headset are significantly different, and based on this, the current audio can be judged as near-field speech. When the distance between the target sound-emitting object and the Bluetooth headset exceeds the first preset distance, that is, the target sound-emitting object is far away from the Bluetooth headset, the times and energy values ​​at which the same audio emitted by the same target sound-emitting object reaches multiple microphones on the Bluetooth headset are not significantly different, and based on this, the current audio can be judged as far-field speech.

[0409] Exemplarily, the energy values ​​of the same audio signal received by multiple microphones on a Bluetooth headset are obtained, and the difference between any two of the multiple energy values ​​satisfies a preset energy difference, for example, is greater than the preset energy difference, and / or the moments of the same audio signal received by multiple microphones on a Bluetooth headset are obtained, and the difference between any two of the multiple moments satisfies a preset time difference, for example, is greater than the preset time difference, then it can be determined that the audio is near-field speech.

[0410] Exemplarily, the energy values ​​of the same audio signal received by multiple microphones on a Bluetooth headset are obtained, and the difference between any two of the multiple energy values ​​does not meet the preset energy difference, for example, is less than the preset energy difference, and / or the moments of the same audio signal received by multiple microphones on a Bluetooth headset are obtained, and the difference between any two of the multiple moments does not meet the preset time difference, for example, is less than the preset time difference, then it can be determined that the audio is far-field voice.

[0411] Based on the above features, near-field speech can be separated from the audio signal collected by the Bluetooth headset.

[0412] Optionally, the step of filtering the audio signal collected by the Bluetooth headset to obtain a near-field voice signal may also be performed by the electronic device 100, which is not limited in this application.

[0413] The specific implementation of how the Bluetooth headset separates the near-field voice from the collected audio signal is similar to the specific implementation of how the electronic device 100 separates the near-field voice from the collected audio signal. For details, please refer to the description of S604 in the embodiment of Figure 6, and this application will not go into details here.

[0414] S907 : The electronic device 100 saves the audio signal in the first application.

[0415] In response to the audio signal sent by the Bluetooth headset, the electronic device 100 saves the audio signal in the first application.

[0416] The electronic device 100 storing the collected audio signal in the first application may include: the electronic device 100 directly storing the audio signal in the first application. Alternatively, the electronic device 100 converts the audio signal into text information and stores the text information in the first application. For example, the first application may be a memo application. For details, please refer to the description of the embodiments in Figures 7B and 4D-4J.

[0417] The electronic device 100 storing the collected audio signal in the first application may also include: the electronic device 100 sending the audio signal to the electronic device 200 via the first application. Alternatively, the electronic device 100 converts the audio signal into text information, and sends the text information to the electronic device 200 via the first application. Exemplarily, the first application may be an instant messaging application. For details, please refer to the description of the embodiments in Figures 7C and 4M-4N.

[0418] S908. The Bluetooth headset continues to monitor whether the Bluetooth headset recognizes the target sound source.

[0419] After the Bluetooth headset sends the collected audio signal to the electronic device 100 , the Bluetooth headset needs to continue to monitor whether the Bluetooth headset recognizes the target sound source.

[0420] If the target sound source is not identified, the Bluetooth headset exits the near-field voice mode and stops sending the audio signal to the electronic device 100 , and then executes S902 .

[0421] When the target sound source continues to be identified, the Bluetooth headset continues to send the collected audio signal to the electronic device 100 and executes S906.

[0422] FIG10 is a flow chart of another audio acquisition method provided in this application.

[0423] S1001. A first electronic device collects a first audio signal output by a target sound-emitting object through a microphone array.

[0424] S1002: The first electronic device determines a first distance between the first electronic device and a target sound-emitting object based on a first audio signal.

[0425] S1003: When the first distance is less than a first preset distance, the first electronic device obtains a second audio signal output by the target sound-emitting object through the microphone array.

[0426] S1004: The first electronic device saves the second audio signal in the first application.

[0427] Optionally, the electronic device may store the first audio signal or may not store the first audio signal.

[0428] Optionally, the first electronic device stores the second audio signal in the first application. The first electronic device may store the second audio signal in the first application directly, or convert the second audio signal into text information and then store the text information in the first application.

[0429] The electronic device can determine a first distance from a target sound-emitting object based on the collected audio signal. When the first distance between the electronic device and the target sound-emitting object is greater than a first preset distance, or when the first distance between the electronic device and the target sound-emitting object is greater than the first preset distance for a certain period of time, the electronic device can determine that near-field speech has been recognized. The electronic device can automatically begin collecting audio and automatically save the audio within the first application. This enables the electronic device to automatically pick up audio, reduces user operations, improves audio collection efficiency, and enhances the user experience.

[0430] In a possible implementation, the first electronic device saves the second audio signal in the first application, specifically including:

[0431] In response to the first distance being less than the first preset distance, the first electronic device opens the first application and saves the second audio signal in the first application, or converts the second audio signal into a first text message and saves the first text message in the first application.

[0432] In this way, after recognizing the near-field voice, the electronic device can automatically save the collected audio signal or the text information corresponding to the collected audio signal in the first application.

[0433] Exemplarily, the first application may be a memo application.

[0434] In one possible implementation, the first electronic device saves the second audio signal in the first application, specifically including: in response to the first distance being less than the first preset distance, the first electronic device opens the first application and sends the second audio signal to the second electronic device through the first application, or the first electronic device converts the second audio signal into a first text message and sends the first text message to the second electronic device through the first application.

[0435] In this way, after recognizing the near-field voice, the electronic device can automatically send the collected audio signal to the second electronic device with which the communication connection is established, or send the text information corresponding to the collected audio signal to the second electronic device with which the communication connection is established.

[0436] Exemplarily, the first application may be an instant messaging application.

[0437] In one possible implementation, the first electronic device is connected to a third electronic device via Bluetooth; the first electronic device saves the second audio signal in the first application, specifically including: the first electronic device sends the second audio signal to the third electronic device via Bluetooth connection; the third electronic device saves the second audio signal in the first application.

[0438] Optionally, the first electronic device may be a Bluetooth headset. After recognizing the near-field voice, the Bluetooth headset may automatically send the collected audio signal to a third electronic device with which a Bluetooth connection is established, or send a text message corresponding to the collected audio signal to the third electronic device with which a Bluetooth connection is established.

[0439] In one possible implementation, while the first electronic device obtains the second audio signal through the microphone array, the first electronic device also obtains the third audio signal output by other sound-emitting objects through the microphone array; the electronic device saves the second audio signal in the first application, specifically including: when the first electronic device determines that the distance between the first electronic device and the other sound-emitting objects exceeds a first preset distance, the electronic device saves the second audio signal in the first application.

[0440] Optionally, the second audio signal and the third audio signal may be audio signals in the same audio file. The electronic device 100 may extract the second audio signal or the third audio signal from the same audio file.

[0441] Optionally, the second audio signal and the third audio signal may also be audio signals in two audio files respectively.

[0442] Near-field speech may refer to audio output from a target sound-emitting object within a first preset distance from the electronic device. Far-field speech may refer to audio output from a target sound-emitting object beyond the first preset distance from the electronic device.

[0443] In this way, after recognizing near-field speech and starting to collect audio signals, the electronic device can determine whether the collected audio signal is a near-field audio signal or a far-field audio signal, and only save the near-field audio signal without saving the far-field audio signal, thereby avoiding interference from the far-field audio signal.

[0444] For details on how electronic devices distinguish between near-field speech and far-field speech, please refer to the detailed descriptions in formulas (1) to (4).

[0445] In one possible implementation, when the first distance is less than a first preset distance, the first electronic device obtains the second audio signal output by the target sound-emitting object through the microphone array, specifically including: the first electronic device determines the probability of the voice signal contained in the first audio signal; when the first distance is less than the first preset distance and the probability of the voice signal contained in the first audio signal is greater than a first threshold, the first electronic device obtains the second audio signal output by the target sound-emitting object through the microphone array.

[0446] In this way, before starting to collect audio, the electronic device can determine the probability of a voice signal contained in the audio signal. When the probability of a voice signal contained in the audio signal is greater than the first threshold, that is, the probability of recognizing that the user is speaking is high, the electronic device can collect and save the second audio signal.

[0447] In one possible implementation, a first electronic device is connected to a third electronic device via Bluetooth; the first electronic device obtains a second audio signal output by a target sound-emitting object through a microphone array, specifically including: when the first electronic device determines that the first electronic device is in a handheld state, the first distance is less than a first preset distance, and the probability of a voice signal contained in the first audio signal is greater than a first threshold, the first electronic device obtains the second audio signal output by the target sound-emitting object through the microphone array.

[0448] Optionally, a sensor is pre-installed on the first electronic device, and whether the device is in a handheld state can be confirmed through a sensor signal collected by the sensor.

[0449] In this way, when the first electronic device is a Bluetooth headset, and the Bluetooth headset is in a handheld state, it can be assumed that the current user has the intention to speak into the Bluetooth headset. When the Bluetooth headset is in a handheld state, the distance between the Bluetooth headset and the target sound-emitting object is within a first preset distance, and the probability of a voice signal contained in the first audio signal is greater than a first threshold, the Bluetooth headset can enter near-field voice mode and begin automatically collecting and saving audio signals, thereby improving the accuracy of the Bluetooth headset entering near-field voice mode.

[0450] In one possible implementation, the first electronic device includes a speaker; before the first electronic device obtains the second audio signal output by the target sound-emitting object through the microphone array, the method also includes: the first electronic device emits a first ultrasonic signal through the speaker; the first electronic device receives the reflected second ultrasonic signal; the first electronic device determines the second distance between the first electronic device and the target sound-emitting object and the vibration frequency of the target sound-emitting part of the target sound-emitting object based on the first ultrasonic signal and the second ultrasonic signal; when the second distance is less than the second preset distance and the vibration frequency of the target sound-emitting part of the target sound-emitting object is within the first range, the first electronic device obtains the second audio signal output by the target sound-emitting object through the microphone array.

[0451] Exemplarily, the first range may be between 20 Hz and 40 Hz.

[0452] Optionally, the second preset distance may be smaller than the first preset distance. Optionally, the second preset distance may be equal to the first preset distance.

[0453] Optionally, the target sound-emitting object may generate sound at a certain vibration frequency. The first electronic device may obtain the vibration frequency of the target sound-emitting part of the target sound-emitting object through the first ultrasonic signal and the second ultrasonic signal, and determine whether the target sound-emitting object is generating sound based on the obtained vibration frequency of the target sound-emitting object. If the vibration frequency of the target sound-emitting object is within a first range, it can be determined that the target sound-emitting object is generating sound.

[0454] In this way, when the first electronic device enters the near-field voice mode, it can be confirmed whether the first electronic device and the target sound-emitting object are close enough and whether the target sound-emitting part of the target sound-emitting object is emitting sound, thereby improving the accuracy of the first electronic device entering the near-field voice mode.

[0455] In one possible implementation, after the first electronic device saves the second audio signal in the first application, the method further includes: when it is determined based on the first ultrasonic signal and the second ultrasonic signal for a continuous first preset time period that the second distance is greater than the second preset distance and / or the vibration frequency of the target sound-emitting part of the target sound-emitting object is not within the first range, the first electronic device stops acquiring the audio signal through the microphone array.

[0456] After the first electronic device enters near-field voice mode, it must continue to monitor whether the first electronic device and the target sound-emitting object are sufficiently close, and whether the target sound-emitting part of the target sound-emitting object is emitting sound. If it is detected that the distance between the first electronic device and the target sound-emitting object exceeds a second preset distance, and / or the target sound-emitting part of the target sound-emitting object is not emitting sound, the first electronic device can automatically exit near-field voice mode and stop collecting audio.

[0457] In one possible implementation, the microphone array includes a first microphone and a second microphone, and the positions of the first microphone and the second microphone are different on the first electronic device; the first electronic device collects a first audio signal output by the target sound-emitting object through the microphone array, specifically including: the first electronic device collects the first audio signal output by the target sound-emitting object through the first microphone and the second microphone respectively; the first electronic device determines a first distance between the first electronic device and the target sound-emitting object based on the first audio signal, specifically including: the first electronic device obtains a first energy value of the first audio signal collected by the first microphone and a second energy value of the first audio signal collected by the second microphone; the first electronic device determines the first distance between the first electronic device and the target sound-emitting object based on the difference between the first energy value and the second energy value; and / or, the first electronic device obtains a first moment of the first audio signal collected by the first microphone and a second moment of the first audio signal collected by the second microphone; the first electronic device determines the first distance between the first electronic device and the target sound-emitting object based on the difference between the first moment and the second moment.

[0458] In this way, the electronic device can determine the distance between the first electronic device and the target sound-emitting object by the difference in energy values ​​of the same audio signal collected by multiple microphones in the microphone array and / or the difference between the times of receiving the same audio signal.

[0459] In one possible implementation, the microphone array includes a first microphone and a second microphone, and the positions of the first microphone and the second microphone are different on the first electronic device. The speaker includes a first speaker and a second speaker, and the first speaker is located near the first microphone and the second speaker is located near the second microphone; the first electronic device sends a first ultrasonic signal through the speaker, specifically including: when the first electronic device determines that the distance between the first microphone and the target sound-emitting object is less than the distance between the second microphone and the target sound-emitting object, the first electronic device sends the first ultrasonic signal through the first speaker, wherein the first speaker is the speaker closest to the target sound-emitting object.

[0460] In combination with the first aspect, in a possible implementation, the first electronic device receives the reflected second ultrasonic signal, which specifically includes: the first electronic device receives the reflected second ultrasonic signal through a first microphone.

[0461] In this way, the first electronic device can emit an ultrasonic signal through a speaker closest to the target sound-emitting object and receive a reflected ultrasonic signal through a microphone closest to the target sound-emitting object, thereby improving the accuracy of identifying the target sound-emitting object.

[0462] In one possible implementation, the first electronic device determines that the distance between the first microphone and the target sound-emitting object is smaller than the distance between the second microphone and the target sound-emitting object, specifically including: the first electronic device obtains a first energy value of the first audio signal collected by the first microphone and a second energy value of the first audio signal collected by the second microphone; when the first energy value is greater than the second energy value, the first electronic device determines that the distance between the first microphone and the target sound-emitting object is smaller than the distance between the second microphone and the target sound-emitting object; and / or, the first electronic device obtains a first moment of the first audio signal collected by the first microphone and a second moment of the first audio signal collected by the second microphone; when the first moment is less than the second moment, the first electronic device determines that the distance between the first microphone and the target sound-emitting object is smaller than the distance between the second microphone and the target sound-emitting object.

[0463] The above are only some of the embodiments and implementations of this application. The scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0464] It is understood that the various user interfaces described in the embodiments of this application are merely exemplary interfaces and do not limit the scope of this application. In other embodiments, the user interface may adopt a different interface layout, include more or fewer controls, and add or remove other functional options. As long as they are based on the same inventive concept provided by this application, they are all within the scope of protection of this application.

[0465] It should be noted that, without causing any contradiction or conflict, any feature in any embodiment of the present application, or any part of any feature, can be combined, and the combined technical solution is also within the scope of the embodiments of the present application.

[0466] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. An audio acquisition method, characterized in that, The first electronic device includes a microphone array, and the microphone array includes at least two microphones. The method includes: The first electronic device acquires a first audio signal output by a target sound - emitting object through the microphone array; The first electronic device determines a first distance between the first electronic device and the target sound - emitting object based on the first audio signal; When the first distance is less than a first preset distance, the first electronic device acquires a second audio signal output by the target sound - emitting object through the microphone array; The first electronic device saves the second audio signal in a first application.

2. The method according to claim 1, characterized in that The first electronic device saves the second audio signal in a first application, specifically including: In response to the first distance being less than the first preset distance, the first electronic device activates the first application and saves the second audio signal in the first application, or the first electronic device converts the second audio signal into first text information and saves the first text information in the first application.

3. The method according to claim 1, wherein The first electronic device saves the second audio signal in a first application, specifically including: In response to the first distance being less than the first preset distance, the first electronic device activates the first application and sends the second audio signal to a second electronic device through the first application, or the first electronic device converts the second audio signal into first text information and sends the first text information to the second electronic device through the first application.

4. The method according to claim 1, wherein The first electronic device is connected to a third electronic device via Bluetooth; the first electronic device saves the second audio signal in a first application, specifically including: The first electronic device sends the second audio signal to the third electronic device through the Bluetooth connection; The third electronic device saves the second audio signal in the first application.

5. The method according to any one of claims 1-4, characterized in that While the first electronic device acquires the second audio signal through the microphone array, the first electronic device also acquires a third audio signal output by other sound - emitting objects through the microphone array; the electronic device saves the second audio signal in a first application, specifically including: When the first electronic device determines that the distance between the first electronic device and the other sound - emitting objects exceeds the first preset distance, the electronic device saves the second audio signal in the first application.

6. The method according to any one of claims 1-5, characterized in that When the first distance is less than a first preset distance, the first electronic device acquires a second audio signal output by the target sound - emitting object through the microphone array, specifically including: The first electronic device determines the probability of the voice signal included in the first audio signal; When the first distance is less than the first preset distance and the probability of the voice signal included in the first audio signal is greater than a first threshold, the first electronic device acquires the second audio signal output by the target sound - emitting object through the microphone array.

7. The method according to claim 6, wherein The first electronic device is connected to a third electronic device via Bluetooth; the first electronic device obtains a second audio signal output by the target sound - emitting object through the microphone array, specifically including: When the first electronic device determines that the first electronic device is in a handheld state, the first distance is less than a first preset distance, and the probability of the voice signal included in the first audio signal is greater than a first threshold, the first electronic device obtains the second audio signal output by the target sound - emitting object through the microphone array.

8. The method according to any one of claims 1 to 7, characterized in that, The first electronic device includes a speaker; before the first electronic device obtains the second audio signal output by the target sound - emitting object through the microphone array, the method further includes: The first electronic device emits a first ultrasonic signal through the speaker; The first electronic device receives a reflected second ultrasonic signal; The first electronic device determines a second distance between the first electronic device and the target sound - emitting object and the vibration frequency of the target sound - emitting part of the target sound - emitting object based on the first ultrasonic signal and the second ultrasonic signal; When the second distance is less than a second preset distance and the vibration frequency of the target sound - emitting part of the target sound - emitting object is within a first range, the first electronic device obtains the second audio signal output by the target sound - emitting object through the microphone array.

9. The method according to claim 8, wherein After the first electronic device saves the second audio signal in a first application, the method further includes: When it is determined that the second distance is greater than the second preset distance and / or the vibration frequency of the target sound - emitting part of the target sound - emitting object is not within the first range based on the first ultrasonic signal and the second ultrasonic signal for a continuous first preset duration, the first electronic device stops obtaining audio signals through the microphone array.

10. The method according to any one of claims 1-9, characterized in that, The microphone array includes a first microphone and a second microphone, and the positions of the first microphone and the second microphone on the first electronic device are different; The first electronic device collects a first audio signal output by the target sound - emitting object through the microphone array, specifically including: The first electronic device collects the first audio signal output by the target sound - emitting object through the first microphone and the second microphone respectively; The first electronic device determines the first distance between the first electronic device and the target sound - emitting object based on the first audio signal, specifically including: The first electronic device obtains a first energy value of the first audio signal collected by the first microphone and a second energy value of the first audio signal collected by the second microphone; The first electronic device determines the first distance between the first electronic device and the target sound - emitting object based on the difference between the first energy value and the second energy value; and / or The first electronic device obtains a first time of the first audio signal collected by the first microphone and a second time of the first audio signal collected by the second microphone; The first electronic device determines the first distance between the first electronic device and the target sound - emitting object based on the difference between the first moment and the second moment.

11. The method according to claim 8 or 9, characterized in that, The microphone array includes a first microphone and a second microphone, and the positions of the first microphone and the second microphone on the first electronic device are different. The speaker includes a first speaker and a second speaker. The first speaker is located near the first microphone, and the second speaker is located near the second microphone; The first electronic device emits a first ultrasonic signal through the speaker, specifically including: When the first electronic device determines that the distance between the first microphone and the target sound - emitting object is less than the distance between the second microphone and the target sound - emitting object, the first electronic device emits the first ultrasonic signal through the first speaker, where the first speaker is the speaker closest to the target sound - emitting object.

12. The method according to claim 11, wherein The first electronic device receives the reflected second ultrasonic signal, specifically including: The first electronic device receives the reflected second ultrasonic signal through the first microphone.

13. The method according to claim 11 or 12, characterized in that, The first electronic device determines that the distance between the first microphone and the target sound - emitting object is less than the distance between the second microphone and the target sound - emitting object, specifically including: The first electronic device obtains a first energy value of the first audio signal collected by the first microphone and a second energy value of the first audio signal collected by the second microphone; When the first energy value is greater than the second energy value, the first electronic device determines that the distance between the first microphone and the target sound - emitting object is less than the distance between the second microphone and the target sound - emitting object; and / or, The first electronic device obtains a first moment of the first audio signal collected by the first microphone and a second moment of the first audio signal collected by the second microphone; When the first moment is less than the second moment, the first electronic device determines that the distance between the first microphone and the target sound - emitting object is less than the distance between the second microphone and the target sound - emitting object.

14. An electronic device, characterized in that, The electronic device includes a microphone array, a memory, and a processor; wherein, the microphone array, the memory, and the processor are coupled. The memory is used to store a computer program. When the processor executes and calls the computer program, the electronic device executes the method according to any one of claims 1 - 13.

15. A computer-readable storage medium, comprising instructions, characterized in that, When the instruction runs on the electronic device, the electronic device executes the method according to any one of claims 1 - 13.

Citation Information

Patent Citations

  • Voice interaction awakening electronic device based on microphone signal, method and medium

    CN110097875A

  • Voice interaction wake-up electronic equipment based on microphone signals, method and medium

    CN110428806A

  • Voice interaction function awakening method and electronic equipment

    CN117119102A

  • Audio acquisition method, electronic equipment and storage medium

    CN119182851A

  • Parametric Spatial Audio Rendering with Near-Field Effect

    US20230362537A1

Cited By

  • Display device and speech recognition method thereof

    CN121171221A