Audio acquisition method, electronic equipment and computer storage medium

By using a microphone array consisting of directional and omnidirectional microphones in electronic devices, combined with multi-level decision technology, the problem of accidental touches during voice wake-up of electronic devices has been solved, achieving higher audio acquisition accuracy.

CN121747602APending Publication Date: 2026-03-27HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing electronic devices are prone to accidental touches when activated by voice, leading to inaccurate audio capture.

Method used

Near-field speech recognition is achieved by combining two electronic devices and using a microphone array consisting of directional and omnidirectional microphones to collect multiple sound signals and perform multi-level judgments, thereby improving the accuracy of near-field speech recognition.

Benefits of technology

It improves the accuracy of automatic audio acquisition by electronic devices, reduces accidental touches, and enhances the precision of audio acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747602A_ABST
    Figure CN121747602A_ABST
Patent Text Reader

Abstract

The invention provides an audio acquisition method, electronic equipment and a computer storage medium. According to the embodiment of the invention, the second electronic equipment can confirm whether the near-field voice is recognized or not through the first sound signal collected by the first electronic equipment; and the second electronic equipment further confirms whether the near-field voice is recognized by combining the third sound signal acquired by the first electronic equipment and the second sound signal acquired by the second electronic equipment, so that the power consumption of the second electronic equipment can be reduced. Moreover, according to the method, whether the near-field voice is recognized or not is jointly confirmed in combination with the sound signals collected by the multiple devices, the accuracy of recognizing the near-field voice by the second electronic device can be improved, and then the accuracy of automatically collecting the audio by the first electronic device and / or the second electronic device is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and in particular to an audio acquisition method, electronic device, and computer storage medium. Background Technology

[0002] With the development of terminal technology, users' functional needs for electronic devices are becoming increasingly diversified. To meet users' needs for sound recording, most electronic devices support audio capture functions, using their own microphones (mics) to capture audio. For example, users can wake up the electronic device to capture audio using voice commands; however, voice wake-up is prone to accidental activation. How to provide a method for accurately controlling the automatic audio capture of electronic devices requires further research. Summary of the Invention

[0003] This application provides an audio acquisition method, an electronic device, and a computer storage medium. This method, by combining two devices to recognize near-field speech, can improve the accuracy of near-field speech recognition by the electronic device, thereby improving the accuracy of automatic audio acquisition by the electronic device.

[0004] In a first aspect, this application provides an audio acquisition method, applied to an audio acquisition system including a first electronic device and a second electronic device, the first electronic device and the second electronic device establishing a communication connection; the method includes: the first electronic device acquiring a first sound signal; the first electronic device sending a first message to the second electronic device, the first message being sent after the first electronic device confirms that the first sound signal meets the first condition; in response to the first message, the second electronic device acquiring a second sound signal; the second electronic device receiving a third sound signal sent by the first electronic device, the acquisition time of the third sound signal being later than the acquisition time of the first sound signal; the first electronic device performing a first functional operation based on a fourth sound signal acquired by the first electronic device and / or a fifth sound signal acquired by the second electronic device, the first functional operation being performed after the first electronic device confirms that the second sound signal and the third sound signal meet the second condition, the acquisition time of the fourth sound signal being later than the acquisition time of the second sound signal, the acquisition time of the fifth sound signal being later than the acquisition time of the third sound signal, the fifth sound signal being sent from the second electronic device to the first electronic device.

[0005] Optionally, after the first electronic device is powered on, it can continuously capture audio.

[0006] The first electronic device can confirm whether near-field speech has been recognized based on the acquired first sound signal. If the first electronic device recognizes near-field speech, it sends a first message to the second electronic device, instructing the second electronic device to start acquiring a second sound signal. The first electronic device also needs to acquire a third sound signal and send it to the second electronic device. The second electronic device can further confirm whether near-field speech has been recognized based on the second and third sound signals. If the second electronic device confirms that near-field speech has been recognized based on the second and third sound signals, it can perform corresponding functional operations based on the sound signals acquired by the first and / or second electronic devices. In other words, this method recognizes near-field speech through a two-level determination, which can improve the accuracy of near-field speech recognition by the electronic device, thereby improving the accuracy of automatic audio acquisition by the electronic device.

[0007] Optionally, the first electronic device performs a first functional operation based on a fourth sound signal collected by the first electronic device and / or a fifth sound signal collected by the second electronic device. This can mean that the first electronic device performs the first functional operation based on the second and fourth sound signals collected by the first electronic device, and / or the first, third, and fifth sound signals collected by the second electronic device. Alternatively, it can mean that the first electronic device performs the first functional operation based solely on the fourth sound signal collected by the first electronic device, and / or the fifth sound signal collected by the second electronic device.

[0008] Optionally, before sending the first message to the second electronic device, the first electronic device can also identify the probability of the speech signal contained in the first sound signal. If the probability of the speech signal contained in the first sound signal is greater than a preset value, the first electronic device then sends the first message to the second electronic device, thus avoiding noise interference. Here, the speech signal can be understood as audio output by a person.

[0009] For example, the energy value of the audio output by a person is much greater than the energy value of noise, and the first electronic device can determine whether the first sound signal contains a speech signal based on the energy value of the first sound signal.

[0010] In conjunction with the first aspect, in one possible implementation, the second and third sound signals are sound signals collected within the same time period.

[0011] In this way, the second electronic device can improve the accuracy of near-field speech recognition by recognizing the second and third sound signals collected within the same time period.

[0012] In conjunction with the first aspect, in one possible implementation, the first sound signal and the third sound signal are sound signals collected by a microphone in the first electronic device.

[0013] In conjunction with the first aspect, in one possible implementation, when the first sound signal and the third sound signal are sound signals collected by an omnidirectional microphone in the first electronic device, the first electronic device further includes an auxiliary recognition unit, and the first message is sent after the first electronic device confirms that the first sound signal meets the first condition and the first signal obtained by the auxiliary recognition unit meets the third condition.

[0014] Thus, when the first electronic device includes an omnidirectional microphone, the first electronic device can use the first signal obtained by the auxiliary recognition unit to assist the first electronic device in recognizing near-field speech based on the sound signal collected by the omnidirectional microphone, thereby improving the accuracy of the first electronic device in recognizing near-field speech.

[0015] Optionally, the first signal may include a first ultrasonic signal output by the auxiliary identification unit and a second ultrasonic signal reflected from the target sound-emitting part received by the auxiliary identification unit.

[0016] In conjunction with the first aspect, in one possible implementation, the second sound signal is a sound signal collected by an omnidirectional microphone in a second electronic device.

[0017] In this way, the directional microphone on the first electronic device and the omnidirectional microphone on the second electronic device can form a microphone array. The second electronic device can recognize near-field speech based on the third sound signal collected by the directional microphone on the first electronic device and the second sound signal collected by the omnidirectional microphone on the second electronic device, which can improve the accuracy of the second electronic device in recognizing near-field speech.

[0018] In conjunction with the first aspect, in one possible implementation, the first condition includes: the energy value of the sound signal in the first frequency band of the first sound signal is greater than a first value, and the difference between the energy value of the first frequency band and the energy values ​​of other frequency points besides the first frequency band is greater than a second value.

[0019] Thus, based on the above judgment conditions, the first electronic device can recognize near-field speech based on the first sound signal collected by the directional microphone.

[0020] In conjunction with the first aspect, in one possible implementation, the first condition includes: the energy value of the first sound signal is greater than the third value; the first signal includes the first transmission time of the first ultrasonic signal and the first reception time of the second ultrasonic signal; the third condition includes: the difference between the first transmission time and the first reception time is less than the sixth value; and / or, the vibration frequency of the target sound-emitting part obtained based on the first ultrasonic signal and the second ultrasonic signal is within a first range.

[0021] Thus, based on the above judgment conditions, the first electronic device can use the auxiliary recognition unit to assist the omnidirectional microphone in recognizing near-field speech based on the collected first sound signal.

[0022] In conjunction with the first aspect, in one possible implementation, when the third sound signal is a sound signal collected by an omnidirectional microphone in the first electronic device, the second condition includes: the difference between the energy value of the second sound signal and the energy value of the third sound signal is greater than a fifth value; and / or, the difference between the collection time of the second sound signal and the collection time of the third sound signal is greater than a first duration.

[0023] Thus, based on the above judgment conditions, the second electronic device can recognize near-field speech based on the third sound signal collected by the omnidirectional microphone in the first electronic device and the second sound signal collected by the omnidirectional microphone in the second electronic device.

[0024] In conjunction with the first aspect, in one possible implementation, when the third sound signal is a sound signal collected by a microphone in the first electronic device, the second condition includes: the difference between the energy value of the sound signal in the second frequency band of the third sound signal and the energy value corresponding to the second frequency band of the second sound signal is greater than the fourth value.

[0025] In one possible implementation, the second condition also includes any one or more of the following: the energy value of the sound signal in the second frequency band of the third sound signal is greater than the first value; the difference between the energy value of the second frequency band and the energy values ​​of other frequency points besides the second frequency band is greater than the second value; and the energy value of the second sound signal is greater than the third value.

[0026] Thus, based on the above judgment conditions, the second electronic device can recognize near-field speech based on the third sound signal collected by the directional microphone in the first electronic device and the second sound signal collected by the omnidirectional microphone in the second electronic device.

[0027] In conjunction with the first aspect, in one possible implementation, the first functional operation includes any one or more of the following: saving the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device; converting the sound signal collected by the first electronic device's acquisition unit and / or the sound signal collected by the second electronic device into text information and saving the text information; recognizing shortcut instructions in the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, and based on the shortcut instructions.

[0028] After the second electronic device also recognizes near-field speech, the second electronic device can perform the first function operation based on the sound signal collected by the first electronic device. The second electronic device can also perform the first function operation based on the sound signal collected by the second electronic device, and the second electronic device can also perform the first function operation based on the sound signal collected by the first electronic device and the sound signal collected by the second electronic device.

[0029] For example, when the second electronic device performs a first function operation based on the sound signal collected by the first electronic device, the sound signal collected by the first electronic device may include a first sound signal, a third sound signal, and a sound signal collected after the third sound signal, or may include a sound signal collected after the third sound signal.

[0030] For example, when the second electronic device performs a first function operation based on the sound signal collected by the second electronic device, the sound signal collected by the second electronic device may include a second sound signal and a sound signal collected after the second sound signal, or may include a sound signal collected after the second sound signal.

[0031] In conjunction with the first aspect, in one possible implementation, the first electronic device performs a first functional operation based on the sound signals collected by the first electronic device and / or the sound signals collected by the second electronic device, specifically including: the first electronic device separating a sixth sound signal and a seventh sound signal from the sound signals collected by the first electronic device and / or the sound signals collected by the second electronic device, wherein the sixth sound signal is the sound signal output by a target sound-emitting object whose distance from the first electronic device and the second electronic device is greater than a first preset distance, and the seventh sound signal is the sound signal output by a target sound-emitting object whose distance from the first electronic device and the second electronic device is less than the first preset distance; and the first electronic device performs the first functional operation based on the seventh sound signal.

[0032] In this way, the second electronic device can improve the accuracy of recognizing the operation intentions included in the near-field voice when performing the first function operation based on near-field voice. The operation intentions included in the near-field voice are used to trigger the second electronic device to perform the first function operation.

[0033] In conjunction with the first aspect, in one possible implementation, the first electronic device is a stylus, and the second electronic device is a tablet.

[0034] Optionally, the first electronic device can also be headphones, and the second electronic device can be a tablet.

[0035] Secondly, this application provides an audio acquisition method, the method comprising: in response to a first message sent by a first electronic device, a second electronic device acquiring a second sound signal, wherein the first message is sent after the first electronic device confirms that the first sound signal meets a first condition, and the first sound signal is the sound signal acquired by the first electronic device; the second electronic device receiving a third sound signal sent by the first electronic device, wherein the acquisition time of the third sound signal is later than the acquisition time of the first sound signal; the first electronic device performing a first functional operation based on a fourth sound signal acquired by the first electronic device and / or a fifth sound signal acquired by the second electronic device, wherein the first functional operation is performed after the first electronic device confirms that the second sound signal and the third sound signal meet a second condition, wherein the acquisition time of the fourth sound signal is later than the acquisition time of the second sound signal, and the acquisition time of the fifth sound signal is later than the acquisition time of the third sound signal, and the fifth sound signal is sent by the second electronic device to the first electronic device.

[0036] In conjunction with the second aspect, in one possible implementation, the second and third sound signals are sound signals collected within the same time period.

[0037] In conjunction with the second aspect, in one possible implementation, the first sound signal and the third sound signal are sound signals collected by a microphone in the first electronic device.

[0038] In conjunction with the second aspect, in one possible implementation, when the first sound signal and the third sound signal are sound signals collected by an omnidirectional microphone in the first electronic device, the first electronic device further includes an auxiliary recognition unit, and the first message is sent after the first electronic device confirms that the first sound signal meets the first condition and the first signal obtained by the auxiliary recognition unit meets the third condition.

[0039] In conjunction with the second aspect, in one possible implementation, the second sound signal is a sound signal collected by an omnidirectional microphone in a second electronic device.

[0040] In conjunction with the second aspect, in one possible implementation, the first condition includes: the energy value of the first sound signal is greater than the third value; the first signal includes the first transmission time of the first ultrasonic signal and the first reception time of the second ultrasonic signal; the third condition includes: the difference between the first transmission time and the first reception time is less than the sixth value; and / or, the vibration frequency of the target sound-emitting part obtained based on the first ultrasonic signal and the second ultrasonic signal is within a first range.

[0041] In conjunction with the second aspect, in one possible implementation, when the third sound signal is a sound signal collected by an omnidirectional microphone in the first electronic device, the second condition includes: the difference between the energy value of the second sound signal and the energy value of the third sound signal is greater than a fifth value; and / or, the difference between the collection time of the second sound signal and the collection time of the third sound signal is greater than a first duration.

[0042] In conjunction with the second aspect, in one possible implementation, when the third sound signal is a sound signal collected by a microphone in the first electronic device, the second condition includes: the difference between the energy value of the sound signal in the second frequency band of the third sound signal and the energy value corresponding to the second frequency band of the second sound signal is greater than the fourth value.

[0043] In conjunction with the second aspect, in one possible implementation, the second condition also includes any one or more of the following: the energy value of the sound signal in the second frequency band of the third sound signal is greater than the first value; the difference between the energy value of the second frequency band and the energy values ​​of other frequency points besides the second frequency band is greater than the second value; and the energy value of the second sound signal is greater than the third value.

[0044] In conjunction with the second aspect, in one possible implementation, the first functional operation includes any one or more of the following: saving the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device; converting the sound signal collected by the first electronic device's acquisition unit and / or the sound signal collected by the second electronic device into text information and saving the text information; recognizing shortcut instructions in the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, and based on the shortcut instructions.

[0045] In conjunction with the second aspect, in one possible implementation, the first electronic device performs a first functional operation based on the sound signals collected by the first electronic device and / or the sound signals collected by the second electronic device. Specifically, the first electronic device separates a sixth sound signal and a seventh sound signal from the sound signals collected by the first electronic device and / or the sound signals collected by the second electronic device. The sixth sound signal is the sound signal output by a target sound-emitting object whose distance from the first electronic device and the second electronic device is greater than a first preset distance, and the seventh sound signal is the sound signal output by a target sound-emitting object whose distance from the first electronic device and the second electronic device is less than the first preset distance. The first electronic device performs the first functional operation based on the seventh sound signal.

[0046] In conjunction with the second aspect, in one possible implementation, the first electronic device is a stylus, and the second electronic device is a tablet.

[0047] Thirdly, this application provides an electronic device, which is a second electronic device. The second electronic device includes a memory and a processor, wherein the memory is used to store a computer program; and the processor is used to invoke the computer program, causing the second electronic device to execute an audio acquisition method provided in any possible implementation of any of the above aspects.

[0048] Fourthly, this application provides an electronic device, namely a first electronic device, which includes a memory and a processor, wherein the memory is used to store a computer program; and the processor is used to invoke the computer program to cause the first electronic device to execute an audio acquisition method provided in any possible implementation of the first aspect.

[0049] Fifthly, this application provides a computationally readable storage medium including instructions that, when executed on a second electronic device, cause the second electronic device to perform an audio acquisition method provided in any possible implementation of any of the above aspects.

[0050] In a sixth aspect, this application provides a computationally readable storage medium including instructions that, when executed on a first electronic device, cause the first electronic device to perform an audio acquisition method provided in any possible implementation of the first aspect.

[0051] In a seventh aspect, this application provides a computer program product comprising computer instructions that, when executed on a second electronic device, cause the second electronic device to perform an audio acquisition method provided in any possible implementation of any of the above aspects.

[0052] Eighthly, this application provides a computer program product comprising computer instructions that, when executed on a first electronic device, cause the first electronic device to perform an audio acquisition method provided in any possible implementation of the first aspect described above.

[0053] Ninthly, this application provides a chip applied to a second electronic device, the chip including one or more processors, the processors being configured to invoke computer instructions to cause the second electronic device to execute an audio acquisition method provided in any possible implementation of any of the above aspects.

[0054] In a tenth aspect, this application provides a chip applied to a first electronic device, the chip including one or more processors, the processors being configured to invoke computer instructions to cause the first electronic device to execute an audio acquisition method provided in any possible implementation of the first aspect described above.

[0055] For the beneficial effects of aspects two through ten, please refer to the description of the beneficial effects in aspect one, which will not be repeated here. Attached Figure Description

[0056] Figure 1A A schematic diagram of a preset pickup direction for a directional microphone is shown.

[0057] Figure 1B A schematic diagram of the preset pickup direction of another directional microphone is shown;

[0058] Figure 1C A schematic diagram of the pickup direction of an omnidirectional microphone is shown;

[0059] Figure 2 A schematic diagram of an audio acquisition system provided in this application;

[0060] Figure 3 An exemplary schematic diagram of the hardware structure of electronic device 100 is shown;

[0061] Figure 4 An exemplary schematic diagram of the hardware structure of electronic device 200 is shown;

[0062] Figure 5 A schematic diagram of a software module for recognizing near-field speech is shown in electronic devices 100 and 200.

[0063] Figures 6A-6E An exemplary diagram illustrates a set of electronic devices 200 automatically acquiring audio;

[0064] Figure 7 This diagram illustrates a method for jointly recognizing near-field speech using electronic devices 100 and 200.

[0065] Figure 8A The diagram shows the frequency response curve of the sound signal collected by the directional microphone when the sound-emitting object is close to the directional microphone and emits sound in the preset pickup direction of the directional microphone or in the direction extending from the preset pickup direction.

[0066] Figure 8B The diagram shows the frequency response curve of the sound signal collected by the directional microphone when the target sound-emitting object is far away from the directional microphone;

[0067] Figure 8C This diagram illustrates the frequency response curve of the sound signal collected by a directional microphone when the target sound-emitting object emits sound in a non-preset pickup direction or in a direction extending beyond the non-preset pickup direction.

[0068] Figure 8D The frequency response curve of the first sound signal acquired by the first audio acquisition device is shown;

[0069] Figure 8EThe frequency response curves of the second and third sound signals are shown in the diagram. Detailed Implementation

[0070] The technical solutions of the embodiments of this application are described below with reference to the accompanying drawings. In the description of the embodiments of this application, the terminology used in the following embodiments is for the purpose of describing specific embodiments only and is not intended to limit the application. As used in the specification and appended claims of this application, the singular expressions "a," "the," "the," "the," and "this" are intended to also include expressions such as "one or more," unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, "at least one" and "one or more" refer to one or more (including two). The term "and / or" is used to describe the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0071] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized. The term "connection" includes direct connections and indirect connections, unless otherwise stated. "First" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated.

[0072] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0073] First, the technical terms used in this application will be explained.

[0074] 1. Directional microphone

[0075] Directional microphones have varying sensitivities to picking up sound signals from different directions. In a preset pickup direction, a directional microphone has higher sensitivity. In non-preset pickup directions, a directional microphone has lower sensitivity.

[0076] Different directional microphones have different structures, and their preset pickup directions also differ.

[0077] For example, the structure of a directional microphone can be a "figure-eight" microphone, and the preset pickup direction of the "figure-eight" microphone can present a "figure-eight" area.

[0078] For example, Figure 1A A schematic diagram of a preset pickup direction for a directional microphone is shown.

[0079] like Figure 1A As shown, the preset pickup direction of a directional microphone can present a "figure-eight" shaped area. For example, the preset pickup direction can be... Figure 1A The solid-lined area is shown. If the target sound-emitting object is in... Figure 1A Within the solid line area shown or Figure 1A The sound originates in the direction of the extended area shown by the solid line. A directional microphone can effectively capture the sound signal output by the target sound-producing object.

[0080] For example, if the structure of a directional microphone is a "cardioid" microphone, then the preset pickup direction of the "cardioid" microphone can present a "cardioid" area.

[0081] For example, Figure 1B A schematic diagram of the preset pickup direction of another directional microphone is shown.

[0082] like Figure 1B As shown, the preset pickup direction of a directional microphone can present a "heart-shaped" area. For example, the preset pickup direction can be... Figure 1B The solid-lined area is shown. If the target sound-emitting object is in... Figure 1B Within the solid line area shown or Figure 1B The sound originates in the direction of the extended area shown by the solid line. A directional microphone can effectively capture the sound signal output by the target sound-producing object.

[0083] Optionally, and not limited to "figure-eight" and "cardioid" microphones, directional microphones may include many other structures. This application only uses "figure-eight" and "cardioid" microphones to explain the default pickup direction of directional microphones, and does not constitute a limitation.

[0084] 2. Omnidirectional microphone

[0085] An omnidirectional microphone has the same sensitivity to pick up sound signals from all directions; in other words, an omnidirectional microphone has the same sensitivity to sound signals from all angles.

[0086] For example, Figure 1C A schematic diagram of the pickup direction of an omnidirectional microphone is shown.

[0087] like Figure 1C As shown, the pickup direction of an omnidirectional microphone can be... Figure 1C The spherical region shown. From Figure 1C It can also be seen that the omnidirectional microphone has the same sensitivity to picking up sound signals from all directions.

[0088] 3. Near-field speech and far-field speech.

[0089] Near-field speech refers to the sound signal output by the target sound-emitting object when the distance between the target sound-emitting object and the microphone on the electronic device is less than a preset distance.

[0090] Far-field speech refers to the sound signal output by the target sound-emitting object when the distance between the target sound-emitting object and the microphone on the electronic device is greater than a preset distance.

[0091] Optionally, the preset distance can be a distance range, such as greater than or equal to distance A and less than or equal to distance B. For example, distance A can be 5 centimeters, distance B can be 10 centimeters, and the preset distance can refer to any distance between 5 centimeters and 10 centimeters.

[0092] Optionally, the electronic device can determine whether near-field speech has been recognized based on the energy value of the sound signal. For details, please refer to... Figure 7 Description in the embodiments.

[0093] This application provides an audio acquisition method that can be applied to an audio acquisition system. The audio acquisition system may include electronic device 100 and electronic device 200, which are connected by a communication link. For example, electronic device 100 may be a mobile phone or similar device, and electronic device 200 may be a stylus, Bluetooth headset, or similar device.

[0094] The electronic device 200 is equipped with a first audio acquisition device, which is used to acquire a first sound signal. The electronic device 200 can confirm whether near-field speech is recognized based on the first sound signal.

[0095] If near-field speech is recognized based on the first sound signal, the electronic device 200 can send a first message to the electronic device 100 through the communication connection. The first message is used to instruct the electronic device 100 to turn on the second audio acquisition device on the electronic device 100 and start acquiring sound signals, such as acquiring the second sound signal.

[0096] Optionally, the electronic device 200 confirms that it has recognized near-field speech based on the first sound signal, which may be that the electronic device 200 confirms that the first sound signal includes near-field speech.

[0097] While the second audio acquisition device on electronic device 100 is acquiring sound signals, the first audio acquisition device on electronic device 200 is also acquiring sound signals in real time, such as acquiring a third sound signal, and sending the third sound signal acquired by the first audio acquisition device to electronic device 100.

[0098] Electronic device 100 can confirm whether near-field speech has been recognized based on the third sound signal collected by the first audio collector of electronic device 200 and the second sound signal collected by the second audio collector of the device itself.

[0099] Optionally, the electronic device 100 confirms the recognition of near-field speech based on the second and third sound signals, or the electronic device 200 confirms that the second and third sound signals include near-field speech.

[0100] If near-field speech is recognized based on the second and third audio signals, in one possible implementation, the first audio acquisition unit on the electronic device 200 can begin acquiring audio signals, for example, acquiring a fifth audio signal, and send the fifth audio signal to the electronic device 100. The electronic device 100 can perform corresponding operations based on the first, third, and fifth audio signals acquired by the first audio acquisition unit on the electronic device 200, or perform corresponding operations based on the fifth audio signal acquired by the first audio acquisition unit on the electronic device 200. The acquisition time of the fifth audio signal is later than the acquisition time of the third audio signal.

[0101] In other possible implementations, the second audio acquisition unit on the electronic device 100 can begin acquiring sound signals, for example, acquiring a fourth sound signal. The electronic device 100 can then perform corresponding operations based on the second and fourth sound signals acquired by the second audio acquisition unit, or perform corresponding operations based on the fourth sound signal acquired by the second audio acquisition unit. The acquisition time of the fourth sound signal is later than the acquisition time of the second sound signal.

[0102] In other possible implementations, the first audio acquisition device on electronic device 200 and the second audio acquisition device on electronic device 100 can also start acquiring sound signals simultaneously. Electronic device 100 can perform corresponding operations based on the sound signals acquired by the first audio acquisition device on electronic device 200 and the sound signals acquired by the second audio acquisition device on electronic device 100, similar to the two implementations mentioned above.

[0103] Optionally, the electronic device 100 may perform corresponding operations based on the sound signal, including but not limited to: saving the sound signal, converting the sound signal into text and saving the text, and performing corresponding shortcut operations based on the instructions contained in the sound signal.

[0104] The first audio acquisition device may include one or more directional microphones, or one or more omnidirectional microphones. For instructions on how the electronic device 200 recognizes that the sound signal acquired by the first audio acquisition device includes near-field speech, please refer to [reference needed]. Figure 7 Description in the embodiments.

[0105] The second audio acquisition device may include one or more omnidirectional microphones. For details on how the electronic device 100 recognizes near-field speech based on the sound signals acquired by the second audio acquisition device and the sound signals acquired by the first audio acquisition device, please refer to [reference needed]. Figure 7 Description in the embodiments.

[0106] Using this method, electronic device 100 can first confirm whether near-field speech is recognized by the first sound signal collected by the first audio collector on electronic device 200. If near-field speech is recognized, the second audio collector on electronic device 100 is then activated, which reduces the power consumption of electronic device 100. After activating the second audio collector on electronic device 100, electronic device 100 can further confirm whether near-field speech is recognized by combining the third sound signal collected by the first audio collector and the sound signal collected by the second audio collector. This method uses sound signals collected by multiple devices to jointly confirm whether near-field speech is recognized, which can improve the accuracy of near-field speech recognition by electronic device 100, thereby improving the accuracy of automatic audio acquisition by electronic device 100 and / or electronic device 200 and reducing the occurrence of accidental touches.

[0107] Figure 2 A schematic diagram of an audio acquisition system provided in this application.

[0108] like Figure 2 As shown, the audio acquisition system includes electronic device 100 and electronic device 200, which are connected by a communication link. For example, electronic device 100 can be a tablet computer, and electronic device 200 can be a stylus.

[0109] Not limited to tablet computers, electronic devices 100 can also be mobile phones, desktop computers, laptop computers, handheld computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cellular phones, personal digital assistants (PDAs), augmented reality (AR) devices, virtual reality (VR) devices, artificial intelligence (AI) devices, wearable devices (such as smart bracelets), in-vehicle devices, smart home devices (such as smart TVs, smart screens, large-screen devices, etc.) and / or smart city devices, etc.

[0110] Not limited to styluses, electronic devices 200 can also include headphones, smartwatches, smart bracelets, and other devices.

[0111] The electronic device 200 is equipped with a first audio acquisition device, which can acquire a first sound signal. If near-field speech is identified based on the first sound signal, the electronic device 200 can send a first message to the electronic device 100 through a communication connection. The first message is used to instruct the electronic device 100 to turn on the second audio acquisition device on the electronic device 100 and start acquiring sound signals.

[0112] The electronic device 100 is equipped with a second audio acquisition device, which can acquire a second sound signal. While the second audio acquisition device on the electronic device 100 is acquiring the sound signal, the first audio acquisition device on the electronic device 200 is also acquiring the sound signal in real time and sending the third sound signal acquired by the first audio acquisition device to the electronic device 100.

[0113] If near-field speech is also recognized based on the second and third sound signals, in one possible implementation, the first audio acquisition device on the electronic device 200 can start to continuously acquire sound signals and send them to the electronic device 100, and the electronic device 100 can perform corresponding operations based on the sound signals acquired by the first audio acquisition device on the electronic device 200.

[0114] In other possible implementations, the second audio acquisition device on the electronic device 100 may also begin to continuously acquire sound signals, and the electronic device 100 may perform corresponding operations based on the sound signals acquired by the second audio acquisition device on the electronic device 100.

[0115] In other possible implementations, the first audio acquisition device on electronic device 200 and the second audio acquisition device on electronic device 100 may also start acquiring sound signals simultaneously, and electronic device 100 may perform corresponding operations based on the sound signals acquired by the first audio acquisition device on electronic device 200 and the sound signals acquired by the second audio acquisition device on electronic device 100.

[0116] Optionally, the electronic device 100 may perform corresponding operations based on the sound signal, including but not limited to: saving the sound signal, converting the sound signal into text and saving the text, and performing corresponding shortcut operations based on the instructions contained in the sound signal.

[0117] Please refer to Figure 3 , Figure 3 An exemplary schematic diagram of the hardware structure of electronic device 100 is shown.

[0118] For example, electronic device 100 may be a tablet. Electronic device 100 may also be other devices, and this application does not limit the type of electronic device 100.

[0119] like Figure 3 As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0120] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0121] Processor 110 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.

[0122] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.

[0123] The processor 110 may also include a memory for storing instructions and data. In some examples, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or is reusing. If the processor 110 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system. In some embodiments, the processor 110 can be used to confirm whether near-field speech has been recognized based on a second sound signal acquired by the electronic device 100 and a third sound signal acquired by the electronic device 200.

[0124] USB interface 130 is an interface that conforms to the USB standard specification. USB interface 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices.

[0125] The charging management module 140 receives charging input from a charger, which can be a wireless charger or a wired charger. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.

[0126] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, external memory, display 194, camera 193, and wireless communication module 160, etc.

[0127] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0128] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network.

[0129] The mobile communication module 150 can provide solutions for wireless communication applications including 2G / 3G / 4G / 5G on the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low-noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1.

[0130] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless-fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near-field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2. In some embodiments, the electronic device 100 can establish a communication connection with the electronic device 200 through the wireless communication module 160.

[0131] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering.

[0132] The display screen 194 is used to display images, videos, etc. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1. In some embodiments, the electronic device 100 may display the user interface of an application through the display screen 194, which may display the text content of the user's voice output.

[0133] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0134] The ISP is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, converting it into an image visible to the naked eye.

[0135] Camera 193 is used to capture still images or videos. In some embodiments, electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0136] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0137] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0138] The external memory interface 120 can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to perform data storage functions.

[0139] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.).

[0140] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0141] Audio module 170 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signal. Audio module 170 can also be used for encoding and decoding audio signals. In some examples, audio module 170 may be located in processor 110, or some functional modules of audio module 170 may be located in processor 110. Speaker 170A, also called a "loudspeaker," is used to convert audio electrical signals into sound signals. Receiver 170B, also called a "handpiece," is used to convert audio electrical signals into sound signals. Microphone 170C, also called a "microphone" or "microphone," is used to convert sound signals into electrical signals. Headphone jack 170D is used to connect wired headphones. The number of microphones 170C can be one or more. In some embodiments, microphone 170C can also be an omnidirectional microphone. Electronic device 100 can simultaneously combine the sound signals collected by the omnidirectional microphone on electronic device 100 and the sound signals collected by the directional / omnidirectional microphone on electronic device 200 to confirm whether near-field speech has been recognized.

[0142] In some embodiments, microphone 170C may also be a directional microphone.

[0143] In some embodiments, microphone 170C may also be referred to as a second audio acquisition device.

[0144] The sensor module 180 may include pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, gravity sensors, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.

[0145] Buttons 190 include a power button, volume buttons, etc. Motor 191 can generate vibration feedback. Indicator 192 can be an indicator light, used to indicate charging status, battery level changes, and also to indicate messages, missed calls, notifications, etc.

[0146] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The electronic device 100 interacts with the network through the SIM card to achieve functions such as making calls and data communication.

[0147] Please refer to Figure 4 , Figure 4 An exemplary schematic diagram of the hardware structure of electronic device 200 is shown.

[0148] For example, electronic device 200 may be a stylus.

[0149] like Figure 4 As shown, the electronic device 200 may include: a processor 401, a memory 402, a Bluetooth communication module 403, a power supply 404, a power management module 405, a microphone 406, and a speaker 407, etc.

[0150] The processor 401 is used to read and execute computer-readable instructions. Specifically, the processor 401 mainly includes a controller, an arithmetic logic unit (ALU), and registers. The controller is primarily responsible for instruction decoding and issuing control signals for the operations corresponding to the instructions. The ALU is primarily responsible for storing register operands and intermediate operation results temporarily stored during instruction execution.

[0151] In some embodiments, the processor 401 can be used to determine whether the sound signal collected by the microphone 406 contains near-field speech. If near-field speech is included, the processor 401 can control the Bluetooth communication module 403 to send a message to the electronic device 100, causing the microphone on the electronic device 100 to start collecting sound signals.

[0152] Memory 402 is coupled to processor 401 and is used to store various software programs and / or multiple sets of instructions. In specific implementations, memory 402 may include high-speed random access memory and may also include non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 402 may store an operating system, such as uCOS, VxWorks, RTLinux, or other embedded operating systems. Memory 402 may also store communication programs that can be used to communicate with electronic device 100 or other devices.

[0153] The Bluetooth communication module 403 supports Bluetooth Low Energy (BLE) protocol communication. Optionally, the Bluetooth communication module 403 can also support classic Bluetooth (also known as Basic Rate / Enhanced Data Rate, BR / EDR) protocol communication. The Bluetooth communication module 403 can be used to establish a Bluetooth connection between electronic device 100 and electronic device 200, and to send messages to electronic device 100 through this Bluetooth connection.

[0154] The power supply 404 can be a rechargeable lithium battery or a replaceable standard battery. The power management module 405 may include an adaptive pulse width modulation (PMW) charging circuit compatible with Universal Serial Bus (USB-compatible), a Buck DC-DC converter, etc. This power management module 405 can provide power to the processor 401, memory 402, Bluetooth communication module 403, microphone 406, speaker 407, and other devices in the electronic device 200.

[0155] The number of microphones 406 can be one or more. Microphone 406 can be a directional microphone or an omnidirectional microphone. Microphone 406 can be used to acquire sound signals and send the sound signals to processor 401, which can confirm whether near-field speech has been recognized based on the sound signals acquired by microphone 406.

[0156] In some embodiments, microphone 406 may also be referred to as a first audio acquisition device.

[0157] The speaker 407 can be used to send and receive ultrasonic signals, and based on the sent and received ultrasonic signals, to confirm whether the target sound-emitting part of the target sound-emitting object has been identified. The speaker 407 can assist the processor 401 in confirming whether near-field speech has been identified based on the sound signals collected by the microphone 406, thereby improving the accuracy of the processor 401 in recognizing near-field speech.

[0158] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 200. In other embodiments of this application, the electronic device 200 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0159] Figure 5A schematic diagram of a software module for recognizing near-field speech is shown in electronic devices 100 and 200.

[0160] like Figure 5 As shown, the electronic device 100 may include a first audio acquisition unit, a near-field speech recognition unit A, and a communication unit A.

[0161] For example, the first audio acquisition device could be Figure 4 The microphone 406 shown. Near-field speech recognition unit A can be... Figure 4 The processing unit in the processor 401 shown. Communication unit A can be... Figure 4 The communication unit in the Bluetooth communication module 403 shown.

[0162] The electronic device 200 may include a second audio acquisition unit, a near-field speech recognition unit B, and a communication unit B.

[0163] For example, the second audio acquisition device could be Figure 3 The microphone shown is 170C. The near-field speech recognition unit B can be... Figure 3 The processor 110 shown is a processing unit. Communication unit B may be a Bluetooth communication module of electronic device 100. Figure 3 The communication unit (not shown in the image) is shown in the image.

[0164] like Figure 5 As shown, the process of near-field speech recognition by electronic devices 100 and 200 can be as follows:

[0165] 1. The first audio acquisition device acquires the first sound signal.

[0166] Optionally, when the electronic device 100 is turned on, the first audio acquisition device may acquire the first sound signal in real time.

[0167] Optionally, the first audio acquisition device can start acquiring the first sound signal after the electronic device 100 and the electronic device 200 have established a communication connection, which can save the power consumption of the electronic device 100.

[0168] Optionally, after electronic device 100 and electronic device 200 establish a communication connection and the electronic device 100 is confirmed to be in a handheld state based on the sensor data collected by electronic device 100, the first audio collector can start collecting the first sound signal, which can save the power consumption of electronic device 100.

[0169] This application does not limit the timing of the first audio acquisition device acquiring the first sound signal.

[0170] 2. The first audio acquisition unit sends the first sound signal to the near-field speech recognition unit A.

[0171] 3. Near-field speech recognition unit A recognizes near-field speech based on the first sound signal.

[0172] After acquiring the first sound signal, the first audio acquisition unit can send the first sound signal to the near-field speech recognition unit A. The near-field speech recognition unit A can then recognize near-field speech based on the first sound signal. For details on how the near-field speech recognition unit A recognizes near-field speech based on the first sound signal, please refer to [link to relevant documentation]. Figure 7 Description of S703 in the embodiments.

[0173] 4. Near-field speech recognition unit A sends message 1 to communication unit A.

[0174] Upon recognizing near-field speech based on the first sound signal, near-field speech recognition unit A can send message 1 to communication unit A. Message 1 instructs communication unit A to send a message to communication unit B in electronic device 200, causing electronic device 200 to begin acquiring sound signals.

[0175] 5. The first audio acquisition device acquires the third sound signal.

[0176] 6. The first audio acquisition unit sends the third sound signal to the communication unit A.

[0177] If near-field speech is recognized based on the first sound signal, the first audio collector can continue to collect sound signals, such as obtaining a third sound signal, and send the third sound signal to the communication unit A. The third sound signal is used by the electronic device 100 in conjunction with the second sound signal collected by the electronic device 100 to further confirm whether near-field speech has been recognized.

[0178] 7. Communication unit A sends message 2 and the third audio signal to communication unit B.

[0179] Communication unit A and communication unit B establish a communication connection, such as a Bluetooth connection, through which communication unit A can send message 2 and a third audio signal to communication unit B.

[0180] Beyond Bluetooth connectivity, electronic devices 100 and 200 can also establish other connections, such as local area network connections and cloud communication connections.

[0181] Message 2 is used to instruct the second audio acquisition device on electronic device 200 to start acquiring sound signals.

[0182] In some embodiments, message 2 may also be referred to as the first message.

[0183] Optionally, message 2 and the third audio signal can be sent simultaneously from communication unit A to communication unit B, or they can be sent to communication unit B in a time-division manner.

[0184] Optionally, communication unit A may choose not to send message 2 to communication unit B.

[0185] 8. Communication unit B sends the third audio signal to near-field speech recognition unit B.

[0186] After receiving the third audio signal from communication unit A, communication unit B can send the third audio signal to near-field speech recognition unit B. The third audio signal is used by electronic device 100 in conjunction with the second audio signal collected by electronic device 100 to further confirm whether near-field speech has been recognized.

[0187] 9. Communication unit B sends message 3 to the second audio acquisition unit.

[0188] After receiving message 2 from communication unit A, communication unit B can send message 3 to the second audio acquisition unit. Message 3 is used to instruct the second audio acquisition unit to start acquiring sound signals.

[0189] Optionally, message 2 and message 3 can be the same or different.

[0190] 10. The second audio acquisition device acquires the second sound signal.

[0191] 11. The second audio acquisition unit sends the second sound signal to the near-field speech recognition unit B.

[0192] 12. The near-field speech recognition unit B combines the second and third sound signals to confirm the recognition of near-field speech.

[0193] After receiving message 3 from communication unit B, the second audio acquisition unit begins collecting sound information, such as a second sound signal. The second audio acquisition unit then sends the second sound signal to near-field speech recognition unit B.

[0194] After acquiring the second and third audio signals, the near-field speech recognition unit B can determine whether near-field speech has been recognized based on the second and third audio signals. If near-field speech is recognized, the electronic device 100 can perform corresponding operations based on the audio signals acquired by the electronic device 100 and / or the audio signals acquired by the electronic device 200. For details on how the near-field speech recognition unit B recognizes near-field speech based on the second and third audio signals, please refer to [reference needed]. Figure 7 Description of S708 in the embodiment.

[0195] It should be noted that, Figure 5 The steps described herein are for illustrative purposes only, and this application does not impose any restrictions on the order in which the above steps are performed.

[0196] This application provides an audio acquisition method, which is applied to electronic devices 100 and 200 that have established a communication connection. After jointly recognizing near-field speech, electronic devices 100 and / or 200 can automatically start acquiring audio and perform corresponding operations based on the acquired audio, such as saving the audio, converting the audio into text and saving it, or performing corresponding operations based on the instructions contained in the audio.

[0197] The following section will introduce the audio capture method provided in this application, using the UI as an example.

[0198] 1. Users can check the weather by using the voice assistant application installed on their electronic devices.

[0199] For example, electronic device 100 can be a tablet, and electronic device 200 can be a stylus. The tablet and the stylus establish a communication connection, and the user can control the tablet through the stylus. For example, the user can input text, graphics, etc. on the tablet through the stylus, which facilitates interaction between the user and the tablet.

[0200] Optionally, the stylus is pre-installed with a first audio collector, which is turned on and can collect audio in real time and confirm whether the user's speech is recognized.

[0201] In some embodiments, when a user needs to check the weather, such as Figure 6A As shown, the user can pick up the stylus and bring it close to their mouth, allowing them to output speech. The stylus can acquire a first sound signal via a first audio acquisition device and determine whether near-field speech is recognized based on the first sound signal. If near-field speech is recognized based on the first sound signal, the stylus can send a first message to the tablet via the communication connection. This first message instructs the tablet to start the sound signal.

[0202] While the second audio acquisition device on the tablet is acquiring the third sound signal, the first audio acquisition device on the stylus is also periodically / irregularly acquiring sound signals and sending the third sound signal acquired by the first audio acquisition device to the tablet.

[0203] The tablet can confirm whether near-field speech has been recognized based on the second and third audio signals. If near-field speech is recognized based on the second and third audio signals, the tablet can recognize shortcut commands included in the audio signals collected by the tablet and / or stylus, and execute corresponding operations based on the shortcut commands. For example, the tablet can recognize the shortcut command "What's the weather like today?" included in the audio signal and display it. Figure 6B The user interface shown.

[0204] like Figure 6BAs shown, the tablet displays a prompt box 6001, which can be displayed by a voice assistant application on the tablet. The prompt box 6001 includes a dialog box 6002, which displays the text "How's the weather today?". This text can be obtained by the tablet from the sound signal collected by the tablet and / or the stylus, and can also be referred to as a shortcut command.

[0205] In response to this shortcut, the voice assistant application on the electronic device can retrieve today's weather information for Shenzhen from the voice assistant application server and display dialog box 6003 and weather details 6004 within prompt box 6001. Dialog box 6003 displays the weather information "Moderate rain in Shenzhen today, 23℃ to 25℃ with an orange rainstorm warning." Weather details 6004 displays the weather information for Shenzhen every hour. For example, at 11:00 AM, Shenzhen's weather is sunny with a temperature of 23℃. At 12:00 PM, Shenzhen's weather is sunny with a temperature of 23℃. At 1:00 PM, Shenzhen's weather is cloudy with a temperature of 25℃. At 2:00 PM, Shenzhen's weather is cloudy with a temperature of 25℃. At 3:00 PM, Shenzhen's weather is cloudy with a temperature of 25℃. At 4:00 PM, Shenzhen's weather is cloudy with a temperature of 25℃.

[0206] 2. The user saves the audio collected by electronic device 100 and / or electronic device 200 through the memo application in electronic device 100.

[0207] For example, electronic device 100 can be a mobile phone, and electronic device 200 can be a Bluetooth headset. The mobile phone and the Bluetooth headset establish a Bluetooth connection.

[0208] Referring to the interaction methods between the tablet and the stylus described above, such as Figure 6C As shown, the user can pick up the Bluetooth headset and bring it close to their mouth. The user can then speak, and the phone and Bluetooth headset can confirm whether the near-field voice has been recognized using the method described above.

[0209] After recognizing near-field speech, the phone can automatically open the Notes app and capture audio. During audio capture, the Notes app can display... Figure 6D The user interface shown is 6100.

[0210] like Figure 6D As shown, the user interface 6100 can display the audio acquisition process, such as waveforms, audio duration, and corresponding text information. This allows users to check if the audio acquired by the memo application matches their output. If the audio acquired by the memo application differs from the user's output, the user can choose to re-record the audio or modify the parts that differ from their output, thus preventing errors in the audio acquisition process.

[0211] After the user stops speaking, the phone can save the captured audio in the Notes app.

[0212] In one possible implementation, the phone can directly save the captured audio in the Notes app and automatically generate voice notes.

[0213] After generating voice notes, the phone can display them. Figure 6E The user interface 6200 shown includes the name of the voice note, the date the voice note was generated, and the duration of the voice note. For example, the name of the voice note could be "Work Notes," or optionally, the name could be automatically extracted and generated by the mobile phone based on the audio content. The date the voice note was generated could be "December 18, 2023." The duration of the voice note could be 1 minute and 26 seconds.

[0214] In other possible implementations, the phone can also convert the captured audio into text, automatically generate text notes, and save the text notes in the Notes app.

[0215] Using the method described above, the phone can automatically open the Notes app and capture audio when it recognizes near-field voice, then save the captured audio within the Notes app. This eliminates the need for users to manually activate the recording function within the Notes app, saving user time and improving the user experience.

[0216] The application scenarios described above are not limited to those described above. These application scenarios are only used to explain this application and do not constitute a limitation.

[0217] The following section explains how electronic devices 100 and 200 jointly recognize near-field speech.

[0218] Figure 7 A schematic diagram of the method for joint recognition of near-field speech by electronic devices 100 and 200 is shown.

[0219] Figure 7 The method shown includes, but is not limited to, the following steps:

[0220] S701, electronic device 100 and electronic device 200 establish a communication connection.

[0221] For example, electronic device 100 and electronic device 200 can establish a Bluetooth connection. Electronic device 100 can be a mobile phone, tablet, or other device, while electronic device 200 can be a Bluetooth headset, stylus, or other device.

[0222] S702, electronic device 200 acquires a first sound signal through a first audio acquisition device.

[0223] Optionally, after electronic device 200 establishes a communication connection with electronic device 100, electronic device 200 can enter a low-power operating mode, and the first audio acquisition device on electronic device 200 can acquire the first sound signal in real time.

[0224] Optionally, the electronic device 200 can also acquire sensor data. If the sensor data confirms that the movement trajectory of the electronic device 200 meets the preset trajectory, such as the upward lifting trajectory, it means that the user has picked up the electronic device 200 and brought it close to their mouth. The electronic device 100 then acquires the first sound signal through the first audio acquisition device.

[0225] This application does not limit the timing of the electronic device 200 acquiring the first sound signal.

[0226] S703, Electronic device 200 confirms whether near-field speech has been recognized based on the first sound signal.

[0227] After acquiring the first sound signal, the electronic device 200 can confirm whether near-field speech has been recognized based on the first sound signal.

[0228] When near-field speech is recognized based on the first sound signal, the electronic device 200 executes S704.

[0229] If no near-field speech is recognized based on the first sound signal, the electronic device 200 may continue to execute S702-S703.

[0230] Optionally, after recognizing near-field speech based on the first sound signal, before executing S704, the electronic device 200 can further confirm whether the first sound signal includes a speech signal, which may refer to speech content output by a person. If the probability of a speech signal in the first sound signal is greater than a first threshold (e.g., 80%), the electronic device 200 executes S704 to avoid noise interference. For example, if the energy value of the speech content output by a person is much greater than the energy value of noise, the electronic device 200 can determine whether the first sound signal contains a speech signal based on the energy value.

[0231] Electronic device 200 can recognize near-field speech based on a first sound signal through, but not limited to, any of the following methods.

[0232] Method 1: When the first audio acquisition device is one or more directional microphones, if the first sound signal meets a first condition, the electronic device 200 can confirm the recognition of near-field speech. The first condition may include: the energy value of a first frequency band in the first sound signal is greater than a first value, and the difference between the energy value of the first frequency band and the energy values ​​of other frequency points besides the first frequency band is greater than a second value. The first frequency band may be a frequency band between the first frequency point and the second frequency point.

[0233] based on Figure 1A and Figure 1B As explained, directional microphones have different sound pickup capabilities in different directions. If the target sound-emitting object is close to the directional microphone, and the target sound-emitting object emits sound in the preset pickup direction or its extension direction, the directional microphone can pick up the sound signal well in the preset pickup direction or its extension direction. That is, the energy value of the sound signal picked up by the directional microphone in the preset pickup direction or its extension direction is also higher, while the energy value of the sound signal picked up by the directional microphone in non-preset pickup directions or their extension directions is lower.

[0234] When the target sound-emitting object emits sound in a direction other than the preset pickup direction or in a direction extending from the preset pickup direction, the energy value of the sound signal collected by the directional microphone in all directions is low.

[0235] When the distance to the directional microphone is far, regardless of whether the target sound-emitting object is emitting sound in the preset pickup direction of the directional microphone or in the direction extending from the preset pickup direction, the energy value of the sound signal collected by the directional microphone in all directions of the directional microphone is low.

[0236] Based on this characteristic, the electronic device 200 can confirm and recognize near-field speech based on the frequency response curve of the first sound signal.

[0237] Figure 8A The diagram shows the frequency response curve of the sound signal collected by the directional microphone when the sound-emitting object is close to the directional microphone and emits sound in the preset pickup direction of the directional microphone or in the direction extending from the preset pickup direction.

[0238] For example, the first frequency band may refer to the frequency band between frequency point a1 and frequency point a2.

[0239] like Figure 8AAs shown, the energy value of the first frequency band is greater than the first value, and the energy values ​​of other frequency points besides the first frequency band are less than the first value. Furthermore, the difference between the energy value of the first frequency band and the energy values ​​of other frequency points besides the first frequency band is also relatively large, for example, the difference is greater than the second value. Therefore, the electronic device 200 can confirm that near-field speech has been recognized based on the first sound signal.

[0240] Figure 8A It also shows that the energy value of the sound signal collected by the pointing microphone varies when the distance between the target sound-emitting object and the pointing microphone is different. Figure 8A The example illustrates the energy differences in sound collected by the directional microphone when the target sound-emitting object is 5cm, 10cm, and 1m away from the microphone. Figure 8A It can be seen that the closer the target sound-emitting object is to the directional microphone, the greater the low-frequency sound energy value collected by the directional microphone.

[0241] Figure 8B The diagram shows the frequency response curve of the sound signal collected by the directional microphone when the target sound-emitting object is far away from the directional microphone.

[0242] In some embodiments, when the distance between the target sound-emitting object and the microphone is far, regardless of whether the target sound-emitting object emits sound in the preset pickup direction or the direction extending from the preset pickup direction, such as... Figure 8B As shown, in the frequency response curve of the sound signal collected by the microphone, the energy values ​​of the sound signal corresponding to different frequency points are not significantly different. For example, the difference between the energy values ​​of the sound signal corresponding to different frequency points is less than a preset value, and the energy values ​​of the sound signal corresponding to different frequency points are also relatively small, for example, less than a first value. When the frequency response curve of the sound signal satisfies... Figure 8B When the characteristics shown are obtained, it can be confirmed that the distance between the target sound-emitting object and the microphone is relatively far. The electronic device 200 continues to collect sound signals and confirms whether the frequency response curve of the sound signal meets the requirements. Figure 8A The features shown.

[0243] Figure 8C This diagram illustrates the frequency response curve of the sound signal collected by a directional microphone when the target sound-emitting object emits sound in a non-preset pickup direction or in a direction extending beyond the non-preset pickup direction.

[0244] In some embodiments, when the target sound-emitting object emits sound in a direction other than a preset pickup direction or an extension of that direction towards the microphone, the frequency response curve of the sound signal collected by the microphone, such as Figure 8CAs shown, the energy values ​​of sound signals at different frequency points do not differ significantly. For example, the difference between the energy values ​​of sound signals at different frequency points is less than a preset value, and the energy values ​​of sound signals at different frequency points are also relatively small, for example, less than a first value. When the frequency response curve of the sound signal satisfies... Figure 8C When the characteristics shown are met, it can be confirmed that although the target sound-emitting object is close to the microphone, it is not emitting sound in the preset pickup direction or the extension direction of the preset pickup direction. The electronic device 200 continues to collect sound signals and confirms whether the frequency response curve of the sound signal meets the requirements. Figure 8A The features shown.

[0245] Method 2: If the first audio acquisition device includes an omnidirectional microphone, the electronic device 200 also includes a near-field speech-assisted recognition unit. If the first sound signal acquired by the first audio acquisition device satisfies a first condition, and the first signal acquired by the near-field speech-assisted recognition unit satisfies a third condition, then the electronic device 200 can confirm the recognition of near-field speech. The first condition may include: the energy value of the first sound signal is greater than the third value. The third condition may include: the difference between the first transmission time and the first reception time is less than a sixth value, and / or, the vibration frequency of the target sound-emitting part obtained based on the first ultrasonic signal and the second ultrasonic signal is within a first range.

[0246] based on Figure 1C The description states that directional microphones have the same sound pickup capability in different directions. For example... Figure 8D As shown, in the frequency response curve of the first sound signal acquired by the first audio acquisition device, the energy values ​​of the sound signal corresponding to different frequency points do not differ significantly. For example, the difference between the energy values ​​of the sound signal corresponding to different frequency points is less than a preset value, and the energy value of the sound signal corresponding to different frequency points is greater than the energy value of the sound signal corresponding to the target sound-emitting object and the electronic device 200. The closer the target sound-emitting object is to the electronic device 200, the greater the energy value of the first sound signal acquired by the first audio acquisition device on the electronic device 200; the farther the target sound-emitting object is from the electronic device 200, the smaller the energy value of the first sound signal acquired by the first audio acquisition device on the electronic device 200.

[0247] Electronic device 200 can combine an omnidirectional microphone and a near-field speech recognition unit to confirm whether near-field speech has been recognized.

[0248] For example, if the energy value of the first sound signal acquired by the omnidirectional microphone is greater than the third value, and the near-field speech-assisted recognition unit identifies the target sound-producing part, then the electronic device 200 can confirm that it has recognized the near-field speech.

[0249] The fact that the energy value of the first sound signal collected by the omnidirectional microphone is greater than the third value indicates that the distance between the target sound-emitting object and the electronic device 200 is within the preset distance. The near-field voice-assisted recognition unit identifies the target sound-emitting part, indicating that the first sound signal is audio output from the target sound-emitting part of the target sound-emitting object.

[0250] The near-field speech-assisted recognition unit identifies a target sound-emitting part through the following method: the near-field speech-assisted recognition unit emits a first ultrasonic signal and receives a reflected second ultrasonic signal. Because the target sound-emitting part (e.g., the mouth) moves with a certain frequency and amplitude when emitting audio, the first ultrasonic signal is reflected off the moving target sound-emitting part, causing a change in the frequency and amplitude of the first ultrasonic signal. The electronic device 200 can determine the vibration frequency of the target sound-emitting part of the target sound-emitting object based on the emitted first ultrasonic signal and the received second ultrasonic signal. When the vibration frequency of the target sound-emitting part of the target sound-emitting object is within a first range, the electronic device 200 can determine that the target sound-emitting part is active and emitting sound, thus confirming that the near-field speech-assisted recognition unit has identified the target sound-emitting part. The first range can be between a first frequency value and a second frequency value. For example, the first range can be 20Hz-40Hz.

[0251] The near-field speech-assisted recognition unit can also identify the target sound-producing part in other ways, which is not limited in this application.

[0252] Optionally, the near-field voice-assisted recognition unit can also detect whether breathing is detected. If breathing is detected, it means that the near-field voice-assisted recognition unit has recognized a person and the person is close to the first audio acquisition unit. The near-field voice-assisted recognition unit can recognize near-field voice by electronic device 200.

[0253] In some embodiments, the near-field speech-assisted recognition unit may also be referred to as an auxiliary recognition unit.

[0254] In other embodiments, if the first audio acquisition device includes multiple omnidirectional microphones, the electronic device 200 can also identify near-field speech based on the sound signals acquired by the electronic device 200 using the multiple omnidirectional microphones. For details, please refer to the description in S708, Method Two, which will not be repeated here.

[0255] The electronic device 200 may also recognize near-field speech in other ways than those described above, and this application does not limit the scope of such recognition.

[0256] S704, Electronic device 200 sends a first message to electronic device 100 via a communication connection.

[0257] Upon recognizing near-field speech based on the first sound signal, electronic device 200 can send a first message to electronic device 100. The first message instructs electronic device 100 to begin acquisition via the second audio acquisition device.

[0258] Using this method, when electronic device 200 does not recognize near-field speech, the second audio acquisition unit on electronic device 100 does not acquire sound signals. When electronic device 200 recognizes near-field speech, electronic device 200 then controls electronic device 100 to acquire sound signals and further confirms whether electronic device 100 has recognized near-field speech, which can reduce the power consumption of electronic device 100.

[0259] S705, Electronic device 100 acquires a second sound signal through a second audio acquisition device.

[0260] The electronic device 100 is equipped with a second audio acquisition device. After receiving the first message sent by the electronic device 200, the electronic device 100 can control the second audio acquisition device to start the sound signal, for example, to acquire the second sound signal.

[0261] S706, Electronic device 200 acquires a third sound signal through a first audio acquisition device.

[0262] S707, electronic device 200 sends a third sound signal to electronic device 100 via a communication connection.

[0263] Optionally, if near-field speech is recognized based on the first sound signal, the electronic device 200 may continue to collect sound signals, such as a third sound signal, and send the third sound signal to the electronic device 100. The third sound signal is used by the electronic device 100 in conjunction with the second sound signal collected by the electronic device 100 to further confirm whether near-field speech has been recognized.

[0264] It should be noted that S706-S707 can be executed before, after, or simultaneously with any step after S703 and before S708.

[0265] S708, electronic device 100 confirms whether near-field speech has been recognized based on the second and third sound signals.

[0266] Optionally, the second and third sound signals can be sound signals collected by electronic devices 200 and 100 within the same time period.

[0267] Optionally, since the distance between the target sound-emitting object and the first audio collector on electronic device 200 and the second audio collector on electronic device 100 are different, the start time of the acquisition of the second sound signal and the third sound signal may be different.

[0268] Electronic device 100 includes a second audio acquisition unit, and electronic device 200 includes a first audio acquisition unit. The second and first audio acquisition units can form a distributed audio acquisition unit array. The first and second audio acquisition units simultaneously acquire sound signals, and electronic device 100 can obtain a distributed sound signal. The distributed sound signal can include a second sound signal acquired by the first audio acquisition unit and a third sound signal acquired by the second audio acquisition unit within the same time period. Electronic device 100 can determine whether near-field speech has been recognized based on the distributed sound signal. In this way, by using electronic device 200, electronic device 100, and electronic device 200 combining two near-field speech recognition operations, this method can improve the accuracy of near-field speech recognition by electronic device 100, improve the accuracy of automatic audio acquisition by electronic device 100, and reduce the occurrence of accidental touches.

[0269] If near-field speech is recognized based on the second and third sound signals, S709 is executed.

[0270] If no near-field speech is recognized based on the second and third audio signals, step S705 is executed. The second audio acquisition unit on electronic device 100 continues to acquire audio signals and receives the audio signals acquired by the first audio acquisition unit on electronic device 200 sent by electronic device 200. It continues to identify the audio signals acquired by the second audio acquisition unit and the audio signals acquired by the first audio acquisition unit to confirm whether near-field speech has been recognized.

[0271] The following describes how the electronic device 100 recognizes near-field speech based on the second and third sound signals.

[0272] Electronic device 100 can recognize near-field speech based on a second sound signal and a third sound signal through, but not limited to, any one or more of the following methods.

[0273] Method 1: When the first audio acquisition device is a directional microphone and the second audio acquisition device is an omnidirectional microphone, and the second and third audio signals meet the second condition, near-field speech can be confirmed. The second condition may include: the difference between the energy value corresponding to the second frequency band in the second audio signal and the energy value corresponding to the second frequency band in the third audio signal is greater than a fourth value. The second frequency band can be the frequency band between the third and fourth frequency points.

[0274] Optionally, the second frequency band and the first frequency band can be the same or different.

[0275] Optionally, the second condition may also include any one or more of the following: the energy value of the second sound signal is greater than the third value; the energy value corresponding to the second frequency band in the third sound signal is greater than the first value; and the difference between the energy value corresponding to the second frequency band in the frequency response curve of the third sound signal and the energy value of the sound signal near other frequency points other than the second frequency band is greater than the second value.

[0276] Figure 8E (a) in the figure shows the frequency response curve of the third sound signal. Figure 8E (b) shows the frequency response curve of the second sound signal.

[0277] The second frequency band can refer to Figure 8E The frequency range between frequency point a3 and frequency point a4 in the text.

[0278] When the target sound-emitting object is close to the electronic device 200, for example, within a first preset distance, and the sound is emitted in a preset pickup direction towards the microphone or in an extension of the preset pickup direction, such as Figure 8E As shown in (a), the energy value of the third sound signal in the second frequency band is greater than the first value, and the difference between the energy value corresponding to the second frequency band in the frequency response curve of the third sound signal and the energy value of the sound signal near other frequency points other than the second frequency band is greater than the second value.

[0279] When the target sound-emitting object is close to the electronic device 100, for example, within a first preset distance, such as Figure 8E As shown in (b), the energy value of the second sound signal is also greater than that of the third.

[0280] Method 2: When both the first and second audio acquisition devices are omnidirectional microphones, near-field speech can be confirmed if the second and third audio signals meet the second condition. The second condition may include: the difference between the acquisition time of the second and third audio signals is greater than the first duration and / or the difference between the energy values ​​of the second and third audio signals is greater than the fifth value.

[0281] The first audio collector on electronic device 200 and the second audio collector on electronic device 100 are located at different distances from the target sound-emitting object. When the target sound-emitting object is close to both electronic devices 100 and 200, for example, within a first preset distance, the time and energy of the same sound signal emitted by the same target sound-emitting object reaching the first and second audio collectors differ significantly. When the target sound-emitting object is far from both electronic devices 100 and 200, for example, beyond the first preset distance, there is no significant difference in the time and energy of the same sound signal emitted by the same target sound-emitting object reaching the first and second audio collectors.

[0282] The electronic device 100 can determine whether near-field speech has been recognized based on the time difference and / or energy difference of the same audio signal arriving at the first audio collector and the second audio collector.

[0283] When the distance between the target sound-emitting object and electronic devices 100 and 200 is within the first preset distance, the sound signal output by the target sound-emitting object is acquired by the second audio collector on electronic device 100 after propagation through the air, thus obtaining the second sound signal. It can also be acquired by the first audio collector on electronic device 200, thus obtaining the third sound signal.

[0284] For example, the second audio acquisition device on electronic device 100 acquires the second sound signal at time 1, and the first audio acquisition device on electronic device 200 acquires the third sound signal at time 2. The acquisition difference between time 1 and time 2 is greater than the first duration.

[0285] For example, the second audio acquisition unit on electronic device 100 acquires the energy value of the second sound signal as energy value 1, and the first audio acquisition unit on electronic device 200 acquires the energy value of the third sound signal as energy value 2. The difference between energy value 1 and energy value 22 is greater than the fifth value.

[0286] It should be noted that the electronic device 100 may confirm whether near-field speech is recognized not only based on the difference between the acquisition time of the second sound signal and the acquisition time of the third sound signal, and / or the difference between the energy value of the second sound signal and the energy value of the third sound signal, but may also confirm whether near-field speech is recognized based on other methods. This application does not limit this.

[0287] S709, Electronic device 100 performs corresponding operations based on the fifth sound signal collected by the first audio collector and / or the fourth sound signal collected by the second audio collector, wherein the fifth sound signal is sent from electronic device 200 to electronic device 100.

[0288] If near-field speech is confirmed to be recognized based on the second and third audio signals, the electronic device 100 may perform corresponding operations based on the fifth audio signal collected by the first audio collector and / or the fourth audio signal collected by the second audio collector, wherein the fifth audio signal is sent from the electronic device 200 to the electronic device 100.

[0289] The electronic device 100 performs corresponding operations based on sound signals, which may include, but are not limited to, any one or more of the following: saving sound signals, converting sound signals into text and saving the text, recognizing shortcut commands in sound signals and executing the shortcut commands, etc.

[0290] Optionally, when near-field speech is confirmed to be recognized based on the second and third sound signals, electronic device 100 and electronic device 200 may collect sound signals simultaneously, or only electronic device 100 may collect sound signals, or only electronic device 200 may collect sound signals.

[0291] Optionally, the first electronic device performs a first functional operation based on a fourth sound signal collected by the first electronic device and / or a fifth sound signal collected by the second electronic device. This can mean that the first electronic device performs the first functional operation based on the second and fourth sound signals collected by the first electronic device, and / or the first, third, and fifth sound signals collected by the second electronic device. Alternatively, it can mean that the first electronic device performs the first functional operation based solely on the fourth sound signal collected by the first electronic device, and / or the fifth sound signal collected by the second electronic device.

[0292] Preferably, when near-field speech is confirmed to be identified based on the second and third sound signals, the electronic device 100 and the electronic device 200 can simultaneously collect sound signals. For example, the electronic device 100 can acquire the fourth sound signal collected by the electronic device 100 and the fifth sound signal collected by the electronic device 200, and separate the near-field speech and far-field speech based on the fourth and fifth sound signals.

[0293] In some embodiments, the electronic device 100 can separate near-field speech from an audio signal in the following manner. An example will be given of separating near-field speech from an audio signal acquired by a distributed audio acquisition array consisting of a second audio acquisition unit on the electronic device 100 and a first audio acquisition unit on the electronic device 200.

[0294] If the near-field speech signal collected by electronic device 100 can be represented as s1(t), with the transmission path of the near-field speech being f1(t,n), and the far-field speech signal can be represented as s2(t), with the transmission path of the far-field speech being f2(t,n), where t represents time, f1(t,n) represents the transmission path of the near-field speech signal to the nth audio collector, and f2(t,n) represents the transmission path of the far-field speech signal to the nth audio collector in the distributed audio collector array, then the sound signal collected by electronic device 100 can be represented by formula (1).

[0295] y(t,n)= s1(t)*f1(t,n) + s2(t)*f2(t,n) Formula (1)

[0296] As shown in formula (1), y(t,n) represents the sound signal collected by the nth audio collector in the distributed audio collector array, s1(t) represents the near-field speech collected by electronic device 100, f1(t,n) represents the transmission path of the near-field speech, s2(t) represents the far-field speech collected by electronic device 100, and f2(t,n) represents the transmission path of the far-field speech.

[0297] The sound signal received by the distributed audio acquisition array can be represented by formula (2).

[0298] Y(t)=[y(t,1),y(t,2),…,y(t,N)] Formula (2)

[0299] As shown in formula (2), Y(t) represents the sound signal received by the distributed audio acquisition array, y(t,1) represents the sound signal received by the first audio acquisition unit in the distributed audio acquisition array, y(t,2) represents the sound signal received by the second audio acquisition unit in the distributed audio acquisition array, and y(t,N) represents the sound signal received by the Nth audio acquisition unit in the distributed audio acquisition array.

[0300] By performing a frequency domain transformation on Y(t), we can obtain Y(t,f) as shown in formula (3).

[0301] Y(t,f)=F(Y(t)) Formula (3)

[0302] As shown in formula (3), Y(t,f) represents the sound signal received by the distributed audio acquisition array in the frequency domain, where f represents the frequency. Y(t) represents the sound signal received by the distributed audio acquisition array in the time domain.

[0303] The electronic device 100 can separate near-field speech from the sound signal collected by the distributed audio acquisition array using the following formula (4).

[0304]

[0305] As shown in Equation (4), s1(t) represents the near-field speech separated from the sound signal collected by the distributed audio acquisition array, Y(t,f) represents the sound signal received by the distributed audio acquisition array in the frequency domain, and W(t,f) represents the solution matrix, which is used to separate the near-field speech from the sound signal collected by the distributed audio acquisition array.

[0306] Optionally, W(t,f) can be different if each audio collector in the distributed audio collector array is of a different type. For example, W(t,f) is different when each audio collector in the distributed audio collector array is an omnidirectional microphone compared to when some audio collectors in the distributed audio collector array are omnidirectional microphones and the rest are directional microphones.

[0307] In some embodiments, when each audio collector in the distributed audio collector array is an omnidirectional microphone, W(t,f) can represent the difference in energy values ​​between the arrival times of the same sound signal emitted by the same sound-emitting object at each audio collector in the distributed audio collector array when the sound signal is near-field speech, and / or the difference in the arrival times of the same sound signal emitted by the same sound-emitting object at each audio collector in the distributed audio collector array. For example, the difference in energy values ​​between the arrival times of the same sound signal emitted by the same sound-emitting object at each audio collector in the distributed audio collector array can be: the difference between any two energy values ​​of the same sound signal received by each audio collector in the distributed audio collector array is greater than a preset energy difference (e.g., a fifth value). The difference in the arrival times of the same sound signal emitted by the same sound-emitting object at each audio collector in the distributed audio collector array can be: the difference between any two moments of the same sound signal received by each audio collector in the distributed audio collector array is greater than a preset time difference (e.g., a first duration).

[0308] Optionally, the solution matrix W(t,f) can be obtained by the electronic device 100 through deep learning, and the solution matrix W(t,f) can also be updated periodically / irregularly.

[0309] It should be noted that the above formulas (1) to (4) are only used to explain how this application separates near-field speech from sound signals. Near-field speech can also be separated from sound signals in other ways, and this application does not limit this.

[0310] In other embodiments, when some audio collectors in the distributed audio collector array are omnidirectional microphones and the remaining audio collectors are directional microphones, W(t,f) can represent the difference between the energy values ​​of the sound signal collected by the directional microphones and the sound signal collected by the omnidirectional microphones when the same sound signal emitted by the same sound-producing object reaches each audio collector in the distributed audio collector array, and when the sound signal is near-field speech. The difference between the energy values ​​of the sound signal collected by the directional microphones and the sound signal collected by the omnidirectional microphones can be such that the difference between the energy value corresponding to the frequency response curve of the sound signal collected by the directional microphones in the second frequency band and the energy value corresponding to the frequency response curve of the sound signal collected by the omnidirectional microphones in the second frequency band is greater than a fourth value.

[0311] Optionally, if the electronic device 100 fails to recognize near-field speech or acquire sound signals for a continuous first duration, or if the electronic device 100 receives user input, the electronic device 100 may stop acquiring sound signals.

[0312] It is understood that the user interfaces described in the embodiments of this application are merely example interfaces and do not constitute a limitation on the solution of this application. In other embodiments, the user interface may adopt different interface layouts, may include more or fewer controls, and may add or remove other functional options, as long as they are based on the same inventive concept provided in this application, they are all within the protection scope of this application.

[0313] It should be noted that, without causing contradictions or conflicts, any feature in any embodiment of this application, or any part of any feature, can be combined, and the combined technical solution is also within the scope of the embodiments of this application.

[0314] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. An audio acquisition method, characterized in that, The method is applied to a system including a first electronic device and a second electronic device, wherein the first electronic device and the second electronic device establish a communication connection; the method includes: The first electronic device acquires the first sound signal; The first electronic device sends a first message to the second electronic device, the first message being sent after the first electronic device confirms that the first sound signal meets the first condition; In response to the first message, the second electronic device acquires the second sound signal; The second electronic device receives a third sound signal sent by the first electronic device, wherein the acquisition time of the third sound signal is later than the acquisition time of the first sound signal; The first electronic device performs a first function operation based on the fourth sound signal collected by the first electronic device and / or the fifth sound signal collected by the second electronic device. The first function operation is performed after the first electronic device confirms that the second sound signal and the third sound signal meet the second condition. The fourth sound signal is collected later than the second sound signal, and the fifth sound signal is collected later than the third sound signal. The fifth sound signal is sent from the second electronic device to the first electronic device.

2. The method according to claim 1, characterized in that, The second sound signal and the third sound signal are sound signals collected within the same time period.

3. The method according to claim 1 or 2, characterized in that, The first sound signal and the third sound signal are sound signals collected by the microphone in the first electronic device.

4. The method according to claim 1 or 2, characterized in that, When the first sound signal and the third sound signal are sound signals collected by the omnidirectional microphone in the first electronic device, the first electronic device further includes an auxiliary recognition unit. The first message is sent after the first electronic device confirms that the first sound signal meets the first condition and the first signal obtained by the auxiliary recognition unit meets the third condition.

5. The method according to claim 3 or 4, characterized in that, The second sound signal is the sound signal collected by the omnidirectional microphone in the second electronic device.

6. The method according to claim 3, characterized in that, The first condition includes: the energy value of the sound signal in the first frequency band of the first sound signal is greater than a first value, and the difference between the energy value of the first frequency band and the energy values ​​of other frequency points other than the first frequency band is greater than a second value.

7. The method according to claim 4, characterized in that, The first condition includes: the energy value of the first sound signal is greater than the third value; The first signal includes a first transmission time of the first ultrasonic signal and a first reception time of the second ultrasonic signal. The third condition includes: the difference between the first transmission time and the first reception time is less than a sixth value, and / or, the vibration frequency of the target sound-emitting part obtained based on the first ultrasonic signal and the second ultrasonic signal is within a first range.

8. The method according to claim 5, characterized in that, When the third sound signal is a sound signal collected by the omnidirectional microphone in the first electronic device, the second condition includes: The difference between the energy value of the second sound signal and the energy value of the third sound signal is greater than the fifth value; And / or, The difference between the acquisition time of the second sound signal and the acquisition time of the third sound signal is greater than the first duration.

9. The method according to claim 5, characterized in that, When the third sound signal is a sound signal collected by the microphone in the first electronic device, the second condition includes: The difference between the energy value of the second frequency band in the third sound signal and the energy value corresponding to the second frequency band in the second sound signal is greater than the fourth value.

10. The method according to claim 9, characterized in that, The second condition also includes any one or more of the following: The energy value of the second frequency band of the third sound signal is greater than the first value, the difference between the energy value of the second frequency band and the energy values ​​of other frequency points besides the second frequency band is greater than the second value, and the energy value of the second sound signal is greater than the third value.

11. The method according to any one of claims 1-10, characterized in that, The first function operation includes any one or more of the following: saving the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device; converting the sound signal collected by the first electronic device's acquisition unit and / or the sound signal collected by the second electronic device into text information and saving the text information; recognizing shortcut instructions in the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, and based on the shortcut instructions.

12. The method according to any one of claims 1-11, characterized in that, The first electronic device performs a first functional operation based on the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, specifically including: The first electronic device separates a sixth sound signal and a seventh sound signal from the sound signals collected by the first electronic device and / or the sound signals collected by the second electronic device. The sixth sound signal is the sound signal output by a target sound-emitting object whose distance from the first electronic device and the second electronic device is greater than a first preset distance. The seventh sound signal is the sound signal output by a target sound-emitting object whose distance from the first electronic device and the second electronic device is less than the first preset distance. The first electronic device performs the first function operation based on the seventh sound signal.

13. The method according to any one of claims 1-12, characterized in that, The first electronic device is a stylus, and the second electronic device is a tablet.

14. An audio acquisition method, characterized in that, The method includes: In response to a first message sent by a first electronic device, a second electronic device acquires a second sound signal. The first message is sent after the first electronic device confirms that the first sound signal meets a first condition. The first sound signal is the sound signal acquired by the first electronic device. The second electronic device receives a third sound signal sent by the first electronic device, wherein the acquisition time of the third sound signal is later than the acquisition time of the first sound signal; The first electronic device performs a first function operation based on the fourth sound signal collected by the first electronic device and / or the fifth sound signal collected by the second electronic device. The first function operation is performed after the first electronic device confirms that the second sound signal and the third sound signal meet the second condition. The fourth sound signal is collected later than the second sound signal, and the fifth sound signal is collected later than the third sound signal. The fifth sound signal is sent from the second electronic device to the first electronic device.

15. The method according to claim 14, characterized in that, The second sound signal and the third sound signal are sound signals collected within the same time period.

16. The method according to claim 14 or 15, characterized in that, The first sound signal and the third sound signal are sound signals collected by the microphone in the first electronic device.

17. The method according to claim 14 or 15, characterized in that, When the first sound signal and the third sound signal are sound signals collected by the omnidirectional microphone in the first electronic device, the first electronic device further includes an auxiliary recognition unit. The first message is sent after the first electronic device confirms that the first sound signal meets the first condition and the first signal obtained by the auxiliary recognition unit meets the third condition.

18. The method according to claim 15 or 16, characterized in that, The second sound signal is the sound signal collected by the omnidirectional microphone in the second electronic device.

19. The method according to claim 16, characterized in that, The first condition includes: the energy value of the first sound signal is greater than the third value; The first signal includes a first transmission time of the first ultrasonic signal and a first reception time of the second ultrasonic signal. The third condition includes: the difference between the first transmission time and the first reception time is less than a sixth value, and / or, the vibration frequency of the target sound-emitting part obtained based on the first ultrasonic signal and the second ultrasonic signal is within a first range.

20. The method according to claim 17, characterized in that, When the third sound signal is a sound signal collected by the omnidirectional microphone in the first electronic device, the second condition includes: The difference between the energy value of the second sound signal and the energy value of the third sound signal is greater than the fifth value; And / or, The difference between the acquisition time of the second sound signal and the acquisition time of the third sound signal is greater than the first duration.

21. The method according to claim 17, characterized in that, When the third sound signal is a sound signal collected by the microphone in the first electronic device, the second condition includes: The difference between the energy value of the second frequency band in the third sound signal and the energy value corresponding to the second frequency band in the second sound signal is greater than the fourth value.

22. The method according to claim 21, characterized in that, The second condition also includes any one or more of the following: The energy value of the second frequency band of the third sound signal is greater than the first value, the difference between the energy value of the second frequency band and the energy values ​​of other frequency points besides the second frequency band is greater than the second value, and the energy value of the second sound signal is greater than the third value.

23. The method according to any one of claims 14-22, characterized in that, The first function operation includes any one or more of the following: saving the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device; converting the sound signal collected by the first electronic device's acquisition unit and / or the sound signal collected by the second electronic device into text information and saving the text information; recognizing shortcut instructions in the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, and based on the shortcut instructions.

24. The method according to any one of claims 14-23, characterized in that, The first electronic device performs a first functional operation based on the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, specifically including: The first electronic device separates a sixth sound signal and a seventh sound signal from the sound signals collected by the first electronic device and / or the sound signals collected by the second electronic device. The sixth sound signal is the sound signal output by a target sound-emitting object whose distance from the first electronic device and the second electronic device is greater than a first preset distance. The seventh sound signal is the sound signal output by a target sound-emitting object whose distance from the first electronic device and the second electronic device is less than the first preset distance. The first electronic device performs the first function operation based on the seventh sound signal.

25. The method according to any one of claims 14-24, characterized in that, The first electronic device is a stylus, and the second electronic device is a tablet.

26. An electronic device, a second electronic device, characterized in that, The second electronic device includes a memory and a processor, wherein the memory is used to store a computer program; and the processor is used to invoke the computer program to cause the second electronic device to perform the method of any one of claims 1-25.

27. A computationally readable storage medium, comprising instructions, characterized in that, When the instructions are executed on the second electronic device, the second electronic device performs the method of any one of claims 1-25.

28. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed on a second electronic device, cause the second electronic device to perform the method of any one of claims 1-25.

29. A system comprising a first electronic device and a second electronic device, the first electronic device storing a first instruction and the second electronic device storing a second instruction, wherein when the first instruction is executed and when the second instruction is executed, the system performs the method as described in any one of claims 1-25.