Audio acquisition method, electronic device and computer storage medium

By establishing communication connections between electronic devices and utilizing a multi-level judgment method to recognize near-field speech, the problem of inaccurate audio acquisition caused by accidental touches during voice wake-up was solved, achieving higher audio acquisition accuracy.

WO2026067360A1PCT designated stage Publication Date: 2026-04-02HUAWEI TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing electronic devices are prone to accidental touches when activated by voice, leading to inaccurate audio capture.

Method used

By combining two devices to recognize near-field speech, a communication connection is established using the first and second electronic devices to collect and analyze sound signals, and a multi-level judgment method is adopted to recognize near-field speech, thereby improving recognition accuracy.

Benefits of technology

This improves the accuracy of electronic devices in recognizing near-field speech, thereby improving the accuracy of audio acquisition and reducing accidental touches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025123138_02042026_PF_FP_ABST
    Figure CN2025123138_02042026_PF_FP_ABST
Patent Text Reader

Abstract

An audio acquisition method, an electronic device, and a computer storage medium. Whether near-field speech is recognized is determined by means of a first sound signal acquired by a first electronic device; and when the near-field speech is recognized, on the basis of a third sound signal acquired by the first electronic device and a second sound signal acquired by a second electronic device, the second electronic device further determines whether the near-field speech is recognized, thereby reducing the power consumption of the second electronic device. In addition, the method combines sound signals acquired by a plurality of devices to jointly determine whether the near-field speech is recognized, improving the accuracy of the second electronic device in recognizing near-field speech, thereby improving the accuracy of the first electronic device and / or the second electronic device in automatically acquiring audio.
Need to check novelty before this filing date? Find Prior Art

Description

Audio acquisition method, electronic device and computer storage medium

[0001] The present application claims priority to the Chinese patent application No. 202411351789.5, filed on September 25, 2024, with the State Intellectual Property Office of China, and entitled "Audio acquisition method, electronic device and computer storage medium", the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the technical field of terminals, and in particular to an audio acquisition method, an electronic device and a computer storage medium. BACKGROUND

[0003] With the development of terminal technology, the functional requirements of users for electronic devices are increasingly diversified. In order to meet the recording requirements of users for sound, most electronic devices support audio acquisition functions, and the electronic devices can use their own microphones (mics) to acquire audio. For example, a user can use voice to wake up an electronic device to acquire audio, but voice wake-up is prone to accidental touch. How to provide a method for accurately controlling an electronic device to automatically acquire audio remains to be further studied. SUMMARY

[0004] The present application provides an audio acquisition method, an electronic device and a computer storage medium. The method can improve the accuracy of the electronic device in recognizing near-field voice by combining two devices to recognize near-field voice, thereby improving the accuracy of the electronic device in automatically acquiring audio.

[0005] In a first aspect, the present application provides an audio acquisition method. The method is applied to an audio acquisition system including a first electronic device and a second electronic device. The first electronic device and the second electronic device have a communication connection. The method includes: the first electronic device acquires a first sound signal; the first electronic device sends a first message to the second electronic device. The first message is sent after the first electronic device confirms that the first sound signal meets a first condition; in response to the first message, the second electronic device acquires a second sound signal; the second electronic device receives a third sound signal sent by the first electronic device. The acquisition time of the third sound signal is later than the acquisition time of the first sound signal; the first electronic device performs a first function operation based on a fourth sound signal acquired by the first electronic device and / or a fifth sound signal acquired by the second electronic device. The first function operation is performed after the first electronic device confirms that the second sound signal and the third sound signal meet a second condition. The acquisition time of the fourth sound signal is later than the acquisition time of the second sound signal. The acquisition time of the fifth sound signal is later than the acquisition time of the third sound signal. The fifth sound signal is sent by the second electronic device to the first electronic device.

[0006] Optionally, after the first electronic device is powered on, the first electronic device can continuously collect audio.

[0007] The first electronic device can determine whether the near-field voice is recognized based on the collected first sound signal. In the case that the first electronic device recognizes the near-field voice, the first electronic device sends a first message to the second electronic device, the first message being used to instruct the second electronic device to start collecting a second sound signal. The first electronic device also needs to collect a third sound signal and send it to the second electronic device. The second electronic device can further determine whether the near-field voice is recognized based on the second sound signal and the third sound signal. In the case that the second electronic device determines that the near-field voice is recognized based on the second sound signal and the third sound signal, the second electronic device can perform a corresponding function operation based on the sound signal collected by the first electronic device and / or the second electronic device. That is, the method is to recognize the near-field voice through two-level determination, which can improve the accuracy of the electronic device in recognizing the near-field voice, and further improve the accuracy of the electronic device in automatically collecting audio.

[0008] Optionally, the first electronic device performs a first function operation based on the fourth sound signal collected by the first electronic device and / or the fifth sound signal collected by the second electronic device. It can mean that the first electronic device performs the first function operation based on the second sound signal collected by the first electronic device, the fourth sound signal, and / or the first sound signal, the third sound signal, and the fifth sound signal collected by the second electronic device. It can also mean that the first electronic device performs the first function operation based only on the fourth sound signal collected by the first electronic device and / or the fifth sound signal collected by the second electronic device. It can also mean that the first electronic device performs the first function operation based only on the second sound signal collected by the first electronic device, the fourth sound signal, and / or the fifth sound signal collected by the second electronic device. It can also mean that the first electronic device performs the first function operation based only on the fourth sound signal collected by the first electronic device and / or the first sound signal, the third sound signal, and the fifth sound signal collected by the second electronic device.

[0009] Optionally, before sending the first message to the second electronic device, the first electronic device can also identify the probability of the voice signal contained in the first sound signal. If the probability of the voice signal contained in the first sound signal is greater than a preset value, the first electronic device sends the first message to the second electronic device, which can avoid the interference of noise. The voice signal can be understood as the audio output by a person.

[0010] For example, the energy value of the audio output by a person is much greater than the energy value of noise, and the first electronic device can determine whether the voice signal is contained in the first sound signal based on the energy value of the first sound signal.

[0011] In a possible implementation manner of the first aspect, the second sound signal and the third sound signal are sound signals collected in a same time period.

[0012] In this way, the second electronic device identifies the near-field voice based on the second sound signal and the third sound signal collected in the same time period, and the accuracy of the second electronic device in identifying the near-field voice can be improved.

[0013] In a possible implementation manner of the first aspect, the first sound signal and the third sound signal are sound signals collected by a directional microphone in the first electronic device.

[0014] In a possible implementation manner of the first aspect, when the first sound signal and the third sound signal are sound signals collected by an omnidirectional microphone in the first electronic device, the first electronic device further includes an auxiliary identification unit, and the first message is sent after the first electronic device confirms that the first sound signal meets the first condition and the auxiliary identification unit obtains that the first signal meets the third condition.

[0015] In this way, when the first electronic device includes the omnidirectional microphone, the first electronic device can assist the first electronic device in identifying the near-field voice based on the sound signal collected by the omnidirectional microphone by using the first signal obtained by the auxiliary identification unit, and the accuracy of the first electronic device in identifying the near-field voice can be improved.

[0016] Optionally, the first signal can include a first ultrasonic signal output by the auxiliary identification unit and a second ultrasonic signal reflected by the target sound-emitting part and received by the auxiliary identification unit.

[0017] In a possible implementation manner of the first aspect, the second sound signal is a sound signal collected by an omnidirectional microphone in the second electronic device.

[0018] In this way, the directional microphone on the first electronic device and the omnidirectional microphone on the second electronic device can form a microphone array, the second electronic device identifies the near-field voice based on the third sound signal collected by the directional microphone on the first electronic device and the second sound signal collected by the omnidirectional microphone on the second electronic device, and the accuracy of the second electronic device in identifying the near-field voice can be improved.

[0019] In a possible implementation manner of the first aspect, the first condition includes that an energy value of a sound signal in a first frequency segment in the first sound signal is greater than a first value, and a difference between the energy value of the first frequency segment and an energy value of a frequency point other than the first frequency segment is greater than a second value.

[0020] In this way, the first electronic device can identify the near-field voice based on the first sound signal collected by the directional microphone by using the above judgment condition.

[0021] With reference to the first aspect, in a possible implementation manner, the first condition comprises that an energy value of the first sound signal is greater than a third value; the first signal comprises a first emission moment of the first ultrasonic signal and a first reception moment of the second ultrasonic signal, and the third condition comprises that a difference between the first emission moment and the first reception moment is less than a sixth value, and / or a vibration frequency of the target sound emitting part obtained based on the first ultrasonic signal and the second ultrasonic signal is within a first range.

[0022] In this way, by using the above judgment condition, the first electronic device can assist the omnidirectional microphone in recognizing the near-field voice based on the collected first sound signal by using the auxiliary recognition unit.

[0023] With reference to the first aspect, in a possible implementation manner, when the third sound signal is a sound signal collected by the omnidirectional microphone in the first electronic device, the second condition comprises that a difference between an energy value of the second sound signal and an energy value of the third sound signal is greater than a fifth value; and / or a difference between a collection moment of the second sound signal and a collection moment of the third sound signal is greater than a first time length.

[0024] In this way, by using the above judgment condition, the second electronic device can recognize the near-field voice based on the third sound signal collected by the omnidirectional microphone in the first electronic device and the second sound signal collected by the omnidirectional microphone in the second electronic device.

[0025] With reference to the first aspect, in a possible implementation manner, when the third sound signal is a sound signal collected by the directional microphone in the first electronic device, the second condition comprises that a difference between an energy value of the sound signal of the second frequency segment in the third sound signal and a corresponding energy value of the second frequency segment in the second sound signal is greater than a fourth value.

[0026] In a possible implementation manner, the second condition further comprises any one or more of the following: the energy value of the sound signal of the second frequency segment in the third sound signal is greater than a first value, a difference between the energy value of the second frequency segment and an energy value of a frequency point other than the second frequency segment is greater than a second value, and the energy value of the second sound signal is greater than a third value.

[0027] In this way, by using the above judgment condition, the second electronic device can recognize the near-field voice based on the third sound signal collected by the directional microphone in the first electronic device and the second sound signal collected by the omnidirectional microphone in the second electronic device.

[0028] With reference to the first aspect, in a possible implementation manner, the first function operation includes any one or more of the following: saving the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, converting the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device into text information and saving the text information, identifying a shortcut instruction in the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, and performing the first function operation based on the shortcut instruction.

[0029] After the second electronic device also identifies the near-field voice, the second electronic device can perform the first function operation based on the sound signal collected by the first electronic device, the second electronic device can also perform the first function operation based on the sound signal collected by the second electronic device, and the second electronic device can also perform the first function operation based on the sound signal collected by the first electronic device and the sound signal collected by the second electronic device.

[0030] For example, when the second electronic device performs the first function operation based on the sound signal collected by the first electronic device, the sound signal collected by the first electronic device can include the first sound signal collected by the first electronic device, the third sound signal, and the sound signal collected after the third sound signal, or can include the sound signal collected after the third sound signal.

[0031] For example, when the second electronic device performs the first function operation based on the sound signal collected by the second electronic device, the sound signal collected by the second electronic device can include the second sound signal and the sound signal collected after the second sound signal, or can include the sound signal collected after the second sound signal.

[0032] With reference to the first aspect, in a possible implementation manner, the first electronic device performs the first function operation based on the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, specifically including: the first electronic device separates the sixth sound signal and the seventh sound signal from the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, the sixth sound signal is a sound signal output by a target sound object with a distance greater than the first preset distance between the first electronic device and the second electronic device, and the seventh sound signal is a sound signal output by a target sound object with a distance less than the first preset distance between the first electronic device and the second electronic device; and the first electronic device performs the first function operation based on the seventh sound signal.

[0033] In this way, the second electronic device performs the first function operation based on the near-field voice, which can improve the accuracy of the second electronic device in identifying the operation intention included in the near-field voice, and the operation intention included in the near-field voice is used to trigger the second electronic device to perform the first function operation.

[0034] With reference to the first aspect, in a possible implementation manner, the first electronic device is a stylus, and the second electronic device is a tablet.

[0035] Optionally, the first electronic device can also be a headset, and the second electronic device can be a tablet.

[0036] With reference to the second aspect, the present application provides an audio acquisition method, which comprises the following steps: in response to a first message sent by a first electronic device, a second electronic device acquires a second sound signal, the first message being sent after the first electronic device confirms that a first sound signal meets a first condition, the first sound signal being a sound signal acquired by the first electronic device; the second electronic device receives a third sound signal sent by the first electronic device, the third sound signal being acquired later than the first sound signal; the first electronic device performs a first function operation based on a fourth sound signal acquired by the first electronic device and / or a fifth sound signal acquired by the second electronic device, the first function operation being performed after the first electronic device confirms that the second sound signal and the third sound signal meet a second condition, the fourth sound signal being acquired later than the second sound signal, the fifth sound signal being acquired later than the third sound signal, and the fifth sound signal being sent by the second electronic device to the first electronic device.

[0037] With reference to the second aspect, in a possible implementation manner, the second sound signal and the third sound signal are sound signals acquired in the same time period.

[0038] With reference to the second aspect, in a possible implementation manner, the first sound signal and the third sound signal are sound signals collected by a directional microphone in the first electronic device.

[0039] With reference to the second aspect, in a possible implementation manner, when the first sound signal and the third sound signal are sound signals collected by an omnidirectional microphone in the first electronic device, the first electronic device further comprises an auxiliary identification unit, and the first message is sent after the first electronic device confirms that the first sound signal meets the first condition and the auxiliary identification unit obtains a first signal meeting a third condition.

[0040] With reference to the second aspect, in a possible implementation manner, the second sound signal is a sound signal collected by an omnidirectional microphone in the second electronic device.

[0041] With reference to the second aspect, in a possible implementation manner, the first condition comprises that an energy value of the first sound signal is greater than a third value; the first signal comprises a first emission time of a first ultrasonic signal and a first reception time of a second ultrasonic signal, the third condition comprises that a difference between the first emission time and the first reception time is less than a sixth value, and / or a vibration frequency of a target sound emitting part obtained based on the first ultrasonic signal and the second ultrasonic signal is within a first range.

[0042] With reference to the second aspect, in a possible implementation manner, when the third sound signal is a sound signal collected by an omnidirectional microphone in the first electronic device, the second condition comprises: a difference between an energy value of the second sound signal and an energy value of the third sound signal is greater than a fifth value; and / or, a difference between a collection time of the second sound signal and a collection time of the third sound signal is greater than a first time length.

[0043] With reference to the second aspect, in a possible implementation manner, when the third sound signal is a sound signal collected by a directional microphone in the first electronic device, the second condition comprises: a difference between an energy value of a sound signal of the second frequency range in the third sound signal and a corresponding energy value of the second frequency range in the second sound signal is greater than a fourth value.

[0044] With reference to the second aspect, in a possible implementation manner, the second condition further comprises any one or more of the following: the energy value of the sound signal of the second frequency range in the third sound signal is greater than a first value, a difference between the energy value of the second frequency range and an energy value of a frequency point other than the second frequency range is greater than a second value, and the energy value of the second sound signal is greater than a third value.

[0045] With reference to the second aspect, in a possible implementation manner, the first function operation comprises any one or more of the following: saving the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, converting the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device into text information and saving the text information, identifying a shortcut instruction in the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, and performing an operation based on the shortcut instruction.

[0046] With reference to the second aspect, in a possible implementation manner, the first electronic device performs the first function operation based on the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, specifically comprising: the first electronic device separates a sixth sound signal and a seventh sound signal from the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, the sixth sound signal is a sound signal output by a target sound-emitting object within a first preset distance from the first electronic device and the second electronic device, and the seventh sound signal is a sound signal output by a target sound-emitting object within a first preset distance from the first electronic device and the second electronic device; and the first electronic device performs the first function operation based on the seventh sound signal.

[0047] With reference to the second aspect, in a possible implementation manner, the first electronic device is a stylus, and the second electronic device is a tablet.

[0048] In a third aspect, the present application provides an electronic device, which is a second electronic device. The second electronic device comprises a memory and a processor, wherein the memory is configured to store a computer program; and the processor is configured to invoke the computer program, so that the second electronic device performs the audio acquisition method in any possible implementation manner of any of the aspects.

[0049] In a fourth aspect, the present application provides an electronic device, which is a first electronic device. The first electronic device comprises a memory and a processor, wherein the memory is configured to store a computer program; and the processor is configured to invoke the computer program, so that the first electronic device performs the audio acquisition method in any possible implementation manner of the first aspect.

[0050] In a fifth aspect, the present application provides a computer readable storage medium, which comprises instructions. When the instructions are run on a second electronic device, the second electronic device performs the audio acquisition method in any possible implementation manner of any of the aspects.

[0051] In a sixth aspect, the present application provides a computer readable storage medium, which comprises instructions. When the instructions are run on a first electronic device, the first electronic device performs the audio acquisition method in any possible implementation manner of the first aspect.

[0052] In a seventh aspect, the present application provides a computer program product, which comprises computer instructions. When the computer instructions are run on a second electronic device, the second electronic device performs the audio acquisition method in any possible implementation manner of any of the aspects.

[0053] In an eighth aspect, the present application provides a computer program product, which comprises computer instructions. When the computer instructions are run on a first electronic device, the first electronic device performs the audio acquisition method in any possible implementation manner of the first aspect.

[0054] In a ninth aspect, the present application provides a chip, which is applied to a second electronic device. The chip comprises one or more processors, which are configured to invoke computer instructions, so that the second electronic device performs the audio acquisition method in any possible implementation manner of any of the aspects.

[0055] In a tenth aspect, the present application provides a chip, which is applied to a first electronic device. The chip comprises one or more processors, which are configured to invoke computer instructions, so that the first electronic device performs the audio acquisition method in any possible implementation manner of the first aspect.

[0056] The beneficial effects of the second aspect to the tenth aspect can refer to the description of the beneficial effects of the first aspect, which will not be repeated here. Attached Figure Description

[0057] Figure 1A shows a schematic diagram of a preset pickup direction of a directional microphone;

[0058] Figure 1B shows a schematic diagram of the preset pickup direction of another directional microphone;

[0059] Figure 1C shows a schematic diagram of the pickup direction of an omnidirectional microphone;

[0060] Figure 2 is a schematic diagram of an audio acquisition system provided in this application;

[0061] Figure 3 illustrates a schematic diagram of the hardware structure of the electronic device 100;

[0062] Figure 4 illustrates a schematic diagram of the hardware structure of the electronic device 200;

[0063] Figure 5 shows a schematic diagram of a software module for recognizing near-field speech in electronic devices 100 and 200;

[0064] Figures 6A-6E illustrate a schematic diagram of a group of electronic devices 200 automatically acquiring audio;

[0065] Figure 7 shows a schematic diagram of the method for joint recognition of near-field speech by electronic devices 100 and 200.

[0066] Figure 8A shows a schematic diagram of the frequency response curve of the sound signal collected by the directional microphone when the target sound-emitting object is close to the directional microphone and emits sound in the preset sound pickup direction or the direction extending from the preset sound pickup direction of the directional microphone.

[0067] Figure 8B shows a schematic diagram of the frequency response curve of the sound signal collected by the directional microphone when the target sound-emitting object is far away from the directional microphone;

[0068] Figure 8C shows a schematic diagram of the frequency response curve of the sound signal collected by the directional microphone when the target sound-emitting object emits sound in a non-preset pickup direction or in a direction extending from the non-preset pickup direction to the microphone.

[0069] Figure 8D shows the frequency response curve of the first sound signal acquired by the first audio acquisition device;

[0070] Figure 8E shows a schematic diagram of the frequency response curves of the second and third sound signals. Detailed Implementation

[0071] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing the specific embodiments and are not intended to be limiting on the present application. As used in the specification and the appended claims of the present application, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “at least one of’ or “one or more of’ as used in the following embodiments means one, two, three, or more than two (including two). The term “and / or” is used to describe the relationship between associated objects, indicating that there can be three relationships; for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character “ / ” generally represents an “or” relationship between the associated objects.

[0072] In this specification, the reference “one embodiment” or “some embodiments” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. Thus, the appearances of the phrases “in one embodiment”, “in some embodiments”, “in other embodiments”, “in additional embodiments”, and so on, in various places in the specification are not necessarily all referring to the same embodiment, unless otherwise specified. The terms “comprise”, “comprising”, “have”, “having”, “include”, “including”, and “contain”, “containing”, or variants thereof, mean “including but not limited to”, unless otherwise specified. The term “connected” includes both direct and indirect connections unless otherwise specified. “First”, “second”, and the like are used only for descriptive purposes and should not be construed as indicating or implying relative importance or an indicated number of technical features.

[0073] In the embodiments of the present application, the words “exemplary” or “for example” are used to mean serving as an example, instance, or illustration. Any embodiment or design presented as “exemplary” or “for example” in the embodiments of the present application is not necessarily to be construed as preferred or advantageous over other embodiments or designs. Rather, use of the words “exemplary” or “for example” is intended to present concepts in a concrete manner.

[0074] First, the technical terms related to the present application are explained.

[0075] 1. Directional microphone

[0076] The pickup sensitivity of the directional microphone to sound signals in different directions is different. The pickup sensitivity of the directional microphone to sound signals in a preset pickup direction is higher. The pickup sensitivity of the directional microphone to sound signals in a non-preset pickup direction is weaker.

[0077] The structure of the directional microphone is different, and the preset pickup direction of the directional microphone is also different.

[0078] For example, the structure of the directional microphone can be an "8-shaped" microphone, and the preset pickup direction of the "8-shaped" microphone can present an "8-shaped" area.

[0079] For example, the structure of the directional microphone can be an "8-shaped" microphone, and the preset pickup direction of the "8-shaped" microphone can present an "8-shaped" area.

[0080] As shown in FIG. 1A, the preset pickup direction of the directional microphone can present an "8-shaped" area, for example, the preset pickup direction can be the solid line area shown in FIG. 1A. If the target sound object sounds in the solid line area shown in FIG. 1A or in the extension direction of the solid line area shown in FIG. 1A, the directional microphone can better collect the sound signal output by the target sound object.

[0081] For example, the structure of the directional microphone can be an "8-shaped" microphone, and the preset pickup direction of the "8-shaped" microphone can present an "8-shaped" area.

[0082] For example, the structure of the directional microphone can be an "8-shaped" microphone, and the preset pickup direction of the "8-shaped" microphone can present an "8-shaped" area.

[0083] As shown in FIG. 1B, the preset pickup direction of the directional microphone can present a "heart-shaped" area, for example, the preset pickup direction can be the solid line area shown in FIG. 1B. If the target sound object sounds in the solid line area shown in FIG. 1B or in the extension direction of the solid line area shown in FIG. 1B, the directional microphone can better collect the sound signal output by the target sound object.

[0084] Optionally, the directional microphone can include other structures in addition to the "8-shaped" microphone and the "heart-shaped" microphone, and the preset pickup direction of the directional microphone is explained by the "8-shaped" microphone and the "heart-shaped" microphone in this application, which does not constitute a limitation.

[0085] 2. Omnidirectional microphone

[0086] The pickup sensitivity of the omnidirectional microphone to sound signals in different directions is the same, that is, the omnidirectional microphone has the same sensitivity to sound signals of all angles.

[0087] For example, FIG. 1C shows a schematic diagram of the pickup direction of an omnidirectional microphone.

[0088] As shown in FIG. 1C, the pickup direction of the omnidirectional microphone can be a spherical region shown in FIG. 1C. As can also be seen from FIG. 1C, the pickup sensitivity of the omnidirectional microphone to the sound signals in each direction is the same.

[0089] 3, near-field speech and far-field speech.

[0090] The near-field speech can refer to a sound signal output by a target sound object when the distance between the target sound object and the microphone on the electronic device is less than a preset distance.

[0091] The far-field speech can refer to a sound signal output by a target sound object when the distance between the target sound object and the microphone on the electronic device is greater than a preset distance.

[0092] Optionally, the preset distance can be a distance range, for example, the distance range can be greater than or equal to distance A and less than or equal to distance B. For example, distance A can be 5 cm, distance B can be 10 cm, and the preset distance can refer to any distance between 5 cm and 10 cm.

[0093] Optionally, the electronic device can confirm whether the near-field speech is recognized based on the energy value of the sound signal. For details, reference can be made to the description in the embodiment of FIG. 7.

[0094] The present application provides an audio acquisition method, which can be applied to an audio acquisition system. The audio acquisition system can include an electronic device 100 and an electronic device 200, and the electronic device 100 and the electronic device 200 have a communication connection. For example, the electronic device 100 can be a mobile phone or the like, and the electronic device 200 can be a stylus, a Bluetooth headset or the like.

[0095] The electronic device 200 is pre-provided with a first audio acquisition device, which is used to acquire a first sound signal. The electronic device 200 can confirm whether the near-field speech is recognized based on the first sound signal.

[0096] If the near-field speech is recognized based on the first sound signal, the electronic device 200 can send a first message to the electronic device 100 through the communication connection. The first message is used to instruct the electronic device 100 to start the second audio acquisition device on the electronic device 100, and start to acquire a sound signal, for example, to acquire a second sound signal.

[0097] Optionally, the electronic device 200 confirms that the near-field speech is recognized based on the first sound signal, which can mean that the electronic device 200 confirms that the first sound signal includes the near-field speech.

[0098] While the second audio collector on the electronic device 100 is collecting the sound signal, the first audio collector on the electronic device 200 is also collecting the sound signal in real time, for example, collecting the third sound signal, and sending the third sound signal collected by the first audio collector to the electronic device 100.

[0099] The electronic device 100 can determine whether the near-field voice is recognized based on the third sound signal collected by the first audio collector of the electronic device 200 and the second sound signal collected by the second audio collector of the electronic device 100.

[0100] Optionally, the electronic device 100 determines that the near-field voice is recognized based on the second sound signal and the third sound signal, which can mean that the electronic device 200 determines that the near-field voice is included in the second sound signal and the third sound signal.

[0101] If the near-field voice is recognized based on the second sound signal and the third sound signal, in one possible implementation, the first audio collector on the electronic device 200 can start collecting the sound signal, for example, collecting the fifth sound signal, and send the fifth sound signal to the electronic device 100. The electronic device 100 can perform corresponding operations based on the first sound signal, the third sound signal, and the fifth sound signal collected by the first audio collector on the electronic device 200, or perform corresponding operations based on the fifth sound signal collected by the first audio collector on the electronic device 200. The collection time of the fifth sound signal is later than the collection time of the third sound signal.

[0102] In other possible implementations, the second audio collector on the electronic device 100 can start collecting the sound signal, for example, collecting the fourth sound signal. The electronic device 100 can perform corresponding operations based on the second sound signal, the fourth sound signal collected by the second audio collector on the electronic device 100, or perform corresponding operations based on the fourth sound signal collected by the second audio collector on the electronic device 100. The collection time of the fourth sound signal is later than the collection time of the second sound signal.

[0103] In other possible implementations, the first audio collector on the electronic device 200 and the second audio collector on the electronic device 100 can also start collecting the sound signal at the same time. The electronic device 100 can perform corresponding operations based on the sound signal collected by the first audio collector on the electronic device 200 and the sound signal collected by the second audio collector on the electronic device 100, which is similar to the above two implementations.

[0104] Optionally, the electronic device 100 performing corresponding operations based on the sound signal can include but is not limited to saving the sound signal, converting the sound signal into text and saving the text, performing corresponding shortcut operations based on the instructions included in the sound signal, and the like.

[0105] The first audio collector can include one or more directional microphones, or one or more omnidirectional microphones. For how the electronic device 200 identifies that the near-field speech is included in the sound signal collected by the first audio collector, reference can be made to the description in the embodiment of FIG. 7.

[0106] The second audio collector can include one or more omnidirectional microphones. For how the electronic device 100 identifies the near-field speech based on the sound signal collected by the second audio collector and the sound signal collected by the first audio collector, reference can be made to the description in the embodiment of FIG. 7.

[0107] Through the method, the electronic device 100 can first confirm whether the near-field speech is identified through the first sound signal collected by the first audio collector on the electronic device 200, and in the case where the near-field speech is identified, the second audio collector on the electronic device 100 is then started, which can reduce the power consumption of the electronic device 100. After the second audio collector on the electronic device 100 is started, the electronic device 100 can further confirm whether the near-field speech is identified by combining the third sound signal collected by the first audio collector and the sound signal collected by the second audio collector, which can improve the accuracy of the electronic device 100 in identifying the near-field speech, and thus improve the accuracy of the electronic device 100 and / or the electronic device 200 in automatically collecting audio, and reduce the occurrence of false touch.

[0108] FIG. 2 is a schematic diagram of an audio collection system provided in the present application.

[0109] As shown in FIG. 2, the audio collection system includes an electronic device 100 and an electronic device 200, and the electronic device 100 and the electronic device 200 are communicatively connected. For example, the electronic device 100 can be a tablet computer, and the electronic device 200 can be a stylus.

[0110] The electronic device 100 is not limited to a tablet computer, and can also be a mobile phone, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a wearable device (for example, a smart bracelet), a vehicle-mounted device, a smart home device (for example, a smart television, a smart screen, a large-screen device, etc.), a smart city device, and / or the like.

[0111] The electronic device 200 can also be a headset, a smart watch, a smart bracelet, etc. without being limited to a stylus.

[0112] The electronic device 200 is pre-installed with a first audio collector, which can collect a first sound signal. If a near-field voice is recognized based on the first sound signal, the electronic device 200 can send a first message to the electronic device 100 through a communication connection, the first message being used to instruct the electronic device 100 to start a second audio collector on the electronic device 100 and start collecting a sound signal.

[0113] The electronic device 100 is pre-installed with a second audio collector, which can collect a second sound signal. While the second audio collector on the electronic device 100 is collecting a sound signal, the first audio collector on the electronic device 200 is also collecting a sound signal in real time, and the third sound signal collected by the first audio collector is sent to the electronic device 100.

[0114] If a near-field voice is also recognized based on the second sound signal and the third sound signal, in a possible implementation, the first audio collector on the electronic device 200 can start collecting a sound signal continuously and send it to the electronic device 100, and the electronic device 100 can perform a corresponding operation based on the sound signal collected by the first audio collector on the electronic device 200.

[0115] In other possible implementations, the second audio collector on the electronic device 100 can also start collecting a sound signal continuously, and the electronic device 100 can perform a corresponding operation based on the sound signal collected by the second audio collector on the electronic device 100.

[0116] In other possible implementations, the first audio collector on the electronic device 200 and the second audio collector on the electronic device 100 can also start collecting a sound signal at the same time, and the electronic device 100 can perform a corresponding operation based on the sound signal collected by the first audio collector on the electronic device 200 and the sound signal collected by the second audio collector on the electronic device 100.

[0117] Optionally, the electronic device 100 performing a corresponding operation based on the sound signal can include but is not limited to saving the sound signal, converting the sound signal into text and saving the text, performing a corresponding shortcut operation based on an instruction contained in the sound signal, etc.

[0118] Please refer to FIG. 3, which exemplarily shows a hardware structure diagram of the electronic device 100.

[0119] Exemplarily, the electronic device 100 can be a tablet. The electronic device 100 can also be other devices, and the type of the electronic device 100 is not limited in the present application.

[0120] As shown in FIG. 3, the electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0121] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than shown, or combine certain components, or split certain components, or different arrangement of components. The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0122] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices, or can be integrated in one or more processors.

[0123] Among them, the controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.

[0124] The processor 110 can also be provided with a memory for storing instructions and data. In some examples, the memory in the processor 110 is a cache memory. The memory can hold instructions or data that the processor 110 has just used or is using in a loop. If the processor 110 needs to use the instructions or data again, it can be called directly from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thereby improving the efficiency of the system. In some embodiments, the processor 110 can be used to confirm whether the near-field voice is recognized based on the second sound signal collected by the electronic device 100 and the third sound signal collected by the electronic device 200.

[0125] The USB interface 130 is an interface that conforms to the USB standard specification. The USB interface 130 can be used to connect a charger to charge the electronic device 100, or to transmit data between the electronic device 100 and a peripheral device.

[0126] The charging management module 140 is used to receive charging input from a charger. The charger can be a wireless charger or a wired charger. The charging management module 140 can also supply power to the electronic device while charging the battery 142 through the power management module 141.

[0127] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to supply power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160, etc.

[0128] The wireless communication function of the electronic device 100 can be realized through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.

[0129] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: the antenna 1 can be multiplexed as a diversity antenna for a wireless local area network.

[0130] The mobile communication module 150 can provide a solution for wireless communication including 2G / 3G / 4G / 5G, etc. applied to the electronic device 100. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive an electromagnetic wave by the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic wave, and transfer the processed electromagnetic wave to the modem processor to be demodulated. The mobile communication module 150 can also amplify a signal modulated by the modem processor, and radiate the amplified signal as an electromagnetic wave through the antenna 1.

[0131] The wireless communication module 160 can provide a solution for wireless communication including wireless local area networks (WLAN) (e.g., wireless-fidelity (Wi-Fi) network), blue tooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrated with at least one communication processing module. The wireless communication module 160 receives an electromagnetic wave via the antenna 2, performs frequency modulation and filtering on the electromagnetic wave signal, and transmits the processed signal to the processor 110. The wireless communication module 160 can also receive a signal to be transmitted from the processor 110, perform frequency modulation and amplification on the received signal, and radiate the processed signal as an electromagnetic wave through the antenna 2. In some embodiments, the electronic device 100 can establish a communication connection with the electronic device 200 through the wireless communication module 160.

[0132] The electronic device 100 can implement a display function through a GPU, a display screen 194, an application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering.

[0133] The display screen 194 is used to display an image, a video, etc. In some embodiments, the electronic device 100 can include one or N display screens 194, N being a positive integer greater than 1. In some embodiments, the electronic device 100 can display a user interface of an application through the display screen 194, and the user interface can display text content of a user's output voice.

[0134] The electronic device 100 can implement a photographing function through an ISP, a camera 193, a video codec, a GPU, a display 194, and an application processor, etc.

[0135] The ISP is used to process data fed back by the camera 193. For example, when taking a photo, the shutter is opened, light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing to convert it into an image visible to the naked eye.

[0136] The camera 193 is used to capture a still image or a video. In some embodiments, the electronic device 100 can include one or N cameras 193, where N is a positive integer greater than 1.

[0137] The digital signal processor is used to process digital signals, in addition to being able to process digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0138] The NPU is a neural-network (NN) calculation processor, which processes input information quickly by drawing on the structure of a biological neural network, such as the transmission mode between human brain neurons, and can also constantly self-learn. Through the NPU, the electronic device 100 can implement intelligent cognitive applications such as image recognition, face recognition, voice recognition, and text understanding, etc.

[0139] The external memory interface 120 can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function.

[0140] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various function applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc. The data storage area can store data created during use of the electronic device 100 (such as audio data, a phonebook, etc.), etc.

[0141] The electronic device 100 can implement an audio function through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, and an application processor, etc. For example, music playing, recording, etc.

[0142] The audio module 170 is configured to convert digital audio information into an analog audio signal output, and to convert an analog audio input into a digital audio signal. The audio module 170 can also be configured to encode and decode audio signals. In some examples, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110. The speaker 170A, also referred to as a "loudspeaker", is configured to convert an audio electrical signal into a sound signal. The receiver 170B, also referred to as a "earpiece", is configured to convert an audio electrical signal into a sound signal. The microphone 170C, also referred to as a "microphone", "microphone", is configured to convert a sound signal into an electrical signal. The earphone interface 170D is configured to connect a wired earphone. The number of microphones 170C can be one or more. In some embodiments, the microphone 170C can also be an omnidirectional microphone. The electronic device 100 can simultaneously combine the sound signals collected by the omnidirectional microphone on the electronic device 100 and the sound signals collected by the directional microphone / omnidirectional microphone on the electronic device 200 to determine whether to recognize the near-field voice.

[0143] In some embodiments, the microphone 170C can also be a directional microphone.

[0144] In some embodiments, the microphone 170C can also be referred to as a second audio collector.

[0145] The sensor module 180 can include a pressure sensor, a gyro sensor, a barometric sensor, a magnetic sensor, an acceleration sensor, a gravity sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.

[0146] The keys 190 include a power key, a volume key, etc. The motor 191 can generate a vibration prompt. The indicator 192 can be an indicator light, which can be used to indicate a charging state, a power change, and can also be used to indicate a message, a missed call, a notification, etc.

[0147] The SIM card interface 195 is configured to connect a SIM card. The SIM card can be inserted into or pulled out of the SIM card interface 195 to achieve contact and separation with the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, N being a positive integer greater than 1. The electronic device 100 interacts with the network through the SIM card to realize functions such as call and data communication.

[0148] Please refer to FIG. 4, which shows a schematic diagram of the hardware structure of the electronic device 200.

[0149] For example, the electronic device 200 can be a stylus.

[0150] As shown in FIG. 4, the electronic device 200 can include a processor 401, a memory 402, a Bluetooth communication module 403, a power supply 404, a power management module 405, a microphone 406, and a speaker 407, etc.

[0151] The processor 401 can be configured to read and execute computer-readable instructions. In a specific implementation, the processor 401 can mainly include a controller, an arithmetic unit, and a register. The controller is mainly responsible for instruction decoding and sending control signals for the corresponding operations of the instructions. The arithmetic unit is mainly responsible for saving the register operands and intermediate operation results temporarily stored during the execution of instructions.

[0152] In some embodiments, the processor 401 can be configured to determine whether the near-field voice is contained in the sound signal collected by the microphone 406. In the case of containing the near-field voice, the processor 401 can control the Bluetooth communication module 403 to send a message to the electronic device 100, so that the microphone on the electronic device 100 starts to collect the sound signal.

[0153] The memory 402 is coupled to the processor 401 and is configured to store various software programs and / or groups of instructions. In a specific implementation, the memory 402 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 402 can store an operating system, such as an embedded operating system uCOS, VxWorks, RTLinux, etc. The memory 402 can also store a communication program that can be used to communicate with the electronic device 100 or other devices.

[0154] The Bluetooth communication module 403 can support communication of the Bluetooth low energy (BLE) protocol. Optionally, the Bluetooth communication module 403 can also support communication of the classic Bluetooth (also known as: basic rate / enhanced data rate (BR / EDR) technology) protocol. The Bluetooth communication module 403 can be used to establish a Bluetooth connection between the electronic device 100 and the electronic device 200, and send a message to the electronic device 100 through the Bluetooth connection.

[0155] The power supply 404 can be a rechargeable lithium battery or a replaceable standard battery, etc. The power management module 405 can include an adaptive USB-compatible pulse width modulation (PMW) charging circuit, a Buck DC-DC converter, etc. The power management module 405 can provide the power required by the processor 401, the memory 402, the Bluetooth communication module 403, the microphone 406, the speaker 407, etc. of the electronic device 200.

[0156] The number of microphones 406 can be one or more. The microphone 406 can be a directional microphone, and the microphone 406 can also be an omnidirectional microphone. The microphone 406 can be used to collect a sound signal and send the sound signal to the processor 401, and the processor 401 can confirm whether the near-field voice is recognized based on the sound signal collected by the microphone 406.

[0157] In some embodiments, the microphone 406 can also be referred to as a first audio collector.

[0158] The speaker 407 can be used to send and receive ultrasonic signals, and based on the sent ultrasonic signals and the received ultrasonic signals, confirm whether the target sound emitting part of the target sound emitting object is recognized. The speaker 407 can assist the processor 401 to confirm whether the near-field voice is recognized based on the sound signal collected by the microphone 406, and improve the accuracy of the processor 401 to recognize the near-field voice.

[0159] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 can include more or fewer components than the illustration, or combine certain components, or split certain components, or different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0160] FIG. 5 illustrates a schematic diagram of a software module for recognizing near-field voice in the electronic device 100 and the electronic device 200.

[0161] As shown in FIG. 5, the electronic device 100 can include a first audio collector, a near-field voice recognition unit A, and a communication unit A.

[0162] For example, the first audio collector can be the microphone 406 shown in FIG. 4. The near-field voice recognition unit A can be a processing unit in the processor 401 shown in FIG. 4. The communication unit A can be a communication unit in the Bluetooth communication module 403 shown in FIG. 4.

[0163] The electronic device 200 can include a second audio collector, a near-field voice recognition unit B, and a communication unit B.

[0164] For example, the second audio collector can be the microphone 170C shown in FIG. 3. The near-field voice recognition unit B can be a processing unit in the processor 110 shown in FIG. 3. The communication unit B can be a communication unit in a Bluetooth communication module (not shown in FIG. 3) of the electronic device 100.

[0165] As shown in FIG. 5, the process of the electronic device 100 and the electronic device 200 recognizing the near-field voice can be as follows:

[0166] 1. The first audio collector collects the first sound signal.

[0167] Optionally, when the electronic device 100 is in the open state, the first audio collector can collect the first sound signal in real time.

[0168] Optionally, the first audio collector can start collecting the first sound signal after the electronic device 100 and the electronic device 200 establish a communication connection, which can save the power consumption of the electronic device 100.

[0169] Optionally, the first audio collector can start collecting the first sound signal after the electronic device 100 and the electronic device 200 establish a communication connection and based on the sensor data collected by the electronic device 100 confirming that the electronic device 100 is in the handheld state, which can save the power consumption of the electronic device 100.

[0170] The present application does not limit the timing of the first audio collector collecting the first sound signal.

[0171] 2. The first audio collector sends the first sound signal to the near-field voice recognition unit A.

[0172] 3. The near-field voice recognition unit A recognizes the near-field voice based on the first sound signal.

[0173] After collecting the first sound signal, the first audio collector can send the first sound signal to the near-field voice recognition unit A. The near-field voice recognition unit A can recognize the near-field voice based on the first sound signal. For how the near-field voice recognition unit A recognizes the near-field voice based on the first sound signal, reference can be made to the description of S703 in the embodiment of FIG. 7.

[0174] 4. The near-field voice recognition unit A sends message 1 to the communication unit A.

[0175] In the case that the near-field speech is recognized based on the first sound signal, the near-field speech recognition unit A can send a message 1 to the communication unit A. The message 1 is used to instruct the communication unit A to send a message to the communication unit B in the electronic device 200, so that the electronic device 200 starts to collect sound signals.

[0176] 5、The first audio collector collects a third sound signal.

[0177] 6、The first audio collector sends the third sound signal to the communication unit A.

[0178] In the case that the near-field speech is recognized based on the first sound signal, the first audio collector can continue to collect sound signals, for example, to obtain a third sound signal, and send the third sound signal to the communication unit A. The third sound signal is used to further confirm whether the near-field speech is recognized in combination with the second sound signal collected by the electronic device 100.

[0179] 7、The communication unit A sends a message 2 and the third sound signal to the communication unit B.

[0180] The communication unit A and the communication unit B have a communication connection, for example, a Bluetooth connection. The communication unit A can send the message 2 and the third sound signal to the communication unit B through the communication connection.

[0181] The connection between the electronic device 100 and the electronic device 200 is not limited to the Bluetooth connection. Other connections, for example, a local area network connection, a cloud communication connection, etc., can also be established.

[0182] The message 2 is used to instruct the second audio collector on the electronic device 200 to start collecting sound signals.

[0183] In some embodiments, the message 2 can also be referred to as a first message.

[0184] Optionally, the message 2 and the third sound signal can be sent to the communication unit B at the same time, or can be sent to the communication unit B at different times.

[0185] Optionally, the communication unit A can also not send the message 2 to the communication unit B.

[0186] 8、The communication unit B sends the third sound signal to the near-field speech recognition unit B.

[0187] After receiving the third sound signal sent by the communication unit A, the communication unit B can send the third sound signal to the near-field speech recognition unit B. The third sound signal is used to further confirm whether the near-field speech is recognized in combination with the second sound signal collected by the electronic device 100.

[0188] 9、The communication unit B sends a message 3 to the second audio collector.

[0189] After receiving the message 2 sent by the communication unit A, the communication unit B can send a message 3 to the second audio collector. The message 3 is used to instruct the second audio collector to start collecting the sound signal.

[0190] Optionally, the message 2 and the message 3 can be the same or different.

[0191] 10. The second audio collector collects the second sound signal.

[0192] 11. The second audio collector sends the second sound signal to the near-field voice recognition unit B.

[0193] 12. The near-field voice recognition unit B combines the second sound signal and the third sound signal, and confirms that the near-field voice is recognized.

[0194] After receiving the message 3 sent by the communication unit B, the second audio collector starts collecting the sound signal, for example, the second sound signal. The second audio collector then sends the second sound signal to the near-field voice recognition unit B.

[0195] After obtaining the second sound signal and the third sound signal, the near-field voice recognition unit B can confirm whether the near-field voice is recognized based on the second sound signal and the third sound signal. In the case that the near-field voice is recognized, the electronic device 100 can perform a corresponding operation based on the sound signal collected by the electronic device 100 and / or the sound signal collected by the electronic device 200. For how the near-field voice recognition unit B recognizes the near-field voice based on the second sound signal and the third sound signal, reference can be made to the description of S708 in the embodiment of FIG. 7.

[0196] It should be noted that the steps in FIG. 5 are only used to explain the present application, and the present application does not limit the execution order of the above steps.

[0197] The present application provides an audio collection method. The method is applied to an electronic device 100 and an electronic device 200 that establish a communication connection. After the electronic device 100 and the electronic device 200 jointly recognize a near-field voice, the electronic device 100 and / or the electronic device 200 can automatically start collecting audio, and perform a corresponding operation based on the collected audio, for example, save the audio, or save the audio after converting the audio into text, or perform a corresponding operation based on the instruction contained in the audio, and the like.

[0198] Next, the audio collection method provided by the present application will be introduced in combination with a UI.

[0199] I. The user queries the weather condition through a voice assistant application installed on the electronic device 100.

[0200] For example, the electronic device 100 can be a tablet, and the electronic device 200 can be a stylus. The tablet and the stylus are communicatively connected, and a user can control the tablet by using the stylus. For example, the user can input text, graphics, and the like on the tablet by using the stylus, which facilitates the interaction between the user and the tablet.

[0201] Optionally, the stylus is pre-installed with a first audio collector, and the first audio collector is in an open state, which can collect audio in real time and confirm whether the user's speech is recognized.

[0202] In some embodiments, when the user needs to query the weather, the user can hold the stylus and approach the mouth, and the user can output speech, as shown in FIG. 6A. The stylus can collect a first sound signal by using the first audio collector, and confirm whether the near-field speech is recognized based on the first sound signal. If the near-field speech is recognized based on the first sound signal, the stylus can send a first message to the tablet by using the communication connection, and the first message is used to instruct the tablet to start collecting sound signals.

[0203] While the second audio collector on the tablet collects a third sound signal, the first audio collector on the stylus also collects sound signals periodically or at an indefinite time, and sends the third sound signal collected by the first audio collector to the tablet.

[0204] The tablet can confirm whether the near-field speech is recognized based on the second sound signal and the third sound signal. If the near-field speech is recognized based on the second sound signal and the third sound signal, the tablet can recognize the shortcut instruction included in the sound signal collected by the tablet and / or the stylus, and perform a corresponding operation based on the shortcut instruction. For example, the tablet can recognize the shortcut instruction "How is the weather today?" included in the sound signal, and display the user interface shown in FIG. 6B.

[0205] As shown in FIG. 6B, the tablet displays a prompt box 6001, and the prompt box 6001 can be displayed by a voice assistant application on the tablet. The prompt box 6001 includes a dialogue box 6002, and the dialogue box 6002 displays the text "How is the weather today?", which can be obtained by the tablet from the sound signal collected by the tablet and / or the stylus, and the text can also be referred to as a shortcut instruction.

[0206] In response to the shortcut instruction, the voice assistant application on the electronic device can obtain the weather condition of Shenzhen today from the voice assistant application server, and display the dialogue box 6003 and the weather details 6004 in the prompt box 6001. The weather content "Shenzhen has rain today, 23-25°C, with orange rainstorm warning" is displayed in the dialogue box 6003, and the weather condition of Shenzhen every hour is displayed in the weather details 6004, for example, at 11 o'clock in the morning, the weather of Shenzhen is sunny, and the temperature is 23°C. At 12 o'clock in the afternoon, the weather of Shenzhen is sunny, and the temperature is 23°C. At 1 o'clock in the afternoon, the weather of Shenzhen is cloudy, and the temperature is 25°C. At 2 o'clock in the afternoon, the weather of Shenzhen is cloudy, and the temperature is 25°C. At 3 o'clock in the afternoon, the weather of Shenzhen is cloudy, and the temperature is 25°C. At 4 o'clock in the afternoon, the weather of Shenzhen is cloudy, and the temperature is 25°C.

[0207] II. The user saves the audio collected by the electronic device 100 and / or the electronic device 200 through the memo application in the electronic device 100.

[0208] For example, the electronic device 100 can be a mobile phone, and the electronic device 200 can be a Bluetooth headset. The mobile phone and the Bluetooth headset have a Bluetooth connection

[0209] Referring to the above interaction method between the tablet and the stylus, as shown in FIG. 6C, the user can pick up the Bluetooth headset and actively approach the user's mouth. The user can output the voice, and the mobile phone and the Bluetooth headset can also confirm whether the near-field voice is recognized according to the above method.

[0210] After recognizing the near-field voice, the mobile phone can automatically start the memo application and collect the audio. During the audio collection process, the memo application can display the user interface 6100 shown in FIG. 6D.

[0211] As shown in FIG. 6D, the user interface 6100 can display the process of audio collection, such as a waveform diagram, an audio duration, and text information corresponding to the audio. In this way, the user can check whether the audio collected by the memo application is the same as the voice output by the user. When the audio collected by the memo application is different from the voice output by the user, the user can choose to re-enter the audio or modify the part different from the audio output by the user, so as to avoid the error of the audio collected by the memo application.

[0212] After the user stops outputting the voice, the mobile phone can save the collected audio in the memo application.

[0213] In one possible implementation, the mobile phone can directly save the collected audio in the memo application and automatically generate a voice note.

[0214] After generating the voice note, the phone can display the user interface 6200 shown in FIG. 6E, which includes the name of the voice note, the generation date of the voice note, the duration of the voice note, etc. For example, the name of the voice note can be "work note", and optionally, the name of the voice note can be automatically extracted and generated by the phone based on the audio content. The generation date of the voice note can be "December 18, 2023". The duration of the voice note can be 1 minute and 26 seconds.

[0215] In other possible implementations, the phone can also convert the collected audio into text and automatically generate a text note, and save the text note in the memo application.

[0216] Through the above method, the phone can automatically start the memo application and collect audio when the near-field voice is recognized, and save the collected audio in the memo application. There is no need for the user to manually start the recording function in the memo application, saving user operations and improving user experience.

[0217] The above application scenarios are not limited to the above, and are only used to explain the present application and do not constitute a limitation.

[0218] Next, how the electronic device 100 and the electronic device 200 jointly recognize the near-field voice will be introduced.

[0219] FIG. 7 shows a method flow diagram for the electronic device 100 and the electronic device 200 to jointly recognize the near-field voice.

[0220] The method shown in FIG. 7 includes but is not limited to the following steps:

[0221] S701, the electronic device 100 and the electronic device 200 establish a communication connection.

[0222] For example, the electronic device 100 and the electronic device 200 can establish a Bluetooth connection. The electronic device 100 can be a phone, a tablet, etc., and the electronic device 200 can be a Bluetooth headset, a stylus, etc.

[0223] S702, the electronic device 200 collects a first sound signal through a first audio collector.

[0224] Optionally, after the electronic device 200 and the electronic device 100 establish a communication connection, the electronic device 200 can enter a low-power consumption working mode, and the first audio collector on the electronic device 200 can collect the first sound signal in real time.

[0225] Optionally, the electronic device 200 can also acquire sensor data, and if it is determined based on the sensor data that the motion trajectory of the electronic device 200 meets a preset trajectory, for example, the preset trajectory can be a motion trajectory of lifting up, it is determined that the user performs the action of picking up the electronic device 200 and approaching the mouth at this time, and the electronic device 100 collects the first sound signal through the first audio collector.

[0226] The application does not limit the timing of the electronic device 200 collecting the first sound signal.

[0227] S703, the electronic device 200 determines whether the near-field voice is recognized based on the first sound signal.

[0228] After acquiring the first sound signal, the electronic device 200 can determine whether the near-field voice is recognized based on the first sound signal.

[0229] In the case where the near-field voice is recognized based on the first sound signal, the electronic device 200 performs S704.

[0230] In the case where the near-field voice is not recognized based on the first sound signal, the electronic device 200 can continue to perform S702-S703.

[0231] Optionally, after recognizing the near-field voice based on the first sound signal, before performing S704, the electronic device 200 can further determine whether the voice signal is included in the first sound signal. The voice signal can refer to the voice content output by the person. In the case where the probability of the voice signal in the first sound signal is greater than a first threshold (for example, 80%), the electronic device 200 performs S704 again, which can avoid the interference of noise. For example, the energy value of the voice content output by the person is much greater than the energy value of the noise, and the electronic device 200 can determine whether the voice signal is included in the first sound signal based on the energy value.

[0232] The electronic device 200 can recognize the near-field voice based on the first sound signal in any one of the following ways, but not limited to.

[0233] Method one, when the first audio collector is one or more directional microphones, if the first sound signal meets a first condition, the electronic device 200 can determine that the near-field voice is recognized. The first condition can include that the energy value of the first frequency segment in the first sound signal is greater than a first value, and the difference between the energy value of the first frequency segment and the energy value of other frequency points except the first frequency segment is greater than a second value. The first frequency segment can be a frequency segment between the first frequency point and the second frequency point.

[0234] Based on the introduction of FIG. 1A and FIG. 1B, the sound pickup ability of the directional microphone is different in different directions. If the target sound object is close to the directional microphone, and the target sound object sounds in the preset sound pickup direction of the directional microphone or the extension direction of the preset sound pickup direction, the directional microphone can better collect the sound signal in the preset sound pickup direction of the directional microphone or the extension direction of the preset sound pickup direction, that is, the energy value of the sound signal collected by the directional microphone in the preset sound pickup direction of the directional microphone or the extension direction of the preset sound pickup direction is also higher, and the energy value of the sound signal collected by the directional microphone in the non-preset sound pickup direction of the directional microphone or the extension direction of the non-preset sound pickup direction is lower.

[0235] When the target sound object sounds in the non-preset sound pickup direction of the directional microphone or the extension direction of the non-preset sound pickup direction, the energy value of the sound signal collected by the directional microphone in each direction of the directional microphone is lower.

[0236] When the target sound object is far away from the directional microphone, whether the target sound object sounds in the preset sound pickup direction of the directional microphone or the extension direction of the preset sound pickup direction, the energy value of the sound signal collected by the directional microphone in each direction of the directional microphone is lower.

[0237] Based on this characteristic, the electronic device 200 can confirm that the near-field voice is recognized based on the frequency response curve of the first sound signal.

[0238] FIG. 8A shows a frequency response curve diagram of the sound signal collected by the directional microphone when the target sound object is close to the directional microphone and sounds in the preset sound pickup direction of the directional microphone or the extension direction of the preset sound pickup direction.

[0239] For example, the first frequency band can refer to a frequency band between the frequency point a1 and the frequency point a2.

[0240] As shown in FIG. 8A, the energy value of the first frequency band is greater than the first value, and the energy value of the frequency point other than the first frequency band is less than the first value. And the difference between the energy value of the first frequency band and the energy value of the frequency point other than the first frequency band is also large, for example, the difference is greater than the second value, then the electronic device 200 can confirm that the near-field voice is recognized based on the first sound signal.

[0241] FIG. 8A also shows that when the distance between the target sound object and the directional microphone is different, the energy value of the sound signal collected by the directional microphone also has a difference. FIG. 8A exemplarily shows the energy difference of the sound collected by the directional microphone when the target sound object is 5 cm, 10 cm and 1 m away from the directional microphone and sounds. As can be seen from FIG. 8A, the closer the distance between the target sound object and the directional microphone, the greater the energy value of the low frequency band of the sound collected by the directional microphone.

[0242] FIG. 8B shows a frequency response curve of the sound signal collected by the directional microphone when the target sound object is far away from the directional microphone.

[0243] In some embodiments, when the target sound object is far away from the directional microphone, whether the target sound object sounds in the preset sound pickup direction or the extension direction of the preset sound pickup direction of the directional microphone, as shown in FIG. 8B, the energy values of the sound signals corresponding to different frequency points in the frequency response curve of the sound signal collected by the directional microphone are not much different, for example, the difference between the energy values of the sound signals corresponding to different frequency points is less than a preset value, and the energy values of the sound signals corresponding to different frequency points are also small, for example, less than a first value. When the frequency response curve of the sound signal satisfies the characteristics shown in FIG. 8B, it can be confirmed that the target sound object is far away from the directional microphone, and the electronic device 200 continues to collect the sound signal and confirms whether the frequency response curve of the sound signal satisfies the characteristics shown in FIG. 8A.

[0244] FIG. 8C shows a frequency response curve of the sound signal collected by the directional microphone when the target sound object sounds in the non-preset sound pickup direction or the extension direction of the non-preset sound pickup direction of the directional microphone.

[0245] In some embodiments, when the target sound object sounds in the non-preset sound pickup direction or the extension direction of the non-preset sound pickup direction of the directional microphone, as shown in FIG. 8C, the energy values of the sound signals corresponding to different frequency points in the frequency response curve of the sound signal collected by the directional microphone are not much different, for example, the difference between the energy values of the sound signals corresponding to different frequency points is less than a preset value, and the energy values of the sound signals corresponding to different frequency points are also small, for example, less than a first value. When the frequency response curve of the sound signal satisfies the characteristics shown in FIG. 8C, it can be confirmed that the target sound object is close to the directional microphone, but the target sound object does not sound in the preset sound pickup direction or the extension direction of the preset sound pickup direction of the directional microphone. The electronic device 200 continues to collect the sound signal and confirms whether the frequency response curve of the sound signal satisfies the characteristics shown in FIG. 8A.

[0246] In the second mode, if the first audio collector includes an omnidirectional microphone, the electronic device 200 further includes a near-field voice auxiliary recognition unit. If the first sound signal collected by the first audio collector satisfies the first condition, and the first signal obtained by the near-field voice auxiliary recognition unit satisfies the third condition, the electronic device 200 can confirm that the near-field voice is recognized. The first condition can include that the energy value of the first sound signal is greater than a third value. The third condition can include that the difference between the first emission time and the first reception time is less than a sixth value, and / or the vibration frequency of the target sound object obtained based on the first ultrasonic signal and the second ultrasonic signal is within a first range.

[0247] Based on the introduction of FIG. 1C, the pickup ability of the pointing microphone in different directions is the same. As shown in FIG. 8D, in the frequency response curve of the first sound signal collected by the first audio collector, the energy values of the sound signals corresponding to different frequency points are not much different, for example, the difference between the energy values of the sound signals corresponding to different frequency points is less than a preset value, and the energy values of the sound signals corresponding to different frequency points are greater than the distance between the target sound object and the electronic device 200. The closer the target sound object is to the electronic device 200, the greater the energy value of the first sound signal collected by the first audio collector on the electronic device 200, and the farther the target sound object is from the electronic device 200, the smaller the energy value of the first sound signal collected by the first audio collector on the electronic device 200.

[0248] The electronic device 200 can confirm whether the near-field voice is recognized in combination with the omnidirectional microphone and the near-field voice auxiliary recognition unit.

[0249] For example, if the energy value of the first sound signal collected by the omnidirectional microphone is greater than a third value, and the near-field voice auxiliary recognition unit recognizes the target sound part, the electronic device 200 can confirm that the near-field voice is recognized.

[0250] Wherein, the energy value of the first sound signal collected by the omnidirectional microphone is greater than the third value, indicating that the distance between the target sound object and the electronic device 200 is within a preset distance. The near-field voice auxiliary recognition unit recognizes the target sound part, indicating that the first sound signal is the audio output by the target sound part of the target sound object.

[0251] The near-field voice auxiliary recognition unit recognizes the target sound part, which can be achieved by the following way: the near-field voice auxiliary recognition unit sends a first ultrasonic signal and receives a reflected second ultrasonic signal. Because the target sound part (such as the mouth) moves at a certain frequency and amplitude when outputting audio, the first ultrasonic signal is reflected when it hits the moving target sound part, causing the frequency and amplitude of the first ultrasonic signal to change. The electronic device 200 can determine the vibration frequency of the target sound part of the target sound object based on the transmitted first ultrasonic signal and the received second ultrasonic signal. When the vibration frequency of the target sound part of the target sound object is within a first range, the electronic device 200 can determine that the target sound part is moving and making a sound, and thus can confirm that the near-field voice auxiliary recognition unit recognizes the target sound part. The first range can be between a first frequency value and a second frequency value. For example, the first range can be 20Hz-40Hz.

[0252] The near-field voice auxiliary recognition unit can also recognize the target sound part in other ways, which are not limited in the present application.

[0253] Optionally, the near-field voice auxiliary recognition unit can also detect whether breathing is recognized. If breathing is detected, it means that the near-field voice auxiliary recognition unit recognizes a person, and the distance between the person and the first audio collector is relatively close. The near-field voice auxiliary recognition unit can recognize the near-field voice by the electronic device 200.

[0254] In some embodiments, the near-field voice auxiliary recognition unit can also be referred to as an auxiliary recognition unit.

[0255] In other embodiments, if the first audio collector includes a plurality of omnidirectional microphones, the electronic device 200 can also confirm that the sound signal collected by the electronic device 200 recognizes the near-field voice based on the plurality of omnidirectional microphones. Specifically, please refer to the description in Mode Two of S708. This application will not be repeated here.

[0256] The electronic device 200 can also recognize the near-field voice in other ways, not limited to the above-mentioned ways.

[0257] S704, the electronic device 200 sends a first message to the electronic device 100 through a communication connection.

[0258] In the case of recognizing the near-field voice based on the first sound signal, the electronic device 200 can send a first message to the electronic device 100. The first message is used to instruct the electronic device 100 to start collecting by the second audio collector.

[0259] Through this method, in the case that the electronic device 200 does not recognize the near-field voice, the second audio collector on the electronic device 100 does not collect the sound signal. In the case that the electronic device 200 recognizes the near-field voice, the electronic device 200 controls the electronic device 100 to collect the sound signal, and further confirms whether the electronic device 100 recognizes the near-field voice, which can reduce the power consumption of the electronic device 100.

[0260] S705, the electronic device 100 collects a second sound signal through the second audio collector.

[0261] The second audio collector is pre-installed on the electronic device 100. After receiving the first message sent by the electronic device 200, the electronic device 100 can control the second audio collector to start collecting the sound signal, for example, collecting the second sound signal.

[0262] S706, the electronic device 200 collects a third sound signal through the first audio collector.

[0263] S707, the electronic device 200 sends the third sound signal to the electronic device 100 through a communication connection.

[0264] Optionally, in a case where the near-field voice is recognized based on the first sound signal, the electronic device 200 further continues to collect a sound signal, for example, a third sound signal, and sends the third sound signal to the electronic device 100. The third sound signal is used by the electronic device 100 to further confirm whether the near-field voice is recognized in combination with the second sound signal collected by the electronic device 100.

[0265] It should be noted that S706-S707 can be executed before any one of the steps after S703 and before S708, after S708, or simultaneously with S708.

[0266] S708, the electronic device 100 confirms whether the near-field voice is recognized based on the second sound signal and the third sound signal.

[0267] Optionally, the second sound signal and the third sound signal can be sound signals collected by the electronic device 200 and the electronic device 100 in the same time period.

[0268] Optionally, since the distance between the target sound-emitting object and the first audio collector on the electronic device 200 and the second audio collector on the electronic device 100 is different, the starting collection time of the second sound signal and the third sound signal can be different.

[0269] The electronic device 100 includes a second audio collector, and the electronic device 200 includes a first audio collector. The second audio collector and the first audio collector can form a distributed audio collector array. The first audio collector and the second audio collector collect sound signals at the same time. The electronic device 100 can obtain a distributed sound signal. The distributed sound signal can include the second sound signal collected by the first audio collector and the third sound signal collected by the second audio collector in the same time period. The electronic device 100 can confirm whether the near-field voice is recognized based on the distributed sound signal. In this way, the method can improve the accuracy of the electronic device 100 in recognizing the near-field voice by using the electronic device 200, the electronic device 100, and the electronic device 200 to recognize the near-field voice twice in succession. The accuracy of the electronic device in automatically collecting audio can also be improved, and the occurrence of false touch can be reduced.

[0270] In a case where the near-field voice is recognized based on the second sound signal and the third sound signal, S709 is executed.

[0271] In a case where the near-field voice is not recognized based on the second sound signal and the third sound signal, S705 is executed. The second audio collector on the electronic device 100 continues to collect a sound signal, receives the sound signal collected by the first audio collector on the electronic device 200 sent by the electronic device 200, and continues to identify whether the near-field voice is recognized based on the sound signal collected by the second audio collector and the sound signal collected by the first audio collector.

[0272] Next, how the electronic device 100 identifies the near-field speech based on the second sound signal and the third sound signal is introduced.

[0273] The electronic device 100 can identify the near-field speech based on the second sound signal and the third sound signal in any one or more of the following ways, but not limited thereto.

[0274] The first way: when the first audio collector is a directional microphone, the second audio collector is an omnidirectional microphone, and the second sound signal and the third sound signal satisfy the second condition, it can be confirmed that the near-field speech is recognized. The second condition can include that the difference between the energy value corresponding to the second frequency segment in the second sound signal and the energy value corresponding to the second frequency segment in the third sound signal is greater than a fourth value. The second frequency segment can be a frequency segment between the third frequency point and the fourth frequency point.

[0275] Optionally, the second frequency segment and the first frequency segment can be the same or different.

[0276] Optionally, the second condition can further include any one or more of the following: the energy value of the second sound signal is greater than a third value, the energy value corresponding to the second frequency segment in the third sound signal is greater than a first value, and the difference between the energy value corresponding to the second frequency segment in the frequency response curve of the third sound signal and the energy value of the sound signal near the frequency points other than the second frequency segment is greater than a second value.

[0277] (a) of FIG. 8E shows the frequency response curve of the third sound signal. (b) of FIG. 8E shows the frequency response curve of the second sound signal.

[0278] The second frequency segment can refer to a frequency segment between the frequency point a3 and the frequency point a4 in FIG. 8E.

[0279] When the distance between the target sound object and the electronic device 200 is close, for example, within the first preset distance, and the sound is emitted in the preset pickup direction of the directional microphone or the extension direction of the preset pickup direction, as shown in (a) of FIG. 8E, the energy value of the third sound signal in the second frequency segment is greater than the first value, and the difference between the energy value corresponding to the second frequency segment in the frequency response curve of the third sound signal and the energy value of the sound signal near the frequency points other than the second frequency segment is greater than the second value.

[0280] When the distance between the target sound object and the electronic device 100 is close, for example, within the first preset distance, as shown in (b) of FIG. 8E, the energy value of the second sound signal is also greater than the third value.

[0281] The second condition can include: a difference between the collection time of the second sound signal and the collection time of the third sound signal is greater than the first time length and / or a difference between the energy value of the second sound signal and the energy value of the third sound signal is greater than the fifth value.

[0282] The first audio collector on the electronic device 200, the second audio collector on the electronic device 100, and the distance from the target sound object are different. When the distance between the target sound object and the electronic device 100 and the electronic device 200 is relatively close, for example, within the first preset distance, the same sound signal emitted by the same target sound object has obvious differences in time and energy when reaching the first audio collector and the second audio collector. When the distance between the target sound object and the electronic device 100 and the electronic device 200 is relatively far, for example, beyond the first preset distance, the same sound signal emitted by the same target sound object has no obvious differences in time and energy when reaching the first audio collector and the second audio collector.

[0283] The electronic device 100 can confirm whether the near-field voice is recognized based on the time difference and / or the energy difference of the same audio signal reaching the first audio collector and the second audio collector.

[0284] When the distance between the target sound object and the electronic device 100 and the electronic device 200 is within the first preset distance, the sound signal output by the target sound object is obtained by the second audio collector on the electronic device 100 after air propagation, obtaining the second sound signal, and can also be obtained by the first audio collector on the electronic device 200, obtaining the third sound signal.

[0285] For example, the second audio collector on the electronic device 100 obtains the second sound signal at time 1, and the first audio collector on the electronic device 200 obtains the third sound signal at time 2. The collection difference between time 1 and time 2 is greater than the first time length.

[0286] For example, the second audio collector on the electronic device 100 obtains the energy value of the second sound signal as energy value 1, and the first audio collector on the electronic device 200 obtains the energy value of the third sound signal as energy value 2. The difference between energy value 1 and energy value 2 is greater than the fifth value.

[0287] It should be noted that the electronic device 100 can confirm whether the near-field voice is recognized based on other manners, not limited to the difference between the collection time of the second sound signal and the collection time of the third sound signal, and / or the difference between the energy value of the second sound signal and the energy value of the third sound signal, and the present application does not make any limitation in this regard.

[0288] In the case where it is confirmed that the near-field voice is recognized based on the second sound signal and the third sound signal, the electronic device 100 can perform a corresponding operation based on the fifth sound signal collected by the first audio collector and / or the fourth sound signal collected by the second audio collector, the fifth sound signal being transmitted by the electronic device 200 to the electronic device 100.

[0289] In the case where it is confirmed that the near-field voice is recognized based on the second sound signal and the third sound signal, the electronic device 100 can perform a corresponding operation based on the fifth sound signal collected by the first audio collector and / or the fourth sound signal collected by the second audio collector, the fifth sound signal being transmitted by the electronic device 200 to the electronic device 100.

[0290] The electronic device 100 can perform a corresponding operation based on the sound signal, which can include but is not limited to any one or several of the following: saving the sound signal, converting the sound signal into text and saving the text, recognizing a shortcut instruction in the sound signal and executing the shortcut instruction, etc.

[0291] Optionally, in the case where it is confirmed that the near-field voice is recognized based on the second sound signal and the third sound signal, the electronic device 100 and the electronic device 200 can simultaneously collect the sound signal, or only the electronic device 100 can collect the sound signal, or only the electronic device 200 can collect the sound signal.

[0292] Optionally, the first electronic device can perform a first function operation based on the fourth sound signal collected by the first electronic device and / or the fifth sound signal collected by the second electronic device, which can mean that the first electronic device performs the first function operation based on the second sound signal collected by the first electronic device, the fourth sound signal, and / or the first sound signal, the third sound signal, the fifth sound signal collected by the second electronic device. It can also mean that the first electronic device performs the first function operation based only on the fourth sound signal collected by the first electronic device, and / or the fifth sound signal collected by the second electronic device. It can also mean that the first electronic device performs the first function operation based only on the second sound signal collected by the first electronic device, the fourth sound signal, and / or the fifth sound signal collected by the second electronic device. It can also mean that the first electronic device performs the first function operation based only on the fourth sound signal collected by the first electronic device, and / or the first sound signal, the third sound signal, the fifth sound signal collected by the second electronic device.

[0293] Preferably, in the case where the near-field voice is identified based on the second sound signal and the third sound signal, the sound signal can be collected by the electronic device 100 and the electronic device 200 at the same time, for example, the electronic device 100 can obtain the fourth sound signal collected by the electronic device 100 and the fifth sound signal collected by the electronic device 200, and separate the near-field voice and the far-field voice based on the fourth sound signal and the fifth sound signal.

[0294] In some embodiments, the electronic device 100 can separate the near-field voice from the sound signal in the following manner. Take the case of separating the near-field voice from the sound signal collected by the distributed audio collector array composed of the second audio collector on the electronic device 100 and the first audio collector on the electronic device 200 as an example.

[0295] If the near-field voice in the sound signal collected by the electronic device 100 can be represented as s1(t), the transmission path of the near-field voice is f1(t, n), and the far-field voice can be represented as s2(t), the transmission path of the far-field voice is f2(t, n). Wherein t represents time, f1(t, n) represents the transmission path of the near-field voice to the nth audio collector, and f2(t, n) represents the transmission path of the far-field voice to the nth audio collector in the distributed audio collector array. The sound signal collected by the electronic device 100 can be represented by formula (1). y(t, n) = s1(t) * f1(t, n) + s2(t) * f2(t, n) Formula (1)

[0296] As shown in formula (1), y(t, n) represents the sound signal collected by the nth audio collector in the distributed audio collector array, s1(t) represents the near-field voice collected by the electronic device 100, f1(t, n) represents the transmission path of the near-field voice, s2(t) represents the far-field voice collected by the electronic device 100, and f2(t, n) represents the transmission path of the far-field voice.

[0297] The sound signal received by the distributed audio collector array can be represented by formula (2). Y(t) = [y(t, 1), y(t, 2), …, y(t, N)] Formula (2)

[0298] As shown in formula (2), Y(t) represents the sound signal received by the distributed audio collector array, y(t, 1) represents the sound signal received by the first audio collector in the distributed audio collector array, y(t, 2) represents the sound signal received by the second audio collector in the distributed audio collector array, and y(t, N) represents the sound signal received by the Nth audio collector in the distributed audio collector array.

[0299] Y(t,f) = F(Y(t)) Equation (3)

[0300] As shown in Equation (3), Y(t,f) represents the sound signal received by the distributed audio collector array in the frequency domain, and f represents the frequency. Y(t) represents the sound signal received by the distributed audio collector array in the time domain.

[0301] The electronic device 100 can separate the near-field voice from the sound signal collected by the distributed audio collector array through Equation (4) as follows.

[0302] As shown in Equation (4), s1(t) represents the near-field voice separated from the sound signal collected by the distributed audio collector array, Y(t,f) represents the sound signal received by the distributed audio collector array in the frequency domain, and W(t,f) represents a solving matrix for separating the near-field voice from the sound signal collected by the distributed audio collector array.

[0303] Optionally, if the types of the audio collectors in the distributed audio collector array are different, W(t,f) can be different. For example, when each audio collector in the distributed audio collector array is an omnidirectional microphone, W(t,f) is different from when part of the audio collectors in the distributed audio collector array are omnidirectional microphones and the remaining audio collectors are directional microphones.

[0304] In some embodiments, when each audio collector in the distributed audio collector array is an omnidirectional microphone, W(t,f) can represent a difference relationship between the energy values of the same sound signal emitted by the same sound object and arriving at each audio collector in the distributed audio collector array and / or a difference relationship between the time instants of the same sound signal emitted by the same sound object and arriving at each audio collector in the distributed audio collector array when the sound signal is a near-field voice. For example, the difference relationship between the energy values of the same sound signal emitted by the same sound object and arriving at each audio collector in the distributed audio collector array can be that the difference between any two of the plurality of energy values of the same sound signal received by each audio collector in the distributed audio collector array is greater than a preset energy difference value (for example, a fifth value). The difference relationship between the time instants of the same sound signal emitted by the same sound object and arriving at each audio collector in the distributed audio collector array can be that the difference between any two of the plurality of time instants of the same sound signal received by each audio collector in the distributed audio collector array is greater than a preset time difference value (for example, a first time length).

[0305] Optionally, the solving matrix W(t, f) can be obtained by the electronic device 100 through deep learning, and the solving matrix W(t, f) can be periodically / unscheduled updated.

[0306] It should be noted that the above formula (1) to formula (4) are only used to explain how the near-field voice is separated from the sound signal in the present application, and the near-field voice can also be separated from the sound signal in other ways, which is not limited in the present application.

[0307] In other embodiments, when part of the audio collectors in the distributed audio collector array are omnidirectional microphones, and the remaining part of the audio collectors are all directional microphones, W(t, f) can represent the difference relationship between the energy value of the sound signal collected by the directional microphone and the energy value of the sound signal collected by the omnidirectional microphone when the same sound signal emitted by the same sound object reaches each audio collector in the distributed audio collector array. The difference relationship between the energy value of the sound signal collected by the directional microphone and the energy value of the sound signal collected by the omnidirectional microphone can be that the difference between the energy value corresponding to the second frequency band of the frequency response curve of the sound signal collected by the directional microphone and the energy value corresponding to the second frequency band of the frequency response curve of the sound signal collected by the omnidirectional microphone is greater than a fourth value.

[0308] Optionally, in the case that the electronic device 100 does not identify the near-field voice or does not obtain the sound signal for a continuous first time length, or the electronic device 100 receives a user input operation, the electronic device 100 can stop collecting the sound signal.

[0309] It can be understood that the various user interfaces described in the embodiments of the present application are only example interfaces, and do not constitute a limitation on the present application scheme. In other embodiments, the user interface can adopt different interface layouts, can include more or fewer controls, can increase or reduce other function options, as long as the same invention idea provided by the present application is based on, it is within the protection scope of the present application.

[0310] It should be noted that any feature in any embodiment of the present application, or any part of any feature, can be combined without contradiction or conflict, and the combined technical solution is also within the scope of the embodiments of the present application.

[0311] The above-described embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. An audio acquisition method, characterized in that, The method is applied to a system comprising a first electronic device and a second electronic device, the first electronic device and the second electronic device being established in communication connection; the method comprises: The first electronic device collects a first sound signal; The first electronic device sends a first message to the second electronic device, the first message being sent after the first electronic device confirms that the first sound signal meets a first condition; In response to the first message, the second electronic device collects a second sound signal; The second electronic device receives a third sound signal sent by the first electronic device, the collection time of the third sound signal being later than that of the first sound signal; The first electronic device performs a first function operation based on a fourth sound signal collected by the first electronic device and / or a fifth sound signal collected by the second electronic device, the first function operation being performed after the first electronic device confirms that the second sound signal and the third sound signal meet a second condition, the collection time of the fourth sound signal being later than that of the second sound signal, the collection time of the fifth sound signal being later than that of the third sound signal, and the fifth sound signal being sent by the second electronic device to the first electronic device.

2. The method of claim 1, wherein, The second sound signal and the third sound signal are sound signals collected in the same time period.

3. The method according to claim 1 or 2, characterized in that, The first sound signal and the third sound signal are sound signals collected by a directional microphone in the first electronic device.

4. The method according to claim 1 or 2, characterized in that, When the first sound signal and the third sound signal are sound signals collected by an omnidirectional microphone in the first electronic device, the first electronic device further comprises an auxiliary identification unit, and the first message is sent after the first electronic device confirms that the first sound signal meets the first condition and the first signal obtained by the auxiliary identification unit meets a third condition.

5. The method according to claim 3 or 4, characterized in that, The second sound signal is a sound signal collected by an omnidirectional microphone in the second electronic device.

6. The method of claim 3, wherein, The first condition comprises that the energy value of the sound signal in the first frequency segment of the first sound signal is greater than a first value, and the difference between the energy value of the first frequency segment and the energy value of other frequency points except the first frequency segment is greater than a second value.

7. The method of claim 4, wherein, The first condition comprises that the energy value of the first sound signal is greater than a third value. The first signal comprises a first emission time of a first ultrasonic signal and a first reception time of a second ultrasonic signal, the third condition comprises that the difference between the first emission time and the first reception time is less than a sixth value, and / or the vibration frequency of the target sound emitting part obtained based on the first ultrasonic signal and the second ultrasonic signal is within a first range.

8. The method of claim 5, wherein, When the third sound signal is a sound signal collected by an omnidirectional microphone in the first electronic device, the second condition comprises: The difference between the energy value of the second sound signal and the energy value of the third sound signal is greater than a fifth value; And / or, The difference between the collection time of the second sound signal and the collection time of the third sound signal is greater than a first time length.

9. The method of claim 5, wherein, When the third sound signal is a sound signal collected by a pointing microphone in the first electronic device, the second condition comprises: a difference between an energy value of a sound signal in a second frequency segment in the third sound signal and an energy value corresponding to the second frequency segment in the second sound signal is greater than a fourth value.

10. The method of claim 9, wherein, The second condition further comprises any one or more of the following: an energy value of a sound signal in a second frequency segment in the third sound signal is greater than the first value, a difference between the energy value of the second frequency segment and an energy value of a frequency point other than the second frequency segment is greater than the second value, and an energy value of the second sound signal is greater than the third value.

11. The method according to any one of claims 1 to 10, characterized in that, The first function operation comprises any one or more of the following: saving the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, converting the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device into text information and saving the text information, identifying a shortcut instruction in the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, and performing an operation based on the shortcut instruction.

12. The method according to any one of claims 1 to 11, characterized in that, The first electronic device performs a first function operation based on the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, specifically comprising: The first electronic device separates a sixth sound signal and a seventh sound signal from the sound signal collected by the first electronic device and / or the sound signal collected by the second electronic device, the sixth sound signal being a sound signal output by a target sound-emitting object within a first preset distance from the first electronic device and the second electronic device, and the seventh sound signal being a sound signal output by a target sound-emitting object within a distance less than the first preset distance from the first electronic device and the second electronic device. The first electronic device performs the first function operation based on the seventh sound signal.

13. The method according to any one of claims 1 to 12, characterized in that, The first electronic device is a handwriting pen, and the second electronic device is a tablet.

14. An audio acquisition method, characterized by, The method comprises: In response to a first message sent by the first electronic device, the second electronic device collects a second sound signal, the first message being sent after the first electronic device confirms that a first sound signal satisfies a first condition, the first sound signal being a sound signal collected by the first electronic device; The second electronic device receives a third sound signal sent by the first electronic device, the third sound signal being collected later than the first sound signal. The first electronic device performs a first function operation based on a fourth sound signal collected by the first electronic device and / or a fifth sound signal collected by the second electronic device, the first function operation being performed after the first electronic device confirms that the second sound signal and the third sound signal meet a second condition, the fourth sound signal being collected at a time later than the second sound signal, the fifth sound signal being collected at a time later than the third sound signal, and the fifth sound signal being sent by the second electronic device to the first electronic device.

15. The method of claim 14, wherein, The second sound signal and the third sound signal are sound signals collected in the same time period.

16. The method according to claim 14 or 15, characterized in that The first sound signal and the third sound signal are sound signals collected by a directional microphone in the first electronic device.

17. The method of claim 14 or 15, wherein, When the first sound signal and the third sound signal are sound signals collected by an omnidirectional microphone in the first electronic device, the first electronic device further comprises an auxiliary identification unit, and the first message is sent after the first electronic device confirms that the first sound signal meets the first condition and the first signal obtained by the auxiliary identification unit meets a third condition.

18. The method of claim 15 or 16, wherein, The second sound signal is a sound signal collected by an omnidirectional microphone in the second electronic device.

19. The method of claim 16, wherein, The first condition comprises that an energy value of the first sound signal is greater than a third value. The first signal comprises a first emission time of a first ultrasonic signal and a first reception time of a second ultrasonic signal, the third condition comprises that a difference between the first emission time and the first reception time is less than a sixth value, and / or a vibration frequency of a target sound emitting part obtained based on the first ultrasonic signal and the second ultrasonic signal is within a first range.

20. The method of claim 17, wherein, When the third sound signal is a sound signal collected by an omnidirectional microphone in the first electronic device, the second condition comprises: A difference between an energy value of the second sound signal and an energy value of the third sound signal is greater than a fifth value. And / or, A difference between a collection time of the second sound signal and a collection time of the third sound signal is greater than a first time length.

21. The method of claim 17, wherein, When the third sound signal is a sound signal collected by a directional microphone in the first electronic device, the second condition comprises: A difference between an energy value of a second frequency segment of the third sound signal and an energy value corresponding to the second frequency segment in the second sound signal is greater than a fourth value.

22. The method of claim 21, wherein, The second condition further comprises any one or more of the following: The energy value of the second frequency segment of the third sound signal is greater than the first value, a difference between the energy value of the second frequency segment and energy values of other frequency points except the second frequency segment is greater than the second value, and the energy value of the second sound signal is greater than the third value.

23. The method according to any one of claims 14-22, characterized by, The first function operation includes any one or more of the following: saving the sound signals collected by the first electronic device and / or the second electronic device, converting the sound signals collected by the first electronic device and / or the second electronic device into text information and saving the text information, identifying a shortcut instruction in the sound signals collected by the first electronic device and / or the second electronic device, and executing the shortcut instruction.

24. The method according to any one of claims 14-23, characterized by, The first electronic device executes a first function operation based on the sound signals collected by the first electronic device and / or the second electronic device, specifically including: The first electronic device separates a sixth sound signal and a seventh sound signal from the sound signals collected by the first electronic device and / or the second electronic device, the sixth sound signal being a sound signal output by a target sound-emitting object at a distance greater than a first preset distance from the first electronic device and the second electronic device, and the seventh sound signal being a sound signal output by a target sound-emitting object at a distance less than the first preset distance from the first electronic device and the second electronic device. The first electronic device executes the first function operation based on the seventh sound signal.

25. The method according to any one of claims 14-24, characterized by, The first electronic device is a stylus, and the second electronic device is a tablet.

26. An electronic device, being a second electronic device, characterized in that The second electronic device includes a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to invoke the computer program to cause the second electronic device to execute the method of any one of claims 1-25.

27. A computer readable storage medium comprising instructions, wherein: The instructions, when executed on the second electronic device, cause the second electronic device to execute the method of any one of claims 1-25.

28. A computer program product, characterised in that, The computer program product includes computer instructions, which, when executed on the second electronic device, cause the second electronic device to execute the method of any one of claims 1-25.

29. A system including a first electronic device and a second electronic device, the first electronic device storing a first instruction, and the second electronic device storing a second instruction, when the first instruction is executed, the second instruction is executed, causing the system to execute the method of any one of claims 1-25.

Citation Information

Patent Citations

  • Far field voice control system based on television device

    CN107566874A

  • Equipment control method and device, electronic equipment and readable storage medium

    CN112581949A

  • Audio acquisition method and system, and related device

    CN113129916A

  • Low-power-consumption standby method, electronic equipment and computer readable storage medium

    CN114816026A

  • Audio acquisition method, electronic equipment and storage medium

    CN115884038A