Audio processing method, computer device, and computer-readable storage medium

By obtaining binaural audio signals in wireless headphones and using binaural time difference effect to determine the location of the sound source, voice enhancement or noise filtering is performed, the speech clarity problem caused by the wireless headphone microphone being far away from the human mouth is solved, and higher speech clarity is achieved.

CN117714942BActive Publication Date: 2025-08-15HONOR DEVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311136007.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-31
Publication Date
2025-08-15
Estimated Expiration
2043-08-31

AI Technical Summary

Technical Problem

The microphone of wireless headphones is far away from the human mouth, resulting in poor speech clarity.

Method used

By obtaining the audio signals collected by the first and second headphones, using binaural time difference effect and voice enhancement technology, it is determined whether the ambient sound source includes the target sound source, and perform voice enhancement or noise filtering processing to improve speech clarity.

Benefits of technology

Improves the voice clarity of the audio signal collected by wireless headphones, enhances user voice and filters out ambient noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117714942B_ABST
    Figure CN117714942B_ABST
Patent Text Reader

Abstract

The present application discloses an audio processing method, a computer device and a computer-readable storage medium, which belong to the field of audio technology. The method comprises: obtaining a first audio signal collected by a first earphone, and obtaining a second audio signal collected by a second earphone. Afterwards, determining whether the ambient sound source includes a target sound source located in a target sound emission area based on the first audio signal and the second audio signal. If the ambient sound source includes the target sound source, speech enhancement processing is performed based on the first audio signal and the second audio signal to obtain a first target audio signal; if the ambient sound source does not include the target sound source, noise filtering processing is performed based on the first audio signal and the second audio signal to obtain a second target audio signal. In the present application, speech enhancement is performed on the sound generated by the target sound source in the target sound emission area, and noise filtering is performed on the sound generated by other sound sources outside the target sound emission area, which can improve the speech clarity of the target audio signal finally obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio technology, and in particular to an audio processing method, a computer device, and a computer-readable storage medium. Background Art

[0002] With the continuous development of wireless communication technology, wireless headsets such as Bluetooth headsets are becoming increasingly popular. During calls or recordings, the microphone of a wireless headset is required to capture audio. Because the microphone of a wireless headset is relatively far from the mouth, voice clarity when wearing wireless headsets has always been a key concern for users. Therefore, a method for processing the audio captured by wireless headsets to improve voice clarity is urgently needed. Summary of the Invention

[0003] This application provides an audio processing method, a computer device, and a computer-readable storage medium that can improve speech clarity. The technical solution is as follows:

[0004] In a first aspect, an audio processing method is provided. In this method, a first audio signal captured by a first headphone is obtained, and a second audio signal captured by a second headphone is obtained. Subsequently, a determination is made based on the first and second audio signals whether the ambient sound source includes a target sound source, where the target sound source is a sound source located in a target sound emission area. If the ambient sound source includes the target sound source, speech enhancement processing is performed on the first and second audio signals to obtain a first target audio signal. If the ambient sound source does not include the target sound source, noise filtering processing is performed on the first and second audio signals to obtain a second target audio signal.

[0005] The environmental sound source may be a single sound source or a mixed sound source including multiple sound sources in the environment.

[0006] The target sound emission area is the user's sound emission area. In other words, the target sound emission area is the area where the mouth of the user using the wireless headset is located, that is, the target sound emission area is the area where the sound emitted by the user using the wireless headset mainly propagates.

[0007] Optionally, the target sound emission area can be a diverging fan-shaped area. For example, the target sound emission area can diverge from the user's position toward the front of the user. In this case, the target sound emission area has a vertex and two boundaries. The vertex is the user's position, and the two boundaries are derived from the vertex diverging toward the front of the user. The central angle of the target sound emission area is the angle between the two boundaries of the target sound emission area.

[0008] If the ambient sound source includes the target sound source, it means that there is currently sound in the target sound emission area, and the collected first audio signal and the second audio signal are likely to contain the user's voice. If the ambient sound source does not include the target sound source, it means that there is currently no sound in the target sound emission area, and the collected first audio signal and the second audio signal are likely to not contain the user's voice.

[0009] In this way, when the ambient sound sources include the target sound source, speech enhancement processing is performed based on the first and second audio signals. This enhances the sound produced by the target sound source within the target sound emission area, thereby enhancing the user's voice. When the ambient sound sources do not include the target sound source, noise filtering processing is performed based on the first and second audio signals. This removes noise from sound sources outside the target sound emission area, thereby filtering out ambient noise. This improves the speech clarity of the resulting target audio signal.

[0010] In one possible implementation, the operation of determining whether the ambient sound source includes the target sound source based on the first audio signal and the second audio signal may be: determining a first audio signal segment in the first audio signal and a second audio signal segment in the second audio signal, and determining whether the ambient sound source includes the target sound source based on a time difference between a start acquisition moment of the first audio signal segment and a start acquisition moment of the second audio signal segment.

[0011] The first audio signal segment is an audio signal segment in the first audio signal, and the second audio signal segment is an audio signal segment in the second audio signal. A similarity between audio features of the first audio signal segment and audio features of the second audio signal segment is greater than or equal to a similarity threshold.

[0012] If the similarity between the audio features of the first audio signal segment and the audio features of the second audio signal segment is greater than or equal to the similarity threshold, it means that the audio content of the first audio signal segment is very similar to the audio content of the second audio signal segment, and the first audio signal segment and the second audio signal segment are very likely produced by the same sound source.

[0013] Since the first audio signal segment and the second audio signal segment are very likely to be generated by the same sound source, the time difference between the start acquisition time of the first audio signal segment and the start acquisition time of the second audio signal segment is the time difference between the sound generated by this sound source reaching the two ears. Therefore, based on the time difference between the start acquisition time of the first audio signal segment and the start acquisition time of the second audio signal segment, it can be determined whether this sound source is located in the target sound emission area, that is, whether this sound source is the target sound source.

[0014] As an example, the operation of determining whether the ambient sound source includes the target sound source based on the time difference between the start acquisition time of the first audio signal segment and the start acquisition time of the second audio signal segment can be: if the time difference between the start acquisition time of the first audio signal segment and the start acquisition time of the second audio signal segment is less than or equal to a time difference threshold, then it is determined that the ambient sound source includes the target sound source; if the time difference between the start acquisition time of the first audio signal and the start acquisition time of the second audio signal segment is greater than the time difference threshold, then it is determined that the ambient sound source does not include the target sound source.

[0015] The time difference threshold is used to determine whether the sound source is located in the target sound emission area. The time difference threshold can be determined based on the central angle of the target sound emission area. In this case, if the time difference between the sound arrival at the two ears is less than or equal to the time difference threshold, it can be determined that the sound source is likely located in the target sound emission area. If the time difference between the sound arrival at the two ears is greater than the time difference threshold, it can be determined that the sound source is likely not located in the target sound emission area.

[0016] Since the target sound emission area is a diverging fan-shaped area radiating from the user position to the front of the user, the distance between the sound source in the target sound emission area and the left and right ears is small, so the time difference between the sound generated by the sound source in the target sound emission area and the two ears is small, and the maximum is the time difference threshold. Therefore, if the time difference between the start acquisition time of the first audio signal segment and the start acquisition time of the second audio signal segment is less than or equal to the time difference threshold, it means that the sound source corresponding to the first audio signal segment and the second audio signal segment is most likely the target sound source, and thus it can be determined that the environmental sound source includes the target sound source. If the time difference between the start acquisition time of the first audio signal segment and the start acquisition time of the second audio signal segment is greater than the time difference threshold, it means that the sound source corresponding to the first audio signal segment and the second audio signal segment is most likely not the target sound source, and thus it can be determined that the environmental sound source does not include the target sound source.

[0017] As another example, the operation of determining whether the ambient sound source includes the target sound source based on the time difference between the start acquisition time of the first audio signal segment and the start acquisition time of the second audio signal segment may be: if the time difference between the start acquisition time of the first audio signal segment and the start acquisition time of the second audio signal segment is less than or equal to the time difference threshold, and the sound pressure level of the first audio signal segment and the sound pressure level of the second audio signal segment are both greater than or equal to the sound pressure level threshold, then it is determined that the ambient sound source includes the target sound source; if the time difference between the start acquisition time of the first audio signal and the start acquisition time of the second audio signal segment is greater than the time difference threshold, and / or if the sound pressure level of the first audio signal segment is less than the sound pressure level threshold, and / or if the sound pressure level of the second audio signal segment is less than the sound pressure level threshold, then it is determined that the ambient sound source does not include the target sound source.

[0018] It should be noted that, in addition to the sound generated by the sound source in the target sound emission area, the time difference of the sound generated by the sound source directly behind or above the user reaching the two ears is also relatively small, and may also be less than or equal to the time difference threshold. In this case, since the sound source in the target sound emission area is the user's mouth, it is closer to the two ears than other sound sources, so the sound pressure level of the sound it produces when it reaches the two ears will be greater. To this end, a sound pressure level threshold can be set according to the distance between the person's mouth and any one ear. The sound pressure level threshold is used to further determine whether the sound source is located in the target sound emission area. In this case, if the sound pressure level when the sound reaches the two ears is greater than or equal to the sound pressure level threshold, it can be determined that the sound source is likely to be located in the target sound emission area; if the sound pressure level when the sound reaches the two ears is less than the sound pressure level threshold, it can be determined that the sound source is likely not located in the target sound emission area.

[0019] In this case, if the time difference between the start acquisition time of the first audio signal segment and the start acquisition time of the second audio signal segment is less than or equal to the time difference threshold, and the sound pressure level of the first audio signal segment and the sound pressure level of the second audio signal segment are both greater than or equal to the sound pressure level threshold, it means that the time difference when the sound generated by the sound source corresponding to the first audio signal segment and the second audio signal segment reaches the two ears is small and the sound pressure level when arriving at the two ears is large, so it can be determined that the sound source is located in the target sound emission area, that is, it is determined that the ambient sound source includes the target sound source.

[0020] If the time difference between the start time of collecting the first audio signal segment and the start time of collecting the second audio signal segment is greater than the time difference threshold, it means that the time difference between the sound generated by the sound source corresponding to the first audio signal segment and the second audio signal segment reaching the two ears is large, so it can be determined that the sound source is not located in the target sound emission area, that is, it is determined that the ambient sound source does not include the target sound source.

[0021] If the sound pressure level of the first audio signal segment is less than the sound pressure level threshold, and / or if the sound pressure level of the second audio signal segment is less than the sound pressure level threshold, it means that the sound pressure levels of the sounds generated by the sound sources corresponding to the first audio signal segment and the second audio signal segment are relatively small when reaching the two ears. Therefore, it can be determined that the sound source is not located in the target sound emission area, that is, it is determined that the ambient sound source does not include the target sound source.

[0022] In one possible implementation, if the ambient sound source includes a target sound source, speech enhancement processing is performed on the first audio signal and the second audio signal to obtain the first target audio signal. The operation may be: if the ambient sound source includes the target sound source, the first audio signal and the second audio signal are merged to obtain a third audio signal; speech recognition is performed on the third audio signal to determine whether the third audio signal contains a human voice; if the third audio signal contains a human voice, speech enhancement processing is performed on the third audio signal to obtain the first target audio signal.

[0023] In the present application, after performing speech recognition on the third audio signal, it is possible to identify whether the third audio signal contains speech content. If the third audio signal contains speech content, it means that the third audio signal contains human voice; if the third audio signal does not contain speech content, it means that the third audio signal does not contain human voice.

[0024] If the third audio signal contains a human voice, speech enhancement processing may be performed on the third audio signal to improve speech clarity to obtain the first target audio signal. If the third audio signal does not contain a human voice, speech enhancement processing is not required and the third audio signal may be directly used as the target audio signal.

[0025] In one possible implementation, if the ambient sound source does not include the target sound source, noise filtering is performed on the first audio signal and the second audio signal to obtain the second target audio signal. The operation may be as follows: if the ambient sound source does not include the target sound source, the first audio signal and the second audio signal are merged to obtain a third audio signal; speech recognition is performed on the third audio signal to determine whether the third audio signal contains a human voice; if the third audio signal does not contain a human voice, it is determined that the sound contained in the third audio signal is ambient noise, and noise filtering is performed on the third audio signal to obtain the second target audio signal.

[0026] In the present application, if the third audio signal does not contain human voice, it can be determined that the sounds contained in the third audio signal are all environmental noise. In this case, the third audio signal can be directly subjected to noise filtering to obtain the second target audio signal.

[0027] Furthermore, speech recognition is performed on the third audio signal to determine whether the third audio signal contains a human voice. If it is determined that the third audio signal contains a human voice, voiceprint recognition is performed on the third audio signal to determine whether the human voice contained in the third audio signal is the voice of the target person; if the human voice contained in the third audio signal is not the voice of the target person, it is determined that the sound contained in the third audio signal is environmental noise, and the third audio signal is noise filtered to obtain a second target audio signal.

[0028] The target person is a user of wireless headphones. If the human voice contained in the third audio signal is not the voice of the target person, it indicates that the human voice contained in the third audio signal is not the voice whose clarity needs to be improved. Therefore, the sound contained in the third audio signal can be determined to be ambient noise. Noise filtering is performed on the third audio signal to obtain the second target audio signal.

[0029] If the human voice contained in the third audio signal is the voice of the target person, the third audio signal can be directly used as the target audio signal; alternatively, the third audio signal can be subjected to voice enhancement processing to obtain a third target audio signal, thereby improving the voice clarity of the target audio signal.

[0030] In a second aspect, an audio processing device is provided, wherein the audio processing device has the function of implementing the audio processing method of the first aspect. The audio processing device includes at least one module, wherein the at least one module is used to implement the audio processing method of the first aspect.

[0031] In a third aspect, a computer device is provided. The computer device includes a processor and a memory, wherein the memory is configured to store a program that supports the computer device in executing the audio processing method provided in the first aspect, as well as data used to implement the audio processing method described in the first aspect. The processor is configured to execute the program stored in the memory. The computer device may also include a communication bus that establishes a connection between the processor and the memory.

[0032] In a fourth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium. When the computer-readable storage medium is run on a computer, the computer executes the audio processing method described in the first aspect.

[0033] In a fifth aspect, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute the audio processing method described in the first aspect.

[0034] The technical effects obtained by the above-mentioned second, third, fourth and fifth aspects are similar to the technical effects obtained by the corresponding technical means in the above-mentioned first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a schematic diagram of a communication system provided in an embodiment of the present application;

[0036] Figure 2 This is a schematic diagram of communication between a wireless headset and a terminal provided in an embodiment of the present application;

[0037] Figure 3 is a schematic diagram of an interaural time difference effect provided by an embodiment of the present application;

[0038] Figure 4 is a schematic diagram of a target sound emission area provided in an embodiment of the present application;

[0039] Figure 5Schematic diagram of a high signal-to-noise ratio sound emission area and a low signal-to-noise ratio sound emission area provided in an embodiment of the present application;

[0040] Figure 6 is a schematic diagram of a first audio signal processing process provided by an embodiment of the present application;

[0041] Figure 7 This is a flowchart of the first audio processing method provided in an embodiment of the present application;

[0042] Figure 8 is a schematic diagram of a second audio signal processing process provided by an embodiment of the present application;

[0043] Figure 9 This is a schematic diagram of a voiceprint registration and voiceprint recognition process provided by an embodiment of the present application;

[0044] Figure 10 is a schematic diagram of a third audio signal processing process provided in an embodiment of the present application;

[0045] Figure 11 is a flowchart of a second audio processing method provided in an embodiment of the present application;

[0046] Figure 12 is a flowchart of a third audio processing method provided in an embodiment of the present application;

[0047] Figure 13 is a flowchart of the fourth audio processing method provided in an embodiment of the present application;

[0048] Figure 14 is a schematic diagram of an audio processing method provided in an embodiment of the present application;

[0049] Figure 15 is a schematic diagram of another audio processing method provided in an embodiment of the present application;

[0050] Figure 16 This is a schematic diagram of the structure of a terminal provided in an embodiment of the present application;

[0051] Figure 17 This is a block diagram of a software system of a terminal provided in an embodiment of the present application;

[0052] Figure 18 This is a schematic structural diagram of a wireless headset provided in an embodiment of the present application;

[0053] Figure 19 It is a structural diagram of an audio processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0055] It should be understood that the “multiple” mentioned in this application refers to two or more. In the description of this application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; “and / or” in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in order to facilitate the clear description of the technical solution of this application, words such as “first” and “second” are used to distinguish between identical or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as “first” and “second” do not limit the quantity and execution order, and words such as “first” and “second” do not necessarily limit them to be different.

[0056] The phrases "one embodiment" or "some embodiments" described in this application mean that the specific features, structures, or characteristics described in that embodiment are included in one or more embodiments of the application. Thus, the phrases "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" that appear in different places in this application do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. In addition, the terms "including," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.

[0057] The application scenarios involved in the embodiments of the present application are described below.

[0058] With the increasing popularity of wireless headsets such as Bluetooth headsets, users are increasingly using them in their daily lives. For example, they are used in voice and video calls, as well as for recording and shooting videos. In these scenarios, wireless headsets need to collect audio signals.

[0059] However, when a user wears a wireless headset, the microphone of the wireless headset is usually located far away from the user's mouth, which will affect the clarity of speech.

[0060] To this end, an embodiment of the present application provides an audio processing method that can utilize binaural time difference effect and speech enhancement technology for directional sound pickup when collecting audio signals through wireless headphones to enhance the user's voice and filter out ambient noise, thereby improving speech clarity.

[0061] The following describes the system architecture involved in the embodiments of the present application.

[0062] Figure 1 Schematic diagram of a communication system provided by an embodiment of the present application. Figure 1 The communication system may include a wireless headset 101 and a terminal 102. The wireless headset 101 and the terminal 102 may communicate via a wireless connection. For example, the wireless headset 101 may be a Bluetooth headset, and the wireless headset 101 and the terminal 102 may communicate via a Bluetooth connection.

[0063] The wireless headset 101 can be any type of true wireless stereo (TWS) headset, neck-worn headset, or wired headset, and the embodiment of the present application is not limited to this.

[0064] The wireless headset 101 may include a first headset and a second headset, one of which is a left headset and the other is a right headset. The left headset is worn on the user's left ear. The right headset is worn on the user's right ear. The first headset has a microphone, and the second headset also has a microphone. The first headset can collect audio signals through its own microphone, and the second headset can also collect audio signals through its own microphone.

[0065] The terminal 102 may be a mobile phone, a PDA, a tablet computer, a laptop computer, a desktop computer or the like, which is not limited in the embodiment of the present application.

[0066] After the terminal 102 is connected to the wireless headset 101, the terminal 102 can collect audio through the wireless headset 101. Figure 2 As shown, the wireless headset 101 is a Bluetooth headset, and the wireless headset 101 is connected to the terminal 102 via Bluetooth. Assume that there are multiple sound sources in the environment where the user is located, such as Figure 2 The sound sources 103, 104, and 105 shown in the figure can all emit sounds, which are collected by the microphone in the wireless headset 101 in the form of audio signals. The wireless headset 101 can transmit the collected audio signals to the terminal 102. The terminal 102 can send the audio signals collected by the wireless headset 101 to a remote device, or generate an audio file based on the audio signals collected by the wireless headset 101 and store it.

[0067] In an embodiment of the present application, when the terminal 102 needs to collect audio through the wireless headset 101, the wireless headset 101 and the terminal 102 can improve the voice clarity of the audio through the following audio processing method.

[0068] The following describes the concepts involved in the embodiments of this application:

[0069] (1) Binaural time difference effect

[0070] The interaural time difference effect refers to the effect of determining the direction of sound by relying on the time difference between the sound reaching the two ears.

[0071] Because there's a certain distance between the left and right ears, sounds arriving from other directions, other than those coming from directly in front, behind, above, or below, arrive at the two ears at different times, creating a time difference. If the sound source is to the right, the sound will arrive at the right ear before the left. If the sound source is to the left, the sound will arrive at the left ear before the right. The further to the side the sound source is, the greater the time difference.

[0072] Next, combine Figure 3 The interaural time difference effect is exemplified.

[0073] Figure 3 Schematic diagram of an interaural time difference effect provided by an embodiment of the present application. Figure 3 As shown, a person 301 and a sound source 302 are in the same environment. Looking down from above the head of the person 301, the sound source 302 is located on the left side of the person 301. Figure 3 The straight and curved lines with arrows roughly represent the propagation paths of the sound generated by sound source 302 to the left and right ears of person 301. As can be seen, because sound source 302 is closer to person 301's left ear, the propagation path of the sound generated by sound source 302 to person 301's left ear is shorter than the propagation path to person 301's right ear. As can be understood, since the speed of sound in air is constant, the sound generated by sound source 302 travels through the air first to person 301's left ear and then to the right ear. The time difference between the arrival of sound at the two ears is an important indicator for determining the direction of sound.

[0074] The difference in the time it takes for a sound to arrive at each ear depends on the location of the sound source relative to the listener. The greater the distance between the sound source and the listener's left and right ears, the greater the time difference. For example, if the sound source is to the listener's left, the sound will arrive at the left ear approximately 0.6 to 0.8 milliseconds earlier than the right ear.

[0075] (2) Target vocalization area

[0076] The target sound emission area is the area where the mouth of the user using the wireless headset is located, that is, the target sound emission area is the area where the sound emitted by the user using the wireless headset is mainly propagated.

[0077] Next, combine Figure 4 A possible implementation of the target vocalization area is exemplified.

[0078] Figure 4Schematic diagram of a target sound area provided in an embodiment of the present application. Figure 4 In Figure (a), the target sound emission area can be a diverging fan-shaped area. For example, the target sound emission area can diverge from the user position to the front of the user. In this case, the target sound emission area has a vertex and two boundaries, the vertex is the user position point, and the two boundaries are obtained by diverging from the vertex to the front of the user. The central angle of the target sound emission area is the angle between the two boundaries of the target sound emission area. For example, the angle between the direction directly in front of the user and any one of the two boundaries of the target sound emission area is δ°, and the angle of the central angle of the target sound emission area is 2δ°.

[0079] For example, see Figure 4 In Figure (b), looking down from above the user's head, the user's position is simplified to a point P. Assuming that the azimuth angle in the direction directly in front of the user is 0°, the entire space can be viewed horizontally as a circular area from 0° to 360° centered on point P. It can be understood that in the direction directly in front of the user, the azimuth angle 0° and the azimuth angle 360° coincide. In the case where the target sound emission area is a divergent fan-shaped area, assuming that δ° is 15°, then the vertex of the target sound emission area is point P, the azimuth angle of one boundary of the target sound emission area is 15°, the azimuth angle of the other boundary is 345°, and the central angle of the target sound emission area is 30°.

[0080] It can be seen from the interaural time difference effect that since the difference between the sound source located in the target sound emission area and the left and right ears is small, the time difference between the sound generated by the sound source located in the target sound emission area and the two ears is small. However, the difference between the sound source located outside the target sound emission area and the left and right ears is large, so the time difference between the sound generated by the sound source located outside the target sound emission area and the two ears is large. For this reason, the embodiment of the present application can determine the time difference threshold based on the angle of the central angle of the target sound emission area. The time difference threshold is used to determine whether the sound source is located in the target sound emission area. Specifically, if the time difference between the sound arriving at the two ears is less than or equal to the time difference threshold, it can be determined that the sound source is located in the target sound emission area; if the time difference between the sound arriving at the two ears is greater than the time difference threshold, it can be determined that the sound source is not located in the target sound emission area.

[0081] like Figure 5 As shown in Figure (a), the user's mouth can be considered a diffuse point sound source. When the user speaks, their voice primarily propagates within the target sound area. Therefore, the audio signal energy of the user's voice within the target sound area is sufficiently large, accounting for approximately 80% or more of the total audio signal energy within the target sound area. Therefore, the target sound area can be defined as a high signal-to-noise ratio sound area, and the area outside the target sound area can be defined as a low signal-to-noise ratio sound area.

[0082] In this case, if Figure 5As shown in Figure (a), the user's sound source is in the high signal-to-noise ratio sound area. Figure 5 As shown in Figure (b), the noise source is in the low SNR region. Speech enhancement can be performed on the sound source in the high SNR region, while noise filtering can be performed on the sound source in the low SNR region, thereby improving speech clarity.

[0083] The audio processing method provided in the embodiment of the present application is explained in detail below.

[0084] The wireless headset includes a first headset and a second headset. One of the first headset and the second headset is a left headset, which can be worn on the user's left ear, and the other headset is a right headset, which can be worn on the user's right ear. For example, the wireless headset described in the embodiments of the present application can be a headset, an earhook headset, a neckhook headset, an earbud headset, etc. The earbud headset can include an in-ear headset (also known as an ear canal headset) or a semi-in-ear headset.

[0085] In an embodiment of the present application, a wireless headset establishes a communication connection with a terminal. Assuming the wireless headset is a Bluetooth headset, the communication connection is a Bluetooth connection. If the Bluetooth switches of both the wireless headset and the terminal are turned on, and if the wireless headset and the terminal have not been previously paired, the wireless headset and the terminal may first perform device discovery and pairing before establishing a Bluetooth connection. If the wireless headset and the terminal have been previously paired, the wireless headset and the terminal may establish a direct Bluetooth connection.

[0086] After the wireless headset is connected to the terminal, the terminal can collect audio through the wireless headset in an audio collection scenario. For example, the audio collection scenario can be a call scenario such as a voice call or video call, or it can be a recording scenario, or it can be a video shooting scenario. Of course, it can also be other scenarios that require audio collection, and the embodiments of the present application are not limited to this.

[0087] In an audio collection scenario, the first earphone and the second earphone need to collect audio signals. In this case, the audio processing method provided in the embodiment of the present application can be used to process the audio signals collected by the first earphone and the second earphone to improve speech clarity.

[0088] In the embodiment of the present application, the audio signal collected by the first earphone may be referred to as a first audio signal, and the audio signal collected by the second earphone may be referred to as a second audio signal.

[0089] It should be noted that the characteristics of audio remain basically unchanged in a short period of time, generally in a short period of time of 10 to 30 milliseconds, that is, it is relatively stable, that is, it has short-term stability. Therefore, the processing of audio can be based on the "short time", that is, the audio can be divided into segments of smaller data signals for processing, where each segment is called a "frame", and the frame length is generally 10 to 30 milliseconds. This not only speeds up the processing speed, but also can well meet the real-time requirements. In an embodiment of the present application, the first audio signal can be the audio frame currently collected by the first earphone, and the second audio signal can be the audio frame currently collected by the second earphone.

[0090] In some embodiments, as Figure 6 As shown in Figure (a), after obtaining a first audio signal captured by a first earphone and a second audio signal captured by a second earphone, the wireless headset can combine the first and second audio signals to obtain a target audio signal, and transmit the target audio signal to the terminal. The terminal then sends the target audio signal to a remote device or generates and stores an audio file based on the target audio signal.

[0091] In other embodiments, Figure 6 As shown in Figure (b), after obtaining the first audio signal collected by the first earphone and the second audio signal collected by the second earphone, the wireless headset can directly transmit the first audio signal and the second audio signal to the terminal. The terminal can combine the first audio signal and the second audio signal to obtain a target audio signal, and then send the target audio signal to the remote device, or generate an audio file based on the target audio signal and store it.

[0092] For example, in a call scenario, the terminal can send the target audio signal to the remote call device. For another example, in a recording scenario or a video shooting scenario, the terminal can generate an audio file based on the target audio signal and store it.

[0093] In some embodiments, the audio processing method provided in the embodiments of the present application can be automatically executed when the terminal collects audio through a wireless headset.

[0094] In other embodiments, the audio processing method provided in the embodiments of the present application can be executed when the terminal collects audio through a wireless headset and the high-definition voice mode is turned on.

[0095] For example, the HD voice mode can be turned on or off by the user in the wireless headset. For example, the user can turn on or off the HD voice mode by long pressing the left earphone, long pressing the right earphone, or operating the relevant button set on the wireless headset. After turning on or off the HD voice mode, the wireless headset can also synchronize the HD voice mode status to the terminal so that the terminal can promptly know whether the HD voice mode is on or off.

[0096] Alternatively, the HD voice mode can be turned on or off by the user in the terminal that is communicating with the wireless headset. For example, after the wireless headset is communicating with the terminal, the user can open the control interface of the wireless headset in the terminal. The control interface may display a voice enhancement control, and the user can turn the HD voice mode on or off by operating the voice enhancement control. After turning the HD voice mode on or off, the terminal can also synchronize the status of the HD voice mode to the wireless headset so that the wireless headset can promptly know whether the HD voice mode is on or off.

[0097] As an optional implementation, the audio processing method provided in the embodiment of the present application can be executed by a wireless headset, as described below. Figure 7 This is described in the examples.

[0098] Figure 7 This is a flowchart of an audio processing method provided in an embodiment of the present application. The method can be executed by a wireless headset, specifically by a processor of the wireless headset. The processor can be provided in the first headset, the second headset, or the neck-wrap portion connecting the first and second headsets. Of course, it can also be provided in other parts of the wireless headset, and this embodiment of the present application is not limited to this.

[0099] See also Figure 7 , the audio processing method may include the following steps:

[0100] Step 701: The wireless headset obtains a first audio signal collected by a first headset, and obtains a second audio signal collected by a second headset.

[0101] The first audio signal is an audio signal obtained by the first earphone from capturing sounds from the external environment. The second audio signal is an audio signal obtained by the second earphone from capturing sounds from the external environment. The first earphone and the second earphone are in the same environment, and the first audio signal and the second audio signal are captured by the first earphone and the second earphone from the same ambient sound source. The ambient sound source can be a single sound source or a mixed sound source comprising multiple sound sources in the environment. For example, while a user is speaking, there may also be the sound of a car horn around the user. In this case, the ambient sound source may include both the user's voice and the car horn.

[0102] In some embodiments, after the wireless headset obtains the first audio signal and the second audio signal, the following steps 702 to 704 may be directly executed.

[0103] In other embodiments, after the wireless headset obtains the first audio signal and the second audio signal, it can first determine whether the high-definition voice mode is turned on; if the high-definition voice mode is turned on, continue to execute the following steps 702 to 704.

[0104] Step 702: The wireless headset determines whether the ambient sound source includes a target sound source based on the first audio signal and the second audio signal. The target sound source is a sound source located in a target sound emission area.

[0105] The target sound emission area is the user's sound emission area. In other words, the target sound emission area is the area where the mouth of the user using the wireless headset is located, that is, the target sound emission area is the area where the sound emitted by the user using the wireless headset mainly propagates.

[0106] If the ambient sound source includes the target sound source, it means that there is currently sound in the target sound emission area, and the collected first audio signal and the second audio signal are likely to contain the user's voice. If the ambient sound source does not include the target sound source, it means that there is currently no sound in the target sound emission area, and the collected first audio signal and the second audio signal are likely to not contain the user's voice.

[0107] In some embodiments, the operation of step 702 may be: the wireless headset determines a first audio signal segment in the first audio signal and a second audio signal segment in the second audio signal, and determines whether the ambient sound source includes the target sound source based on the time difference between the start acquisition time of the first audio signal segment and the start acquisition time of the second audio signal segment.

[0108] The first audio signal segment is an audio signal segment in the first audio signal, and the second audio signal segment is an audio signal segment in the second audio signal. A similarity between audio features of the first audio signal segment and audio features of the second audio signal segment is greater than or equal to a similarity threshold.

[0109] Audio features are features used to characterize audio content. For example, an audio feature may be a feature vector including sub-feature data of one or more dimensions. For example, an audio feature may include sub-feature data of different dimensions such as time domain and frequency domain.

[0110] The similarity threshold can be pre-set and can be set to a larger value, for example, the similarity threshold can be 80%, 90%, etc., which is not limited in the embodiment of the present application.

[0111] If the similarity between the audio features of the first audio signal segment and the audio features of the second audio signal segment is greater than or equal to the similarity threshold, it means that the audio content of the first audio signal segment is very similar to the audio content of the second audio signal segment, and the first audio signal segment and the second audio signal segment are very likely produced by the same sound source.

[0112] The start time of collecting the first audio signal segment is the time when the first earphone begins collecting the first audio signal segment, and the start time of collecting the second audio signal segment is the time when the second earphone begins collecting the second audio signal segment. Since the first and second audio signal segments are likely generated by the same sound source, the time difference between the start time of collecting the first and second audio signal segments (i.e., the duration between the start time of collecting the first and second audio signal segments) is the time difference between the arrival of the sound generated by the same sound source at both ears. Therefore, based on the time difference between the start time of collecting the first and second audio signal segments, it is possible to determine whether the sound source is located in the target sound emission area, that is, to determine whether the sound source is the target sound source. In other words, based on the time difference between the start time of collecting the first and second audio signal segments, it is possible to determine whether the ambient sound sources include the target sound source.

[0113] In some embodiments, the wireless headset may determine whether the ambient sound source includes the target sound source based on the time difference between the start time of collecting the first audio signal segment and the start time of collecting the second audio signal segment in the following two ways:

[0114] In the first method, if the time difference between the start time of collecting the first audio signal segment and the start time of collecting the second audio signal segment is less than or equal to a time difference threshold, the wireless headset determines that the ambient sound source includes the target sound source. If the time difference between the start time of collecting the first audio signal segment and the start time of collecting the second audio signal segment is greater than the time difference threshold, the wireless headset determines that the ambient sound source does not include the target sound source.

[0115] The time difference threshold is used to determine whether the sound source is located in the target sound emission area. The time difference threshold can be determined based on the angle of the center angle of the target sound emission area. In this case, if the time difference between the sound reaching the two ears is less than or equal to the time difference threshold, it can be determined that the sound source is likely located in the target sound emission area. If the time difference between the sound reaching the two ears is greater than the time difference threshold, it can be determined that the sound source is likely not located in the target sound emission area. For example, if the angle of the center angle of the target sound emission area is 30 degrees, the time difference threshold can be 0.5 milliseconds.

[0116] Since the target sound emission area is a diverging fan-shaped area radiating from the user position to the front of the user, the distance between the sound source in the target sound emission area and the left and right ears is small, so the time difference between the sound generated by the sound source in the target sound emission area and the two ears is small, and the maximum is the time difference threshold. Therefore, if the time difference between the start acquisition time of the first audio signal segment and the start acquisition time of the second audio signal segment is less than or equal to the time difference threshold, it means that the sound source corresponding to the first audio signal segment and the second audio signal segment is most likely the target sound source, and thus it can be determined that the environmental sound source includes the target sound source. If the time difference between the start acquisition time of the first audio signal segment and the start acquisition time of the second audio signal segment is greater than the time difference threshold, it means that the sound source corresponding to the first audio signal segment and the second audio signal segment is most likely not the target sound source, and thus it can be determined that the environmental sound source does not include the target sound source.

[0117] Second method: If the time difference between the start time of collecting the first audio signal segment and the start time of collecting the second audio signal segment is less than or equal to a time difference threshold, and the sound pressure level of the first audio signal segment and the sound pressure level of the second audio signal segment are both greater than or equal to a sound pressure level threshold, then the wireless headset determines that the ambient sound source includes the target sound source. If the time difference between the start time of collecting the first audio signal and the start time of collecting the second audio signal segment is greater than the time difference threshold, and / or if the sound pressure level of the first audio signal segment is less than the sound pressure level threshold, and / or if the sound pressure level of the second audio signal segment is less than the sound pressure level threshold, then the wireless headset determines that the ambient sound source does not include the target sound source.

[0118] It should be noted that, in addition to the sound generated by the sound source in the target sound emission area, the time difference of the sound generated by the sound source directly behind or above the user reaching the two ears is also relatively small, and may also be less than or equal to the time difference threshold. In this case, since the sound source in the target sound emission area is the user's mouth, it is closer to the two ears than other sound sources, so the sound pressure level of the sound it produces when it reaches the two ears will be greater. To this end, a sound pressure level threshold can be set according to the distance between the person's mouth and any one ear. The sound pressure level threshold is used to further determine whether the sound source is located in the target sound emission area. In this case, if the sound pressure level when the sound reaches the two ears is greater than or equal to the sound pressure level threshold, it can be determined that the sound source is likely to be located in the target sound emission area; if the sound pressure level when the sound reaches the two ears is less than the sound pressure level threshold, it can be determined that the sound source is likely not located in the target sound emission area.

[0119] In this case, if the time difference between the start acquisition time of the first audio signal segment and the start acquisition time of the second audio signal segment is less than or equal to the time difference threshold, and the sound pressure level of the first audio signal segment and the sound pressure level of the second audio signal segment are both greater than or equal to the sound pressure level threshold, it means that the time difference when the sound generated by the sound source corresponding to the first audio signal segment and the second audio signal segment reaches the two ears is small and the sound pressure level when arriving at the two ears is large, so it can be determined that the sound source is located in the target sound emission area, that is, it is determined that the ambient sound source includes the target sound source.

[0120] If the time difference between the start time of collecting the first audio signal segment and the start time of collecting the second audio signal segment is greater than the time difference threshold, it means that the time difference between the sound generated by the sound source corresponding to the first audio signal segment and the second audio signal segment reaching the two ears is large, so it can be determined that the sound source is not located in the target sound emission area, that is, it is determined that the ambient sound source does not include the target sound source.

[0121] If the sound pressure level of the first audio signal segment is less than the sound pressure level threshold, and / or if the sound pressure level of the second audio signal segment is less than the sound pressure level threshold, it means that the sound pressure levels of the sounds generated by the sound sources corresponding to the first audio signal segment and the second audio signal segment are relatively small when reaching the two ears. Therefore, it can be determined that the sound source is not located in the target sound emission area, that is, it is determined that the ambient sound source does not include the target sound source.

[0122] When the ambient sound source includes the target sound source, the wireless headset may continue to execute step 703. When the ambient sound source does not include the target sound source, the wireless headset may continue to execute step 704.

[0123] Step 703: If the ambient sound source includes the target sound source, the wireless headset performs speech enhancement processing on the first audio signal and the second audio signal to obtain a first target audio signal.

[0124] The speech enhancement processing is used to improve speech clarity. The speech enhancement processing is used to weaken (e.g., remove) non-speech content (e.g., noise, background sound) in the audio signal and / or to enhance the speech content in the audio signal. As a result, the first target audio signal obtained by the wireless headset through speech enhancement processing based on the first audio signal and the second audio signal has higher speech clarity.

[0125] The speech enhancement process can be implemented using one or more speech enhancement algorithms. The one or more speech enhancement algorithms may include one or more speech noise reduction algorithms, and may also include one or more voice enhancement algorithms. The speech noise reduction algorithm is used to remove noise components from the audio signal. The voice enhancement algorithm is used to weaken the background sound components in the audio signal and / or strengthen the effective speech components in the audio signal. The voice enhancement algorithm may use spectral subtraction, minimum mean square error method, etc., which is not limited in the embodiments of the present application.

[0126] When the ambient sound source includes a target sound source, speech enhancement processing is performed based on the first audio signal and the second audio signal, that is, speech enhancement is performed on the sound generated by the target sound source in the target sound emission area, thereby enhancing the user's voice and improving the speech clarity of the overall audio.

[0127] In some embodiments, as Figure 8 As shown, the operation of step 703 may be: if the ambient sound source includes the target sound source, merging the first audio signal and the second audio signal to obtain a third audio signal; performing speech recognition on the third audio signal to determine whether the third audio signal contains a human voice; if the third audio signal contains a human voice, performing speech enhancement processing on the third audio signal to obtain the first target audio signal.

[0128] For example, the wireless headset can perform speech recognition on the third audio signal through automatic speech recognition (ASR) technology. Of course, speech recognition on the third audio signal can also be performed through other speech recognition technologies, which is not limited in this embodiment of the present application.

[0129] After performing speech recognition on the third audio signal, the wireless headset can identify whether the third audio signal contains speech content. If the third audio signal contains speech content, it indicates that the third audio signal contains human voice; if the third audio signal does not contain speech content, it indicates that the third audio signal does not contain human voice.

[0130] If the third audio signal contains a human voice, the wireless headset may perform voice enhancement processing on the third audio signal to improve voice clarity to obtain a first target audio signal, which may then be transmitted to the terminal. If the third audio signal does not contain a human voice, voice enhancement processing is not required, and the wireless headset may directly transmit the third audio signal as the target audio signal to the terminal.

[0131] Step 704: If the ambient sound source does not include the target sound source, the wireless headset performs noise filtering processing on the first audio signal and the second audio signal to obtain a second target audio signal.

[0132] The noise filtering process is used to filter out noise. The noise in the second target audio signal obtained after the wireless headset performs noise filtering on the first audio signal and the second audio signal has been filtered out, thereby helping to improve the overall speech clarity.

[0133] When the ambient sound source does not include the target sound source, noise filtering is performed based on the first audio signal and the second audio signal, that is, noise filtering is performed on the sound generated by other sound sources outside the target sound emission area, thereby filtering out the ambient noise and improving the speech clarity of the overall audio.

[0134] In some embodiments, the operation of step 704 may be: if the ambient sound source does not include the target sound source, merging the first audio signal and the second audio signal to obtain a third audio signal; performing speech recognition on the third audio signal to determine whether the third audio signal contains a human voice; if the third audio signal does not contain a human voice, determining that the sound contained in the third audio signal is ambient noise, and performing noise filtering on the third audio signal to obtain a second target audio signal.

[0135] If the third audio signal does not contain human voice, it can be determined that the sounds contained in the third audio signal are all environmental noise. In this case, the third audio signal can be directly subjected to noise filtering to obtain the second target audio signal.

[0136] For example, the third audio signal can be subjected to noise filtering using an inverse filter. Specifically, an inverse signal of the third audio signal can be generated using an inverse filter, and then the inverse signal can be superimposed on the third audio signal to cancel out the noise component, thereby obtaining a second target audio signal. The wireless headset can then transmit the second target audio signal to the terminal.

[0137] As an example, if the third audio signal includes human voice, the wireless headset may directly transmit the third audio signal as the target audio signal to the terminal.

[0138] As another example, if the third audio signal contains a human voice, the wireless headset can perform voiceprint recognition on the third audio signal to determine whether the human voice contained in the third audio signal is the voice of the target person; if the human voice contained in the third audio signal is not the voice of the target person, the sound contained in the third audio signal is determined to be ambient noise, and the third audio signal is noise filtered to obtain the second target audio signal.

[0139] The target person is a user using a wireless headset. Optionally, the wireless headset may perform voiceprint recognition on the third audio signal by obtaining a voiceprint feature of the third audio signal, inputting the voiceprint feature into a voiceprint recognition model, and obtaining a voiceprint recognition result output by the voiceprint recognition model, wherein the voiceprint recognition result indicates whether the voiceprint feature belongs to the target person. If the voiceprint recognition result indicates that the voiceprint feature belongs to the target person, the voice contained in the third audio signal may be determined to be the voice of the target person; if the voiceprint recognition result indicates that the voiceprint feature does not belong to the target person, the voice contained in the third audio signal may be determined to be not the voice of the target person.

[0140] The voiceprint recognition model is used to identify whether the input voiceprint features belong to the target person. The voiceprint recognition model can be pre-trained. For example, the voiceprint recognition model can be trained by the terminal based on the voice of the target person.

[0141] For example, see Figure 9 In Figure (a), the target person can register a voiceprint on the terminal. Specifically, the terminal can obtain the human voice signal input by the user (i.e., the target person), then identify the voiceprint phonemes in the human voice signal, perform voice enhancement on the human voice signal based on the voiceprint phonemes, and then extract the voiceprint features of the human voice signal after voice enhancement, and use the voiceprint features to train and generate a voiceprint recognition model. Afterwards, see Figure 9 In Figure (b), during the voiceprint recognition stage, the wireless headset can extract the voiceprint features of the third audio signal, and then input the voiceprint features into the voiceprint recognition model to obtain the voiceprint recognition result output by the voiceprint recognition model. The voiceprint recognition result is used to indicate whether the human voice contained in the third audio signal is the voice of the owner.

[0142] After the terminal obtains the voiceprint recognition model through training, it can transmit the voiceprint recognition model to the wireless headset after device pairing and communication connection with the wireless headset.

[0143] It should be noted that once a wireless headset is paired and connected to a terminal, it can receive the voiceprint recognition model of the terminal's owner from that terminal. If the wireless headset subsequently unpairs with that terminal, the wireless headset can delete the voiceprint recognition model until it is paired and connected to another terminal, at which point it can receive a new fingerprint recognition model from that terminal.

[0144] If the human voice contained in the third audio signal is not the voice of the target person, it means that the human voice contained in the third audio signal is not the voice whose clarity needs to be improved. Therefore, it can be determined that the sound contained in the third audio signal is environmental noise, and the third audio signal is subjected to noise filtering to obtain the second target audio signal.

[0145] If the human voice contained in the third audio signal is the voice of the target person, the wireless headset can directly transmit the third audio signal as the target audio signal to the terminal; or, the wireless headset can perform voice enhancement processing on the third audio signal to obtain a third target audio signal and transmit it to the terminal, thereby improving the voice clarity of the target audio signal transmitted to the terminal.

[0146] It should be noted that if Figure 10As shown, in an embodiment of the present application, it is possible to determine whether a target sound source is currently present based on the time difference between the collected two-way audio signals (i.e., the first audio signal and the second audio signal). If the target sound source is currently absent, it can be determined that the currently collected audio signal is likely a noise signal, and noise filtering can then be performed; if the target sound source is currently present, it can be determined that the currently collected audio signal is likely a human voice signal, and speech enhancement can then be performed. In this way, the speech clarity of the processed target audio signal is higher.

[0147] In an embodiment of the present application, the wireless headset obtains a first audio signal collected by the first headset, and obtains a second audio signal collected by the second headset. Afterwards, it is determined whether the ambient sound source includes a target sound source located in the target sound emission area based on the first audio signal and the second audio signal. If the ambient sound source includes the target sound source, speech enhancement processing is performed based on the first audio signal and the second audio signal to obtain a first target audio signal; if the ambient sound source does not include the target sound source, noise filtering processing is performed based on the first audio signal and the second audio signal to obtain a second target audio signal. In this way, the wireless headset performs speech enhancement on the sound generated by the target sound source in the target sound emission area, and performs noise filtering on the sound generated by other sound sources outside the target sound emission area, so as to enhance the user's voice and filter out the ambient noise, thereby improving the speech clarity of the target audio signal finally obtained.

[0148] As an optional implementation, the audio processing method provided in the embodiment of the present application can be executed by the terminal, as described below. Figure 11 This is described in the examples.

[0149] Figure 11 This is a flowchart of an audio processing method provided by an embodiment of the present application. This method can be executed by a terminal. Figure 11 , the audio processing method may include the following steps:

[0150] Step 1101: The terminal obtains a first audio signal collected by a first earphone, and obtains a second audio signal collected by a second earphone.

[0151] After the wireless headset obtains a first audio signal collected by the first headset and a second audio signal collected by the second headset, the first audio signal and the second audio signal can be transmitted to the terminal. In this case, both the first audio signal and the second audio signal have timestamps. The timestamp of the first audio signal is used to indicate the time when the first audio signal was collected, and the timestamp of the second audio signal is used to indicate the time when the second audio signal was collected.

[0152] In some embodiments, after the terminal obtains the first audio signal and the second audio signal, it may directly execute the following steps 1102 to 1104 .

[0153] In other embodiments, after obtaining the first audio signal and the second audio signal, the terminal may first determine whether the high-definition voice mode is enabled; if the high-definition voice mode is enabled, the following steps 1102 to 1104 are continued to be executed.

[0154] Step 1102: The terminal determines whether the ambient sound source includes a target sound source based on the first audio signal and the second audio signal. The target sound source is a sound source located in a target sound emission area.

[0155] The operation of step 1102 is similar to that of the above-mentioned step 702, and will not be described in detail in this embodiment of the present application.

[0156] When the ambient sound source includes the target sound source, the terminal may continue to execute the following step 1103. When the ambient sound source does not include the target sound source, the terminal may continue to execute the following step 1104.

[0157] Step 1103: If the ambient sound source includes the target sound source, the terminal performs speech enhancement processing according to the first audio signal and the second audio signal to obtain a first target audio signal.

[0158] The operation of step 1103 is similar to that of the above-mentioned step 703, and will not be described in detail in this embodiment of the present application.

[0159] Step 1104: If the ambient sound source does not include the target sound source, the terminal performs noise filtering processing on the first audio signal and the second audio signal to obtain a second target audio signal.

[0160] The operation of step 1104 is similar to the operation of the above-mentioned step 704, and will not be repeated in this embodiment of the present application.

[0161] It should be noted that after the terminal obtains the target audio signal (including the first target audio signal, the second target audio signal, and the third target audio signal) through step 1103 or step 1104 above, the target audio signal can be transmitted to the remote device, or an audio file can be generated and stored based on the target audio signal. For example, in a call scenario, the terminal can send the target audio signal to the remote call device. For another example, in a recording scenario or a video shooting scenario, the terminal can generate an audio file based on the target audio signal and store it.

[0162] In an embodiment of the present application, the terminal obtains a first audio signal collected by a first earphone, and obtains a second audio signal collected by a second earphone. Afterwards, it is determined whether the ambient sound source includes a target sound source located in the target sound emission area based on the first audio signal and the second audio signal. If the ambient sound source includes the target sound source, speech enhancement processing is performed based on the first audio signal and the second audio signal to obtain a first target audio signal; if the ambient sound source does not include the target sound source, noise filtering processing is performed based on the first audio signal and the second audio signal to obtain a second target audio signal. In this way, the terminal performs speech enhancement on the sound generated by the target sound source in the target sound emission area, and performs noise filtering on the sound generated by other sound sources outside the target sound emission area, so as to enhance the user's voice and filter out the ambient noise, thereby improving the speech clarity of the target audio signal finally obtained.

[0163] As an optional implementation, the audio processing method provided in the embodiment of the present application can be executed by the wireless headset and the terminal in cooperation, as described below. Figure 12 This is described in the examples.

[0164] Figure 12 This is a flow chart of an audio processing method provided by an embodiment of the present application. Figure 12 , the audio processing method may include the following steps:

[0165] Step 1201: The wireless headset obtains a first audio signal collected by a first headset, and obtains a second audio signal collected by a second headset.

[0166] The operation of step 1201 is similar to that of the above-mentioned step 701, and will not be described in detail in this embodiment of the present application.

[0167] In some embodiments, after the wireless headset obtains the first audio signal and the second audio signal, the following step 1202 may be directly executed.

[0168] In other embodiments, after the wireless headset obtains the first audio signal and the second audio signal, it can first determine whether the high-definition voice mode is turned on; if the high-definition voice mode is turned on, continue to execute the following step 1202.

[0169] Step 1202: The wireless headset determines whether the ambient sound source includes a target sound source based on the first audio signal and the second audio signal. The target sound source is a sound source located in a target sound emission area.

[0170] The operation of step 1202 is similar to the operation of the above-mentioned step 702, and will not be repeated in this embodiment of the present application.

[0171] When the ambient sound source includes the target sound source, the wireless headset may transmit a first audio signal, a second audio signal, and first indication information to the terminal, where the first indication information is used to indicate that the ambient sound source includes the target sound source. Upon receiving the first audio signal, the second audio signal, and the first indication information, the terminal learns that the ambient sound source includes the target sound source and may then execute step 1203.

[0172] When the ambient sound source does not include the target sound source, the wireless headset may transmit a first audio signal, a second audio signal, and second indication information to the terminal, where the second indication information is used to indicate that the ambient sound source includes the target sound source. Upon receiving the first audio signal, the second audio signal, and the second indication information, the terminal learns that the ambient sound source does not include the target sound source and may then execute step 1204.

[0173] Step 1203: If the ambient sound source includes the target sound source, the terminal performs speech enhancement processing according to the first audio signal and the second audio signal to obtain a first target audio signal.

[0174] The operation of step 1203 is similar to that of the above-mentioned step 703, and will not be described in detail in this embodiment of the present application.

[0175] Step 1204: If the ambient sound source does not include the target sound source, the terminal performs noise filtering processing on the first audio signal and the second audio signal to obtain a second target audio signal.

[0176] The operation of step 1204 is similar to the operation of the above-mentioned step 704, and will not be repeated in this embodiment of the present application.

[0177] It should be noted that after the terminal obtains the target audio signal (including the first target audio signal, the second target audio signal, and the third target audio signal) through the above step 1203 or step 1204, it can transmit the target audio signal to the remote device, or generate an audio file based on the target audio signal and store it. For example, in a call scenario, the terminal can send the target audio signal to the remote call device. For another example, in a recording scenario or a video shooting scenario, the terminal can generate an audio file based on the target audio signal and store it.

[0178] In an embodiment of the present application, the wireless headset obtains a first audio signal collected by the first headset, and obtains a second audio signal collected by the second headset, and then determines whether the ambient sound source includes a target sound source located in the target sound emission area based on the first audio signal and the second audio signal. If the ambient sound source includes the target sound source, the terminal performs voice enhancement processing based on the first audio signal and the second audio signal to obtain a first target audio signal; if the ambient sound source does not include the target sound source, the terminal performs noise filtering processing based on the first audio signal and the second audio signal to obtain a second target audio signal. In this way, the sound generated by the target sound source in the target sound emission area is voice enhanced, and the sound generated by other sound sources outside the target sound emission area is noise filtered to enhance the user's voice and filter out the ambient noise, thereby improving the voice clarity of the target audio signal finally obtained.

[0179] As an optional implementation, the audio processing method provided in the embodiment of the present application can be executed by the wireless headset and the terminal in cooperation, as described below. Figure 13 This is described in the examples.

[0180] Figure 13 This is a flow chart of an audio processing method provided by an embodiment of the present application. Figure 13 , the audio processing method may include the following steps:

[0181] Step 1301: The wireless headset obtains a first audio signal collected by a first headset, and obtains a second audio signal collected by a second headset.

[0182] The operation of step 1301 is similar to that of the above-mentioned step 701, and will not be described in detail in this embodiment of the present application.

[0183] In some embodiments, after the wireless headset obtains the first audio signal and the second audio signal, the following step 1302 may be directly executed.

[0184] In other embodiments, after the wireless headset obtains the first audio signal and the second audio signal, it can first determine whether the high-definition voice mode is turned on; if the high-definition voice mode is turned on, continue to execute the following step 1302.

[0185] Step 1302: The wireless headset determines whether the ambient sound source includes a target sound source based on the first audio signal and the second audio signal. The target sound source is a sound source located in a target sound emission area.

[0186] The operation of step 1302 is similar to that of the above-mentioned step 702, and will not be described in detail in this embodiment of the present application.

[0187] After determining whether the ambient sound source includes the target sound source, the wireless headset may execute step 1303 .

[0188] Step 1303: The wireless headset combines the first audio signal and the second audio signal to obtain a third audio signal.

[0189] When the ambient sound source includes the target sound source, the wireless headset may transmit a third audio signal and first indication information to the terminal, where the first indication information is used to indicate that the ambient sound source includes the target sound source. Upon receiving the third audio signal and the first indication information, the terminal learns that the ambient sound source includes the target sound source and may then execute step 1304.

[0190] When the ambient sound source does not include the target sound source, the wireless headset may transmit a third audio signal and second indication information to the terminal, where the second indication information is used to indicate that the ambient sound source includes the target sound source. After receiving the third audio signal and the second indication information, the terminal learns that the ambient sound source does not include the target sound source and may then execute step 1305.

[0191] Step 1304: If the ambient sound source includes the target sound source, the terminal performs speech enhancement processing on the third audio signal to obtain a first target audio signal.

[0192] The operation of step 1304 is similar to that of the above-mentioned step 703, and will not be described in detail in this embodiment of the present application.

[0193] Step 1305: If the ambient sound source does not include the target sound source, the terminal performs noise filtering processing on the third audio signal to obtain a second target audio signal.

[0194] The operation of step 1305 is similar to that of the above-mentioned step 704, and will not be described in detail in this embodiment of the present application.

[0195] It should be noted that after the terminal obtains the target audio signal (including the first target audio signal, the second target audio signal, and the third target audio signal) through the above step 1304 or step 1305, it can transmit the target audio signal to the remote device, or generate an audio file based on the target audio signal and store it. For example, in a call scenario, the terminal can send the target audio signal to the remote call device. For another example, in a recording scenario or a video shooting scenario, the terminal can generate an audio file based on the target audio signal and store it.

[0196] In an embodiment of the present application, the wireless headset obtains a first audio signal collected by the first headset, and obtains a second audio signal collected by the second headset, and then determines whether the ambient sound source includes a target sound source located in the target sound emission area based on the first audio signal and the second audio signal. Thereafter, the first audio signal and the second audio signal are merged to obtain a third audio signal. If the ambient sound source includes the target sound source, the terminal performs voice enhancement processing based on the third audio signal to obtain a first target audio signal; if the ambient sound source does not include the target sound source, the terminal performs noise filtering processing based on the third audio signal to obtain a second target audio signal. In this way, the sound generated by the target sound source in the target sound emission area is voice enhanced, and the sound generated by other sound sources outside the target sound emission area is noise filtered to enhance the user's voice and filter out the ambient noise, thereby improving the voice clarity of the target audio signal finally obtained.

[0197] The following combination Figure 14 and Figure 15 The above audio processing method is exemplified below.

[0198] Figure 14 This is a schematic diagram of an audio processing method provided in an embodiment of the present application.

[0199] See also Figure 14 ,During the signal acquisition stage, you can turn on the high-definition voice mode in the audio acquisition scenario, and then collect two-way audio signals.

[0200] In the threshold judgment stage, the time difference between the start acquisition moments of the audio signal segments whose audio feature similarity in the two-way audio signal is greater than or equal to the similarity threshold is compared with the time difference threshold to determine whether the time difference is less than or equal to the time difference threshold or greater than the time difference threshold.

[0201] In the signal processing stage, if the time difference is less than or equal to the time difference threshold, the merged audio signal of the two-way audio signal is subjected to speech recognition and speech enhancement to obtain a target audio signal. If the time difference is greater than the time difference threshold, speech recognition is performed on the merged audio signal of the two-way audio signal to determine whether the merged audio signal contains a human voice; if the merged audio signal does not contain a human voice, noise is filtered out of the merged audio signal; if the merged audio signal contains a human voice, voiceprint recognition is performed on the merged audio signal to determine whether the human voice contained in the merged audio signal is the voice of the target person; if the human voice contained in the merged audio signal is not the voice of the target person, noise is filtered out of the merged audio signal; if the human voice contained in the merged audio signal is the voice of the target person, the merged audio signal is used as the target audio signal.

[0202] In this way, the above-mentioned method uses the binaural time difference effect and speech enhancement technology for directional sound pickup, which can enhance the user's voice and filter out environmental noise, thereby improving speech clarity.

[0203] Figure 15 This is a schematic diagram of an audio processing method provided in an embodiment of the present application.

[0204] See also Figure 15 ,During the signal acquisition stage, you can turn on the high-definition voice mode in the audio acquisition scenario, and then collect two-way audio signals.

[0205] In the threshold judgment stage, the time difference between the start acquisition moments of the audio signal segments whose audio feature similarity in the two-way audio signal is greater than or equal to the similarity threshold is compared with the time difference threshold to determine whether the time difference is less than or equal to the time difference threshold or greater than the time difference threshold.

[0206] During the signal processing phase, if the time difference is less than or equal to the time difference threshold, speech recognition and speech enhancement are performed on the merged audio signal of the two-channel audio signals to obtain a target audio signal. If the time difference is greater than the time difference threshold, speech recognition is performed on the merged audio signal of the two-channel audio signals to determine whether the merged audio signal contains a human voice. If the merged audio signal does not contain a human voice, noise is filtered out of the merged audio signal. If the merged audio signal contains a human voice, the merged audio signal is used as the target audio signal.

[0207] In this way, the above-mentioned method uses the binaural time difference effect and speech enhancement technology for directional sound pickup, which can enhance the user's voice and filter out environmental noise, thereby improving speech clarity.

[0208] The terminal involved in the embodiments of the present application is described below.

[0209] Figure 16 This is a schematic diagram of the structure of a terminal provided by an embodiment of the present application. Figure 16The terminal 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0210] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the terminal 100. In other embodiments of the present application, the terminal 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0211] The processor 110 may include one or more processing units, for example, an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0212] The controller may be the nerve center and command center of the terminal 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0213] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0214] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0215] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the terminal 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created by the terminal 100 during use (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0216] The USB interface 130 is an interface that complies with USB standards and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. The USB interface 130 can be used to connect a charger to charge the terminal 100 and to transfer data between the terminal 100 and peripheral devices. It can also be used to connect headphones to play audio. The USB interface 130 can also be used to connect other devices, such as AR devices.

[0217] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the terminal 100. While charging the battery 142, the charging management module 140 can also provide power to the terminal 100 via the power management module 141.

[0218] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.

[0219] The wireless communication function of the terminal 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0220] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0221] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied on the terminal 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.

[0222] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. applied on the terminal 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0223] Terminal 100 implements display functions through a GPU, display screen 194, and an application processor. The GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0224] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLED, or a quantum dot light-emitting diode (QLED). In some embodiments, terminal 100 may include one or N display screens 194, where N is an integer greater than one.

[0225] The terminal 100 can realize the shooting function through the ISP, camera 193, video codec, GPU, display screen 194 and application processor.

[0226] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and transformed into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.

[0227] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the terminal 100 may include 1 or N cameras 193, where N is an integer greater than 1.

[0228] The terminal 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D and the application processor.

[0229] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.

[0230] The buttons 190 include a power button, a volume button, etc. The buttons 190 may be mechanical buttons or touch buttons. The terminal 100 may receive button inputs and generate key signal inputs related to user settings and function control of the terminal 100.

[0231] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. Touch operations on different areas of the display screen 194 can also correspond to different vibration feedback effects. Different application scenarios (such as: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0232] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.

[0233] The SIM card interface 195 is used to connect a SIM card. The SIM card can be connected to or removed from the terminal 100 by inserting it into or removing it from the SIM card interface 195. The terminal 100 can support 1 or N SIM card interfaces, where N is an integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The terminal 100 interacts with the network through the SIM card to implement functions such as calls and data communications. In some embodiments, the terminal 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the terminal 100 and cannot be separated from the terminal 100.

[0234] Next, the software system of the terminal 100 will be described.

[0235] The software system of the terminal 100 may adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present application, the Android system with a layered architecture is used as an example to exemplify the software system of the terminal 100.

[0236] Figure 17 This is a block diagram of a software system of a terminal 100 provided in an embodiment of the present application. Figure 17 The layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other via software interfaces. In some embodiments, the Android system is divided into an application layer, an application framework layer, an Android runtime (Android Runtime), a system layer, and a kernel layer.

[0237] The application layer can include a series of application packages. Figure 17As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.

[0238] The application framework layer provides application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions. Figure 17 As shown, the application framework layer may include a window manager, content provider, view system, telephony manager, resource manager, and notification manager. The window manager manages window programs. It can obtain the display size, determine whether a status bar exists, lock the screen, and take screenshots. The content provider stores and retrieves data and makes it accessible to applications. This data may include video, images, audio, incoming and outgoing calls, browsing history and bookmarks, and the phone book. The view system includes visual controls, such as those for displaying text and images. The view system is used to construct the application's display interface, which may consist of one or more views, such as a view that displays a text message notification icon, a view that displays text, and a view that displays images. The telephony manager provides terminal 100 communication functions, such as managing call status (including connected and disconnected calls). The resource manager provides applications with various resources, such as localized strings, icons, images, layout files, and video files. The notification manager enables applications to display notifications in the status bar. These notifications can be used to convey informational messages and can disappear automatically after a short pause, without requiring user interaction. For example, a notification manager is used to notify users of completed downloads, message alerts, and more. A notification manager can also appear as an icon or scrolling text bar in the system's top status bar, such as notifications from background applications. A notification manager can also appear as a dialog window on the screen, such as a text message in the status bar, a beep, a vibration on an electronic device, or a flashing indicator light.

[0239] The Android Runtime consists of core libraries and a virtual machine (VM). The Android Runtime is responsible for scheduling and management of the Android system. The core libraries consist of two parts: one containing the Java language's callable functions and the other the Android core library. The application layer and the application framework layer run in the VM. The VM executes the Java files in the application layer and application framework layer as binary files. The VM is responsible for performing functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0240] The system layer can include multiple functional modules, such as: surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc. The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications. The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc. The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing. The 2D graphics engine is a drawing engine for 2D drawing.

[0241] The kernel layer is the layer between hardware and software. The kernel layer may include display drivers, camera drivers, audio drivers, sensor drivers, etc.

[0242] The wireless headset involved in the embodiment of the present application is described below.

[0243] Figure 18 This is a schematic diagram of the structure of a wireless headset provided by an embodiment of the present application. Figure 18 The wireless headset may include at least one processor 1801 , a communication bus 1802 , a memory 1803 and at least one communication interface 1804 .

[0244] The processor 1801 may be a microprocessor (including a central processing unit (CPU), etc.), an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present application.

[0245] The communication bus 1802 may include a pathway for transmitting information between the aforementioned components.

[0246] The memory 1803 may be a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an optical disc (including a compact disc read-only memory (CD-ROM), a compact disc, a laser disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1803 may exist independently and be connected to the processor 1801 via the communication bus 1802. The memory 1803 may also be integrated with the processor 1801.

[0247] The communication interface 1804 uses any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area network (WLAN), etc.

[0248] In a specific implementation, as an embodiment, the processor 1801 may include one or more CPUs, such as Figure 18 CPU0 and CPU1 are shown in the figure.

[0249] In a specific implementation, as an embodiment, the wireless headset may include multiple processors, such as Figure 18 1 and 1805. Each of these processors can be a single-core processor or a multi-core processor. A processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0250] The memory 1803 is used to store the program code 1810 for executing the solution of the present application, and the processor 1801 is used to execute the program code 1810 stored in the memory 1803. The wireless headset can implement the audio processing method provided in the above embodiment through the processor 1801 and the program code 1810 in the memory 1803.

[0251] Figure 19 This is a structural diagram of an audio processing device provided by an embodiment of the present application. The device can be implemented as part or all of a computer device by software, hardware, or a combination of both. The computer device can be Figure 17 and Figure 18 The terminal shown, or can be Figure 19 Wireless headphones shown. Figure 19 The device includes: an acquisition module 1901, a determination module 1902, a speech enhancement module 1903 and a noise filtering module 1904.

[0252] An acquisition module 1901 is configured to acquire a first audio signal acquired by a first earphone and a second audio signal acquired by a second earphone;

[0253] A determination module 1902 is configured to determine whether the ambient sound source includes a target sound source based on the first audio signal and the second audio signal, where the target sound source is a sound source located in a target sound emission area;

[0254] a speech enhancement module 1903 configured to perform speech enhancement processing on the first audio signal and the second audio signal to obtain a first target audio signal if the ambient sound source includes a target sound source;

[0255] The noise filtering module 1904 is configured to perform noise filtering processing on the first audio signal and the second audio signal to obtain a second target audio signal if the ambient sound source does not include the target sound source.

[0256] In an embodiment of the present application, a first audio signal collected by a first earphone is obtained, and a second audio signal collected by a second earphone is obtained. Afterwards, it is determined whether the ambient sound source includes a target sound source located in the target sound emission area based on the first audio signal and the second audio signal. If the ambient sound source includes the target sound source, speech enhancement processing is performed based on the first audio signal and the second audio signal to obtain a first target audio signal; if the ambient sound source does not include the target sound source, noise filtering processing is performed based on the first audio signal and the second audio signal to obtain a second target audio signal. In this way, speech enhancement is performed on the sound generated by the target sound source in the target sound emission area, and noise filtering is performed on the sound generated by other sound sources outside the target sound emission area, so as to enhance the user's voice and filter out the ambient noise, thereby improving the speech clarity of the target audio signal finally obtained.

[0257] It should be noted that: the audio processing device provided in the above embodiment only uses the division of the above-mentioned functional modules as an example to illustrate audio processing. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0258] The functional units and modules in the above embodiments may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The above integrated units may be implemented in the form of hardware or software functional units. In addition, the specific names of the functional units and modules are only for the purpose of distinguishing them from each other and are not intended to limit the scope of protection of the embodiments of this application.

[0259] The audio processing device and the audio processing method provided in the above embodiments belong to the same concept. The specific working process of the units and modules in the above embodiments and the technical effects brought about can be found in the method embodiment part and will not be repeated here.

[0260] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (such as a coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more available media integrations. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital versatile disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).

[0261] The above are optional embodiments provided for this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the technical scope disclosed in this application should be included in the scope of protection of this application.

Claims

1. An audio processing method, characterized in that: The method comprises: Acquire a first audio signal collected by the first earphone, and acquire a second audio signal collected by the second earphone; determining a first audio signal segment in the first audio signal and a second audio signal segment in the second audio signal, wherein a similarity between an audio feature of the first audio signal segment and an audio feature of the second audio signal segment is greater than or equal to a similarity threshold; If the time difference between the start acquisition time of the first audio signal segment and the start acquisition time of the second audio signal segment is less than or equal to a time difference threshold, and the sound pressure level of the first audio signal segment and the sound pressure level of the second audio signal segment are both greater than or equal to a sound pressure level threshold, it is determined that the ambient sound source includes a target sound source, and the target sound source is a sound source located in a target sound emission area; if the ambient sound source includes the target sound source, the first audio signal and the second audio signal are merged to obtain a third audio signal; speech recognition is performed on the third audio signal to determine whether the third audio signal contains a human voice; if the third audio signal contains a human voice, speech enhancement processing is performed on the third audio signal to obtain a first target audio signal; if the third audio signal does not contain a human voice, the third audio signal is used as the target audio signal; If the time difference between the start acquisition time of the first audio signal and the start acquisition time of the second audio signal segment is greater than the time difference threshold, and / or if the sound pressure level of the first audio signal segment is less than the sound pressure level threshold, and / or if the sound pressure level of the second audio signal segment is less than the sound pressure level threshold, it is determined that the ambient sound source does not include the target sound source; if the ambient sound source does not include the target sound source, the first audio signal and the second audio signal are combined to obtain a third audio signal; speech recognition is performed on the third audio signal to determine whether the third audio signal contains human voice; if the third audio signal does not contain human voice, the third audio signal is determined to be If the sound contained in the signal is environmental noise, the third audio signal is subjected to noise filtering processing to obtain a second target audio signal; if the third audio signal contains a human voice, the third audio signal is subjected to voiceprint recognition to determine whether the human voice contained in the third audio signal is the voice of the target person; if the human voice contained in the third audio signal is not the voice of the target person, the sound contained in the third audio signal is determined to be environmental noise, the third audio signal is subjected to noise filtering processing to obtain a second target audio signal; if the human voice contained in the third audio signal is the voice of the target person, the third audio signal is subjected to voice enhancement processing to obtain a third target audio signal.

2. The method according to claim 1, wherein The target sound emission area is a diverging fan-shaped area, and the target sound emission area diverges from the user position to the front of the user. The time difference threshold is determined according to the central angle of the target sound emission area.

3. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 2 when executed by the processor.

4. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the method according to any one of claims 1 to 2.

Citation Information

Patent Citations

  • Audio processing method and device, electronic equipment and storage medium

    CN114333886A