Audio generation method and system

The audio generation method dynamically adjusts frequency intervals based on evaluation indices to enhance audio quality by selecting and fusing signals from air and bone conduction microphones, addressing inconsistent quality issues.

JP7738068B2Active Publication Date: 2025-09-11SHENZHEN SHOKZ CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023533791
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-12-31
Publication Date
2025-09-11
Estimated Expiration
2040-12-31

AI Technical Summary

Technical Problem

Existing audio generation methods using a combination of air conduction and bone conduction microphones fail to dynamically adjust frequency splicing points based on environmental noise and speaker variability, resulting in inconsistent audio quality.

Method used

An audio generation method that dynamically adjusts frequency intervals based on evaluation indices in the frequency domain of audio signals from air and bone conduction microphones to maximize sound quality by selecting and fusing audio signals with higher quality at each frequency.

Benefits of technology

Improves audio quality by ensuring that all frequencies in the target audio signal have the highest possible sound quality, adapting to different environments and speakers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007738068000001
    Figure 0007738068000001
  • Figure 0007738068000002
    Figure 0007738068000002
  • Figure 0007738068000003
    Figure 0007738068000003
Patent Text Reader

Abstract

The audio generation method and system presented in this specification dynamically selects a frequency junction point of the audio signals based on the sound quality corresponding to each frequency in the frequency domain of the first audio signal and the second audio signal, thereby dividing the frequency domain into a first frequency range and a second frequency range, and splicing together audio signals with higher sound quality corresponding to each frequency range to obtain a target audio after fusion of the first audio signal and the second audio signal, so that the sound quality of the target audio within each frequency range at the said frequency is all of the highest quality, thereby improving the sound quality of the target audio after fusion.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This document relates to the field of audio signal processing, and more particularly to methods and systems for audio generation. [Background technology]

[0002] In various situations in our lives, we are surrounded by noise, and for a better hearing experience, we need to enhance the sound. The so-called sound enhancement can also be called noise suppression, that is, reducing or suppressing noise to a certain extent, and improving the quality and clarity of the sound surrounded by noise. In conventional methods, the signal source collection device is usually an air conduction component, i.e., an air conduction microphone. In noisy environments, the effective sound signal output by the air conduction microphone is almost completely drowned by noise.

[0003] Currently, bone conduction microphones are increasingly used in electronic products such as earphones to receive audio signals. Bone conduction components differ from air conduction microphones in that they directly pick up vibration signals from the vocal tract and can reduce the effects of environmental noise to a certain extent. In the future, electronic devices will increasingly combine air conduction and bone conduction microphones, which have different characteristics, using the air conduction microphone to pick up external audio signals and the bone conduction microphone to pick up vibration signals from the vocal tract, and then performing audio enhancement processing and fusion on the picked-up signals. This can optimize audio quality in situations with high wind noise or other loud noises.

[0004] In the case of combining an air conduction microphone and a bone conduction microphone, the high frequency band of the signal picked up by the air conduction microphone and the low frequency band of the signal picked up by the bone conduction microphone are usually cut out and combined to obtain the final output audio signal. Currently, in the majority of the air conduction microphone and bone conduction microphone combination schemes, the bone conduction microphone signal corresponding to a frequency lower than the frequency junction point and the air conduction microphone signal corresponding to a frequency higher than the frequency junction point are combined to obtain the combined audio signal.

[0005] However, if different speakers use the same bone conduction microphone or air conduction microphone under the same environment and noise, the signal strength and signal characteristics will be different. Even if the speaker is the same, if the speaker uses the same bone conduction microphone or air conduction microphone under different environment and noise, the signal strength and signal characteristics will be different. Therefore, it is unreasonable to use the same frequency splice point to splice audio signals under different environment and noise conditions or by different speakers, and the resulting audio quality will be poor.

[0006] Therefore, there is a need to provide a new audio generation method and system that selects frequency splicing points based on the audio signals of the environment, noise, or speakers, and splices and fuses the audio signals to obtain better voice quality. Summary of the Invention

[0007] This description provides a novel audio generation method and system that selects frequency joint points based on the audio signals of the environment, noise, or speakers, and splices and fuses the audio signals to obtain better voice quality.

[0008] In a first aspect, the present description provides a method for generating audio, comprising: obtaining a first audio signal and a second audio signal; and generating target audio based on the first audio signal and the second audio signal, wherein a frequency domain of the target audio comprises a first frequency interval and a second frequency interval, an audio signal in the first frequency interval of the target audio signal comprises an audio signal in the first frequency interval of the first audio signal, and an audio signal in the second frequency interval of the target audio signal comprises an audio signal in the second frequency interval of the second audio signal, and the ranges of the first frequency interval and the second frequency interval are dynamically adjusted based on at least a dynamic variation of a first evaluation index in the frequency domain of the first audio signal and a dynamic variation of a second evaluation index in the frequency domain of the second audio signal.

[0009] In some embodiments, the first evaluation index and the voice quality of the first audio signal are positively correlated, the second evaluation index and the voice quality of the second audio signal are positively correlated, and in the first frequency range, the voice quality of the first audio signal is higher than the voice quality of the second audio signal, and in the second frequency range, the voice quality of the first audio signal is lower than the voice quality of the second audio signal.

[0010] In some embodiments, the first evaluation index corresponding to each frequency in the first frequency range is higher than the second evaluation index.

[0011] In some embodiments, the first evaluation metric comprises a first SNR corresponding to the first audio signal, and the second evaluation metric comprises a second SNR corresponding to the second audio signal.

[0012] In some embodiments, generating the target audio based on the first audio signal and the second audio signal includes determining and comparing the first evaluation index and the second evaluation index in the frequency domain, and determining at least one target frequency based on a comparison result between at least the first evaluation index and the second evaluation index, thereby determining the first frequency range and the second frequency range, wherein each target frequency of the at least one target frequency corresponds to a continuous portion of the first frequency range and the second frequency range; and generating the target audio based on the first frequency range and the second frequency range and the first audio signal and the second audio signal.

[0013] In some embodiments, the first frequency interval comprises at least one contiguous frequency interval and the second frequency interval comprises at least one contiguous frequency interval.

[0014] In some embodiments, determining the first frequency interval and the second frequency interval includes determining the at least one target frequency based on a frequency corresponding when the first SNR and the second SNR are equal, and determining that, with the at least one target frequency as a critical point, a corresponding frequency interval when the first SNR exceeds the second SNR is the first frequency interval, and a frequency interval outside the first frequency interval is the second frequency interval.

[0015] In some embodiments, each of the at least one target frequencies includes any frequency within a frequency interval of a given width around a corresponding frequency when the first SNR and the second SNR are equal.

[0016] In some embodiments, determining the first frequency interval and the second frequency interval includes obtaining an SNR threshold; comparing the first SNR and the second SNR, and determining a corresponding frequency when the first SNR and the second SNR are equal as at least one first target frequency; comparing the first SNR and the SNR threshold, and determining a corresponding frequency when the first SNR and the SNR threshold are equal as at least one second target frequency; comparing the first SNR, the second SNR, and the SNR threshold corresponding to each frequency in the at least one first target frequency and the at least one second target frequency, and determining a corresponding frequency when the first SNR is not lower than the second SNR and the SNR threshold as the at least one target frequency; and regarding the at least one target frequency as a critical point, a corresponding frequency interval when the first SNR is higher than the second SNR is the first frequency interval, and a frequency interval outside the first frequency interval is the second frequency interval.

[0017] In some embodiments, generating the target audio based on the first frequency range and the second frequency range and the first audio signal and the second audio signal includes performing a smoothing process on the first audio signal and the second audio signal corresponding to frequencies within a given range of each of the at least one target frequency, thereby achieving a smooth transition between the audio signal of the first audio signal corresponding to frequencies within the given range and the audio signal of the second audio signal corresponding to frequencies within the given range; and acquiring the target audio by connecting the audio signal of the first audio signal that is in the first frequency range and the audio signal of the second audio signal that is in the second frequency range that have been subjected to the smoothing process based on frequency distributions.

[0018] In some embodiments, the first audio signal is an audio signal output by at least one first-class microphone, and the second audio signal is an audio signal output by at least one second-class microphone.

[0019] In some embodiments, the at least one first-class microphone comprises a bone conduction microphone used to collect human body vibration signals, and the at least one second-class microphone comprises an air conduction microphone used to collect air vibration signals.

[0020] In some embodiments, the first audio signal comprises an audio signal directly output by the at least one first-class microphone, and the second audio signal comprises an audio signal directly output by the at least one second-class microphone.

[0021] In some embodiments, the first audio signal includes an audio signal obtained by applying noise reduction processing to an audio signal directly output by the at least one first-class microphone, and the second audio signal includes an audio signal obtained by applying noise reduction processing to an audio signal directly output by the at least one second-class microphone.

[0022] In a second aspect, the present disclosure provides a system for audio generation, comprising at least one storage medium and at least one processor, the memory of the at least one storage medium having at least one instruction set for audio generation, the at least one processor and the at least one storage medium being in communication with each other, wherein, when the system for audio generation is running, the at least one processor reads the at least one instruction set and performs the method for audio generation described in the first aspect of the present disclosure based on instructions in the at least one instruction set.

[0023] As can be seen from the above technical solutions, the audio generation method and system presented in this specification can obtain and compare evaluation indicators corresponding to each frequency in the frequency domain of the first audio signal and the second audio signal to compare the sound quality corresponding to each frequency in the frequency domain of the first audio signal and the second audio signal. By selecting a frequency junction point of the audio signals based on the sound quality dynamics, each frequency in the frequency domain is divided into regions, and audio signals with higher sound quality corresponding to each frequency range are combined to obtain a target audio after fusion of the first audio signal and the second audio signal. The sound quality of each frequency range in the frequency domain of the target audio is maximized, thereby improving the sound quality of the fused target audio. In different scenes, such as scenes with different speaker voice signals or different environmental noises, the method and system can select a crossover point based on the sound quality dynamics of the first audio signal and the second audio signal in the current scene, divide the frequencies into dynamic ranges, and combine the audio signals to obtain a higher sound quality of the fused target audio.

[0024] Other features of the audio generation method and system presented in this document are partially listed in the following description. Based on the description, the following figures and examples are basic for those of ordinary skill in the art. The inventive aspects of the audio generation method and system presented in this document can be fully appreciated by practicing or using the methods, devices and combinations shown in the detailed examples below. [Brief explanation of the drawings]

[0025] In order to more clearly explain the technical solutions of the embodiments of this specification, the following briefly introduces the drawings that need to be used in the description of the embodiments. The drawings in the following description are merely examples of this specification, and general technical personnel in this field can obtain other drawings based on these drawings without providing creative work. [Figure 1]1 shows a schematic diagram of the equipment of an audio generation system presented based on an embodiment of this manual. [Figure 2] 1 shows a flow diagram of a method for audio generation presented in accordance with an embodiment of the present disclosure. [Figure 3] 1 shows graphs of frequency spectra of a first audio signal and a second audio signal presented in accordance with an embodiment of the present disclosure; [Figure 4] 1 shows graphs of first SNR and second SNR presented based on an embodiment of the present instructions. [Figure 5] 1 shows a flow diagram of the first and second frequency intervals presented based on an embodiment of the present manual. [Figure 6] 1 shows graphs of a first frequency range and a second frequency range presented based on an embodiment of the present instructions. [Figure 7] 1 shows another flow chart for determining the first and second frequency ranges proposed according to an embodiment of the present disclosure. [Figure 8] 10 shows another graph of the first and second frequency ranges presented in accordance with an embodiment of the present disclosure. [Figure 9] 10 shows a graph of target audio presented according to an embodiment of the present instructions. [Figure 10] 10 shows another target audio graph presented according to an embodiment of the present instructions. DETAILED DESCRIPTION OF THE INVENTION

[0026] The following description presents specific application scenarios and requirements of this specification, and aims to assist those skilled in the art in their manufacturing using the contents of this specification. For those skilled in the art, various local modifications of the disclosed embodiments are fundamental, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification does not limit the embodiments shown, but rather adopts the broadest scope consistent with the scope of the claims.

[0027] The terms used herein are used for the purpose of describing particular examples and embodiments and are not limiting. For example, unless expressly stated otherwise in the context below, the odd number forms "one," "one," and "the" as used herein can also include the complex number forms. As used in this description, the terms "comprise," "include," and / or "contain" refer to the presence of associated integers, steps, operations, elements, and / or components, but do not preclude the presence of one or more other features, integers, steps, operations, elements, components, and / or sets, or that other features, integers, steps, operations, elements, components, and / or sets can be added to the system / method.

[0028] In view of the following description, these and other features of this specification and the operation and function of the associated parts of the structure, and the economy of assembly and manufacture of parts, have been clearly improved. While the referenced drawings are an integral part of this specification, it is to be expressly understood that the drawings are for illustrative purposes only and do not limit the scope of this specification. It is further to be understood that the drawings are not drawn to scale.

[0029] The flow diagrams used in this specification illustrate the operations of a system implemented according to the embodiments of this specification. It should be clearly understood that the operations in the flow diagrams may be implemented out of order. Conversely, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more operations may be added to a flow diagram, or one or more operations may be removed from a flow diagram.

[0030] To improve the audio quality of the synthesized audio signal, the audio generation method and system presented in this document synthesizes a bone conduction microphone signal and an air conduction microphone signal to generate a target audio signal based on the audio quality of the bone conduction microphone signal and the air conduction microphone signal in different application scenarios, thereby selecting an audio signal with better audio quality at every frequency in the frequency domain, and combining the selected audio signals to obtain the target audio signal, thereby ensuring that all audio signals at every frequency in the frequency domain of the target audio are audio signals of the highest quality.

[0031] 1 shows a schematic diagram of an audio generation system 100 (hereinafter referred to as "system 100"). System 100 can be used in an electronic device 200.

[0032] In some embodiments, electronic device 200 may be a wireless earphone, a wired earphone, a smart wearable device (such as a device with audio processing capabilities, such as smart glasses, a smart helmet, or a smart watch), a mobile device, a tablet PC, a laptop PC, an in-car device, or similar content, or any combination thereof. In some embodiments, the mobile device may comprise a smart home device, a smart mobile device, or similar devices, or any combination thereof. For example, the smart mobile device may comprise a mobile phone, a personal digital assistant, a gaming device, a navigation device, an ultra-mobile personal computer (UMPC), or the like, or any combination thereof. In some embodiments, the smart home device may comprise a smart television, a desktop PC, or the like, or any combination thereof. In some embodiments, the in-car device may comprise an in-car PC, an in-car TV, or the like.

[0033] The electronic device 200 can store data or instructions for implementing the audio generation methods described herein and can execute the data and / or instructions. The electronic device 200 can receive audio signals ready for processing and execute the data or instructions for the audio generation methods described herein to perform synthesis processing on the ready audio signals and generate target audio. The audio generation methods are described elsewhere in this document. For example, the description of Figures 2-10 introduces the audio generation methods.

[0034] The audio signals waiting to be processed include at least two different audio signals. The audio generation method is used to combine the at least two different audio signals to obtain a target audio based on the sound quality within the frequency domain of the at least two different audio signals, thereby improving the sound quality of the target audio. Specifically, the electronic device 200 can compare the sound quality corresponding to each frequency within the frequency domain of the at least two different audio signals, and select and combine the audio signal with the better sound quality within each frequency to obtain the target audio. The sound quality of all audio signals corresponding to all frequencies within the frequency domain of the target audio is the highest quality.

[0035] The audio signal waiting to be processed may be an audio signal stored locally in the electronic device 200, an audio signal output by an audio collecting device of the electronic device 200, or an audio signal sent to the electronic device 200 by another device. The audio collecting device may be integrated into the electronic device 200 or may be an external device communicatively connected to the electronic device 200. The audio signal waiting to be processed may be an audio signal that has undergone noise reduction processing or an audio signal that has not undergone noise reduction processing. For convenience, the following description will be given taking an example in which the audio signal waiting to be processed is an audio signal output by the audio collecting device of the electronic device 200.

[0036] 1, electronic device 200 may include at least one storage medium 230 and at least one processor 220. In some embodiments, electronic device 200 may also include a communication port 250 and an internal communication bus 210, while electronic device 200 may also include I / O components 260. In some embodiments, electronic device 200 may also include a microphone module 240.

[0037] The internal communication bus 210 can connect different system components, including a storage medium 230, a processor 220, and a microphone module 240.

[0038] The I / O component 260 corresponds to input / output between the electronic device 200 and other components. For example, the electronic device 200 can obtain the audio signal waiting to be processed through the I / O component 260.

[0039] The communication port 250 is used for data communication between the electronic device 200 and the outside world. For example, the electronic device 200 can also receive the audio signal waiting to be processed through the communication port 250.

[0040] The at least one storage medium 230 may comprise a data memory device. The data memory device may be either a non-transitory storage medium or a transitory storage medium. For example, the data memory device may comprise one or more of a disk 232, a read-only memory (ROM) 234, or a random access memory (RAM) 236. The storage medium 230 may further comprise at least one of the data memory devices including a set of instructions for audio generation. The instructions may be program code, and the program code may include programs, routines, objects, components, data structures, processes, modules, etc., of the audio generation methods presented herein. The at least one storage medium 230 may store the audio signal awaiting processing.

[0041] The at least one processor 220 can be communicatively connected to the at least one storage medium 230 via the internal communication bus 210. The communicative connection refers to any type of connection that directly or indirectly receives information. The at least one processor 220 is used to execute the at least one instruction set. When the system 100 is running, the at least one processor 220 reads the at least one instruction set and performs the method of audio generation presented herein based on the instructions of the at least one instruction set. The processor 220 can perform all steps included in the method of audio generation. Processor 220 may take the form of one or more processors, and in some embodiments, processor 220 may comprise one or more hardware processors, such as a microcontroller, microprocessor, reduced instruction set computer (RISC), application specific integrated circuit (ASIC), application specific instruction set processor (ASIP), central processing unit (CPU), graphics processing unit (GPU), physics processing unit (PPU), microcontroller unit, digital signal processor (DSP), field programmable gate array (FPGA), advanced RISC machine (ARM), programmable logic device (PLD), or any combination thereof, capable of performing one or more functions. For purposes of illustration only, this description will describe electronic device 200 with only one processor 220; however, it should be noted that this description may also include multiple processors. Accordingly, the operations and / or method steps described herein may be performed by either one processor or multiple processors as described herein. For example, in this description, when a processor 220 of electronic device 200 performs procedure A and procedure B, it should be understood that procedure A and procedure B may be performed jointly or separately by two different processors 220 (e.g., a first processor may perform procedure A and a second processor may perform procedure B, or the first and second processors may perform procedures A and B jointly).

[0042] In some embodiments, the electronic device 200 may also include a microphone module 240. The microphone module 240 may be an audio collection device of the electronic device 200. The microphone module 240 may be configured to acquire local audio signals and output microphone signals, i.e., electronic signals carrying audio information. The audio signals awaiting processing may be the microphone signals output by the microphone module 240. The microphone module 240 may be communicatively coupled to at least one processor 220 and at least one storage medium 230. If the audio signals awaiting processing are microphone signals, when the system 100 is running, the at least one processor 220 may read the at least one instruction set and acquire the microphone signals according to the instructions of the at least one instruction set to perform the audio generation method presented herein. The microphone module 240 may be integrated into the electronic device 200 or may be an external device to the electronic device 200.

[0043] The microphone module 240 can be configured to receive local audio signals and output microphone signals, i.e., electronic signals carrying audio information. The microphone module 240 can be either an extra-aural canal microphone module or an intra-aural canal microphone module. For example, the microphone module 240 can be either an extra-aural canal microphone or an intra-aural canal microphone. The microphone module 240 can include at least one first-class microphone 242 and at least one second-class microphone 244. The first-class microphone 242 is different from the second-class microphone 244. The first-class microphone 242 is a microphone capable of directly collecting human body vibration signals, such as a bone conduction microphone. The second-class microphone 244 is a microphone capable of directly collecting air vibration signals, such as an air conduction microphone. Of course, the microphone module 240 can also be other types of microphones. For example, the first-class microphone 242 can be an optical microphone, and the second-class microphone 244 can be a microphone that receives surface myoelectric potential signals. For convenience, in this presentation, the first type microphone 242 will be explained as a bone conduction microphone, and the second type microphone 244 will be explained as an air conduction microphone.

[0044] A bone conduction microphone may include a vibration sensor, such as an optical vibration sensor or an acceleration sensor. The vibration sensor can collect mechanical vibration signals (e.g., signals generated by vibrations in the skin or bones when a user speaks) and convert the mechanical vibration signals into electrical signals. Here, mechanical vibration signals primarily refer to vibrations transmitted through solid objects. The bone conduction microphone collects vibration signals generated in the bones or skin when the user speaks by contacting the user's skin or bones through the vibration sensor or a vibrating component connected to the vibration sensor, and converts the vibration signals into electrical signals. In some embodiments, the vibration sensor may be a device that is sensitive to mechanical vibrations but not airborne vibrations (i.e., the response capability of the vibration sensor to mechanical vibrations exceeds the response capability of the vibration sensor to airborne vibrations). Because the bone conduction microphone can directly pick up vibration signals from the speaking area, the bone conduction microphone can reduce the influence of environmental noise.

[0045] An air conduction microphone collects air vibration signals caused when a user speaks, and converts the air vibration signals into electrical signals. An air conduction microphone can be either a single air conduction microphone or a microphone array consisting of two or more air conduction microphones. A microphone array can be a beamforming microphone array or other similar microphone array. A microphone array can collect sounds from different directions and positions in a space.

[0046] The audio signal output by the bone conduction microphone can effectively reduce the influence of noise at low frequencies. Therefore, the audio quality of the audio signal output by the bone conduction microphone at low frequencies is higher than that of the audio signal output by the air conduction microphone at low frequencies. In the high frequency band, the audio quality of the audio signal output by the bone conduction microphone is inferior to that of the audio signal output by the air conduction microphone. In contrast, the audio signal output by the air conduction microphone is relatively stable in all frequency bands.

[0047] The first microphone 242 can output a first audio signal, and the second microphone 244 can output a second audio signal, and the audio signals waiting to be processed can include the first audio signal and the second audio signal.

[0048] The audio generation method presented in this specification can generate a target audio by combining the first audio signal and the second audio signal. The first audio signal may be either an audio signal directly output by the first microphone 242 or an audio signal obtained by applying noise reduction processing to the audio signal directly output by the first microphone 242. The second audio signal may be either an audio signal directly output by the second microphone 244 or an audio signal obtained by applying noise reduction processing to the audio signal directly output by the second microphone 244. It should be noted that if the first audio signal is an audio signal directly output by the first microphone 242, the second audio signal is an audio signal directly output by the second microphone 244. If the first audio signal is an audio signal directly output by the first microphone 242 and applying noise reduction processing to the audio signal directly output by the second microphone 244, the second audio signal is an audio signal obtained by applying noise reduction processing to the audio signal directly output by the second microphone 244. The noise reduction processing methods for the first audio signal and the second audio signal may be the same or different.

[0049] When the number of first-class microphones 242 is plural, the first audio signal is an audio signal obtained by combining the audio signals of the single microphones output from the plurality of first-class microphones 242. When the number of second-class microphones 244 is plural, the second audio signal is an audio signal obtained by combining the audio signals of the single microphones output from the plurality of second-class microphones 244.

[0050] For example, when there is one first-class microphone 242 and one second-class microphone 244, the first audio signal may be the audio signal directly output by the first-class microphone 242, and the second audio signal may be the audio signal directly output by the second-class microphone 244. Alternatively, the first audio signal may be the audio signal directly output by the first-class microphone 242 that has been subjected to noise reduction processing, and the second audio signal is the audio signal directly output by the second-class microphone 244 that has been subjected to noise reduction processing.

[0051] For example, when there is one first-class microphone 242 and multiple second-class microphones 244, the first audio signal may be the audio signal directly output by the first-class microphone 242, and the second audio signal may be an audio signal obtained by fusing the audio signals directly output by the multiple microphones among the second-class microphones 244 after performing noise reduction processing on the audio signals directly output by a single microphone. Alternatively, the first audio signal may be an audio signal obtained by performing noise reduction processing on the audio signal directly output by the first-class microphone 242, and the second audio signal may be an audio signal obtained by performing noise reduction processing on the audio signals directly output by the multiple microphones among the second-class microphones 244 after performing noise reduction processing on the audio signals directly output by a single microphone, fusing the signals, and then performing noise reduction processing on the multiple microphones. The noise reduction processing algorithm may be a conventional speech noise reduction algorithm, such as one or any combination of spectral subtraction, Wiener filter, MMSE algorithm, and MMSE-based improvement algorithm.

[0052] In particular, in the case of the second-class microphone 244 consisting of a plurality of air conduction microphones, the audio quality is significantly improved by applying noise reduction processing to the audio signal directly output by the second-class microphone 244. Therefore, if the audio signal obtained after applying noise reduction processing to the audio signal directly output by the second-class microphone 244 is used as the second audio signal, the efficiency of audio generation increases, the audio quality of the target audio increases, and the amount of calculation and calculation cost can be reduced.

[0053] The system 100 can further perform noise reduction processing on the target audio to improve the sound quality of the target audio. The system 100 may either perform noise reduction processing on the first audio signal and the second audio signal first, then synthesize the audio to generate the target audio, or perform noise reduction processing on the first audio signal and the second audio signal first, then generate the target audio.

[0054] FIG. 2 shows a flow diagram of an audio generation method P100 according to an embodiment of the present disclosure. The method P100 combines the first audio signal and the second audio signal to obtain an audio signal with higher sound quality. Specifically, the method P100 selects and combines the audio signal with the higher sound quality based on the sound quality in the frequency domain of the first audio signal and the second audio signal to obtain a target audio signal. As shown in FIG. 2, the method P100 may include the following:

[0055] S120: The electronic device 200 acquires a first audio signal and a second audio signal.

[0056] As described above, the first audio signal and the second audio signal are different audio signals. The first audio signal and the second audio signal have different characteristics. The first audio signal and the second audio signal have different sound qualities in the frequency domain. If the first audio signal is an audio signal output by a bone conduction microphone and the second audio signal is an audio signal output by an air conduction microphone, the first audio signal has relatively high sound quality in the low frequency band, and the sound quality of the second audio signal in the high frequency band is higher than the sound quality of the first audio signal in the low frequency band. Of course, the first audio signal and the second audio signal may be other types of audio signals, such as an audio signal output by an optical microphone or an audio signal output by a microphone that receives surface electromyography signals.

[0057] S140: The electronic device 200 generates the target audio based on the first audio signal and the second audio signal. Specifically, step S140 may include the following steps:

[0058] S142: The electronic device 200 determines and compares a first evaluation index in the frequency domain of the first audio signal and a second evaluation index in the frequency domain of the second audio signal.

[0059] When synthesizing the first audio signal and the second audio signal, the audio signal with better audio quality can be selected and spliced ​​together by comparing the audio qualities of the first audio signal and the second audio signal. Specifically, the electronic device 200 can indicate the audio quality of the audio signals waiting to be processed through an evaluation index. The first evaluation index can indicate the audio quality of the first audio signal, and the first evaluation index and the audio quality of the first audio signal are positively correlated, and the second evaluation index can indicate the audio quality of the second audio signal, and the second evaluation index and the audio quality of the second audio signal are positively correlated.

[0060] The voice quality of the waiting audio signal can be evaluated based on the signal strength of an effective audio signal included in the waiting audio signal. The effective audio signal may be an important audio signal accompanying the audio signal. The noise signal may be an audio signal other than the effective audio signal. For example, when making a voice call, the effective audio signal may be a voice signal of a call user speaking, and the noise signal may be environmental noise such as the sound of a car or a horn. When collecting a special sound, for example, when collecting the sound of birds chirping, the effective audio signal may be an audio signal of birds chirping, and the noise signal may be the sound of wind or water. For convenience, the following description will be given taking a voice call as an example. The effective audio signal may be a voice signal of a call user speaking, and the noise signal may be environmental noise. The voice quality of the waiting audio signal can be evaluated based on the signal strength of an effective audio signal included in the waiting audio signal. For example, when the effective audio signal is a human voice signal, the higher the strength of the effective voice signal and the clarity of the effective voice signal, the higher the voice quality of the waiting audio signal.

[0061] Here, it should be noted that the noise signal and the effective audio signal are both signals obtained through an estimation algorithm, and the effective audio signal and the noise signal are not accurate. The noise signal can be estimated through a noise estimation algorithm. The effective audio signal is obtained by estimating the noise signal from the audio signal waiting to be first processed.

[0062] Specifically, the strength of the effective audio signal can be evaluated through the evaluation index. The evaluation index may be the SNR of the audio signal waiting to be processed. The first evaluation index may be a first SNR corresponding to the first audio signal, and the second evaluation index may be a second SNR corresponding to the second audio signal. The first SNR may be a ratio of the effective audio signal to the noise signal in the first audio signal. The second SNR may be a ratio of the effective audio signal to the noise signal in the second audio signal. A higher first SNR of the first audio signal indicates a higher ratio of the effective audio signal at the current frequency, which indicates a higher audio quality of the first audio signal. A higher second SNR of the second audio signal indicates a higher ratio of the effective audio signal at the current frequency, which indicates a higher audio quality of the second audio signal. If the first evaluation index is higher than the second evaluation index, the value of the first SNR is higher than the value of the second SNR.

[0063] Naturally, the audio quality of the waiting audio signal can be evaluated directly through the effective audio signal in the waiting audio signal. That is, the evaluation index may be the effective audio signal. If the first evaluation index is higher than the second evaluation index corresponding to the second audio signal, a numerical value indicating the intensity of the first effective audio signal corresponding to the first audio signal will be higher than a numerical value indicating the intensity of the second effective audio signal corresponding to the second audio signal. Naturally, the evaluation index may be a noise signal in the waiting audio signal. If the first evaluation index is higher than the second evaluation index corresponding to the second audio signal, a numerical value indicating the intensity of the first noise signal corresponding to the first audio signal will be lower than a numerical value indicating the intensity of the noise signal corresponding to the second audio signal. Naturally, the evaluation index may be the noise signal intensity in the waiting audio signal. For convenience, in the following description, the evaluation index is SNR, the first evaluation index is a first SNR corresponding to the first audio signal, and the second evaluation index is a second SNR corresponding to the second audio signal. Those skilled in the art should understand that all other parameters that can evaluate voice quality can be the first evaluation index and the second evaluation index.

[0064] The SNR is a parameter related to frequency. Different frequencies result in different SNRs for the audio signals. Specifically, in step S142, by determining a first evaluation index in the frequency domain of the first audio signal and an evaluation index in the frequency domain of the second audio signal, a first SNR corresponding to each frequency in the frequency domain of the first audio signal and a second SNR corresponding to each frequency in the frequency domain of the second audio signal can be determined.

[0065] To obtain the first evaluation index of the first audio signal and the second evaluation index of the second audio signal, the system 100 may first perform frame processing on the first audio signal and the second audio signal, respectively. A frame is a basic unit constituting an audio signal. In data processing of audio signals, frames are often used as basic units for calculations. The first audio signal and the second audio signal each include one or more audio frames. The audio frame includes an audio signal of a given time span. The audio signal in each audio frame is stable. Adjacent audio frames may partially overlap. The given time span may be 20 to 50 milliseconds, such as 20 milliseconds, 25 milliseconds, 30 milliseconds, 40 milliseconds, or 50 milliseconds. Of course, the given time span may be longer or shorter. The lengths of different audio frames may be the same or different.

[0066] Each audio frame is made up of overlapping signals of multiple frequencies. To obtain a first evaluation index corresponding to each frequency in the frequency domain of the first audio signal and an evaluation index corresponding to each frequency in the frequency domain of the second audio signal, the system 100 can perform a Fourier transform on the audio frame to obtain a signal distribution of each frequency in the audio frame. The signal distribution of each frequency may be the intensity of the audio signal corresponding to each frequency in the audio frame.

[0067] FIG. 3 shows frequency spectrum diagrams of a first audio signal and a second audio signal according to an embodiment of the present disclosure. FIG. 3 shows frequency spectrum diagrams corresponding to corresponding audio frames in the first audio signal and the second audio signal. The frequency spectrum diagram shows the relationship between frequency and audio signal strength in one audio frame. As shown in FIG. 3, the horizontal axis represents frequency and the vertical axis represents signal width. Curve 1 is the frequency spectrum diagram corresponding to the first audio signal, and Curve 2 is the frequency spectrum diagram corresponding to the second audio signal. FIG. 3 is merely a sample explanation, and those skilled in the art should understand that Curve 1 and Curve 2 corresponding to different audio frames may be different, Curve 1 and Curve 2 may be subject to dynamic fluctuations, and Curve 1 and Curve 2 may be any type of frequency spectrum curve.

[0068] 4 shows a graph of the first SNR and the second SNR according to an embodiment of the present disclosure. In FIG. 4, the vertical axis represents SNR and the horizontal axis represents frequency f. Curve 5 represents the first SNR curve corresponding to each frequency of the first audio signal. Curve 6 represents the second SNR curve corresponding to each frequency of the second audio signal.

[0069] 4, comparing curve 5 and curve 6, it can be seen that in the low frequency band, the first SNR of the first audio signal is higher than the second SNR of the second audio signal, and in the high frequency band, the first SNR of the first audio signal is lower than the second SNR of the second audio signal. That is, in the low frequency band, the sound quality of the first audio signal is higher than the sound quality of the second audio signal, and in the high frequency band, the sound quality of the first audio signal is lower than the sound quality of the second audio signal.

[0070] Different audio frames may have different first and second SNRs, and the first and second SNRs may vary dynamically. Similarly, the first and second evaluation indices may vary dynamically.

[0071] It should be noted that FIG. 4 is an example. Curves 5 and 6 in FIG. 4 are illustrated by taking the first audio signal as the output signal of a bone conduction microphone and the second audio signal as the output signal of an air conduction microphone as an example. The output signal of a bone conduction microphone has a relatively high SNR in the low frequency band and relatively good sound quality, while the SNR in the high frequency band is relatively low and relatively poor sound quality. On the other hand, the output signal of an air conduction microphone is relatively stable in each frequency band. Those skilled in the art should understand that when the first audio signal and the second audio signal are audio signals output by other types of microphones, the relative relationship between Curves 5 and 6 may be different. Those skilled in the art should also understand that all types of graphs of the first SNR and the second SNR are within the scope of protection of this specification.

[0072] Step S140 may further include the following:

[0073] S144: The electronic device 200 determines at least one target frequency based on the comparison result between at least the first evaluation index and the second evaluation index, thereby determining a first frequency interval 001 and a second frequency interval 002.

[0074] As described above, when synthesizing the first audio signal and the second audio signal, the method P100 can combine audio signals with higher sound quality corresponding to each frequency in the frequency domain. Therefore, the method P100 can compare the sound qualities of the first audio signal and the second audio signal in the frequency domain by comparing evaluation metrics in the frequency domain of the first audio signal and the second audio signal. Specifically, in step S144, the electronic device 200 can divide the frequency domain into the first frequency interval 001 and the second frequency interval 002 based on changes in sound quality of the first audio signal in the frequency domain and changes in sound quality of the second audio signal in the frequency domain, so that the sound quality of the first audio signal in the first frequency interval 001 is higher than the sound quality of the second audio signal, and the sound quality of the first audio signal in the second frequency interval 002 is lower than the sound quality of the second audio signal. The ranges of the first frequency interval 001 and the second frequency interval 002 can be dynamically adjusted based on a dynamic fluctuation of a first evaluation index in the frequency domain of the first audio signal and a dynamic fluctuation of a second evaluation index in the frequency domain of the second audio signal. The frequency domain includes the first frequency interval 001 and the second frequency interval 002. Each target frequency in the at least one target frequency is a frequency corresponding to a boundary between the first frequency interval 001 and the second frequency interval 002.

[0075] In some embodiments, the method P100 can divide frequencies in the frequency domain into the first frequency interval 001 and the second frequency interval 002 based on a relative result of comparing the first evaluation index of the first audio signal with the second evaluation index of the second audio signal. When the first evaluation index of the first audio signal exceeds the second evaluation index of the second audio signal, it indicates that the sound quality of the first audio signal is superior to that of the first audio signal, and the corresponding frequency interval when the first evaluation index exceeds the second evaluation index is divided into the first frequency interval 001. Frequencies outside the first frequency interval 001 are divided into the second frequency interval 002.

[0076] In another embodiment, the method P100 may divide the frequencies in the frequency domain into the first frequency range 001 and the second frequency range 002 based on a relative result of comparing the first evaluation index with the second evaluation index and a result of comparing the first evaluation index with an absolute threshold value of the evaluation index. When the first evaluation index exceeds the evaluation index, it does not necessarily indicate that the sound quality of the first audio signal is superior to the first audio signal. For example, when the SNR of the audio signal output by the bone conduction microphone is higher than the SNR of the audio signal output by the air conduction microphone and the SNR of the audio signal output by the bone conduction microphone is relatively low and below an SNR threshold, the sound quality of the audio signal output by the bone conduction microphone may be lower than the audio quality of the audio signal output by the air conduction microphone. Therefore, in some embodiments, particularly when the first audio signal is an audio signal output by a bone conduction microphone, the method P100 divides frequencies in the frequency domain into the first frequency range 001 and the second frequency range 002 based on a relative result of comparing the first evaluation index with the evaluation index and a result of comparing the first evaluation index with an absolute threshold value of the evaluation index, thereby improving the accuracy of band division and improving the sound quality of the target audio. As described above, the first evaluation index may be a first SNR, and the second evaluation index may be a second SNR. The absolute threshold value of the evaluation index may be an SNR threshold.

[0077] 5 shows a flow chart for determining a first frequency interval 001 and a second frequency interval 002 according to an embodiment of the present disclosure. In the graph shown in FIG. 5, the method P100 can divide frequencies in the frequency domain into the first frequency interval 001 and the second frequency interval 002 based on a relative result of comparing the first SNR and the second SNR. As shown in FIG. 5, step S144 can include the following:

[0078] S144-2: The electronic device 200 determines the at least one target frequency based on a corresponding frequency when the first SNR and the second SNR are equal.

[0079] S144-3: The electronic device 200 determines that the corresponding frequency interval when the first SNR exceeds the second SNR is the first frequency interval 001, and the frequency interval outside the first frequency interval 001 is the second frequency interval 002, using the at least one target frequency as a critical point.

[0080] FIG. 6 shows a graph of a first frequency interval 001 and a second frequency interval 002 presented based on an embodiment of the present specification. FIG. 6 shows a graph in which frequency intervals are divided based on FIG. 4. FIG. 6 corresponds to FIG. 5. As shown in FIG. 6, for convenience, we define the frequency corresponding to the intersection of curve 5 and curve 6 as a first target frequency f1. That is, the first target frequency f1 is the corresponding frequency when the first SNR and the second SNR are equal.

[0081] In some embodiments, each target frequency in the at least one target frequency may be a first target frequency f1. In other embodiments, each target frequency in the at least one target frequency may be any frequency within a frequency interval of a given width around the first target frequency f1. That is, any frequency within a frequency interval of a given width around the corresponding frequency when the first SNR and the second SNR are equal. The given width may be a preset frequency width.

[0082] The electronic device 200 may determine, using the at least one target frequency as a critical point, that the corresponding frequency interval when the first SNR exceeds the second SNR is the first frequency interval 001, and that the frequency interval outside the first frequency interval 001 is the second frequency interval 002. As shown in FIG. 6 , in a band below the first target frequency f1, the first SNR exceeds the second SNR. That is, the sound quality of the first audio signal exceeds that of the second audio signal. In a band above the first target frequency f1, the first SNR is lower than the second SNR. That is, the sound quality of the first audio signal is lower than that of the second audio signal. We define the band below the first target frequency f1 as the first frequency interval 001, and the band above the first target frequency f1 as the second frequency interval 002.

[0083] The first frequency interval 001 may include at least one continuous frequency interval. The second frequency interval 002 may include at least one continuous frequency interval. Although only one first target frequency f1 is shown in FIG. 6, those skilled in the art should understand that there may be multiple first target frequencies f1 depending on the differences between the first audio signal and the second audio signal. If there are multiple first target frequencies f1, there will also be multiple corresponding target frequencies, and the first frequency interval 001 may include multiple continuous frequency intervals, and the second frequency interval 002 may also include multiple continuous frequency intervals.

[0084] 7 shows another flow diagram for determining a first frequency interval 001 and a second frequency interval 002 according to an embodiment of the present disclosure. In the graph of FIG. 7, the method P100 can divide frequencies in the frequency domain into the first frequency interval 001 and the second frequency interval 002 based on a relative result of comparing the first SNR with the second SNR and a result of comparing the first SNR with the SNR threshold. As shown in FIG. 7, step S144 can include the following:

[0085] S144-4: Obtain the SNR threshold.

[0086] S144-5: The electronic device 200 compares the first SNR with the second SNR, and determines a corresponding frequency when the first SNR is equal to the second SNR as at least one first target frequency f1.

[0087] S144-6: The electronic device 200 compares the first SNR with the SNR threshold, and determines a corresponding frequency when the first SNR is equal to the SNR threshold as at least one second target frequency f2.

[0088] S144-8: The electronic device 200 compares the first SNR, the second SNR and the SNR threshold corresponding to each frequency in the at least one first target frequency f1 and the at least one second target frequency f2, and sets the corresponding frequency when the first SNR is not lower than the second SNR and the SNR threshold as the at least one target frequency.

[0089] S144-9: The electronic device 200 determines the at least one target frequency as a critical point, and determines the corresponding frequency range when the first SNR exceeds all of the second SNRs as the first frequency range, and determines the frequency range outside the first frequency range as the second frequency range.

[0090] FIG. 8 shows another graph of a first frequency range and a second frequency range presented based on an embodiment of the present specification. FIG. 8 is a graph in which frequency ranges are divided based on FIG. 4. FIG. 8 corresponds to FIG. 7. As shown in FIG. 8, for convenience, we define SNR0 as the SNR threshold. A first target frequency f1 is the corresponding frequency when the first SNR and the second SNR are equal. That is, it is the frequency corresponding to the intersection of curve 5 and curve 6. A second target frequency f2 is the corresponding frequency when the first SNR and the SNR threshold SNR0 are equal. That is, it is the frequency corresponding to the intersection of curve 5 and the SNR threshold SNR0.

[0091] The SNR threshold SNR0 may be any value and may be pre-stored in at least one storage medium 230. The SNR threshold SNR0 may be manually set or changed. Alternatively, the SNR threshold SNR0 may be obtained through machine learning, for example, the SNR threshold SNR0 may be 3 dB or 6 dB or other values, etc. Depending on the type of the first audio signal, the SNR threshold SNR0 may be different.

[0092] The electronic device 200 may compare the first SNR, the second SNR, and the SNR threshold SNR0 corresponding to each frequency of the at least one first target frequency f1 and the at least one second target frequency f2, and determine the corresponding frequency when the first SNR is not lower than the second SNR and the SNR threshold SNR0 as the at least one target frequency. Taking FIG. 8 as an example, FIG. 8 shows one first target frequency f1 and one second target frequency f2. The electronic device 200 may compare the first SNR, the second SNR, and the SNR threshold SNR0 corresponding to the first target frequency f1 and one second target frequency f2, respectively. The first SNR corresponding to the first target frequency f1 is equal to the second SNR corresponding to the first target frequency f1 but is lower than the SNR threshold SNR0. The first SNR corresponding to the second target frequency f2 is higher than the second SNR corresponding to the second target frequency f2 and equal to the SNR threshold SNR0. Therefore, we define the second target frequency f2 as the target frequency. In a band below the second target frequency f2, the first SNR exceeds the second SNR and also exceeds the SNR threshold SNR0, which proves that the sound quality of the first audio signal is better than that of the second audio signal. At this time, we define the frequency band corresponding to the frequency band below the second target frequency f2 as a first frequency band 001, and define the band above the second target frequency f2 as a second frequency band 002.

[0093] The first frequency interval 001 may include at least one continuous frequency interval. The second frequency interval 002 may include at least one continuous frequency interval. Although FIG. 8 shows only one first target frequency f1 and one second target frequency f2, those skilled in the art should understand that there may be multiple first target frequencies f1 and second target frequencies f2, and multiple corresponding target frequencies, depending on the difference between the first audio signal and the second audio signal. When there are multiple target frequencies, the first frequency interval 001 may include multiple continuous frequency intervals, and the second frequency interval 002 may also include multiple continuous frequency intervals.

[0094] As shown in Figures 4 to 8, the first SNR and the second SNR may fluctuate within a small range. That is, the first SNR and the second SNR corresponding to multiple frequencies may be equal within a small range. The width of the frequency range may be preset so that the accuracy of audio generation is not affected by the fluctuations in SNR. If the gap between the multiple frequencies is within the width of the frequency range, the target frequency may be any one of the multiple frequencies, the frequency with the largest corresponding first SNR among the multiple frequencies, or the average value of the multiple frequencies.

[0095] Step S140 may further include the following:

[0096] S146: The electronic device 200 generates the target audio based on the first frequency interval 001 and the second frequency interval 002, and the first audio signal and the second audio signal.

[0097] Specifically, in step S146, the electronic device 200 combines the audio signal in the first frequency interval 001 of the first audio signal with the audio signal in the second frequency interval 002 of the second audio signal to obtain the target audio. Specifically, in the frequency domain, the audio signal in the first frequency interval 001 of the target audio includes the audio signal in the first frequency interval of the first audio signal, and the audio signal in the second frequency interval 002 of the target audio signal includes the audio signal in the second frequency interval of the second audio signal.

[0098] In some embodiments, the intensities of the first audio signal and the second audio signal at the target frequency portion may be different. When the audio signal in the first frequency range 001 of the first audio signal and the audio signal in the second frequency range 002 of the second audio signal are spliced ​​together, the signal at the target frequency portion may become discontinuous. To avoid the signal discontinuity, step S146 may include the following:

[0099] S146-2: The electronic device 200 performs a smoothing process on the first audio signal and the second audio signal corresponding to frequencies within a given range of each target frequency in the at least one target frequency, thereby smoothly transitioning between the audio signal of the first audio signal corresponding to frequencies within the given range and the audio signal of the second audio signal corresponding to frequencies within the given range.

[0100] S146-4: The electronic device 200 combines the audio signal in the first frequency range 001 of the first audio signal that has been smoothed and the audio signal in the second frequency range 002 of the second audio signal based on their frequency distributions to obtain the target audio.

[0101] The given range may include a frequency interval of a given width that includes the target frequency. The smoothing process may be an amplification process, which amplifies the audio signal within the given range through an amplification factor.

[0102] FIG. 9 shows a graph of a target audio provided according to an embodiment of the present disclosure. FIG. 10 shows a graph of another target audio provided according to an embodiment of the present disclosure. FIG. 9 corresponds to FIG. 6, and the target frequency of the target audio shown in FIG. 9 is the first target frequency f1. FIG. 10 corresponds to FIG. 8, and the target frequency of the target audio shown in FIG. 10 is the second target frequency f2.

[0103] To summarize, the method P100 and system 100 can compare the sound qualities of the first audio signal and the second audio signal in the frequency domain based on the evaluation indexes of the first audio signal and the second audio signal. A frequency range where the sound quality of the first audio signal exceeds that of the second audio signal is defined as a first frequency range 001, and a frequency range where the sound quality of the first audio signal falls below that of the second audio signal is defined as a second frequency range 002. The audio signal located in the first frequency range 001 of the first audio signal and the audio signal located in the second frequency range 002 of the second audio signal are combined to obtain the target audio, thereby improving the audio generation effect and the sound quality of the target audio. The method P100 and system 100 dynamically select a target frequency based on the sound qualities of the first audio signal and the second audio signal, and distinguish between the first frequency range 001 and the second frequency range 002 based on the dynamics of the target frequency, ensuring that the method P100 and system 100 can be applied to any scenario. That is, the method P100 and the system 100 can maximize the quality of the target audio in all frequency ranges in all scenes.

[0104] This document also presents a non-transitory storage medium, the memory of which contains at least one set of executable instructions for audio generation, which, when executed by a processor, direct the processor to perform the steps of the audio generation method P100 described herein. In some embodiments, aspects of this document can be embodied in the form of a program product, which includes program code. When the program product is run on an electronic device 200, the program code causes the electronic device 200 to perform the steps of the audio generation described herein. A program product used to implement the method can be a compact disc read-only memory (CD-ROM) containing program code and run on the electronic device 200. However, the program product of this document is not limited thereto, and in this document, a readable storage medium can be any tangible medium for an embedded or stored program, which can be used by or in conjunction with an execution command system (e.g., processor 220). The program product can employ any combination of one or more readable media. The computer-readable medium may be either a signal-readable medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or any combination thereof. More specific examples of computer-readable storage media include an electrical connection having one or more conductors, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), fiber optics, a compact disc read-only memory (CD-ROM), an optical memory device, a magnetic memory device, or any suitable combination of the above. The computer-readable storage medium may include a propagating data signal, in baseband or as part of a carrier wave, carrying the computer-readable program code.Such propagated data signals may take several forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may be any readable medium other than a readable storage medium, which may transmit, propagate, or transmit a program for use by or in conjunction with an execution command system, apparatus, or device. The program code embodied in the readable storage medium may be transmitted using any suitable medium, including, but not limited to, wireless, wired, optical cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of this description may be written in any combination of one or more programming languages, including target-specific programming languages ​​such as Java, C++, etc., as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may execute entirely on the electronic device 200, partially on the electronic device 200, as a separate software package, partially on the electronic device 200, partially on a remote computing device, or entirely on a remote computing device.

[0105] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in the order of the respective embodiments and still achieve the desired results. Also, even if the processes depicted in the figures require a particular order or sequential order, the desired results will not necessarily be achieved. In some implementations, both multitasking and parallel processing may be advantageous or effective.

[0106] In summary, after reading the detailed disclosure, those skilled in the art will understand that the detailed disclosure above is provided for illustrative purposes only and is not limiting. Although not explicitly stated herein, those skilled in the art will understand that the requirements of this description encompass various reasonable modifications, improvements, or alterations to the embodiments. The spirit of these modifications, improvements, or alterations is presented in this description and falls within the spirit and scope of the exemplary embodiments shown in this description.

[0107] Furthermore, the technical terms used in this specification are used to describe the embodiments of this specification. For example, "one embodiment," "embodiment," and / or "some embodiments" mean that a particular feature, structure, or characteristic described in the embodiment can be included in at least one embodiment of this specification. Therefore, it should be emphasized and understood that two or more references to "an embodiment," "one embodiment," or "alternative embodiments" in various parts of this specification do not necessarily refer to the same embodiment. Furthermore, particular features, structures, or characteristics can be combined as appropriate in one or more embodiments of this specification.

[0108] In the above description of the embodiments of this specification, it should be understood that in order to facilitate understanding of a feature and to simplify this specification, this specification combines various features in one embodiment, drawing, or description. However, this does not mean that the combination of these features is essential. When reading this specification, those skilled in the art may extract some features and understand them as an independent embodiment. In other words, the embodiments of this specification can be understood as a combination of multiple secondary embodiments. This also applies when the content of each secondary embodiment contains less than all the features of one of the above-mentioned disclosed embodiments.

[0109] Each patent, patent application, patent application publication, or other material, such as articles, books, manuals, publications, documents, articles, etc., cited herein may be linked by citation. All content for all purposes is the history of any related prosecuting documents, excluding the history of any related prosecuting documents that may be inconsistent with or in conflict with this document, or that may have a limiting effect on the broadest scope of any claim, now or hereafter related to this document. By way of example, in the event of any inconsistency or conflict between the terminology, explanations, definitions, and / or terms associated with any included material and / or the terminology, explanations, definitions, and / or terms associated with this file, the terminology in this document shall control.

[0110] Finally, it should be understood that the implementation plan of the application disclosed herein is merely an explanation of the principles of the implementation plan of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are provided as examples and are not limiting. Those skilled in the art can make alternative arrangements based on the embodiments in this specification to realize the application in this specification. Therefore, the embodiments in this specification are not limited to the embodiments exactly described in the application.

Claims

1. 1. A method of audio generation, comprising: obtaining a first audio signal and a second audio signal; and generating a target audio based on the first audio signal and the second audio signal; the frequency domain of the target audio signal includes a first frequency section in which a first evaluation index of the first audio signal is higher than a second evaluation index of the second audio signal, and a second frequency section in which the first evaluation index of the first audio signal is lower than the second evaluation index of the second audio signal; In the first frequency range, a target audio signal is generated based on a first audio signal corresponding to the first frequency range; In the second frequency range, a target audio signal is generated based on a second audio signal corresponding to the second frequency range; ranges of the first frequency range and the second frequency range are dynamically adjusted based on at least a dynamic variation of the first evaluation index in the frequency domain of the first audio signal and a dynamic variation of the second evaluation index in the frequency domain of the second audio signal; the first evaluation metric represents a speech quality of the first audio signal; the second evaluation metric represents a speech quality of the second audio signal; the first audio signal is an audio signal output by at least one first-class microphone including a bone conduction microphone, and the second audio signal is an audio signal output by at least one second-class microphone including an air conduction microphone; 1. A method of audio generation, comprising:

2. 2. The audio generation method of claim 1, the first evaluation index and the speech quality of the first audio signal are positively correlated, and the second evaluation index and the speech quality of the second audio signal are positively correlated; In the first frequency range, the sound quality of the first audio signal is higher than the sound quality of the second audio signal; 2. A method of generating audio, wherein in the second frequency range, the sound quality of the first audio signal is lower than the sound quality of the second audio signal.

3. 2. The audio generation method of claim 1, A method of generating audio, wherein the first evaluation index corresponding to each frequency within the first frequency range is higher than the second evaluation index.

4. 4. The audio generation method of claim 3, 10. A method of audio generation, wherein the first evaluation metric comprises a first SNR corresponding to the first audio signal, and the second evaluation metric comprises a second SNR corresponding to the second audio signal.

5. 5. The audio generation method of claim 4, and generating the target audio based on the first audio signal and the second audio signal, the method comprising: determining and comparing the first evaluation index and the second evaluation index in the frequency domain; and determining at least one target frequency based on a comparison result between the first evaluation index and the second evaluation index, thereby determining the first frequency range and the second frequency range, wherein each target frequency in the at least one target frequency range corresponds to a continuous portion of the first frequency range and the second frequency range; and generating the target audio based on the first frequency range and the second frequency range and the first audio signal and the second audio signal.

6. 6. The audio generation method of claim 5, 10. A method of audio generation, wherein the first frequency interval comprises at least one contiguous frequency interval and the second frequency interval comprises at least one contiguous frequency interval.

7. 6. The audio generation method of claim 5, and determining that a corresponding frequency range when the first SNR exceeds the second SNR, using the at least one target frequency as a critical point, is the first frequency range, and that a frequency range outside the first frequency range is the second frequency range, wherein the determining of the first frequency range and the second frequency range includes determining the at least one target frequency based on a corresponding frequency when the first SNR is equal to the second SNR, with the at least one target frequency as a critical point.

8. 8. The audio generation method of claim 7, A method for generating audio, wherein each target frequency in the at least one target frequency includes an arbitrary frequency within a frequency interval of a given width around a corresponding frequency when the first SNR and the second SNR are equal.

9. 6. The audio generation method of claim 5, determining the first frequency interval and the second frequency interval by obtaining an SNR threshold; comparing the first SNR with the second SNR, and determining a corresponding frequency when the first SNR is equal to the second SNR as at least one first target frequency; comparing the first SNR with the SNR threshold, and determining a corresponding frequency when the first SNR is equal to the SNR threshold as at least one second target frequency; comparing the first SNR and the second SNR corresponding to each of the at least one first target frequency and the at least one second target frequency with the SNR threshold, and determining a corresponding frequency when the first SNR is not lower than the second SNR and the SNR threshold as the at least one target frequency; and determining, with the at least one target frequency as a critical point, that a corresponding frequency interval when the first SNR is higher than the second SNR is the first frequency interval, and that a frequency interval outside the first frequency interval is the second frequency interval.

10. 6. The audio generation method of claim 5, generating the target audio based on the first frequency range and the second frequency range and the first audio signal and the second audio signal, the target audio including: performing a smoothing process on the first audio signal and the second audio signal corresponding to frequencies within a given range of each of the at least one target frequency, thereby creating a smooth transition between the audio signal corresponding to the frequency within the given range among the first audio signal and the audio signal corresponding to the frequency within the given range among the second audio signal; and acquiring the target audio by connecting the audio signal in the first frequency range among the first audio signal that has been subjected to the smoothing process and the audio signal in the second frequency range among the second audio signal that has been subjected to the smoothing process based on a frequency distribution.

11. 2. The audio generation method of claim 1, 10. A method of audio generation, characterized in that said at least one first-class microphone collects body vibration signals and said at least one second-class microphone collects air vibration signals.

12. 2. The audio generation method of claim 1, 1. A method for generating audio, characterized in that the first audio signal comprises an audio signal directly output by the at least one first-class microphone, and the second audio signal comprises an audio signal directly output by the at least one second-class microphone.

13. 2. The audio generation method of claim 1, 1. A method for generating audio, wherein the first audio signal includes an audio signal obtained by performing noise reduction processing on an audio signal directly output by the at least one first-class microphone, and the second audio signal includes an audio signal obtained by performing noise reduction processing on an audio signal directly output by the at least one second-class microphone.

14. In an audio generation system, 14. An audio generation system comprising: at least one storage medium having at least one instruction set for audio generation; and at least one processor in communication with said at least one storage medium, wherein, when executing said audio generation system, said at least one processor reads said at least one instruction set and performs the audio generation method of any one of claims 1 to 13 based on instructions in said at least one instruction set.

Citation Information

Patent Citations

  • Earphone signal processing method, earphone signal processing system and earphone

    CN111131947A

  • Telephone transmitter

    JP1996223677A

  • Handset

    JP2000261534A

  • Voice correction device, voice correction program, and voice correction method

    JP2014239346A

  • NOISE SIGNAL DETERMINATION METHOD AND APPARATUS AND AUDIO NOISE REMOVAL METHOD AND APPARATUS

    JP2018534618A