Audio signal restoration methods, apparatus, equipment, storage media, and computer program products

By combining bone conduction microphones, a first air conduction microphone, and a second air conduction microphone, and utilizing low-frequency and full-frequency repair network models, low-frequency and mid-to-high-frequency signals are generated and fused, solving the problem of missing mid-to-high frequencies in the audio signal acquired by headphones, thereby improving audio signal quality and user experience.

CN118214970BActive Publication Date: 2026-01-30HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211622597.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-16
Publication Date
2026-01-30
Estimated Expiration
2042-12-16

AI Technical Summary

Technical Problem

During the process of acquiring audio signals, the headphones may experience weak or missing mid-to-high frequency signals, which affects the audio signal quality and leads to a poor user experience.

Method used

By combining a bone conduction microphone, a first air conduction microphone, and at least one second air conduction microphone, and utilizing low-frequency and acoustic features, a low-frequency repair network model and a full-frequency repair network model are used to generate and fuse low-frequency and mid-to-high-frequency signals, thereby improving the audio signal quality.

Benefits of technology

It restored mid-to-high frequency signals, improved the quality and clarity of audio signals, and enhanced the user's call experience and video recording quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118214970B_ABST
    Figure CN118214970B_ABST
Patent Text Reader

Abstract

This application discloses an audio signal restoration method, apparatus, device, storage medium, and computer program, belonging to the field of audio processing technology. Applied to headphones, the method includes: determining low-frequency characteristics of a bone conduction audio signal acquired by a bone conduction microphone; determining a low-frequency restoration signal based on the low-frequency characteristics and a first air conduction audio signal acquired by a first air conduction microphone; determining acoustic characteristics based on the low-frequency restoration signal and the low-frequency characteristics; and determining a target audio signal based on the low-frequency restoration signal, a second air conduction audio signal acquired by at least one second air conduction microphone, and the acoustic characteristics. The embodiments of this application improve the quality and clarity of the target audio signal by simultaneously combining the bone conduction audio signal acquired by the bone conduction microphone, the first air conduction audio signal acquired by the first air conduction microphone, and the second air conduction audio signal acquired by at least one second air conduction microphone.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio processing, and in particular to an audio signal repairing method and device, equipment, storage medium and computer program. BACKGROUND

[0002] With the popularity of true wireless stereo (TWS) earphones, more and more users use earphones for calls, live broadcasts or video recordings. Therefore, the quality of the audio signal collected by the earphone becomes one of the key factors affecting the user experience. In the process of collecting the audio signal, the energy of the medium and high frequencies is weak or even missing due to the microphone, the environment in which the user is located, the wearing posture and other reasons, thereby affecting the quality of the audio signal. Therefore, there is an urgent need for an audio signal repairing method to restore the damaged or missing medium and high frequency signals, improve the sound quality of the audio signal, and improve the user's earphone experience. SUMMARY

[0003] The present application provides an audio signal repairing method, device, equipment, storage medium and computer program, which can improve the quality and clarity of the target audio signal. The technical solution is as follows:

[0004] In a first aspect, an audio signal repairing method is provided, applied to an earphone, the earphone including a bone conduction microphone, a first air conduction microphone and at least one second air conduction microphone, the first air conduction microphone being used to collect an air conduction signal inside an ear canal, and the at least one second air conduction microphone being used to collect an air conduction signal of an external environment. In the method, a low frequency feature of a bone conduction audio signal collected by the bone conduction microphone is determined. A low frequency repairing signal is determined based on the low frequency feature and a first air conduction audio signal collected by the first air conduction microphone. An acoustic feature is determined based on the low frequency repairing signal and the low frequency feature. A target audio signal is determined based on the low frequency repairing signal, a second air conduction audio signal collected by the at least one second air conduction microphone, and the acoustic feature, the target audio signal including a low frequency signal and a medium and high frequency signal.

[0005] Since the bone conduction microphone collects the bone conduction audio signal through bone vibration, it can shield environmental noise to some extent, so that the low-frequency characteristics of the bone conduction audio signal can be more accurately extracted from the bone conduction audio signal. The first air conduction audio signal has a high signal-to-noise ratio, so the low-frequency signal repair effect can be improved by using the low-frequency characteristics and the first air conduction audio signal. Moreover, since the acoustic characteristics generally include the characteristics of the audio signal in each frequency range, and the second air conduction audio signal collected by the at least one second air conduction microphone includes a mid-high frequency signal. Therefore, by combining the second air conduction audio signal collected by the at least one second air conduction microphone and the acoustic characteristics, the full-frequency signal repair can be better guided, thereby improving the quality and clarity of the target audio signal. That is, the combination of the bone conduction audio signal collected by the bone conduction microphone, the first air conduction audio signal collected by the first air conduction microphone, and the second air conduction audio signal collected by the at least one second air conduction microphone can improve the quality and clarity of the target audio signal.

[0006] Due to the propagation medium and device, the bone conduction audio signal lacks a mid-high frequency signal. Therefore, in order to better fuse the bone conduction audio signal with the second air conduction audio signal later, the low-frequency signal in the bone conduction audio signal is taken as a reference, and a harmonic is generated according to a related algorithm to make the bone conduction audio signal include both low-frequency signals and mid-high frequency signals.

[0007] There are various ways to determine the low-frequency repair signal based on the low-frequency characteristics and the first air conduction audio signal. Next, two of the ways will be introduced.

[0008] In the first way, the low-frequency characteristics and the first air conduction audio signal are taken as inputs of a low-frequency repair network model to obtain a low-frequency repair signal output by the low-frequency repair network model.

[0009] Since the first air conduction microphone is blocked by the earphone and the pinna, the first air conduction audio signal lacks a mid-high frequency signal. Therefore, in order to better fuse the first air conduction audio signal with the second air conduction audio signal later, the low-frequency signal in the first air conduction audio signal is taken as a reference, and a harmonic is generated according to a related algorithm to make the first air conduction audio signal include both low-frequency signals and mid-high frequency signals.

[0010] In the second way, the low-frequency repair signal is determined based on the low-frequency characteristics, the bone conduction audio signal, the first air conduction audio signal, and the second air conduction audio signal collected by some or all of the at least one second air conduction microphone.

[0011] That is, the bone conduction audio signal, the first air conduction audio signal, the second air conduction audio signal collected by part or all of the at least one second air conduction microphone, and the low frequency feature are used to determine the low frequency repair signal, so as to further improve the repair effect of the low frequency signal.

[0012] Optionally, the bone conduction audio signal, the first air conduction audio signal, and the second air conduction audio signal collected by part or all of the at least one second air conduction microphone are fused to obtain a first fused signal. Then, the low frequency feature and the first fused signal are used as inputs of a low frequency repair network model to obtain a low frequency repair signal output by the low frequency repair network model.

[0013] Optionally, before determining the low frequency repair signal based on the low frequency feature, the bone conduction audio signal, the first air conduction audio signal, and the second air conduction audio signal collected by part or all of the at least one second air conduction microphone, the bone conduction audio signal, the first air conduction audio signal, and the second air conduction audio signal collected by part or all of the at least one second air conduction microphone can also be low-pass filtered respectively, so as to further improve the repair effect of the low frequency signal. That is, the middle and high frequency signals included in the bone conduction audio signal, the first air conduction audio signal, and the second air conduction audio signal collected by part or all of the at least one second air conduction microphone are blocked or weakened by low-pass filtering, so as to obtain low frequency signals included in each of the signals. Then, the low frequency signals included in the signals are fused to obtain a fused signal, and the low frequency repair signal is determined more accurately based on the low frequency feature and the fused signal.

[0014] The first fusion coefficient and the second fusion coefficient are determined, the first fusion coefficient being a fusion coefficient of the low frequency repair signal, and the second fusion coefficient including a fusion coefficient of the second air conduction audio signal collected by the at least one second air conduction microphone. The target audio signal is determined based on the low frequency repair signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature.

[0015] The first fusion coefficient and the second fusion coefficient are determined, the first fusion coefficient being a fusion coefficient of the low frequency repair signal, and the second fusion coefficient including a fusion coefficient of the second air conduction audio signal collected by the at least one second air conduction microphone. The target audio signal is determined based on the low frequency repair signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature.

[0016] The first fusion coefficient and the second fusion coefficient are determined, the first fusion coefficient being a fusion coefficient of the low frequency repair signal, and the second fusion coefficient including a fusion coefficient of the second air conduction audio signal collected by the at least one second air conduction microphone. The target audio signal is determined based on the low frequency repair signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature.

[0017] The first fusion coefficient and the second fusion coefficient are determined, the first fusion coefficient being a fusion coefficient of the low frequency repair signal, and the second fusion coefficient including a fusion coefficient of the second air conduction audio signal collected by the at least one second air conduction microphone. The target audio signal is determined based on the low frequency repair signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature.

[0018] The computer device determines a target scene currently located according to a related algorithm and displays a first user interface, and the first user interface includes an identification of the target scene. When the computer device detects a confirmation operation of the user, based on the identification of the target scene, the first fusion coefficient and the second fusion coefficient corresponding to the target scene are obtained from the corresponding relationship between the stored scene identification and the fusion coefficient. When the computer device detects a cancel operation of the user, a second user interface is displayed, and the second user interface includes a plurality of scene identifications. The user selects a target scene identification from the plurality of scene identifications, and when the computer device detects a confirmation operation of the user, the scene identification selected by the user is taken as the target scene identification, and then based on the identification of the target scene, the first fusion coefficient and the second fusion coefficient corresponding to the target scene are obtained from the corresponding relationship between the stored scene identification and the fusion coefficient.

[0019] The above is an example of the computer device setting the corresponding relationship between the scene identification and the fusion coefficient in advance. Of course, in actual application, the user can also adjust the first fusion coefficient and the second fusion coefficient involved in the repair of the audio signal in the target scene in real time. For example, when the computer device detects an adjustment operation of the user, the computer device displays a third user interface, and the third user interface includes an adjustment bar corresponding to the fusion coefficient. The user can adjust the size of the first fusion coefficient and the second fusion coefficient by sliding the adjustment bar up and down. When the computer device detects a confirmation operation of the user, the first fusion coefficient and the second fusion coefficient adjusted by the user are determined as the first fusion coefficient and the second fusion coefficient corresponding to the target scene.

[0020] It should be noted that whether the user adjusts the fusion coefficient in real time for the target scene or the computer device sets the corresponding relationship between the scene identification and the fusion coefficient in advance, the first fusion coefficient and the second fusion coefficient are personalized settings of the user. That is, for the same target scene, different users can adaptively adjust the fusion coefficient according to their own needs, so as to subsequently perform personalized and differentiated audio signal repair according to the fusion coefficient adjusted by the user. Moreover, different fusion coefficients can also be set for different scenes such as conference rooms and outdoor sports.

[0021] The third way is to perform environment detection on the target scene currently located to obtain an environment detection result. Based on the environment detection result, the first fusion coefficient and the second fusion coefficient are determined.

[0022] In the third way, the fusion coefficient is determined based on the environment detection result by performing environment detection on the target scene currently located, so that the target audio signal can be better adapted to the current environment.

[0023] Optionally, the low-frequency repaired signal, the second air channel audio signal collected by the at least one second air channel microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature are input into the full-frequency repair network model as inputs of the full-frequency repair network model to obtain the target audio signal output by the full-frequency repair network model.

[0024] It should be noted that directly inputting the low-frequency repaired signal, the second air channel audio signal collected by the at least one second air channel microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature into the full-frequency repair network model to determine the target audio signal is only an example. Alternatively, the target audio signal can also be determined in other manners. For example, based on the first fusion coefficient and the second fusion coefficient, the low-frequency repaired signal and the second air channel audio signal collected by the at least one second air channel microphone are fused to obtain a second fusion signal. Then, the second fusion signal and the acoustic feature are input into the full-frequency repair network model as inputs of the full-frequency repair network model to obtain the target audio signal output by the full-frequency repair network model. In this way, the calculation amount of the full-frequency repair network model can be reduced, thereby improving the efficiency of audio signal repair.

[0025] In a second aspect, an audio signal repair apparatus is provided, which has functions to implement behaviors of the audio signal repair method in the first aspect. The audio signal repair apparatus includes at least one module for implementing the audio signal repair method provided in the first aspect.

[0026] In a third aspect, a computer device is provided, which includes a processor and a memory for storing a computer program for executing the audio signal repair method provided in the first aspect. The processor is configured to execute the computer program stored in the memory to implement the audio signal repair method in the first aspect.

[0027] Optionally, the computer device can further include a communication bus for establishing a connection between the processor and the memory.

[0028] In a fourth aspect, a computer readable storage medium is provided, which stores instructions therein, and when the instructions are run on a computer, the computer is caused to perform the steps of the audio signal repair method in the first aspect.

[0029] In a fifth aspect, a computer program product including instructions is provided, and when the instructions are run on a computer, the computer is caused to perform the steps of the audio signal repair method in the first aspect. Alternatively, a computer program is provided, and when the computer program is run on a computer, the computer is caused to perform the steps of the audio signal repair method in the first aspect.

[0030] The technical effects obtained by the second aspect to the fifth aspect are similar to the technical effects obtained by the corresponding technical means in the first aspect, and will not be described here again. BRIEF DESCRIPTION OF DRAWINGS

[0031] Figure 1 is a structural schematic diagram of an earphone provided by an embodiment of the present application.

[0032] Figure 2 is a flowchart of an audio signal repair method provided by an embodiment of the present application.

[0033] Figure 3 is a schematic diagram of low-frequency repair network model training provided by an embodiment of the present application.

[0034] Figure 4 is a structural schematic diagram of a low-frequency repair network model provided by an embodiment of the present application.

[0035] Figure 5 is another schematic diagram of low-frequency repair network model training provided by an embodiment of the present application.

[0036] Figure 6 is a schematic diagram of obtaining a target audio signal by a full-frequency repair network model provided by an embodiment of the present application.

[0037] Figure 7 is a schematic diagram of an audio signal repair process provided by an embodiment of the present application.

[0038] Figure 8 is a structural schematic diagram of an audio signal repair device provided by an embodiment of the present application.

[0039] Figure 9 is a structural schematic diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0040] To make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0041] Before the audio signal repair method provided by the embodiments of the present application is explained in detail, the terms and business scenarios involved in the embodiments of the present application are introduced.

[0042] To facilitate understanding, the terms involved in the embodiments of the present application are first explained.

[0043] For example, please refer to Figure 1 , Figure 1 is a structural schematic diagram of an earphone provided by an embodiment of the present application. As shown in Figure 1As shown in the middle left view, the earphone 100 comprises a bone conduction microphone 101, a first air conduction microphone 102, and at least one second air conduction microphone 103. Figure 1 As shown in the middle left view, the earphone 100 comprises a bone conduction microphone 101, a first air conduction microphone 102, and at least one second air conduction microphone 103.

[0044] Bone conduction microphone: a microphone for collecting an audio signal propagating through the bone, which can be referred to as a bone conduction audio signal. That is, in the process of collecting the bone conduction audio signal by the bone conduction microphone, the audio emitted by the sound source is transmitted to the bone conduction microphone through the vibration of the bone, and after the bone conduction microphone receives the vibration signal, the vibration signal is converted into an electrical signal, thereby collecting the bone conduction audio signal.

[0045] Since the bone conduction microphone collects the bone conduction audio signal through bone vibration, it can shield environmental noise to a certain extent. However, due to the propagation medium and device, etc., the bone conduction audio signal usually lacks mid-high frequency signals. Generally, the frequency range of the bone conduction audio signal is a first frequency range.

[0046] First air conduction microphone: a microphone for collecting an audio signal propagating through the air, which can be referred to as a first air conduction audio signal. As shown in the middle left view, the first air conduction microphone is disposed inside the earphone, and after the earphone is worn on the human ear, the first air conduction microphone is located inside the human ear, so the first air conduction microphone can also be referred to as an intra-aural air conduction microphone. The frequency range of the first air conduction audio signal is a second frequency range. In the process of collecting the first air conduction audio signal by the first air conduction microphone, the first air conduction microphone is shielded by the earphone and the pinna to a certain extent, which can shield environmental noise, but also causes the first air conduction audio signal to lack mid-high frequency signals. Figure 1 Second air conduction microphone: a microphone for collecting an audio signal propagating through the air, which can be referred to as a second air conduction audio signal. As shown in the middle left view, the second air conduction microphone is disposed outside the earphone, and after the earphone is worn on the human ear, the second air conduction microphone is located outside the human ear, so the second air conduction microphone can also be referred to as an extra-aural air conduction microphone. The frequency range of the second air conduction audio signal is a third frequency range. Generally, at least one second air conduction microphone is disposed outside the earphone, and the second air conduction audio signal usually includes environmental noise.

[0047] Figure 1 It should be noted that the first frequency range is less than the second frequency range and the third frequency range. That is, the frequency range of the bone conduction audio signal is the lowest, the frequency range of the first air conduction audio signal is higher, and the frequency range of the second air conduction audio signal is the highest. Moreover, the sampling rate of the bone conduction microphone, the sampling rate of the first air conduction microphone, and the sampling rate of the second air conduction microphone are all the same.

[0048] It should be noted that the first frequency range is less than the second frequency range and the third frequency range. That is, the frequency range of the bone conduction audio signal is the lowest, the frequency range of the first air conduction audio signal is higher, and the frequency range of the second air conduction audio signal is the highest. Moreover, the sampling rate of the bone conduction microphone, the sampling rate of the first air conduction microphone, and the sampling rate of the second air conduction microphone are all the same. ​

[0049] Secondly, the business scenarios related to the embodiments of the present application are introduced.

[0050] The audio signal repairing method provided by the embodiments of the present application can be applied to various scenarios. For example, in the case that the audio signal collected by the earphone lacks mid-high frequency signals, resulting in that the audio signal is low-pitched and unnatural, and lacks recognition and emotional expression, the audio signal collected by the earphone is repaired according to the method provided by the embodiments of the present application, so as to restore the mid-high frequency signals in the audio signal. The repaired audio signal is closer to the audio signal emitted by the sound source, and can effectively improve the quality, recognition and emotional expression of the audio signal. In this way, not only the user's call experience can be improved, but also the wrong word rate and wrong character rate of audio recognition in video recording can be reduced, and the user's video editing efficiency can be improved.

[0051] For another example, due to the incorrect posture of the user wearing the earphone, the audio signal collected by one or more of the bone conduction microphone, the first air conduction microphone and the second air conduction microphone deployed on the earphone may be abnormal. At this time, the audio signal collected by the earphone is repaired according to the method provided by the embodiments of the present application, which can improve the abnormal signal.

[0052] The execution subject of the audio signal repairing method provided by the embodiments of the present application is a computer device. The computer device includes a low-frequency feature extraction module, a low-frequency repairing module, an acoustic feature extraction module, a fusion coefficient acquisition module and a full-frequency repairing module. The low-frequency feature extraction module is configured to extract features of the bone conduction audio signal by an artificial intelligence (AI) algorithm to obtain low-frequency features. The low-frequency repairing module is configured to determine a low-frequency repairing signal based on the low-frequency features and the first air conduction audio signal. The acoustic feature extraction module is configured to extract features by an AI algorithm in combination with the low-frequency repairing signal and the low-frequency features to obtain acoustic features. The fusion coefficient acquisition module is configured to determine a first fusion coefficient and a second fusion coefficient. The full-frequency repairing module is configured to fuse the bone conduction audio signal, the first air conduction audio signal and the second air conduction audio signal collected by the at least one second air conduction microphone by the first fusion coefficient and the second fusion coefficient to determine a target audio signal including low-frequency signals and mid-high frequency signals.

[0053] The computer device can be any kind of electronic product that can interact with the user through one or more ways such as keyboard, touchpad, touch screen, remote control, voice interaction or handwriting device, for example, personal computer (PC), mobile phone, smart phone, personal digital assistant (PDA), wearable device, pocket pc (PPC), tablet computer, smart screen, car audio, etc.

[0054] Those skilled in the art shall understand that the computer device described above is only an example, and other existing or future computer devices can also be applicable to the embodiments of the present application and shall be included in the protection scope of the embodiments of the present application, and are hereby included by reference.

[0055] It should be noted that the business scenarios described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0056] Figure 2 is a flowchart of an audio signal repairing method provided by an embodiment of the present application. Please refer to Figure 2 The method comprises the following steps.

[0057] Step 201: Based on the bone conduction audio signal collected by the bone conduction microphone, the low-frequency feature of the bone conduction audio signal is determined.

[0058] Based on the above description, the computer device comprises a low-frequency feature extraction module. The low-frequency feature extraction module extracts the low-frequency feature from the bone conduction audio signal through an AI algorithm. Of course, in actual application, the low-frequency feature can also be determined through other ways, such as cepstrum method or short-time fourier transfer (STFT), etc., which are not limited by the embodiments of the present application.

[0059] The low-frequency feature refers to the fundamental frequency and its multiple harmonic frequencies included in the audio signal.

[0060] Based on the above description, due to the reasons of propagation medium and devices, etc., the bone conduction audio signal lacks mid-high frequency signals. Therefore, in order to make the bone conduction audio signal better fused with the second air conduction audio signal subsequently, the low-frequency signal in the bone conduction audio signal is taken as a reference, and the bone conduction audio signal is generated in accordance with relevant algorithms, so that the bone conduction audio signal includes both low-frequency signals and mid-high frequency signals.

[0061] Step 202: Based on the low-frequency feature and the first air conduction audio signal collected by the first air conduction microphone, the low-frequency repairing signal is determined.

[0062] There are various ways to determine the low-frequency repairing signal based on the low-frequency feature and the first air conduction audio signal, and next two of them will be introduced respectively.

[0063] The first mode is to take the low-frequency feature and the first air-conducted audio signal as inputs of the low-frequency restoration network model to obtain a low-frequency restoration signal output by the low-frequency restoration network model.

[0064] The low-frequency restoration network model is obtained by training a to-be-trained low-frequency restoration network model based on a first sample data set. The first sample data set includes a plurality of groups of sample data, and each group of sample data includes a first sample air-conducted audio signal, a sample low-frequency feature, and an actual sample low-frequency signal.

[0065] For example, refer to Figure 3 , Figure 3 is a schematic diagram of training of a low-frequency restoration network model provided by an embodiment of the present application. In Figure 3 , the sample low-frequency feature and the first sample air-conducted audio signal are input into the to-be-trained low-frequency restoration network model to obtain a low-frequency restoration signal output by the to-be-trained low-frequency restoration network model. Based on the low-frequency restoration signal and the actual sample low-frequency signal, a first loss function is calculated according to a related algorithm. Then, the low-frequency restoration network model is trained by means of back propagation based on the first loss function.

[0066] The loss function is also called the cost function. The loss function is an iteration basis in the training of the network model. The loss function is used to evaluate the difference between the estimated value and the actual value output by the network model. Different network models generally use different loss functions. In the embodiment of the present application, the loss function used by the low-frequency restoration network model is the mean squared error (MSE) between the estimated value and the actual value. Of course, the low-frequency restoration network model can also use other loss functions, which are not limited in the embodiment of the present application.

[0067] For example, the structure of the low-frequency restoration network model is a neural network structure. Refer to Figure 4 , Figure 4 is a structural schematic diagram of a low-frequency restoration network model provided by an embodiment of the present application. In Figure 4 , the low-frequency restoration network model includes N restoration blocks, and N is an integer greater than or equal to 1. Each restoration block in the N restoration blocks includes at least one convolutional layer and at least one activation layer. It should be noted that Figure 4 the low-frequency restoration network model shown in the figure is only an example. In actual applications, the low-frequency restoration network model can also have other structures, the low-frequency restoration network model can also include other modules, and each restoration block can also include other functional layers, which are not limited in the embodiment of the present application.

[0068] Based on the above description, due to the shielding of the earphone and the auricle to the first air conduction microphone, the first air conduction audio signal lacks mid-high frequency signals. Therefore, in order to enable the first air conduction audio signal to be better fused with the second air conduction audio signal subsequently, in the process of determining the low frequency repair signal by using the low frequency repair network model, the low frequency signal in the first air conduction audio signal is taken as a reference, and the first air conduction audio signal is generated in accordance with a relevant algorithm, so that the first air conduction audio signal includes both low frequency signals and mid-high frequency signals.

[0069] In the second mode, the low frequency repair signal is determined based on the low frequency feature, the bone conduction audio signal, the first air conduction audio signal, and the second air conduction audio signals collected by some or all of the at least one second air conduction microphone.

[0070] That is, the low frequency repair signal is determined jointly based on the bone conduction audio signal, the first air conduction audio signal, the second air conduction audio signals collected by some or all of the at least one second air conduction microphone, and the low frequency feature, so as to further improve the repair effect of the low frequency signal.

[0071] In some embodiments, the bone conduction audio signal, the first air conduction audio signal, and the second air conduction audio signals collected by some or all of the at least one second air conduction microphone are fused to obtain a first fused signal. Then, the low frequency feature and the first fused signal are taken as inputs of the low frequency repair network model to obtain the low frequency repair signal output by the low frequency repair network model.

[0072] In the case of determining the low frequency repair signal by using the low frequency repair network model based on the low frequency feature and the plurality of signals, the low frequency repair network model is obtained by training a to-be-trained low frequency repair network model based on a second sample data set. The second sample data set includes a plurality of groups of sample data, and each group of sample data includes a first sample fused signal, a sample low frequency feature, and an actual sample low frequency signal.

[0073] For example, referring to Figure 5 , Figure 5 is another schematic diagram of training a low frequency repair network model provided by the embodiments of the present application. In Figure 5 , the sample low frequency feature and the first sample fused signal are input into the to-be-trained low frequency repair network model to obtain a low frequency repair signal output by the to-be-trained low frequency repair network model. Based on the low frequency repair signal and the actual sample low frequency signal, a second loss function is calculated in accordance with a relevant algorithm. Then, the low frequency repair network model is trained by means of back propagation based on the second loss function.

[0074] Optionally, before the low-frequency repair signal is determined based on the low-frequency feature, the bone conduction audio signal, the first air conduction audio signal, and the second air conduction audio signals collected by the at least one second air conduction microphone, low-pass filtering can be performed on the bone conduction audio signal, the first air conduction audio signal, and the second air conduction audio signals collected by the at least one second air conduction microphone, respectively, to further improve the repair effect of the low-frequency signal. That is, the low-pass filtering method is used to block and weaken the medium and high frequency signals included in the bone conduction audio signal, the first air conduction audio signal, and the second air conduction audio signals collected by the at least one second air conduction microphone, respectively, to obtain the low-frequency signals included in each of the signals. Then, the low-frequency signals included in the signals are fused to obtain a fusion signal, and the low-frequency repair signal is more accurately determined based on the low-frequency feature and the fusion signal.

[0075] Step 203: determining the acoustic feature based on the low-frequency repair signal and the low-frequency feature.

[0076] Based on the above description, the computer device includes an acoustic feature extraction module. The acoustic feature extraction module extracts features by combining the low-frequency repair signal and the low-frequency feature through an AI algorithm to obtain an acoustic feature, which includes the features of the audio signal in each frequency range.

[0077] For example, the low-frequency repair signal and the low-frequency feature are taken as inputs of the acoustic feature extraction module to obtain multiple features such as the fundamental frequency, the harmonic, the pitch, and the loudness output by the acoustic feature extraction module. Then, the multiple features are vectorized to obtain a target vector, which is used to indicate the acoustic feature. The target vector includes multiple elements that correspond to the multiple features one by one. Of course, in actual applications, the acoustic feature can also be determined by other methods, which are not limited by the embodiments of the present application.

[0078] Step 204: determining the target audio signal based on the low-frequency repair signal, the second air conduction audio signals collected by the at least one second air conduction microphone, and the acoustic feature, the target audio signal including the low-frequency signal and the medium and high frequency signal.

[0079] The first fusion coefficient and the second fusion coefficient are determined, the first fusion coefficient being the fusion coefficient of the low-frequency repair signal, and the second fusion coefficient including the fusion coefficient of the second air conduction audio signals collected by the at least one second air conduction microphone. The target audio signal is determined based on the low-frequency repair signal, the second air conduction audio signals collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature.

[0080] The manner of determining the first fusion coefficient and the second fusion coefficient includes multiple manners, and three manners will be described in the following.

[0081] In a first manner, the first fusion coefficient is determined based on an acoustic feature. The second fusion coefficient is determined based on the second air conduction audio signal collected by the at least one second air conduction microphone.

[0082] The computer device stores a corresponding relationship between acoustic features and fusion coefficients, so after the acoustic feature is determined according to the above step 103, the fusion coefficient corresponding to the acoustic feature is obtained from the stored corresponding relationship between acoustic features and fusion coefficients based on the acoustic feature, so as to obtain the first fusion coefficient.

[0083] Based on the above description, at least one second air conduction microphone is usually arranged outside the earphone, each of the at least one second air conduction microphone can collect a second air conduction audio signal to obtain at least one second air conduction audio signal, each of the at least one second air conduction audio signal corresponds to a second fusion coefficient. Since the process of determining the second fusion coefficient corresponding to each second air conduction audio signal is the same, in the following, any second air conduction audio signal is selected from the at least one second air conduction audio signal as a target air conduction audio signal, and the process of determining the second fusion coefficient corresponding to the target air conduction audio signal is introduced. The determination process of the second fusion coefficient of other air conduction audio signals in the at least one second air conduction audio signal can refer to the determination process of the second fusion coefficient corresponding to the target air conduction audio signal.

[0084] The target air conduction audio signal is pre-processed such as beamforming, gain control and double-microphone filtering to obtain multiple features such as signal-to-noise ratio and loudness, and then the multiple features are vectorized to obtain a feature vector, and the feature vector is used to indicate the features of the target air conduction audio signal. Then, based on the feature vector, the fusion coefficient corresponding to the feature vector is obtained from the stored corresponding relationship between the feature vector and the fusion coefficient, so as to obtain the second fusion coefficient of the target air conduction audio signal.

[0085] In a second manner, a target scene currently located is determined. Based on the target scene, the first fusion coefficient and the second fusion coefficient are determined.

[0086] The computer device determines a target scene according to a related algorithm and displays a first user interface, the first user interface including an identification of the target scene. When the computer device detects a confirmation operation of the user, the first fusion coefficient and the second fusion coefficient corresponding to the target scene are obtained from the stored corresponding relationship between the scene identification and the fusion coefficient based on the identification of the target scene. When the computer device detects a cancel operation of the user, a second user interface is displayed, the second user interface including a plurality of scene identifications. The user selects a target scene identification from the plurality of scene identifications, and when the computer device detects a confirmation operation of the user, the scene identification selected by the user is taken as the target scene identification, and then the first fusion coefficient and the second fusion coefficient corresponding to the target scene are obtained from the stored corresponding relationship between the scene identification and the fusion coefficient based on the identification of the target scene.

[0087] That is, for different target scenes, the computer device stores a corresponding relationship between the scene identification and the fusion coefficient in advance. In a case where the target scene determined by the computer device is the same as the scene actually located by the user, the first fusion coefficient and the second fusion coefficient corresponding to the target scene are obtained from the stored corresponding relationship between the scene identification and the fusion coefficient based on the identification of the target scene determined by the computer device. In a case where the target scene determined by the computer device is not the same as the scene actually located by the user, the first fusion coefficient and the second fusion coefficient corresponding to the target scene are obtained from the stored corresponding relationship between the scene identification and the fusion coefficient based on the identification of the target scene selected by the user.

[0088] The identification of the target scene is used to uniquely identify the target scene, and the identification of the target scene can be a number, a type, an area size, etc. of the target scene, or obtained by combining these information.

[0089] The confirmation operation of the user can be triggered by voice interaction, and can also be triggered by a click operation on a submit button in the first user interface or the second user interface.

[0090] The above is an example in which the computer device sets the corresponding relationship between the scene identification and the fusion coefficient in advance. Of course, in actual application, the user can also adjust the first fusion coefficient and the second fusion coefficient involved in repairing the audio signal in the target scene in real time. For example, when the computer device detects an adjustment operation of the user, the computer device displays a third user interface, the third user interface including an adjustment bar corresponding to the fusion coefficient. The user can adjust the size of the first fusion coefficient and the second fusion coefficient by sliding the adjustment bar up and down. When the computer device detects a confirmation operation of the user, the first fusion coefficient and the second fusion coefficient adjusted by the user are determined as the first fusion coefficient and the second fusion coefficient corresponding to the target scene.

[0091] The adjustment operation of the user is triggered in a voice interaction manner or triggered by a click operation on an adjustment button. For example, the user triggers the adjustment operation by voice inputting "adjust the fusion coefficient".

[0092] It should be noted that, whether the user adjusts the fusion coefficient in real time for the target scene or the computer device sets the correspondence between the scene identifier and the fusion coefficient in advance, the first fusion coefficient and the second fusion coefficient are both user personalized settings. That is, for the same target scene, different users can adaptively adjust the fusion coefficient according to their own needs, so as to subsequently perform personalized and differentiated audio signal repair according to the fusion coefficient adjusted by the user. Moreover, different fusion coefficients can also be set for different scenes such as conference rooms and outdoor sports.

[0093] In the third mode, the target scene currently present is detected to obtain an environment detection result. Based on the environment detection result, the first fusion coefficient and the second fusion coefficient are determined.

[0094] The computer device performs environment detection on the target scene currently present according to a related algorithm to obtain an environment detection result. Since the computer device stores a correspondence between environment detection results and fusion coefficients, based on the environment detection result, the fusion coefficient corresponding to the environment detection result is obtained from the stored correspondence between environment detection results and fusion coefficients to obtain the first fusion coefficient and the second fusion coefficient.

[0095] The environment detection result includes a signal-to-noise ratio detection result and a wind noise detection result. Of course, in actual application, the environment detection result can also include other information, which is not limited in the embodiments of the present application.

[0096] In the third mode, the target scene currently present is detected, and then the fusion coefficient is determined based on the environment detection result, so that the target audio signal can be better adapted to the environment currently present.

[0097] In some embodiments, the low-frequency repair signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature are taken as inputs of the full-frequency repair network model to obtain the target audio signal output by the full-frequency repair network model.

[0098] The full-frequency repair network model is obtained by training a to-be-trained full-frequency repair network model based on a third sample data set. The third sample data set includes a plurality of groups of sample data, and each group of sample data in the plurality of groups of sample data includes a sample low-frequency repair signal, a second sample air conduction audio signal, a first sample fusion coefficient, a second sample fusion coefficient, a sample acoustic feature, and an actual sample audio signal.

[0099] The training process of the full-frequency restoration network model is similar to the training process of the low-frequency restoration network model in step 202, so the relevant content of step 202 can be referred to, and details are not repeated here.

[0100] The structure of the full-frequency restoration network model is a convolution recurrent neural network (CRNN) structure, and the full-frequency restoration network model includes an encoding part and a decoding part. The encoding part includes at least one encoding layer, and the decoding part includes at least one decoding layer.

[0101] Next, taking an example in which the encoding part includes a first encoding layer, a second encoding layer, and a third encoding layer, and the decoding part includes a first decoding layer, a second decoding layer, and a third decoding layer. At this time, the implementation process of taking the low-frequency restoration signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature as the input of the full-frequency restoration network model to obtain the target audio signal output by the full-frequency restoration network model includes the following steps (1)-(6).

[0102] (1) Input the low-frequency restoration signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature to the first encoding layer to obtain the first encoding result output by the first encoding layer.

[0103] (2) Input the first encoding result to the second encoding layer to obtain the second encoding result output by the second encoding layer.

[0104] (3) Input the second encoding result to the third encoding layer to obtain the third encoding result output by the third encoding layer.

[0105] (4) Input the third encoding result output by the third encoding layer and the second encoding result output by the second encoding layer to the first decoding layer to obtain the first decoding result output by the first decoding layer.

[0106] (5) Input the first decoding result output by the first decoding layer and the first encoding result output by the first encoding layer to the second decoding layer to obtain the second decoding result output by the second decoding layer.

[0107] (6) Input the second decoding result output by the second decoding layer, the low-frequency restoration signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature to the third decoding layer to obtain the third decoding result output by the third decoding layer, and then determine the third decoding result as the target audio signal.

[0108] For example, please refer to Figure 6 ,Figure 6 is a schematic diagram of obtaining a target audio signal through a full-frequency repair network model provided by an embodiment of the present application. In Figure 6 , the low-frequency repair signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature are input into the full-frequency repair network model, and then the target audio signal is determined through the three encoding layers and the three decoding layers included in the full-frequency repair network model.

[0109] It should be noted that directly inputting the low-frequency repair signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature into the full-frequency repair network model to determine the target audio signal is only an example. In other embodiments, the target audio signal can also be determined in other ways. For example, based on the first fusion coefficient and the second fusion coefficient, the low-frequency repair signal and the second air conduction audio signal collected by the at least one second air conduction microphone are fused to obtain a second fusion signal. Then, the second fusion signal and the acoustic feature are input into the full-frequency repair network model to obtain the target audio signal output by the full-frequency repair network model. In this way, the calculation amount of the full-frequency repair network model can be reduced, thereby improving the efficiency of audio signal repair.

[0110] wherein the low-frequency repair signal includes signal frequencies at multiple time points, and the second air conduction audio signal collected by the at least one second air conduction microphone also includes signal frequencies at multiple time points. The implementation process of fusing the low-frequency repair signal and the second air conduction audio signal collected by the at least one second air conduction microphone based on the first fusion coefficient and the second fusion coefficient to obtain a fusion signal includes: obtaining the signal frequencies at the same time point in the low-frequency repair signal and the second air conduction audio signal collected by the at least one second air conduction microphone respectively to obtain a first frequency and at least one second frequency. Multiplying the first frequency by the first fusion coefficient to obtain a first reference frequency. Multiplying the at least one second frequency by the second fusion coefficient corresponding to the at least one second air conduction audio signal respectively to obtain at least one second reference frequency. Then, adding the first reference frequency and the at least one second reference frequency to obtain a target frequency, and then taking the target frequency as the signal frequency at the time point in the fusion signal. In this way, each time point in the multiple time points is traversed in turn to obtain multiple target frequencies, and the multiple target frequencies correspond to the multiple time points one by one. The multiple target frequencies are arranged and combined in time sequence to obtain the fusion signal.

[0111] Next, the audio signal repair process provided by an embodiment of the present application is completely described by taking Figure 7 as an example. In Figure 7In the embodiments of the present application, since the bone conduction microphone collects the bone conduction audio signal through bone vibration, it can shield environmental noise to some extent, so that the low-frequency feature of the bone conduction audio signal can be more accurately extracted from the bone conduction audio signal. The signal-to-noise ratio of the first air conduction audio signal is high, so the low-frequency signal repair effect can be improved by using the low-frequency feature and the first air conduction audio signal. Moreover, the acoustic feature usually includes the features of the audio signal in each frequency range, and the second air conduction audio signal collected by the at least one second air conduction microphone includes a mid-high frequency signal. Therefore, the combination of the second air conduction audio signal collected by the at least one second air conduction microphone and the acoustic feature can better guide the full-frequency signal repair, thereby improving the quality and clarity of the target audio signal. That is, the combination of the bone conduction audio signal collected by the bone conduction microphone, the first air conduction audio signal collected by the first air conduction microphone, and the second air conduction audio signal collected by the at least one second air conduction microphone can improve the quality and clarity of the target audio signal. In addition, by combining the current scene and the fusion coefficient adjusted by the user according to the user's own needs, personalized and differentiated audio signal repair can be achieved. Alternatively, by detecting the target scene, the fusion coefficient can be determined based on the detection result, so that the target audio signal can be better adapted to the current environment.

[0112] In the embodiments of the present application, since the bone conduction microphone collects the bone conduction audio signal through bone vibration, it can shield environmental noise to some extent, so that the low-frequency feature of the bone conduction audio signal can be more accurately extracted from the bone conduction audio signal. The signal-to-noise ratio of the first air conduction audio signal is high, so the low-frequency signal repair effect can be improved by using the low-frequency feature and the first air conduction audio signal. Moreover, the acoustic feature usually includes the features of the audio signal in each frequency range, and the second air conduction audio signal collected by the at least one second air conduction microphone includes a mid-high frequency signal. Therefore, the combination of the second air conduction audio signal collected by the at least one second air conduction microphone and the acoustic feature can better guide the full-frequency signal repair, thereby improving the quality and clarity of the target audio signal. That is, the combination of the bone conduction audio signal collected by the bone conduction microphone, the first air conduction audio signal collected by the first air conduction microphone, and the second air conduction audio signal collected by the at least one second air conduction microphone can improve the quality and clarity of the target audio signal. In addition, by combining the current scene and the fusion coefficient adjusted by the user according to the user's own needs, personalized and differentiated audio signal repair can be achieved. Alternatively, by detecting the target scene, the fusion coefficient can be determined based on the detection result, so that the target audio signal can be better adapted to the current environment.

[0113] Figure 8 FIG. 1 is a structural schematic diagram of an audio signal repair device provided by an embodiment of the present application. The audio signal repair device can be realized as part or all of a computer device by software, hardware, or a combination of both. Referring to FIG. 1, Figure 8 The device includes a first determination module 801, a second determination module 802, a third determination module 803, and a fourth determination module 804.

[0114] The first determining module 801 is configured to determine a low-frequency feature of the bone conduction audio signal based on the bone conduction audio signal collected by the bone conduction microphone. For details, refer to the corresponding content in the above embodiments, which will not be described here.

[0115] The second determining module 802 is configured to determine a low-frequency repair signal based on the low-frequency feature and the first air conduction audio signal collected by the first air conduction microphone. For details, refer to the corresponding content in the above embodiments, which will not be described here.

[0116] The third determining module 803 is configured to determine an acoustic feature based on the low-frequency repair signal and the low-frequency feature. For details, refer to the corresponding content in the above embodiments, which will not be described here.

[0117] The fourth determining module 804 is configured to determine a target audio signal based on the low-frequency repair signal, the second air conduction audio signal collected by the at least one second air conduction microphone, and the acoustic feature, the target audio signal including a low-frequency signal and a medium-high frequency signal. For details, refer to the corresponding content in the above embodiments, which will not be described here.

[0118] Optionally, the second determining module 802 is specifically configured to:

[0119] The low-frequency feature and the first air conduction audio signal are taken as inputs of a low-frequency repair network model to obtain a low-frequency repair signal output by the low-frequency repair network model.

[0120] Optionally, the second determining module 802 includes:

[0121] A determining unit is configured to determine a low-frequency repair signal based on the low-frequency feature, the bone conduction audio signal, the first air conduction audio signal, and the second air conduction audio signal collected by part or all of the at least one second air conduction microphone.

[0122] Optionally, the determining unit is specifically configured to:

[0123] The bone conduction audio signal, the first air conduction audio signal, and the second air conduction audio signal collected by part or all of the at least one second air conduction microphone are fused to obtain a first fusion signal.

[0124] The low-frequency feature and the first fusion signal are taken as inputs of a low-frequency repair network model to obtain a low-frequency repair signal output by the low-frequency repair network model.

[0125] Optionally, the fourth determining module 804 includes:

[0126] The first determining unit is configured to determine a first fusion coefficient and a second fusion coefficient, the first fusion coefficient being a fusion coefficient of the low-frequency restoration signal, and the second fusion coefficient including a fusion coefficient of the second binaural audio signal collected by the at least one second binaural microphone.

[0127] The second determining unit is configured to determine the target audio signal based on the low-frequency restoration signal, the second binaural audio signal collected by the at least one second binaural microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature.

[0128] Optionally, the second determining unit is specifically configured to:

[0129] input the low-frequency restoration signal, the second binaural audio signal collected by the at least one second binaural microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature into the full-frequency restoration network model to obtain the target audio signal output by the full-frequency restoration network model.

[0130] Optionally, the second determining unit is specifically configured to:

[0131] fuse the low-frequency restoration signal and the second binaural audio signal collected by the at least one second binaural microphone based on the first fusion coefficient and the second fusion coefficient to obtain a second fusion signal;

[0132] input the second fusion signal and the acoustic feature into the full-frequency restoration network model to obtain the target audio signal output by the full-frequency restoration network model.

[0133] Optionally, the first determining unit is specifically configured to:

[0134] determine the first fusion coefficient based on the acoustic feature, and determine the second fusion coefficient based on the second binaural audio signal collected by the at least one second binaural microphone; or

[0135] determine a target scene currently located in, and determine the first fusion coefficient and the second fusion coefficient based on the target scene; or

[0136] perform environment detection on the target scene currently located in to obtain an environment detection result, and determine the first fusion coefficient and the second fusion coefficient based on the environment detection result.

[0137] In the embodiments of the present application, since the bone conduction microphone collects bone conduction audio signals in the manner of bone vibration, environmental noise can be shielded to a certain extent, so that the low-frequency characteristics of the bone conduction audio signals can be more accurately extracted from the bone conduction audio signals. The first air conduction audio signal has a high signal-to-noise ratio, so that the repair effect of the low-frequency signal can be improved by the low-frequency characteristics and the first air conduction audio signal. Moreover, since the acoustic characteristics generally include the characteristics of the audio signal in each frequency range, and the second air conduction audio signal collected by the at least one second air conduction microphone includes a mid-high frequency signal. Therefore, the combination of the second air conduction audio signal collected by the at least one second air conduction microphone and the acoustic characteristics can better guide the full-frequency signal repair, thereby improving the quality and clarity of the target audio signal. That is, the combination of the bone conduction audio signal collected by the bone conduction microphone, the first air conduction audio signal collected by the first air conduction microphone, and the second air conduction audio signal collected by the at least one second air conduction microphone can improve the quality and clarity of the target audio signal. In addition, by combining the current scene and the fusion coefficient adaptively adjusted by the user according to the user's own needs, personalized and differentiated audio signal repair can be achieved. Alternatively, by detecting the target scene currently, the fusion coefficient is determined based on the environmental detection result, so that the target audio signal can be better adapted to the current environment.

[0138] It should be noted that the audio signal repair device provided in the above embodiments is only used as an example for the division of the above functional modules during audio signal repair. In actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the audio signal repair device and the audio signal repair method provided in the above embodiments belong to the same concept, and the specific implementation process is described in the method embodiments, which will not be repeated here.

[0139] For reference Figure 9 , Figure 9 is a structural schematic diagram of a computer device according to an embodiment of the present application. The computer device includes at least one processor 901, a communication bus 902, a memory 903, and at least one communication interface 904.

[0140] The processor 901 can be a general central processing unit (CPU), a network processing unit (NP), a microprocessor, or can be one or more integrated circuits used to implement the schemes of the present application, such as an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0141] The communication bus 902 is used to transmit information between the above-mentioned components. The communication bus 902 can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, only one thick line is shown in the figure, but it does not mean that there is only one bus or only one type of bus.

[0142] The memory 903 can be a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an optical disk (including a compact disc read-only memory (CD-ROM), a compressed disk, a laser disk, a digital versatile disk, a Blu-ray disk, and the like), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but is not limited thereto. The memory 903 can exist independently and be connected to the processor 901 through the communication bus 902. The memory 903 can also be integrated with the processor 901.

[0143] The communication interface 904 is configured to communicate with other devices or communication networks using any transceiver-like mechanism. The communication interface 904 includes a wired communication interface and can also include a wireless communication interface. For example, the wired communication interface can be an Ethernet interface. The Ethernet interface can be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface can be a wireless local area networks (WLAN) interface, a cellular network communication interface, or a combination thereof.

[0144] In some embodiments, the processor 901 can include one or more central processing units (CPUs) as an example. For example, the processor 901 can include CPU0 and CPU1 as shown in FIG. 9. Figure 9

[0145] In some embodiments, the computer device can include multiple processors as an example. For example, the computer device can include the processor 901 and the processor 905 as shown in FIG. 9. Each of the processors can be a single-core processor or a multi-core processor. The processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions). Figure 9

[0146] In some embodiments, the computer device can further include an output device 906 and an input device 907 as an example. The output device 906 is configured to communicate with the processor 901 and display information in various ways. For example, the output device 906 can be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, a projector, or the like. The input device 907 is configured to communicate with the processor 901 and receive user input in various ways. For example, the input device 907 can be a mouse, a keyboard, a touch screen device, a sensor device, or the like.

[0147] In some embodiments, the memory 903 is configured to store program code 910 for implementing the solutions of the present application. The processor 901 can execute the program code 910 stored in the memory 903. The program code 910 can include one or more software modules. The computer device can implement the above-mentioned solutions by means of the processor 901 and the program code 910 in the memory 903. Figure 2 The audio signal repairing method provided by the embodiments.

[0148] ​​In the embodiments described above, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (for example: coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example: infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium accessible by a computer, or a data storage device such as a server, data center, etc. integrated with one or more available media. The available medium can be a magnetic medium (for example: floppy disk, hard disk, magnetic tape), an optical medium (for example: digital versatile disc (DVD)) or a semiconductor medium (for example: solid state disk (SSD)) and the like. It is worth noting that the computer readable storage medium mentioned in the embodiments of the present application can be a non-volatile storage medium, in other words, it can be a non-transitory storage medium.

[0149] That is, the embodiments of the present application also provide a computer readable storage medium, which stores instructions, when the instructions are executed on a computer, the computer executes the steps of the above audio signal repair method.

[0150] The embodiments of the present application also provide a computer program product including instructions, when the instructions are executed on a computer, the computer executes the steps of the above audio signal repair method. Alternatively, a computer program is provided, when the computer program is executed on a computer, the computer executes the steps of the above audio signal repair method.

[0151] It should be understood that the "multiple" mentioned herein refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B; "and / or" herein only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist together, and B exists alone. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, "first", "second" and the like are used to distinguish the same items or similar items with basically the same function and role. The skilled in the art can understand that "first", "second" and the like do not limit the quantity and execution order, and "first", "second" and the like do not necessarily mean different.

[0152] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the bone conduction microphone collected bone conduction audio signal, the first air conduction microphone collected first air conduction audio signal, and the second air conduction microphone collected second air conduction audio signal in the embodiments of the present application are all obtained under sufficient authorization.

[0153] The above is the embodiment provided by the present application, which does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method of audio signal restoration, characterized by, The method is applied to a headset, the headset comprising a bone conduction microphone, a first air conduction microphone and at least one second air conduction microphone, the first air conduction microphone being configured to collect an air conduction signal inside an ear canal, and the at least one second air conduction microphone being configured to collect an air conduction signal of an external environment; the method comprising: determining a low-frequency feature of a bone conduction audio signal collected by the bone conduction microphone; determining a low-frequency restoration signal based on the low-frequency feature and a first air conduction audio signal collected by the first air conduction microphone; determining an acoustic feature based on the low-frequency restoration signal and the low-frequency feature; determining a target audio signal based on the low-frequency restoration signal, a second air conduction audio signal collected by the at least one second air conduction microphone, and the acoustic feature, the target audio signal comprising a low-frequency signal and a medium-high frequency signal.

2. The method of claim 1, wherein, The method of determining a low-frequency restoration signal based on the low-frequency feature and a first air conduction audio signal collected by the first air conduction microphone comprises: inputting the low-frequency feature and the first air conduction audio signal into a low-frequency restoration network model to obtain the low-frequency restoration signal output by the low-frequency restoration network model.

3. The method of claim 1, wherein, The method of determining a low-frequency restoration signal based on the low-frequency feature and a first air conduction audio signal collected by the first air conduction microphone comprises: determining the low-frequency restoration signal based on the low-frequency feature, the bone conduction audio signal, the first air conduction audio signal, and a second air conduction audio signal collected by some or all of the at least one second air conduction microphone.

4. The method of claim 3, wherein, The method of determining the low-frequency restoration signal based on the low-frequency feature, the bone conduction audio signal, the first air conduction audio signal, and a second air conduction audio signal collected by some or all of the at least one second air conduction microphone comprises: fusing the bone conduction audio signal, the first air conduction audio signal, and the second air conduction audio signal collected by some or all of the at least one second air conduction microphone to obtain a first fused signal; inputting the low-frequency feature and the first fused signal into a low-frequency restoration network model to obtain the low-frequency restoration signal output by the low-frequency restoration network model.

5. The method according to any one of claims 1 to 4, characterized in that, The method of determining a target audio signal based on the low-frequency restoration signal, a second air conduction audio signal collected by the at least one second air conduction microphone, and an acoustic feature comprises: determining a first fusion coefficient and a second fusion coefficient, the first fusion coefficient being a fusion coefficient of the low-frequency restoration signal, and the second fusion coefficient comprising a fusion coefficient of the second air conduction audio signal collected by the at least one second air conduction microphone; determining the target audio signal based on the low-frequency restoration signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature.

6. The method of claim 5, wherein, The method of determining the target audio signal based on the low-frequency restoration signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature comprises: The low-frequency repair signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature are taken as inputs of a full-frequency repair network model to obtain the target audio signal output by the full-frequency repair network model.

7. The method of claim 5, wherein, The target audio signal is determined based on the low-frequency repair signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature, and the target audio signal includes a low-frequency signal and a medium-high frequency signal. The low-frequency repair signal and the second air conduction audio signal collected by the at least one second air conduction microphone are fused based on the first fusion coefficient and the second fusion coefficient to obtain a second fusion signal. The second fusion signal and the acoustic feature are taken as inputs of a full-frequency repair network model to obtain the target audio signal output by the full-frequency repair network model.

8. The method of claim 5, wherein, The first fusion coefficient and the second fusion coefficient are determined in the following manners: The first fusion coefficient is determined based on the acoustic feature, and the second fusion coefficient is determined based on the second air conduction audio signal collected by the at least one second air conduction microphone; or A target scene currently located is determined, and the first fusion coefficient and the second fusion coefficient are determined based on the target scene; or An environment detection is performed on the target scene currently located to obtain an environment detection result, and the first fusion coefficient and the second fusion coefficient are determined based on the environment detection result.

9. An audio signal restoration apparatus characterized by comprising: The device is applied to an earphone, and the earphone includes a bone conduction microphone, a first air conduction microphone, and at least one second air conduction microphone. The first air conduction microphone is configured to collect an air conduction signal inside an ear canal, and the at least one second air conduction microphone is configured to collect an air conduction signal of an external environment. The device includes: A first determination module configured to determine a low-frequency feature of a bone conduction audio signal collected by the bone conduction microphone. A second determination module configured to determine a low-frequency repair signal based on the low-frequency feature and a first air conduction audio signal collected by the first air conduction microphone. A third determination module configured to determine an acoustic feature based on the low-frequency repair signal and the low-frequency feature. A fourth determination module configured to determine a target audio signal based on the low-frequency repair signal, a second air conduction audio signal collected by the at least one second air conduction microphone, and the acoustic feature. The target audio signal includes a low-frequency signal and a medium-high frequency signal.

10. The apparatus of claim 9, wherein, The second determination module is specifically configured to: Take the low-frequency feature and the first air conduction audio signal as inputs of a low-frequency repair network model to obtain the low-frequency repair signal output by the low-frequency repair network model.

11. The apparatus of claim 9, wherein, The second determination module includes: A determination unit configured to determine the low-frequency repair signal based on the low-frequency feature, the bone conduction audio signal, the first air conduction audio signal, and a second air conduction audio signal collected by part or all of the at least one second air conduction microphone.

12. The apparatus of claim 11, wherein, The determination unit is specifically configured to: fuse the bone conduction audio signal, the first air conduction audio signal, and the second air conduction audio signal collected by the at least one second air conduction microphone to obtain a first fused signal; input the low frequency feature and the first fused signal into a low frequency restoration network model to obtain the low frequency restoration signal output by the low frequency restoration network model.

13. The apparatus of any one of claims 9-12, wherein, The fourth determination module comprises: a first determination unit configured to determine a first fusion coefficient and a second fusion coefficient, the first fusion coefficient being a fusion coefficient of the low frequency restoration signal, and the second fusion coefficient including a fusion coefficient of the second air conduction audio signal collected by the at least one second air conduction microphone; a second determination unit configured to determine the target audio signal based on the low frequency restoration signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature.

14. The apparatus of claim 13, wherein, The second determination unit is specifically configured to: input the low frequency restoration signal, the second air conduction audio signal collected by the at least one second air conduction microphone, the first fusion coefficient, the second fusion coefficient, and the acoustic feature into a full frequency restoration network model to obtain the target audio signal output by the full frequency restoration network model.

15. The apparatus of claim 13, wherein, The second determination unit is specifically configured to: fuse the low frequency restoration signal and the second air conduction audio signal collected by the at least one second air conduction microphone based on the first fusion coefficient and the second fusion coefficient to obtain a second fused signal; input the second fused signal and the acoustic feature into a full frequency restoration network model to obtain the target audio signal output by the full frequency restoration network model.

16. The apparatus of claim 13, wherein, The first determination unit is specifically configured to: determine the first fusion coefficient based on the acoustic feature and determine the second fusion coefficient based on the second air conduction audio signal collected by the at least one second air conduction microphone; or determine a target scene currently located in, and determine the first fusion coefficient and the second fusion coefficient based on the target scene; or perform environment detection on the target scene currently located in to obtain an environment detection result, and determine the first fusion coefficient and the second fusion coefficient based on the environment detection result.

17. A computer device, comprising: The computer device comprises a memory and a processor, the memory is configured to store a computer program, and the processor is configured to execute the computer program stored in the memory to implement the method in any one of claims 1-8.

18. A computer-readable storage medium, characterized in that, The storage medium has instructions stored therein, and when the instructions are executed on the computer, the computer executes the steps of the method in any one of claims 1-8.

19. A computer program product, characterised in that, The computer program product comprises instructions, and when the instructions are executed on the computer, the computer executes the method in any one of claims 1-8.

Citation Information

Patent Citations

  • Sound source control method and loudspeaker equipment

    CN110782912A

  • Voice control method and device

    CN115132212A