Method executed by electronic equipment, electronic equipment and storage medium

By dynamically adjusting the processing delay and performing signal correction through the cascade neural network, the problem of inaccurate sound amplification at low delays in the prior art is solved, and efficient and smooth audio signal processing is achieved.

CN120220716APending Publication Date: 2025-06-27BEIJING SAMSUNG TELECOM R&D CENT +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311824210.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to accurately amplify and compensate sound under low latency conditions, resulting in poor user auditory experience, especially when fine frequency control is required.

Method used

The processing delay and amplification signal of the audio signal are determined through multiple cascading neural networks, the processing delay is dynamically adjusted to adapt to signal changes, and signal correction and compensation are performed through the neural network to improve amplification accuracy and fluency.

Benefits of technology

It enables efficient amplification and compensation of sound under low latency conditions, improving the user's auditory experience, especially when processing rapidly changing audio signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220716A_ABST
    Figure CN120220716A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method executed by electronic equipment, the electronic equipment and a storage medium, and relates to the field of artificial intelligence. The method comprises the following steps: determining a first processing time delay for performing amplification processing on a first audio signal to be processed and a second audio signal obtained through the amplification processing; if the first processing time delay is different from the processing time delay for amplifying the previous first audio signal, determining a third audio signal based on the second audio signal; and outputting the third audio signal. Optionally, the method performed by the electronic device may be performed using an artificial intelligence model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of signal processing. Specifically, the present application relates to a method for processing an audio signal executed by an electronic device, an electronic device, and a storage medium. Background Art

[0002] Personal Sound Amplification Products (PSAP), as wearable electronic products, are designed to amplify sound for non-deaf or hearing-impaired people. Wide Dynamic Range Compression (WDRC) is a sound amplification technology that can amplify or reduce sounds of different frequencies or frequency bands according to the user's audiogram. WDRC is an important module in PSAP and can be used in True Wireless Stereo (TWS) headphones with an "ambient sound amplification" function to improve the auditory perception quality of people with normal hearing or hearing impairment. PSAP requires low latency (<= 3 ms). If the latency of the sound played by TWS exceeds 3 ms compared to the sound directly heard by the ear, the sound heard by the user will be unclear. WDRC can be used to provide hearing compensation for frequency bands with hearing defects. The more sub-bands in the WDRC algorithm, the finer the hearing compensation and the better the sound amplification effect.

[0003] How to accurately amplify and compensate sound to meet user needs is a technical problem that those skilled in the art have been constantly striving to study. Summary of the Invention

[0004] To at least solve the above problems existing in the prior art, the present invention provides a method executed by an electronic device, an electronic device, and a storage medium.

[0005] According to a first aspect of an embodiment of the present application, there is provided a method executed by an electronic device, including: determining a first processing delay for amplifying a first audio signal to be processed and a second audio signal obtained through the amplification processing; if the first processing delay is different from the processing delay for amplifying a previous first audio signal, determining a third audio signal based on the second audio signal; and outputting the third audio signal.

[0006] Optionally, the step of determining a first processing delay for amplifying a first audio signal to be processed and a second audio signal obtained through the amplification processing includes: obtaining a processed audio signal and its corresponding prediction probability based on the first audio signal by using each first neural network in a cascade of multiple first neural networks; and determining the first processing delay and the second audio signal based on the prediction probability.

[0007] Optionally, the step of determining the first processing delay and the second audio signal based on the prediction probability includes: for each first neural network among the multiple first neural networks: if the processing delay corresponding to the first neural network is determined as the first processing delay based on the prediction probability obtained by the first neural network, then the audio signal obtained by the first neural network is determined as the second audio signal; if the processing delay corresponding to the first neural network is not determined as the first processing delay based on the prediction probability obtained by the first neural network, then the first processing delay and the second audio signal are determined according to the prediction probability obtained by the next first neural network.

[0008] Optionally, the step of determining the first processing delay and the second audio signal based on the prediction probability further includes: for each first neural network among the multiple first neural networks, if the processing delay corresponding to the first neural network is not determined as the first processing delay based on the prediction probability obtained by the first neural network, then the processed audio signal obtained by the first neural network and the output features of at least one network layer other than the last network layer in the first neural network are input into the next first neural network.

[0009] Optionally, the step of obtaining the processed audio signal and its corresponding prediction probability based on the first audio signal by using each first neural network in a cascade of multiple first neural networks includes: for the first first neural network among the multiple first neural networks, based on the target amplification gain and the first audio signal to be processed corresponding to the first neural network, use the first neural network to obtain the processed audio signal and its corresponding prediction probability; for each of the other first neural networks among the multiple first neural networks, based on at least one of the following, use the first neural network to obtain the processed audio signal and its corresponding prediction probability: the processed audio signal obtained by the previous first neural network, the output features of the previous first neural network, and the audio signal other than the first audio signal to be processed corresponding to the previous first neural network among the first audio signals to be processed corresponding to the first neural network.

[0010] Optionally, the step of determining the first processing delay and the second audio signal based on the prediction probability includes: for each first neural network among the multiple first neural networks, if the prediction probability obtained by the first neural network is greater than a first predetermined threshold, then the processing delay corresponding to the first neural network is determined as the first processing delay, and the processed audio signal obtained by the first neural network is determined as the second audio signal.

[0011] Optionally, the processing delay corresponding to each first neural network increases sequentially in the order of cascading of the multiple first neural networks; the time length of the first audio signal to be processed corresponding to each first neural network corresponds to the processing delay corresponding to the corresponding first neural network.

[0012] Optionally, the steps of determining a first processing delay for amplifying the first audio signal to be processed and a second audio signal obtained by the amplifying process include: by using each of a plurality of fourth neural networks, respectively obtaining a processed audio signal and its corresponding prediction probability based on the first audio signal to be processed corresponding to the fourth neural network and a target amplification gain, wherein each fourth neural network corresponds to a different processing delay; determining the first processing delay and the second audio signal based on the prediction probability.

[0013] Optionally, the steps of determining the first processing delay and the second audio signal based on the prediction probability include: if the prediction probability obtained by at least one fourth neural network is greater than a second predetermined threshold, determining the audio signal obtained by the fourth neural network corresponding to the lowest processing delay among the at least one fourth neural network as the second audio signal, and determining the processing delay corresponding to the fourth neural network as the first processing delay; if the prediction probabilities obtained by the plurality of fourth neural networks are all less than or equal to the second predetermined threshold, determining the audio signal obtained by the fourth neural network corresponding to the highest processing delay among the plurality of fourth neural networks as the second audio signal, and determining the processing delay corresponding to the fourth neural network as the first processing delay.

[0014] Optionally, the target amplification gain is obtained by the following operations: dividing an audio signal of a preset time length into a plurality of sub-band signals; respectively calculating the amplification gain of each sub-band signal in the plurality of sub-band signals according to the hearing information of the user; selecting the amplification gain corresponding to the first audio signal to be processed from the calculated plurality of amplification gains as the target amplification gain.

[0015] Optionally, if the first processing delay is greater than the processing delay for amplifying the previous first audio signal, the method further includes: obtaining, by a second neural network, at least one fourth audio signal to be filled before the third audio signal based on the audio signal output for the previous first audio signal and the first audio signal to be processed; and outputting the at least one fourth audio signal.

[0016] Optionally, the step of determining a third audio signal based on the second audio signal includes: correcting the second audio signal by a third neural network based on the output fourth audio signal to obtain the third audio signal.

[0017] Optionally, the method further includes: for each first neural network among the multiple first neural networks, if the prediction probability obtained based on the first neural network does not determine the processing delay corresponding to the first neural network as the first processing delay, then based on at least one of the following, determine, through a second neural network, a fourth audio signal that needs to be filled before the third audio signal: the previously output audio signal, the output features of at least one network layer in the first neural network except for the last network layer, and the first audio signal to be processed corresponding to the first first neural network among the multiple first neural networks; output the fourth audio signal.

[0018] Optionally, if the first processing delay is less than the processing delay for amplifying the previous first audio signal, the method further includes: based on the previously output audio signal, correct the second audio signal through a third neural network to obtain the third audio signal.

[0019] Optionally, the method further includes: if the first processing delay is the same as the processing delay for amplifying the previous first audio signal, output the second audio signal.

[0020] According to a second aspect of the embodiments of the present application, there is provided an electronic device, including: at least one processor; and at least one memory storing computer-executable instructions, wherein when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the method performed by the electronic device as described above.

[0021] According to a third aspect of the embodiments of the present application, there is provided a computer-readable storage medium storing instructions, wherein when the instructions are run by at least one processor, the at least one processor is caused to execute the method performed by the electronic device as described above.

[0022] The beneficial effects brought by the technical solutions provided by the embodiments of the present application will be described in combination with specific optional embodiments later, or can be learned from the description of the embodiments, or can be learned through the implementation of the embodiments. Description of the Drawings

[0023] To more clearly, easily illustrate and understand the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.

[0024] Figure 1 It is a schematic structural diagram showing a WDRC system.

[0025] Figure 2AIt is a schematic diagram showing the influence of time delay on the gain of each frequency.

[0026] Figure 2B It is a schematic diagram showing the distortion generated when an audio signal is amplified using a neural network;

[0027] Figure 3 It is a flowchart showing a method executed by an electronic device according to an exemplary embodiment of the present application.

[0028] Figure 4 It is a schematic diagram showing the process of a method executed by an electronic device according to an exemplary embodiment of the present application.

[0029] Figure 5 It is a flowchart showing the process of determining a first processing time delay for amplifying a first audio signal to be processed and a second audio signal according to an exemplary embodiment of the present application.

[0030] Figure 6 It is a diagram showing the cascaded structure of multiple first neural networks according to an exemplary embodiment of the present application.

[0031] Figure 7 It is a flowchart showing the process of determining a first processing time delay for amplifying a first audio signal to be processed and a second audio signal according to another exemplary embodiment of the present application.

[0032] Figure 8 It is a diagram showing an example structure of multiple fourth neural networks according to an exemplary embodiment of the present application.

[0033] Figure 9 It is shown that the above is referred to Figure 5 A schematic diagram showing an example where the first processing time delay increases when a second audio signal is obtained by applying the process described above.

[0034] Figure 10 It is a schematic diagram showing the influence caused by the increase in the processing time delay for amplifying the first audio signal.

[0035] Figure 11 It is shown that the above is referred to Figure 5 A schematic diagram showing an example where the first processing time delay decreases when a second output signal is obtained by applying the process described above.

[0036] Figure 12 It is a schematic diagram showing the influence caused by the decrease in the processing time delay for amplifying the first audio signal.

[0037] Figure 13It is a flowchart showing a process of determining a third audio signal based on a second audio signal when a first processing delay for amplifying a first audio signal to be processed is different from a processing delay for amplifying a previous first audio signal according to an exemplary embodiment of the present application.

[0038] Figure 14 It is a schematic diagram showing a process of generating a fourth audio signal through a second neural network according to an exemplary embodiment of the present application.

[0039] Figure 15 It is a schematic diagram showing a process of obtaining a third audio signal by correcting a second audio signal through a third neural network according to an exemplary embodiment of the present application.

[0040] Figure 16 It is a schematic diagram showing a process of applying a method executed by an electronic device according to an exemplary embodiment of the present application.

[0041] Figure 17 A schematic diagram of a structure of an electronic device applicable according to an exemplary embodiment of the present application is shown. Detailed Description

[0042] The following description with reference to the accompanying drawings is provided to facilitate a thorough understanding of various embodiments of the present disclosure defined by the claims and their equivalents. This description includes various specific details to facilitate understanding but should be considered merely exemplary. Thus, those of ordinary skill in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, descriptions of well-known functions and structures may be omitted for clarity and conciseness.

[0043] The terms and phrases used in the following specification and claims are not limited to their dictionary meanings but are used solely by the inventors to enable a clear and consistent understanding of the present disclosure. Thus, it should be apparent to those skilled in the art that the following description of the various embodiments of the present disclosure is provided for illustrative purposes only and not for the purpose of limiting the present disclosure as defined by the appended claims and their equivalents.

[0044] It should be understood that the singular forms "a", "an", and "the" may also include plural referents unless the context clearly indicates otherwise. Thus, for example, a reference to "a component surface" includes a reference to one or more such surfaces. When we say that one element is "connected" or "coupled" to another element, that one element may be directly connected or coupled to the other element, or it may mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include a wireless connection or wireless coupling.

[0045] The term "comprises" or "may comprise" refers to the presence of the corresponding disclosed function, operation, or component that can be used in various embodiments of the present disclosure, and does not limit the presence of one or more additional functions, operations, or features. In addition, the term "comprises" or "has" may be interpreted to mean that certain characteristics, numbers, steps, operations, components, components, or combinations thereof are present, but should not be construed as excluding the possibility of the presence of one or more other characteristics, numbers, steps, operations, components, components, or combinations thereof.

[0046] The term "or" used in various embodiments of the present disclosure includes any of the listed terms and all combinations thereof. For example, "A or B" may include A, may include B, or may include both A and B. When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items may refer to one, multiple, or all of the multiple items. For example, for the description of "parameter A includes A1, A2, A3", it may be implemented such that parameter A includes A1 or A2 or A3, or it may also be implemented such that parameter A includes at least two of the three items of parameter A1, A2, and A3.

[0047] Unless otherwise defined, all terms (including technical terms or scientific terms) used in the present disclosure have the same meaning as understood by those skilled in the art to which the present disclosure pertains. Commonly used terms defined in a dictionary are interpreted to have a meaning consistent with the context in the relevant technical field, and should not be interpreted idealistically or overly formally unless explicitly defined as such in the present disclosure.

[0048] At least some of the functions in the apparatus or electronic device provided in the embodiments of the present disclosure can be implemented by an AI model. For example, at least one of the multiple modules of the apparatus or electronic device can be implemented by an AI model. The functions associated with AI can be executed by a non-volatile memory, a volatile memory, and a processor.

[0049] The processor may include one or more processors. At this time, the one or more processors may be general-purpose processors, such as a central processing unit (CPU), an application processor (AP), etc., or a pure graphics processing unit, such as a graphics processing unit (GPU), a vision processing unit (VPU), and / or an AI dedicated processor, such as a neural processing unit (NPU).

[0050] The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence (AI) models stored in the non-volatile memory and the volatile memory. The predefined operation rules or artificial intelligence models are provided through training or learning.

[0051] Here, providing through learning refers to obtaining a predefined operation rule or an AI model with desired characteristics by applying a learning algorithm to multiple learning data. This learning can be performed in the device or electronic device itself where the AI according to the embodiment is executed, and / or can be implemented by a separate server / system.

[0052] The AI model can include multiple neural network layers. Each layer has multiple weight values, and each layer performs neural network calculations through the calculation between the input data of the layer (such as the calculation result of the previous layer and / or the input data of the AI model) and the multiple weight values of the current layer. Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial network (GAN), and deep Q-network.

[0053] A learning algorithm is a method of using multiple learning data to train a predetermined target device (e.g., a robot) to enable, allow, or control the target device to make a determination or prediction. Examples of this learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0054] The method provided by the present disclosure may relate to one or more of the technical fields such as speech, language, image, video, or data intelligence.

[0055] Optionally, when involving the field of speech or language, in the method executed by an electronic device according to the present disclosure, a speech signal as an analog signal can be received via a speech input device (e.g., a microphone), and the speech part can be converted into computer-readable text using an automatic speech recognition (ASR) model. The user's utterance intention can be obtained by interpreting the converted text using a natural language understanding (NLU) model. The ASR model or the NLU model can be an AI model. The AI model can be processed by an AI dedicated processor designed in a hardware structure specified for AI model processing. Language understanding is a technology for recognizing and applying / processing human language / text, including, for example, natural language processing, machine translation, dialogue systems, question answering, or speech recognition / synthesis.

[0056] Optionally, when involving the field of image or video, in the method executed by an electronic device according to the present disclosure, output data can be obtained by using image data as input data of an AI model. The method of the present disclosure may relate to the field of visual understanding of AI technology, and visual understanding is a technology for recognizing and processing things like human vision, and includes, for example, object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / positioning, or image enhancement.

[0057] Optionally, when it comes to the field of data intelligent processing, in the method executed by an electronic device according to the present disclosure, during the inference or prediction phase, an artificial intelligence model can be used to perform predictions by using real-time input data. The processor of the electronic device can perform preprocessing operations on the data to convert it into a form suitable for use as input to the artificial intelligence model. Inference and prediction are techniques for logical reasoning and prediction by determining information, including, for example, knowledge-based reasoning, optimization prediction, preference-based planning, or recommendation.

[0058] In this application, the artificial intelligence model can be obtained through training. Here, "obtained through training" means obtaining a predefined operation rule or artificial intelligence model configured to perform expected features (or purposes) by training a basic artificial intelligence model with multiple pieces of training data using a training algorithm. The artificial intelligence model can include multiple neural network layers. Each of the multiple neural network layers includes multiple weight values and performs neural network calculations through calculations between the calculation results of the previous layer and the multiple weight values.

[0059] The WDRC system in this application mainly adopts filter design, as Figure 1 shown. The WDRC system mainly includes four modules: a subband analysis module, a gain calculation module, an amplifier module, and a subband synthesis module. The subband analysis module obtains multiple subband signals by downsampling the input audio signal, then uses auditory filters to perform filter analysis on the multiple subband signals with multiple analysis frequencies as the center frequencies, and then performs a transformation (such as performing a Hilbert transform) on each filtered subband signal to extract the amplitude of each subband signal. The gain calculation module calculates the subband gain (also referred to as "amplification gain") of each subband signal respectively according to the amplitude of each subband signal obtained from the subband analysis module and the audiogram of the user corresponding to the above audio signal by using a hearing aid prescription fitting formula. The amplifier module applies the subband gain calculated by the gain calculation module to the corresponding subband signal based on a filtering method, thereby obtaining the amplified subband signals. The subband synthesis module upsamples the amplified subband signals, that is, by padding zeros to the amplified subband signals, and then filtering the padded subband signals through an interpolation filter, so as to synthesize all the subband signals into a complete full-band signal, that is, the amplified signal.

[0060] However, the delay of the existing WDRC system is fixed and relatively low, which means that the filter is relatively wide, as Figure 2A shown by curve 1 in. However, a relatively wide filter cannot accurately apply the calculated gain to each frequency point or frequency band. As Figure 2AAs shown, the operating range (i.e., frequency resolution) of the filter indicated by curves 1 and 2 is from f_a to f_b. However, if the user wants to only amplify the frequency band f_0 without affecting the frequencies near f_a and f_b, obviously the filter indicated by curves 1 and 2 cannot meet this requirement. The user hopes that the filter can accurately distinguish and act on each frequency band. When the frequency band is narrower, the required filter is also narrower, and the frequency resolution of the filter is higher. Since the frequency resolution of traditional filters is restricted by time delay, such a requirement cannot be met. Figure 2A It shows the influence of time delay on the applied gain. In the case of ideal infinite time delay, the frequency-gain curve is indicated by curve 3, which applies gain to the frequency f_0. At this time, f_a and f_b near f_0 will not be affected by the filter; when the time delay is high, the frequency-gain curve is indicated by curve 2, and the filter will have a slight influence on f_a and f_b; when the time delay is low, the frequency-gain curve is indicated by curve 1, and the filter will have a greater influence on f_a and f_b. At this time, the influence on f_a and f_b cannot be ignored. The contradiction between low time delay and the operating range of the filter often makes it impossible to accurately control the frequency band to be amplified, resulting in some frequencies of the original voice being over-amplified, bringing an unnatural listening experience to the user.

[0061] This application proposes that a neural network can be used to replace Figure 1 the amplifier module and the subband synthesis module in, and directly use the neural network to predict the amplified signal, that is, predict the amplified signal according to the historical signal information and the currently input audio signal, then the problem of insufficient frequency resolution of the filter described above with reference to Figure 2A can be solved under low time delay conditions, and more subbands can be supported. For the situation where the audio signal input in the actual application scenario changes significantly within a short period of time (for example, a sudden voice, a knock on the door, or other short and high-energy sounds), this application can not only avoid the distortion of the amplified signal under low time delay (such as 3ms) through the neural network method, but also bring a good auditory experience to the user. For example, during a sudden speech, at the initial stage of the signal, the energy suddenly becomes significantly larger. Using the method proposed in this application can effectively avoid the distortion of the predicted amplified signal at the initial stage as shown in the yellow virtual circle part in Figure 2B and bring a good auditory experience to the user.

[0062] Based on this, the present application proposes a method executed by an electronic device. This method can determine a first processing delay for amplifying a first audio signal to be processed and a second audio signal obtained through the amplification processing by means of multiple first neural networks or multiple fourth neural networks. If the first processing delay is different from the processing delay for amplifying the previous first audio signal, then a third audio signal is determined based on the second audio signal, and then the third audio signal is output. Among them, if the first processing delay is greater than the processing delay for amplifying the previous first audio signal, then the method can obtain at least one fourth audio signal to be filled before the third audio signal based on the audio signal output for the previous first audio signal and the first audio signal to be processed through a second neural network, and output the at least one fourth audio signal. If the first processing delay is less than the processing delay for amplifying the previous first audio signal, then the second audio signal is corrected through a third neural network based on the previously output audio signal to obtain a third audio signal. Through the above operations, the present application can use ultra-low latency for sound amplification in some cases, and appropriately increase the latency in some cases. It can adaptively adjust the processing delay according to the signal situation. Moreover, using neural networks solves the limitations of amplification accuracy and the number of subbands, and when the output delay changes, neural networks are used to compensate the output, thereby obtaining a smooth listening experience.

[0063] The technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure will be described below through the description of several alternative embodiments. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.

[0064] Figure 3 It is a flowchart showing a method executed by an electronic device according to an exemplary embodiment of the present application. Figure 4 It is a schematic diagram showing the process of a method executed by an electronic device according to an exemplary embodiment of the present application.

[0065] As Figure 3 shown, in step S310, a first processing delay for amplifying a first audio signal to be processed and a second audio signal obtained through the amplification processing are determined.

[0066] Specifically, the above step S310 can be performed by Figure 4It is executed by the adaptive time-delay amplification module in []. The first audio signal to be processed mentioned in step S310 may be a sampled signal in an input audio signal of a predetermined time length. Before executing step S310, it is first necessary to determine multiple sub-band signals of the input audio signal of a predetermined time length and calculate the sub-band gains of the multiple sub-band signals. The following will refer to Figure 4 for a detailed description of this.

[0067] As Figure 4 shown in [], the sub-band analysis module can divide the input audio signal of a predicted time length into multiple sub-band signals. For example, the audio signal can be downsampled at the sampling rate of the audio signal to obtain multiple sub-band signals. For example, the sampling rates can be fs, fs / 2, fs / 4, fs / 8, fs / 16. Correspondingly, 5 sub-band signals can be obtained. Then, the sub-band analysis module can perform filtering analysis on the multiple sub-band signals obtained above with multiple analysis frequencies (such as the above sampling rates fs, fs / 2, fs / 4, fs / 8, fs / 16) as the center frequencies. Among them, the center frequency is related to the frequency that the user needs to compensate. For example, if the user is more sensitive to low-frequency loss, more filtering needs to be allocated in the frequency band where the user is damaged. After that, the sub-band analysis module performs a transformation (such as a Hilbert transform) on each filtered sub-band signal to extract the amplitude of each sub-band signal.

[0068] The gain calculation module can calculate the amplification gain (also called "sub-band gain") of each sub-band signal in the multiple sub-band signals according to the user's hearing information (such as the user's audiogram). Specifically, the gain calculation module can obtain the amplitude of each sub-band signal extracted by the sub-band analysis module, receive the user's audiogram corresponding to the audio signal, and then use a hearing aid prescription fitting formula (such as NAL, NAL-R, NAL-RP, NAL-NL1, NAL-NL2, etc.) to calculate the amplification gain of each sub-band signal according to the amplitude of each sub-band signal and the audiogram. Among them, the audiogram is a table that can vividly "describe" the user's hearing condition. It records the user's hearing situation and is one of the important ways to judge whether the user's hearing is healthy. The audiogram is usually represented by a block diagram. Among them, the abscissa represents the frequency of the sound, which can also be called the pitch, and the unit is Hertz (Hz). From left to right, it means that the sound changes from low to high-pitched, and the range is usually 250 Hz to 8000 Hz. The ordinate represents the intensity, that is, the size of the sound, and the unit is decibel (dB). The larger the value, the louder the sound heard.

[0069] The process of obtaining the amplification gain of each sub-band signal is described above. The following specifically describes step S310.

[0070] In an exemplary embodiment of the present application, the first processing delay and the second audio signal can be determined by cascading multiple first neural networks. That is, the first processing delay and the second audio signal can be determined by multiple first neural networks connected in series, where the number of first neural networks is set according to the user's delay requirement for the finally output audio signal. In addition, each first neural network can adopt the same neural network, but their weight parameters are different from each other. In the present application, the first neural network can adopt, but is not limited to, RNN, CNN, etc. The following refers to Figure 5 to describe in detail the process of determining the first processing delay and the second audio signal by cascading multiple first neural networks.

[0071] Figure 5 is a flowchart showing the process of determining the first processing delay for amplifying the first audio signal to be processed and the second audio signal according to an exemplary embodiment of the present application.

[0072] In step S510, based on the first audio signal, by using each first neural network in the cascaded multiple first neural networks, the processed audio signal and its corresponding prediction probability are obtained. As described above, the number of first neural networks can be set according to the user's delay requirement for the finally output audio signal. For example, Figure 6 shows the cascaded structure of multiple first neural networks according to an exemplary embodiment of the present application. In Figure 6 it, the cascaded structure includes 3 first neural networks, but the present application is not limited thereto. If the user has a higher delay requirement for the finally output audio signal and hopes for more refined delay control, the number of cascaded first neural networks can be set larger.

[0073] Specifically, for the first first neural network among the multiple first neural networks, based on the target amplification gain and the first audio signal to be processed corresponding to this first neural network, this first neural network is used to obtain the processed audio signal and its corresponding prediction probability. At this time, the first audio signal to be processed corresponding to the first first neural network is the sampling signal of the input audio signal with a predetermined time length at time t, and the target amplification gain is the amplification gain corresponding to the first audio signal to be processed among the multiple amplification gains of the multiple sub-band signals calculated by the gain calculation module. As Figure 6As shown in [figure], assume that the first audio signal to be processed (i.e., the sampling signal of the input audio signal with a predetermined time length at time t = 1) is x(1). Then, in addition to the sampling signal x(1), the input of the first first neural network also includes the amplification gain corresponding to the sampling signal x(1) among the multiple amplification gains of the multiple subband signals. Based on these two input quantities, the first first neural network can obtain the processed audio signal and its corresponding prediction probability, that is, predict the prediction signal y(1) corresponding to the sampling signal x(1) at time t and the prediction probability p corresponding to the prediction signal y(1).

[0074] For each of the other first neural networks among the multiple first neural networks, based on at least one of the following, use the first neural network to obtain the processed audio signal and its corresponding prediction probability: the processed audio signal obtained by the previous first neural network, the output features of at least one network layer except the last network layer in the previous first neural network, and the audio signal in the first audio signal to be processed corresponding to this first neural network except the first audio signal to be processed corresponding to the previous first neural network. As Figure 6 As shown in [figure], for the second first neural network, the first audio signal to be processed corresponding to it includes the sampling signal x(1) of the input audio signal with a predetermined time length at time t = 1 and the sampling signal x(2) at time t + 1. At this time, the input of the second first neural network includes the sampling signal x(2) at time t + 1, and also includes the processed audio signal (i.e., the prediction signal) obtained by the first first neural network, and the output features f(1) of at least one network layer except the last network layer in the first first neural network (which can also be called "intermediate layer features"). Similarly, for the third first neural network, the first audio signal to be processed corresponding to it includes the sampling signal x(1) of the input audio signal with a predetermined time length at time t = 1, the sampling signal x(2) at time t + 1, and the sampling signal x(3) at time t + 2. At this time, the input of the third first neural network includes the sampling signal x(3) at time t + 2, and also includes the processed audio signal (i.e., the prediction signal) obtained by the second first neural network, and the output features f(2) of at least one network layer except the last network layer in the second first neural network. In the above description, the first neural networks cascaded and located behind the first first neural network all use the intermediate layer features of the previous first neural network (i.e., the output features of at least one network layer except the last network layer). These intermediate layer features imply the information of historical data, avoid the trouble of recalculating from scratch, and save computing resources.

[0075] In step S520, determine the first processing delay and the second audio signal based on the prediction probability.

[0076] As described above, multiple first neural networks are connected and operated in a cascaded structure (i.e., connected and operated in a serial manner). Therefore, in step S510, when a processed audio signal and its corresponding prediction probability are obtained through a first neural network, it is necessary to determine whether the currently obtained processed audio signal can be used as the second audio signal according to the prediction probability. If it cannot be used as the second audio signal, then in a similar manner, the processed audio signal and the prediction probability corresponding to the processed audio signal are obtained through the next first neural network in the cascade; if it can be used as the second audio signal, the operation can be ended in advance, that is, the next first neural network in the cascade is no longer used to perform the prediction.

[0077] Specifically, the step of determining the first processing delay and the second audio signal based on the prediction probability may include: for each first neural network among the multiple first neural networks, if the processing delay corresponding to the first neural network is determined as the first processing delay based on the prediction probability obtained through the first neural network, then the audio signal obtained through the first neural network is determined as the second audio signal. If the processing delay corresponding to the first neural network is not determined as the first processing delay based on the prediction probability obtained through the first neural network, then the first processing delay and the second audio signal are determined according to the prediction probability obtained through the next first neural network. In addition, the step of determining the first processing delay and the second audio signal based on the prediction probability may further include: for each first neural network among the multiple first neural networks, if the processing delay corresponding to the first neural network is not determined as the first processing delay based on the prediction probability obtained through the first neural network, then the processed audio signal obtained through the first neural network and the output features of at least one network layer except the last network layer in the first neural network are input into the next first neural network.

[0078] In an exemplary embodiment of the present application, it is possible to determine whether to determine the processing delay corresponding to the first neural network as the first processing delay by comparing the prediction probability obtained by a first neural network with a first predetermined threshold, and determine the corresponding processed audio signal as the second audio signal, where the first predetermined threshold can be set according to experience based on the actual situation. Specifically, the steps of determining the first processing delay and the second audio signal based on the prediction probability may include: for each of the multiple first neural networks, if the prediction probability obtained by the first neural network is greater than the first predetermined threshold, then determine the processing delay corresponding to the first neural network as the first processing delay, and determine the processed audio signal obtained by the first neural network as the second audio signal; otherwise, compare the prediction probability obtained by the next first neural network with the first predetermined threshold, and determine whether to determine the processing delay corresponding to the next first neural network as the first processing delay according to the comparison result, and determine the processed audio signal obtained by the next first neural network as the second audio signal.

[0079] As Figure 6 shown, when the prediction probability p of the amplified signal corresponding to the first audio signal to be processed (i.e., the sampling signal at time t = 1) obtained by the first first neural network is greater than the first predetermined threshold T, determine the processing delay corresponding to the first first sub - neural network as the first processing delay, and determine the amplified signal obtained by the first first neural network as the second audio signal y(1).

[0080] When the corresponding prediction probability p of the amplified signal y(1) obtained by the first first neural network based on the corresponding first audio signal to be processed (i.e., the sampling signal at time t = 1) is less than or equal to the first predetermined threshold T, the processed audio signal obtained by the first first neural network (i.e., the amplified signal), the output features f(1) of at least one network layer in the first first neural network except the last network layer (i.e., the intermediate layer features), and the first audio signal to be processed corresponding to the second first neural network (i.e., the sampling signal x(1) at time t = 1 and the sampling signal x(2) at time t + 1) except the first audio signal to be processed corresponding to the first first neural network (i.e., the sampling signal x(1) at time t = 1) (i.e., the sampling signal x(2) at time t + 1) are input into the second first neural network, and then the processed audio signal and its corresponding prediction probability are obtained by the second first neural network, that is, the amplified signal corresponding to the sampling signal x(1) at time t = 1 and the prediction probability p corresponding to the amplified signal. When the prediction probability p obtained by the second first neural network is greater than the first predetermined threshold T, the processing delay corresponding to the second first neural network is determined as the first processing delay, and the amplified signal obtained by the second first neural network is determined as the second audio signal y(1).

[0081] When the prediction probability p obtained by the second first neural network is less than or equal to the first predetermined threshold T, the processed audio signal (i.e., the amplified signal) obtained by the second first neural network, the output feature f(2) of at least one network layer except the last network layer in the second first neural network, and the first audio signal to be processed corresponding to the third first neural network (i.e., the sampling signal x(1) at time t = 1, the sampling signal x(2) at time t + 1, and the sampling signal x(3) at time t + 2) except the first audio signal to be processed corresponding to the second first neural network (i.e., the sampling signal x(1) at time t = 1 and the sampling signal x(2) at time t + 1) (i.e., the sampling signal x(3) at time t + 2) are input into the third first neural network. Then, the processed audio signal and its corresponding prediction probability are obtained by the third first neural network, that is, the amplified signal corresponding to the sampling signal x(1) at time t = 1 and the prediction probability p corresponding to the amplified signal. Then, it is determined whether the prediction probability p obtained by the third first neural network is greater than the first predetermined threshold T. When the prediction probability p obtained by the third first neural network is greater than the first predetermined threshold T, the processing delay corresponding to the third first neural network is determined as the first processing delay, and the amplified signal obtained by the third first neural network is determined as the second audio signal y(1). When the prediction probability p obtained by the third first neural network is less than or equal to the first predetermined threshold T, the processing delay corresponding to the third first neural network is also determined as the first processing delay, and the amplified signal obtained by the third first neural network is determined as the second audio signal y(1).

[0082] That is to say, when the prediction probability obtained by each of the multiple first neural networks is less than or equal to the first predetermined threshold, the processing delay corresponding to the last first neural network is determined as the first processing delay, and the processed audio signal obtained by the last first neural network is determined as the second audio signal.

[0083] In the above reference Figure 5 In the process of determining the first processing delay and the second audio signal for amplifying the first audio signal to be processed according to an exemplary embodiment of the present application described above, the processing delay corresponding to each first neural network increases sequentially in the order of cascading of the multiple first neural networks, and the time length of the first audio signal to be processed corresponding to each first neural network corresponds to the processing delay corresponding to the corresponding first neural network. In the present application, among the cascaded multiple first neural networks, the processing delay corresponding to a first neural network refers to the time spent by the first neural network and all the first neural networks before the first neural network in processing the audio signal.

[0084] In the process described above, the time delay of the output signal is adaptively adjusted using the prediction probability corresponding to the amplified signal. Taking 16-bit data as an example, the dynamic range of the signal amplitude is -32768 to 32767. Correspondingly, the amplitude of the amplified signal is divided into 65536 categories. The first neural network can predict the prediction probability of each of these 65536 categories. The maximum prediction probability among the 65536 predicted prediction probabilities is output as the prediction probability corresponding to the amplified signal, and the amplitude of the category corresponding to this prediction probability is output as the amplified signal (i.e., the amplitude of this amplified signal). For example, assume that a certain first neural network among multiple first neural networks outputs a vector of [1, 63356], and this vector represents the prediction probability of each category. If the prediction probability of the first category is the largest at this time, for example, 0.9, then the prediction probability of the amplified signal finally output by this first neural network is 0.9, and at the same time, the amplified signal corresponding to this prediction probability of 0.9 can be obtained (i.e., the amplitude corresponding to the first category).

[0085] In the above reference Figure 5 and Figure 6 In the process of determining the first processing time delay for amplifying the first audio signal to be processed and the second audio signal according to an exemplary embodiment of the present application described above, a plurality of cascaded first neural networks are used to determine the first processing time delay and the second audio signal. However, the present application is not limited thereto. In another exemplary embodiment of the present application, a plurality of mutually independent fourth neural networks can be used to determine the first processing time delay for amplifying the first audio signal to be processed and the second audio signal. The following will refer to Figure 7 for a detailed description of this.

[0086] Figure 7 is a flowchart showing the process of determining the first processing time delay for amplifying the first audio signal to be processed and the second audio signal according to another exemplary implementation of the present application.

[0087] In step S710, by using each of the plurality of fourth neural networks, a processed audio signal and its corresponding prediction probability are respectively obtained based on the first audio signal to be processed corresponding to the fourth neural network and the target amplification gain, where each fourth neural network corresponds to a different processing time delay. As described above, the number of fourth neural networks can be set according to the user's time delay requirement for the finally output audio signal. For example, Figure 8 shows an example structure of a plurality of fourth neural networks according to an exemplary embodiment of the present application. In Figure 8Among them, three mutually independent fourth neural networks are shown. However, the present application is not limited thereto. If the user has higher requirements for the time delay of the finally output audio signal and hopes for more refined time delay control, the number of fourth neural networks can be set larger.

[0088] Specifically, the input of the first fourth neural network includes the first audio signal to be processed corresponding to the first fourth neural network and the target amplification gain. As Figure 8 shown in the figure, assuming that the sampled signal of the input audio signal with a predetermined time length at time t = 1 is x(1), the input of the first fourth neural network includes the sampled signal x(1) and the target amplification gain corresponding to the sampled signal x(1) among the multiple amplification gains of multiple subbands. Based on these two input quantities, the first fourth neural network can predict the processed audio signal and its corresponding prediction probability, that is, the amplified signal corresponding to the sampled signal x(1) at time t and the prediction probability p corresponding to this amplified signal.

[0089] For the second fourth neural network, its input includes the first audio signal to be processed corresponding to the second fourth neural network and the target amplification gain. As Figure 8 shown in the figure, the input of the second fourth neural network includes two sampled signals x(1) and x(2) of the input audio signal with a predetermined time length from time t to time t + 1, and the target amplification gain corresponding to the sampled signal x(1) among the multiple amplification gains of multiple subbands. Based on these input quantities, the second fourth neural network can predict the processed audio signal and its corresponding prediction probability, that is, the amplified signal corresponding to the sampled signal x(1) at time t and the prediction probability p corresponding to this amplified signal.

[0090] Similarly, for the third fourth neural network, its input includes the first audio signal to be processed corresponding to the third fourth neural network and the target amplification gain. As Figure 8 shown in the figure, the input of the third fourth neural network includes three sampled signals x(1), x(2) and x(3) of the input audio signal with a predetermined time length from time t to time t + 2, and the target amplification gain corresponding to the sampled signal x(1) among the multiple amplification gains of multiple subbands. Based on these input quantities, the third fourth neural network can predict the processed audio signal and its corresponding prediction probability, that is, the amplified signal y(1) corresponding to the sampled signal x(1) at time t and the prediction probability p corresponding to this amplified signal y(1).

[0091] In the above description, multiple fourth neural networks can run simultaneously in parallel, and then each fourth neural network respectively obtains the amplified signal corresponding to the sampled signal at time t and the corresponding prediction probability.

[0092] At step S720, determine the first processing delay and the second audio signal based on the predicted probability.

[0093] As described above, the multiple fourth neural networks operate independently of each other. Therefore, at step S710, each fourth neural network will obtain an amplified signal corresponding to the sampled signal at time t and a predicted probability corresponding to the amplified signal. In this case, it is necessary to determine the second audio signal according to the predicted probability obtained by each fourth neural network.

[0094] Specifically, if the predicted probability obtained by at least one fourth neural network is greater than a second predetermined threshold, then determine the audio signal obtained by the fourth neural network corresponding to the lowest processing delay in the at least one fourth neural network as the second audio signal, and determine the processing delay corresponding to this fourth neural network as the first processing delay. Among them, the second predetermined threshold can be set according to experience according to the actual situation. In one embodiment, the second predetermined threshold may be the same as the first predetermined threshold. In another embodiment, the second predetermined threshold may be different from the first predetermined threshold. As Figure 8 shown, if only the first fourth neural network and the second fourth neural network among the 3 fourth neural networks obtain a predicted probability p greater than the second predetermined threshold, it is necessary to compare the processing delay corresponding to the first fourth neural network and the processing delay corresponding to the second fourth neural network, determine the lower processing delay of these two processing delays as the first processing delay, and determine the processed audio signal obtained by the fourth neural network corresponding to this processing delay as the second audio signal.

[0095] If the predicted probabilities obtained by the multiple fourth neural networks are all less than or equal to the second predetermined threshold, then determine the processed audio signal obtained by the fourth neural network corresponding to the highest processing delay in the multiple fourth sub-neural networks as the second audio signal, and determine the processing delay corresponding to this fourth neural network as the first processing delay. As Figure 8 shown, if the predicted probabilities p obtained by the 3 fourth neural networks are all less than or equal to the second predetermined threshold, it is necessary to compare the processing delay corresponding to each of the 3 fourth neural networks, select the highest processing delay from these 3 processing delays as the first processing delay, and determine the processed audio signal obtained by the fourth neural network corresponding to this processing delay as the second audio signal.

[0096] The above refers to Figure 5 the network structure formed by cascading multiple first neural networks described above and refers to Figure 7The described network structure formed by multiple independent fourth neural networks uses multiple neural networks to implement a multi-stage amplification network, enabling the time delay to increase as the number of neural network stages increases and being able to adaptively adjust the time delay of the second audio signal according to the prediction probability corresponding to the amplified signal. In addition, referring to Figure 5 the described network structure formed by cascading multiple first neural networks and referring to Figure 7 the described network structure formed by multiple independent fourth neural networks can be used for different types of hardware. Among them, the network structure formed by cascading multiple first neural networks can be used for hardware platforms with poor parallel computing capabilities, such as Advanced RISC Machine (ARM) and Intel CPUs. However, the network structure formed by multiple independent fourth neural networks can be used for hardware platforms with parallel computing capabilities, such as Neural Processing Unit (NPU), etc.

[0097] Return reference Figure 3 , in step S320, it is determined whether the first processing time delay is different from the processing time delay for amplifying the previous first audio signal. If the first processing time delay is different from the processing time delay for amplifying the previous first audio signal, then in step S330, a third audio signal is determined based on the second audio signal and the third audio signal is output.

[0098] Since the time delay may increase or decrease when predicting the first output signal using the first neural network or the fourth neural network as described above, the Figure 4 signal compensation module in it determines whether to adjust the second audio signal based on whether the processing time delays of the two consecutive amplification processes are the same, and then obtains the adjusted audio signal, that is, the third audio signal. First, the reason for the time delay change in the processing time delays of the two consecutive amplification processes is described below with reference to Figures 9 to 12 the description.

[0099] Specifically, Figure 9 is a schematic diagram showing an example where the first processing time delay increases when obtaining the second audio signal by applying the process described above with reference to Figure 5 . As shown in Figure 9As shown, for the sampling signal ① at time t = 1 and the sampling signal ② at time t = 2, the second audio signal 1 and the second audio signal 2 can be obtained through the first first neural network respectively. For the sampling signal ③ at time t = 3, the prediction probability obtained by the first first neural network is less than the first predetermined threshold. Therefore, it is necessary to wait for the sampling signal ④ at time t + 1 = 4 as the input of the second first neural network to perform the prediction operation. The prediction probability obtained by the second first neural network is also less than the first predetermined threshold. Therefore, it is necessary to further wait for the sampling signal ⑤ at time t + 2 = 5 as the input of the third first neural network to perform the prediction operation, and finally obtain the second audio signal 3. Compared with the second audio signal 1 and the second audio signal 2, in the process of obtaining the second audio signal 3, it is necessary to use the sampling signal ③ at time t = 3 as the input of the first first neural network to perform the prediction operation, then use the sampling signal ④ at time t + 1 = 4 as the input of the second first neural network to perform the prediction operation, and then use the sampling signal ⑤ at time t + 2 = 5 as the input of the third first neural network to perform the prediction operation to obtain the second audio signal 3. Therefore, the first processing delay for amplifying the first audio signal to be processed (i.e., the sampling signal ③ at time t = 3, the sampling signal ④ at time 4, and the sampling signal ⑤ at time 5, which can also be called "the samples processed once") is different from (greater than) the first processing delay for amplifying the sampling signal ② at time t = 2, and thus there will be a silent segment between the second audio signal 2 and the second audio signal 3, that is, there is discontinuity.

[0100] Therefore, for Figure 9 the situation where the first processing delay increases in the process of obtaining the second audio signal 3 as shown, the present application generates a filling signal through a neural network based on the previously output audio signal (i.e., the historical audio signal) and the first audio signal to be processed, and fills the silent segment (as shown in Figure 10 (b) of Figure 10 ) reflected in the time domain waveform of the second audio signal based on the generated filling signal, and correspondingly eliminates the noise covering the entire frequency band (as shown in Figure 10 (c) of

[0101] Figure 11 ), and obtains a harmonic structure close to the original harmonic structure of the audio signal (as shown in Figure 5 (a) of Figure 11 ).

[0101] Figure 11 FIG. Figure 5 shows an example of the reduction of the first processing delay when obtaining the second output signal by applying the process described above with reference to Figure 11 As shown in Figure 11 , the sampling signals ①, ②, and ③ all need to be delayed by 4 sampling signals to obtain the second audio signals 1, 2, and 3 through the third first neural network. The sampling signal ⑦ is delayed by 1 sampling signal to obtain the second audio signal 7 through the first first neural network. The second audio signal appears in the time domain waveformFigure 12 The situation shown in (b) of

[0102] For this reason, in view of the situation where the first processing delay is reduced during the process of obtaining the second audio signal 7 shown in Figure 11 , this application proposes to correct the obtained second audio signal through a neural network based on the previously output audio signal (i.e., the historical audio signal) to obtain a corrected audio signal, thereby eliminating the numerical discontinuity in the time-domain waveform of the second audio signal (as shown in (b) of Figure 12 ), and correspondingly eliminating the noise covering the entire frequency band that appears in the spectrogram (as shown in (c) of Figure 12 ), and obtaining a harmonic structure close to the original harmonic structure of the audio signal (as shown in (a) of Figure 12 ).

[0103] Therefore, it is necessary to determine whether to adjust (also referred to as "compensate") the second audio signal by comparing the first processing delay with the processing delay for amplifying the previous first audio signal to obtain a third output signal. The following will refer to Figure 13 to describe in detail the process of determining the third audio signal based on the second audio signal when the first processing delay for amplifying the first audio signal to be processed is different from the processing delay for amplifying the previous first audio signal.

[0104] In step S1310, it is judged whether the first processing delay D(t) for amplifying the first audio signal to be processed is greater than the processing delay D(t - 1) for amplifying the previous first audio signal. If the first processing delay is greater than the processing delay for amplifying the previous first audio signal, then in step S1320, at least one fourth audio signal to be filled before the third audio signal is obtained through a second neural network based on the audio signal output for the previous first audio signal and the first audio signal to be processed, and the at least one fourth audio signal is output.

[0105] Specifically, according to the above reference to Figure 9As can be seen from the described process, when the predicted probability obtained based on a certain first neural network does not determine the processing delay corresponding to the first neural network as the first processing delay, the first processing delay will be greater than the processing delay for amplifying the previous first audio signal. As a result, there will be a silent segment between two consecutive second audio signals. In this regard, the present application can fill this silent segment with at least one fourth audio signal generated by the above step S1320. Therefore, in the present application, for each first neural network among multiple first neural networks, if the predicted probability obtained based on the first neural network does not determine the processing delay corresponding to the first neural network as the first processing delay, then based on at least one of the following, the second neural network determines the fourth audio signal that needs to be filled before the third audio signal: the previously output audio signal, the output features (i.e., intermediate layer features) of at least one network layer in the first neural network except for the last network layer, and the first audio signal to be processed corresponding to the first first neural network among the multiple first neural networks; and outputs the fourth audio signal. For example, as Figure 14As shown, for the sampled signal ③ at time t = 3, the prediction probability obtained based on the first first neural network does not determine the processing delay corresponding to this first first neural network as the first processing delay. Therefore, based on the previously output audio signal (i.e., the second audio signal 2), the output features of at least one network layer in the first first neural network except for the last network layer, and the first audio signal to be processed corresponding to the first first neural network (i.e., the sampled signal ③ at time t = 3), the first fourth audio signal (i.e., the first padding signal) is determined through the second neural network, and this first fourth audio signal is output. In addition, for the first audio signal to be processed corresponding to the second first neural network (i.e., the sampled signal ③ at time t = 3 and the sampled signal ④ at time t = 3 + 1), the prediction probability obtained based on the second first neural network does not determine the processing delay corresponding to this second first neural network as the first processing delay. Therefore, based on the previously output audio signal (i.e., the first fourth audio signal), the output features of at least one network layer in the second first neural network except for the last network layer, and the first audio signal to be processed corresponding to the first first neural network (i.e., the sampled signal ③ at time t = 3), the second fourth audio signal (i.e., the second padding signal) is determined through the second neural network, and this second fourth audio signal is output. Since for the first audio signal to be processed corresponding to the third first neural network (i.e., the sampled signal ③ at time t = 3, the sampled signal ④ at time t = 3 + 1, and the sampled signal ⑤ at time t = 3 + 2), the prediction probability obtained based on the third first neural network determines the processing delay corresponding to this third first neural network as the first processing delay, no further fourth audio signal is generated. At this time, by successively outputting the two fourth audio signals generated above, the filling of the silent segment can be achieved, thereby eliminating the discontinuity of the output signal.

[0106] After determining and outputting at least one fourth audio signal, the step of determining the third audio signal based on the second audio signal may include: based on the output fourth audio signal, the second audio signal is corrected through the third neural network to obtain the third audio signal. That is, after performing step S1320, it proceeds to step S1330. Based on the last fourth audio signal among the at least one fourth audio signals, the second audio signal is corrected through the third neural network (which can also be called the correction model) to eliminate the numerical discontinuity between this last fourth audio signal and the second audio signal, thereby eliminating the possible noise in hearing and ensuring the smoothness of hearing.

[0107] For example, as Figure 14As shown in the figure, based on the second fourth audio signal, the third neural network corrects the second audio signal obtained for the first audio signal to be processed corresponding to the third first neural network (i.e., the sampling signal ③ at time t = 3, the sampling signal ④ at time t = 3 + 1, and the sampling signal ⑤ at time t = 3 + 2), so as to obtain the third audio signal 3.

[0108] If it is determined in step S1310 that the first processing delay is less than the processing delay for amplifying the previous first audio signal, then in step S1330, based on the previously output audio signal, the third neural network corrects the second audio signal to obtain the third audio signal. The above-described steps S1310 to S1330 can be performed by Figure 4 the signal compensation module shown in the figure.

[0109] Specifically, when the situation described above with reference to Figure 11 occurs, the first processing delay will be less than the processing delay for amplifying the previous first audio signal, and then a problem of numerical discontinuity will occur. In this regard, the present application can correct the second audio signal based on the previously output audio signal through the third neural network, and then obtain the third audio signal.

[0110] For example, as Figure 15 shown in the figure, first, based on the sampling signals ①, ②, ③, and ④, the second audio signal 1 is obtained through the first four first neural networks in the cascaded multiple first neural networks. Based on the sampling signals ②, ③, ④, and ⑤, the second audio signal 2 is obtained through the first four first neural networks in the cascaded multiple first neural networks. Then, based on the sampling signals ③, ④, ⑤, and ⑥, the second audio signal 3 is obtained through the first four first neural networks in the cascaded multiple first neural networks. After that, based on the sampling signal ⑦, the second audio signal (i.e., Figure 15 the amplified signal of the sampling signal ⑦ in the figure) is obtained through the first first neural network in the cascaded multiple first neural networks. At this time, the problem of numerical discontinuity caused by the reduction of the processing delay described above with reference to Figure 11 will occur. In this case, the second audio signal 3 and the amplified signal of the sampling signal ⑦ are input into the third neural network. The third neural network corrects the amplified signal of the sampling signal ⑦ based on the second audio signal 3, and then obtains the third audio signal 7, ensuring the numerical continuity of the output signal and the smoothness of hearing.

[0111] The second neural network and the third neural network mentioned in the above description may adopt a Generative Adversarial Network (GAN). Among them, the generators adopted by these two neural networks may be, but are not limited to, CNN, RNN, etc., but the weight parameters of the generators adopted by these two neural networks are different. In addition, these two neural networks are obtained through offline training, and the discriminators adopted in the training stage may be the same.

[0112] In addition, return to reference Figure 3 , if it is determined in step S320 that the first processing delay is the same as the processing delay for amplifying the previous first audio signal, then in step S340, the second audio signal is output. That is to say, if the processing delays of the two consecutive amplification processes are the same, the second audio signal can be directly output. The above steps S320 to S340 can be executed by Figure 4 the signal compensation module shown in

[0113] The following gives the test results in the case of adjusting the second audio signal and / or generating the fourth audio signal and in the case of not adjusting the second audio signal and generating the fourth audio signal. Specifically, the test results in the case of not adjusting the second audio signal and generating the fourth audio signal are used as the benchmark control group Benchmark. By using a preset first predetermined threshold to determine whether to increase or decrease the processing delay for amplifying the first processing signal to be processed, the processing delay is preset to the following categories: 0.5ms, 1ms, 1.5ms, 2ms, 2.5ms, 3ms. In the case of adjusting the second audio signal and / or generating the fourth audio signal, both the second neural network and the third neural network adopt a Dual-Path RNN (DPRNN) to implement the adjustment of the second audio signal and the generation of the fourth audio signal (the generators in the DPRNNs of the second neural network and the third neural network adopt different weight parameters). In the test, 2000 test data from the standard test dataset WSJ0 are used for testing, and the DNSMOS, a commonly used method for speech quality evaluation without reference, is used to evaluate the speech quality. The higher the DNSMOS value, the better the speech quality. It can be seen from the results that the DNSMOS of the method proposed in this application is 9.4% higher than that of the benchmark control group Benchmark, and the average delay is reduced to 1.16ms, that is, reduced by 61.3%. Obviously, the method proposed in this application can make the sound more natural even when the average delay is reduced.

[0114] The detailed process of the method executed by the electronic device has been described above. This method can be applied to various scenarios. For example, in the scenario of hearing a distant sound (such as listening to a report by a keynote speaker in a very large venue or chatting with a person at a distance), that is, by amplifying the sound to help the user hear the clear voice of the distant speaker. The following will refer to Figure 16 for a brief description.

[0115] Figure 16 is a schematic process diagram showing the process of the method executed by the electronic device according to an exemplary embodiment of the present application.

[0116] First, the electronic device (such as headphones) picks up an audio signal (including, for example, the user's own voice and / or ambient sound); then, the user's own audiogram is imported; the amplification gain of each sub-band signal is calculated using a fitting formula based on the user's audiogram, etc.; the picked-up audio signal is processed (such as amplified) by an adaptive delay WDRC module according to the calculated amplification gain of each sub-band signal to obtain a processed audio signal. Specifically, the adaptive delay WDRC module first predicts a second audio signal based on the audio signal and the amplification gain of each sub-band signal, and then, according to the change (such as increase or decrease) of the processing delay of the second audio signal, generates a third audio signal based on the second audio signal. Among them, when the processing delay increases, at least one fourth audio signal (i.e., a filling signal) is obtained through a second neural network based on the audio signal output for the previous first audio signal and the current first audio signal to be processed, the at least one fourth audio signal is output, and based on the output fourth audio signal, the second audio signal is corrected through a third neural network to obtain the third audio signal; when the processing delay decreases, the second audio signal is corrected through the third neural network based on the previously output audio signal to obtain the third audio signal, and the processed third audio signal is output through a speaker; when the processing delay does not change, the second audio signal is directly output. Through the above processing, the user can clearly hear the distant sound.

[0117] An electronic device is also provided in an embodiment of the present disclosure. The electronic device includes at least one processor. Optionally, it may further include at least one transceiver and / or at least one memory coupled to the at least one processor. The at least one processor is configured to execute the steps of the method provided in any optional embodiment of the present disclosure.

[0118] Figure 17 shows a schematic structural diagram of an electronic device applicable to an embodiment of the present invention, as Figure 17 shown, Figure 17The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as being connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between this electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, each of the processor 4001, the memory 4003, and the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present disclosure. Optionally, this electronic device may be a first network node, a second network node, or a third network node.

[0119] The processor 4001 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the content disclosed in the present disclosure. The processor 4001 may also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0120] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard structure) bus, etc. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 17 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0121] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited here.

[0122] The memory 4003 is used to store the computer programs or executable instructions for implementing the embodiments of the present disclosure and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the computer programs or executable instructions stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.

[0123] The embodiments of the present disclosure provide a computer-readable storage medium with computer programs or instructions stored thereon. When the computer programs or instructions are executed by at least one processor, they can execute or implement the steps and corresponding contents of the foregoing method embodiments.

[0124] The embodiments of the present disclosure also provide a computer program product, including a computer program. When the computer program is executed by a processor, it can implement the steps and corresponding contents of the foregoing method embodiments.

[0125] Terms such as "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims, and the above-mentioned drawings of the present disclosure are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than the illustrated or textually described order.

[0126] It should be understood that although the flowcharts of the embodiments of the present disclosure indicate various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present disclosure do not limit this.

[0127] The above text and drawings are provided only as examples to assist the reader in understanding the present disclosure. They are not intended and should not be construed as limiting the scope of the present disclosure in any way. Although certain embodiments and examples have been provided, it will be apparent to those skilled in the art based on the content disclosed herein that changes can be made to the illustrated embodiments and examples without departing from the scope of the present disclosure, and other similar implementation means based on the technical idea of the present disclosure can be adopted, which also fall within the protection scope of the embodiments of the present disclosure.

Claims

1. A method performed by an electronic device, comprising: Determining a first processing delay for amplifying a first audio signal to be processed, and a second audio signal obtained through the amplification processing; If the first processing delay is different from the processing delay for amplifying a previous first audio signal, determining a third audio signal based on the second audio signal; and Outputting the third audio signal.

2. The method according to claim 1, wherein The step of determining a first processing delay for amplifying a first audio signal to be processed, and a second audio signal obtained through the amplification processing, comprises: Based on the first audio signal, obtaining a processed audio signal and its corresponding prediction probability by using each first neural network in a plurality of cascaded first neural networks; Determining the first processing delay and the second audio signal based on the prediction probability.

3. The method according to claim 2, wherein The step of determining the first processing delay and the second audio signal based on the prediction probability comprises, for each first neural network in the plurality of first neural networks: If the prediction probability obtained based on the first neural network determines the processing delay corresponding to the first neural network as the first processing delay, determining the audio signal obtained through the first neural network as the second audio signal; If the prediction probability obtained based on the first neural network does not determine the processing delay corresponding to the first neural network as the first processing delay, determining the first processing delay and the second audio signal according to the prediction probability obtained through the next first neural network.

4. The method according to claim 3, wherein The step of determining the first processing delay and the second audio signal based on the prediction probability further comprises: For each first neural network in the plurality of first neural networks, if the prediction probability obtained based on the first neural network does not determine the processing delay corresponding to the first neural network as the first processing delay, inputting the processed audio signal obtained through the first neural network, and output features of at least one network layer other than the last network layer in the first neural network, into the next first neural network.

5. The method according to claim 4, wherein The step of obtaining a processed audio signal and its corresponding prediction probability by using each first neural network in a plurality of cascaded first neural networks based on the first audio signal, comprises: For the first first neural network in the plurality of first neural networks, based on the target amplification gain and the first audio signal to be processed corresponding to the first neural network, using the first neural network to obtain a processed audio signal and its corresponding prediction probability; For each of the other first neural networks in the plurality of first neural networks, based on at least one of the following, using the first neural network to obtain a processed audio signal and its corresponding prediction probability: the processed audio signal obtained through the previous first neural network, the output features of the previous first neural network, and the audio signal other than the first audio signal to be processed corresponding to the previous first neural network in the first audio signal to be processed corresponding to the first neural network.

6. The method according to claim 2, wherein, The step of determining the first processing delay and the second audio signal based on the prediction probability comprises: For each of the multiple first neural networks, if the prediction probability obtained by the first neural network is greater than a first predetermined threshold, the processing delay corresponding to the first neural network is determined as the first processing delay, and the processed audio signal obtained through the first neural network is determined as the second audio signal.

7. The method according to claim 2, wherein: The processing delays corresponding to each of the first neural networks increase sequentially in the order of cascading of the multiple first neural networks; The time length of the first audio signal to be processed corresponding to each of the first neural networks corresponds to the processing delay corresponding to the corresponding first neural network.

8. The method according to claim 1, wherein The steps of determining the first processing delay for amplifying the first audio signal to be processed and the second audio signal obtained through the amplification processing include: By using each of the multiple fourth neural networks, respectively based on the first audio signal to be processed corresponding to the fourth neural network and the target amplification gain, obtaining the processed audio signal and its corresponding prediction probability, wherein each fourth neural network corresponds to a different processing delay; Determining the first processing delay and the second audio signal based on the prediction probability.

9. The method according to claim 8, wherein, The steps of determining the first processing delay and the second audio signal based on the prediction probability include: If the prediction probability obtained by at least one fourth neural network is greater than a second predetermined threshold, the audio signal obtained by the fourth neural network corresponding to the lowest processing delay among the at least one fourth neural network is determined as the second audio signal, and the processing delay corresponding to the fourth neural network is determined as the first processing delay; If the prediction probabilities obtained by the multiple fourth neural networks are all less than or equal to the second predetermined threshold, the audio signal obtained by the fourth neural network corresponding to the highest processing delay among the multiple fourth neural networks is determined as the second audio signal, and the processing delay corresponding to the fourth neural network is determined as the first processing delay.

10. The method according to any one of claims 5, 8, and 9, wherein The target amplification gain is obtained through the following operations: Dividing the audio signal of a preset time length into multiple subband signals; According to the user's hearing information, respectively calculating the amplification gain of each subband signal among the multiple subband signals; Selecting the amplification gain corresponding to the first audio signal to be processed from the calculated multiple amplification gains as the target amplification gain.

11. The method according to claim 1, wherein If the first processing delay is greater than the processing delay for amplifying the previous first audio signal, the method further includes: Based on the audio signal output for the previous first audio signal and the first audio signal to be processed, obtaining, through a second neural network, at least one fourth audio signal to be filled before the third audio signal; and Outputting the at least one fourth audio signal.

12. The method according to claim 11, wherein, The steps of determining the third audio signal based on the second audio signal include: Based on the output fourth audio signal, correcting the second audio signal through a third neural network to obtain the third audio signal.

13. The method according to claim 3, further includes: For each of the multiple first neural networks, if the prediction probability obtained based on the first neural network does not determine the processing delay corresponding to the first neural network as the first processing delay, then based on at least one of the following, a fourth audio signal to be filled before the third audio signal is determined by a second neural network: the previously output audio signal, the output features of at least one network layer in the first neural network except for the last network layer, and the first audio signal to be processed corresponding to the first first neural network among the multiple first neural networks; Output the fourth audio signal.

14. The method according to claim 1, wherein If the first processing delay is less than the processing delay for amplifying the previous first audio signal, the method further includes: Based on the previously output audio signal, the second audio signal is corrected by a third neural network to obtain the third audio signal.

15. The method according to claim 1, further including: If the first processing delay is the same as the processing delay for amplifying the previous first audio signal, output the second audio signal.

16. An electronic device, including: At least one processor; And At least one memory storing computer-executable instructions, wherein, when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the method according to any one of claims 1 to 15.

17. A computer-readable storage medium for storing instructions, wherein, When the instructions are run by at least one processor, the at least one processor is caused to execute the method according to any one of claims 1 to 15.