Audio super-resolution model training, audio super-resolution processing method, device and electronic equipment

Through Schrödinger Bridge Model Training and Distillation Optimization, the problem of low accuracy in generated audio in existing audio overscore processing is solved, and high-frequency details and fidelity are improved. It is suitable for edge computing and lightweight audio overscore processing on mobile terminals.

CN120319255BActive Publication Date: 2025-08-26BEIJING SHENGSHU TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510800704.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-08-26
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

In the existing audio oversubscription processing scheme, the audio generated by noise based on silent information is long, resulting in low-frequency partial signals being unfavorable for maintenance and the generated audio accuracy is low.

Method used

By obtaining multiple super-segment data pairs, using the Schrödinger Bridge model to establish a solutionable path, training the Schrödinger Bridge model, and distillation optimization, obtaining the distillation-optimized Schrödinger Bridge audio super-segment model, using the linear interpolation algorithm to generate prior information, keeping the low-frequency information unchanged, and focusing on the generation of high-frequency information.

Benefits of technology

The high-frequency details and fidelity in audio overscore processing are optimized, computing consumption is reduced, practical feasibility is enhanced in low-latency application scenarios, and the efficiency of audio overscore processing and fidelity of high-frequency details are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120319255B_ABST
    Figure CN120319255B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure disclose an audio super-resolution model training, audio super-resolution processing method, device and electronic device, including: obtaining multiple super-resolution data pairs, each super-resolution data pair including a low-resolution waveform signal and a high-resolution waveform signal representing the same audio; for each super-resolution data pair, using a Schrödinger bridge model to establish a solvable path from the low-resolution waveform signal to the high-resolution waveform signal; based on the solvable path, training the Schrödinger bridge model to obtain a trained Schrödinger bridge audio super-resolution model; using multiple super-resolution data pairs, performing at least one round of distillation optimization on the trained Schrödinger bridge audio super-resolution model to obtain a distillation-optimized Schrödinger bridge audio super-resolution model. The distillation-optimized Schrödinger bridge audio super-resolution model obtained by the present disclosure greatly reduces the computational consumption in the audio super-resolution processing process, thereby enhancing the Schrödinger bridge audio super-resolution model in low-latency application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of computer technology and artificial intelligence technology, and in particular to an audio super-resolution model training, an audio super-resolution processing method, a device, and an electronic device. Background Art

[0002] Speech signal processing is widely used in many fields, such as game development, smart homes, and autonomous driving. Speech super-resolution technology is a key branch of speech signal processing. Through speech super-resolution processing, high-sampling-rate speech signals can be generated from low-sampling-rate speech signals, thereby improving the quality and effectiveness of speech signals.

[0003] Related technologies use a diffusion model to gradually generate high-resolution audio from standard Gaussian noise, achieving the goal of generating audio from noise without any cues. However, this audio super-resolution solution generates audio from noise without any cues, resulting in a long trace and poor preservation of low-frequency signals, leading to low accuracy in the generated audio. Summary of the Invention

[0004] Embodiments of the present disclosure provide an audio super-resolution model training, an audio super-resolution processing method, an apparatus, and an electronic device.

[0005] According to a first aspect of an embodiment of the present disclosure, a method for training an audio super-resolution model is provided, the method comprising: obtaining a plurality of super-resolution data pairs, each super-resolution data pair comprising a low-resolution waveform signal and a high-resolution waveform signal representing the same audio; for each super-resolution data pair, establishing a solvable path from the low-resolution waveform signal to the high-resolution waveform signal using a Schrödinger bridge model; training the Schrödinger bridge model based on the solvable path to obtain a trained Schrödinger bridge audio super-resolution model; and performing at least one round of distillation optimization on the trained Schrödinger bridge audio super-resolution model using the plurality of super-resolution data pairs to obtain a distillation-optimized Schrödinger bridge audio super-resolution model.

[0006] According to the second aspect of the embodiment of the present disclosure, a method for audio super-resolution processing is provided, which includes: obtaining a low-resolution signal to be processed; using a linear interpolation algorithm to interpolate the low-resolution signal to be processed to obtain generated prior information; based on the generated prior information, using a distilled and optimized Schrödinger bridge audio super-resolution model to generate a high-resolution target waveform, and the distilled and optimized Schrödinger bridge audio super-resolution model keeps the low-frequency information in the generated prior information unchanged in each sampling step of generating the high-resolution target waveform.

[0007] According to a third aspect of an embodiment of the present disclosure, an audio super-resolution model training device is provided, comprising: a sample acquisition module for acquiring a plurality of super-resolution data pairs, each super-resolution data pair comprising a low-resolution waveform signal and a high-resolution waveform signal representing the same audio; a path generation module for establishing, for each super-resolution data pair, a solvable path from the low-resolution waveform signal to the high-resolution waveform signal using a Schrödinger bridge model; an audio super-resolution model training module for training the to-be-trained Schrödinger bridge model based on the solvable path to obtain a trained Schrödinger bridge audio super-resolution model; and a model distillation module for performing at least one round of distillation optimization on the trained Schrödinger bridge audio super-resolution model using a plurality of super-resolution data pairs to obtain a distilled-optimized Schrödinger bridge audio super-resolution model.

[0008] According to the fourth aspect of the embodiment of the present disclosure, an audio super-resolution processing device is provided, including: an acquisition module for acquiring a low-resolution signal to be processed; an adjustment module for interpolating the low-resolution signal to be processed using a linear interpolation algorithm to obtain generated prior information; a generation module for generating a high-resolution target waveform based on the generated prior information and using a distilled and optimized Schrödinger bridge audio super-resolution model, wherein the distilled and optimized Schrödinger bridge audio super-resolution model keeps the low-frequency information in the generated prior information unchanged in each sampling step of generating the high-resolution target waveform.

[0009] According to a fifth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which stores a computer program for executing the above-mentioned audio super-resolution processing method or audio super-resolution model training method.

[0010] According to the sixth aspect of an embodiment of the present disclosure, an electronic device is provided, which includes: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the above-mentioned audio super-resolution processing method or audio super-resolution model training method.

[0011] According to the seventh aspect of the embodiments of the present disclosure, a computer program product is provided, including computer program instructions. When the computer program instructions are executed by a processor, the above-mentioned audio super-resolution processing method or audio super-resolution model training method is implemented.

[0012] Based on the embodiments of the present disclosure, when training an audio super-resolution model, multiple super-resolution data pairs are first obtained, and the super-resolution data pairs include low-resolution waveform signals and high-resolution waveform signals representing unified audio. For each super-resolution data pair, a Schrödinger bridge model is used to establish a solvable path from the low-resolution waveform signal to the high-resolution waveform signal. Based on the solvable path, the Schrödinger bridge model is trained to obtain a trained Schrödinger bridge audio super-resolution model. Then, the above-mentioned multiple super-resolution data pairs are used to perform at least one round of distillation optimization on the trained Schrödinger bridge audio super-resolution model. The technical solution disclosed herein utilizes super-resolution data to train a Schrödinger bridge audio super-resolution model from a low-resolution waveform signal to a high-resolution waveform signal. The model has a small number of parameters and is lightweight. Moreover, through multiple rounds of distillation optimization, the sampling path of the Schrödinger bridge audio super-resolution model can be compressed, the reasoning efficiency can be improved, and the computational consumption in the audio super-resolution processing process can be greatly reduced, thereby enhancing the practical feasibility of the Schrödinger bridge audio super-resolution model in low-latency application scenarios, such as edge computing and mobile terminals. It is suitable for a wide range of practical application scenarios. In addition, the Schrödinger bridge audio super-resolution model can directly optimize the low-resolution audio in the audio waveform space to obtain high-resolution audio. While maintaining the low-frequency data, it can restore the audio details of the high-frequency part, which can greatly improve the efficiency of the audio super-resolution processing and the fidelity of the high-frequency details.

[0013] Based on the embodiment of the present disclosure, when audio super-resolution processing is required, a low-resolution signal to be processed is obtained; a linear interpolation algorithm is used to interpolate the low-resolution signal to be processed to obtain generation prior information; based on the generation prior information, a high-resolution target waveform is generated using the Schrödinger bridge audio super-resolution model optimized by distillation, and the Schrödinger bridge audio super-resolution model optimized by distillation keeps the low-frequency information in the generation prior information unchanged in each sampling step of generating the high-resolution target waveform. In the present disclosure, the Schrödinger bridge audio super-resolution model optimized by distillation is used to keep the low-frequency information in the generation prior information unchanged in each sampling step of generating the corresponding high-resolution target waveform based on the generation prior information, focusing on the generation of high-frequency information, optimizing the high-frequency details and fidelity in the audio super-resolution processing, making the high-resolution signal after super-resolution processing more natural, and performing audio super-resolution processing directly in the waveform space avoids the problems of cascade errors and data space loss.

[0014] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The above and other purposes, features, and advantages of the present disclosure will become more apparent through a more detailed description of the embodiments of the present disclosure in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and are not intended to limit the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.

[0016] The present application can be more clearly understood from the following detailed description with reference to the accompanying drawings, in which:

[0017] Figure 1 is a system framework diagram to which the present disclosure is applicable;

[0018] Figure 2 1 is a flow chart of an audio super-resolution processing method provided by an exemplary embodiment of the present disclosure;

[0019] Figure 3 is a training flow chart of a Schrödinger bridge model provided by an exemplary embodiment of the present disclosure;

[0020] Figure 4 is a flow chart of distillation optimization training of a Schrödinger bridge model provided by an exemplary embodiment of the present disclosure;

[0021] Figure 5 is a training flow chart of a Schrödinger bridge model provided by another exemplary embodiment of the present disclosure;

[0022] Figure 6 is a training flow chart of a Schrödinger bridge model provided by another exemplary embodiment of the present disclosure;

[0023] Figure 7 is a structural diagram of an audio super-resolution processing device provided by an exemplary embodiment of the present disclosure;

[0024] Figure 8 1 is a structural diagram of an audio super-resolution model training device provided by an exemplary embodiment of the present disclosure;

[0025] Figure 9 is a structural diagram of an audio super-resolution model training device provided by another exemplary embodiment of the present disclosure;

[0026] Figure 10 is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] Below, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.

[0028] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.

[0029] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meanings, nor do they indicate a necessary logical order between them.

[0030] It should also be understood that in the embodiments of the present disclosure, “a plurality of” may refer to two or more than two, and “at least one” may refer to one, two, or more than two.

[0031] It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.

[0032] In addition, the term "and / or" in this disclosure is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this disclosure generally indicates that the related objects are in an "or" relationship.

[0033] It should also be understood that the description of the various embodiments in this disclosure focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced with each other. For the sake of brevity, they will not be described one by one.

[0034] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0035] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.

[0036] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0037] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0038] The embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate in conjunction with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, among others.

[0039] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system-executable instructions (such as program modules) executed by the computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, and the like that perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are performed by remote processing devices linked via a network. In distributed cloud computing environments, program modules can be located on local or remote computer system storage media, including storage devices.

[0040] Overview of the Disclosure

[0041] Current audio super-resolution solutions typically use diffusion models to gradually generate high-resolution waveforms from standard Gaussian noise, achieving audio generation from noise without any cues. However, this noise-based audio generation process results in long traces and is not conducive to preserving low-frequency signals, resulting in low audio accuracy.

[0042] The present invention obtains generative prior information by interpolating a low-resolution signal, and then uses a distilled and optimized Schrödinger bridge audio super-resolution model to generate a high-resolution signal based on the generative prior information, keeping the low-frequency information in the generative prior information unchanged and focusing on the generation of high-frequency information, thereby optimizing the high-frequency details and fidelity in the audio super-resolution processing. In addition, the distilled and optimized Schrödinger bridge audio super-resolution model has fewer inference steps and lower inference efficiency, thereby reducing the computational cost of the audio super-resolution processing and enhancing the practical feasibility of the distilled and optimized Schrödinger bridge audio super-resolution model in low-latency application scenarios.

[0043] Exemplary Systems

[0044] Figure 1 An exemplary system architecture is shown to which the audio super-resolution processing method or audio super-resolution processing device according to the embodiments of the present disclosure can be applied.

[0045] like Figure 1As shown, the system architecture may include a terminal device 11, a network 12, an audio input device 13, and an audio output device 14. The network 12 is used to provide a medium for a communication link between the terminal device 11 and the audio input device 13. The network 12 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0046] A user can use a terminal device 11 to interact with an audio input device 13 and an audio output device 14 via a network 12 to receive or send audio, etc. Various communication client applications can be installed on the terminal device 11, such as applications for generating artificial intelligence (AI) content (such as AI videos or AI images), multimedia applications, search applications, web browser applications, shopping applications, instant messaging tools, etc.

[0047] The terminal device 11 can be any electronic device that can deploy the Schrödinger bridge model, including but not limited to mobile terminals such as mobile phones, laptops, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), etc., as well as fixed terminals such as digital televisions, desktop computers, smart home appliances, etc.

[0048] The audio input device 13 may be a device that provides various low-resolution audios. The audio output device 14 is a device that receives high-resolution audios.

[0049] It should be noted that the audio super-resolution processing method provided in the embodiment of the present disclosure can be executed by the terminal device 11. Accordingly, the audio super-resolution processing device can be set in the audio input device 13 or in the terminal device 11.

[0050] The audio super-resolution model training method provided in the embodiment of the present disclosure can be executed by the terminal device 11 or by another server, and after the Schrödinger bridge model is trained, the Schrödinger bridge model is deployed on the terminal device 11.

[0051] It should be understood that Figure 1 The number of terminal devices 11, network 12, audio input devices 13, and audio output devices 14 in the above description is merely illustrative. Any number of terminal devices 11, network 12, audio input devices 13, and audio output devices 14 may be provided as needed. For example, if audio super-resolution processing does not require remote processing, the above system architecture may include only the terminal devices 11, excluding the network and server.

[0052] Exemplary Methods

[0053] Figure 2 FIG. 1 is a flow chart of an audio super-resolution processing method provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to electronic devices such as Figure 1 On mobile devices, such as Figure 2 As shown, the following steps are included:

[0054] Step 201: Obtain a low-resolution signal to be processed.

[0055] The low-resolution signal to be processed is a low-resolution signal. For example, audio with a resolution between 2kHz and 16kHz is considered low-resolution audio. The low-resolution signal to be processed in this disclosure indicates an audio signal that requires super-resolution processing. Super-resolution processing indicates the process of upsampling an audio signal to a higher resolution, such as 48kHz.

[0056] Step 202: Use a linear interpolation algorithm to interpolate the low-resolution signal to be processed to generate prior information.

[0057] In this embodiment, the sampling rate of the low-resolution signal to be processed is increased to the target sampling rate through linear interpolation (e.g., sinc interpolation), so that the number of sampling points used to generate the prior information is the same as the number of sampling points used to generate the high-resolution target waveform. For example, interpolation is performed between any two sampling points of the low-resolution signal using a sine wave paradigm, thereby increasing the sampling rate and increasing the data space.

[0058] In step 203, based on the above-mentioned generated prior information, a high-resolution target waveform is generated using the distilled and optimized Schrödinger bridge audio super-resolution model, wherein the distilled and optimized Schrödinger bridge audio super-resolution model maintains the low-frequency information in the generated prior information unchanged in each sampling step of generating the high-resolution target waveform.

[0059] Among them, the high-resolution target waveform is a high-resolution waveform corresponding to the generated prior information generated by the distilled and optimized Schrödinger bridge audio super-resolution model based on the generated prior information, for example, an audio waveform with a resolution of 24kHz to 48kHz.

[0060] Among them, the distilled optimized Schrödinger bridge audio super-resolution model is used to indicate a model that can determine and remove the noise to be removed corresponding to each sampling step in the generated prior information based on the generated prior information, and thus obtain a high-resolution target waveform.

[0061] In the disclosed embodiment, the Schrödinger bridge audio super-resolution model after distillation optimization keeps the low-frequency information used to generate prior information unchanged at each sampling step when performing audio super-resolution processing, and predicts the super-resolution audio waveform, thereby realizing a sampling path from data to data and avoiding the diffusion model's generation process from noise to data. It can better utilize low-resolution audio as prior information for generating super-resolution audio, thereby improving the generation efficiency of super-resolution audio signals and the fidelity of high-frequency information.

[0062] In the present disclosure, the distilled optimized Schrödinger bridge audio super-resolution model can be obtained by performing at least one round of distillation optimization on the trained Schrödinger bridge audio super-resolution model obtained by the forward process training of the Schrödinger bridge model. It can be trained using multiple super-resolution data pairs. The training process can be found in Figure 3 The reverse process of the distillation-optimized Schrödinger bridge audio super-resolution model is used to indicate the process of gradually removing noise to obtain a high-resolution target waveform.

[0063] Through the above steps 201 to 203, when audio super-resolution processing is required, a low-resolution signal to be processed is obtained; a linear interpolation algorithm is used to interpolate the low-resolution signal to be processed to obtain generation prior information; based on the generation prior information, a high-resolution target waveform is generated using the Schrödinger bridge audio super-resolution model optimized by distillation, and the Schrödinger bridge audio super-resolution model optimized by distillation keeps the low-frequency information in the generation prior information unchanged in each sampling step of generating the high-resolution target waveform. In the present disclosure, the Schrödinger bridge audio super-resolution model optimized by distillation keeps the low-frequency information in the generation prior information unchanged in each sampling step of generating the corresponding high-resolution target waveform based on the generation prior information, focuses on the generation of high-frequency information, optimizes the high-frequency details and fidelity in the audio super-resolution processing, makes the high-resolution signal after super-resolution processing more natural, and directly performs audio super-resolution processing in the waveform space to avoid cascade errors and data space loss problems.

[0064] like Figure 3 As shown in FIG, the training process of the distillation-optimized Schrödinger bridge audio super-resolution model may include the following steps.

[0065] Step 301: Acquire multiple super-resolution data pairs, each super-resolution data pair including a low-resolution waveform signal and a high-resolution waveform signal representing the same audio.

[0066] Among them, the low-resolution waveform signal is the waveform of low-resolution audio, for example, the waveform of audio with a resolution of 2kHz; the high-resolution waveform signal is the waveform of high-resolution audio, for example, the waveform of audio with a resolution of 48kHz.

[0067] In this embodiment, the high-resolution waveform signal can be copied using a signal processing filter to obtain a low-resolution waveform signal corresponding to the high resolution. Therefore, the low-resolution waveform signal and the high-resolution waveform signal are audio waveforms of different resolutions representing the same audio.

[0068] In some optional implementations, after using a signal processing filter to replicate the high-resolution waveform signal to obtain a corresponding low-resolution waveform signal, a scaling factor can be used to amplify the low-resolution waveform signal and the high-resolution waveform signal in each super-resolution data pair to obtain an amplified low-resolution waveform signal and a high-resolution waveform signal. The amplified low-resolution waveform signal and the high-resolution waveform signal are used to train the Schrödinger bridge model. In this implementation, by amplifying the low-resolution waveform signal and the high-resolution waveform signal, the trained Schrödinger bridge model can better perceive the difference between the low-resolution waveform and the high-resolution waveform, ensuring that the loss function remains sensitive, thereby avoiding the problem of vanishing gradients during training, and helping to improve the training efficiency and stability of the model.

[0069] Step 302: For each super-resolved data pair, a solvable path from the low-resolution waveform signal to the high-resolution waveform signal is established using a Schrödinger bridge model.

[0070] Among them, the Schrödinger bridge model uses a stochastic differential equation to generate a solvable path, and the stochastic differential equation is defined in the form of an asymmetric noise scheduling strategy.

[0071] Specifically, see Figure 4 , which illustrates the process of using the Schrödinger bridge model to train the trained Schrödinger bridge audio super-resolution model. First, a forward stochastic differential equation (SDE) is defined, as shown in Equation (1); the stochastic differential equation is defined in the form of an asymmetric noise scheduling strategy, as shown in Equation (2).

[0072] Formula (1)

[0073] Formula (2)

[0074] In formula (1), is the drift term, is the diffusion coefficient, is a standard Wiener process, Used to represent a probability distribution, t is the sampling step, represents a random process, is the sampling step tThe waveform of the audio signal.

[0075] Formula (2) is used to determine the noise in the audio waveform at each sampling step, t is the sampling step, is the noise of each sampling step, and is a function of the Schrödinger bridge sampling step determined by predefined hyperparameters, represents the integration variable.

[0076] According to formula (2), the asymmetric noise scheduling strategy of the Schrödinger bridge model can be realized.

[0077] According to formula (1), we can construct a system that satisfies the boundary condition “p(0)= , p(T)= ", according to which the T The diffusion process of sampling steps gradually adds noise to the high-resolution waveform signal to obtain the final low-resolution waveform signal.

[0078] In the disclosed embodiment, the sampling step is used to indicate the discrete time unit of the evolution of the simulated system, and can be used to control the progress of the noise addition process of the Schrödinger bridge model. Usually, each sampling step corresponds to a specific noise. As the sampling step increases, the noise level gradually increases. After increasing to a certain noise level, the added noise will offset part of the noise in the audio, causing the noise in the audio to gradually decrease.

[0079] In the embodiment of the present disclosure, the trained Schrödinger bridge audio super-resolution model uses an asymmetric noise scheduling strategy to denoise the audio signal to be processed. The number of sampling steps for adding noise is different from the number of sampling steps for reducing noise, where the noise peak Located to the right of the noise curve, the number of sampling steps for noise addition is less than the number of sampling steps for noise reduction. Therefore, more sampling steps are used for noise reduction to obtain the super-resolution audio data generation process. Using this asymmetric noise scheduling strategy to denoise the audio signal can improve the super-resolution effect of the audio and the quality of the super-resolution audio.

[0080] The sampling step can be the time basis used by the Schrödinger bridge model to add noise to the high-resolution waveform during the forward pass, or it can be the time basis used by the Schrödinger bridge model to determine and remove noise from the low-resolution audio waveform during the reverse pass. The noise to be removed corresponding to each sampling step can refer to the noise to be removed from the audio waveform. In practical applications, each sampling step can correspond to a noise to be removed.

[0081] Step 303: Based on the solvable path, the Schrödinger bridge model is trained to obtain a trained Schrödinger bridge audio super-resolution model.

[0082] Among them, based on the solvable path, the Schrödinger bridge model loss and multi-scale auxiliary loss can be used to guide the training of the Schrödinger bridge model to obtain the trained Schrödinger bridge audio super-resolution model.

[0083] In the disclosed embodiment, the Schrödinger bridge model loss is used to characterize the error between the predicted audio waveform signal and the true high-frequency waveform signal of the Schrödinger bridge model at each sampling step, and the multi-scale auxiliary loss is used to characterize the multi-scale auxiliary loss of the Schrödinger bridge model at each sampling step.

[0084] In an embodiment of the present disclosure, the training process of the Schrödinger bridge model may include a pre-training process, a fine-tuning process and a distillation optimization process: based on a solvable path, the Schrödinger bridge model is pre-trained using the Schrödinger bridge model loss to obtain a pre-trained Schrödinger bridge model; then, the pre-trained Schrödinger bridge model is fine-tuned using the Schrödinger bridge model loss and the multi-scale auxiliary loss to finally obtain a trained Schrödinger bridge audio super-resolution model; and then the trained Schrödinger bridge audio super-resolution model is distilled and optimized using the distillation optimization method to obtain a distilled optimized Schrödinger bridge audio super-resolution model.

[0085] In the disclosed embodiment, the Schrödinger bridge model loss may be obtained based on the error between the predicted audio waveform signal and the true high-frequency waveform signal at each sampling step.

[0086] Specifically, in the process of using the Schrödinger bridge model to establish a solvable path from a low-resolution waveform signal to a high-resolution waveform signal, for the t For each sampling step, the Schrödinger bridge model is responsible for predicting the waveform state of the predicted audio waveform signal at the current sampling step based on the forward stochastic differential equation and the waveform state at the previous sampling step, and determining the Schrödinger bridge model loss based on the real high-frequency waveform signal, as shown in Equation (3).

[0087] Formula (3)

[0088] In formula (3), is the Schrödinger bridge model loss, is the sampling step predicted by the Schrödinger bridge model t The waveform, is the Schrödinger bridge model, is a low-resolution sample waveform, is a high-resolution sample waveform, E Express expectations, represents the Schrödinger bridge prior probability distribution, represents the Schrödinger bridge target probability distribution.

[0089] Through the loss function of the above formula (3), the error between the predicted audio waveform signal and the real high-frequency waveform signal at each sampling step can be obtained, that is, the Schrödinger bridge model loss. According to the Schrödinger bridge model loss, the forward stochastic differential equation of the Schrödinger bridge model is optimized to guide the training of the Schrödinger bridge model, so that the Schrödinger bridge model is at each sampling step. t Both have the ability to predict high-resolution waveforms.

[0090] Step 304 : Using multiple super-resolution data pairs, perform at least one round of distillation optimization on the trained Schrödinger bridge audio super-resolution model to obtain a distilled and optimized Schrödinger bridge audio super-resolution model.

[0091] In the disclosed embodiments, a trained Schrödinger bridge audio super-resolution model can be used as the initial teacher model and student model. The trained Schrödinger bridge audio super-resolution model is then distilled and optimized through rounds of distillation, reducing the number of inference sampling steps in the Schrödinger bridge audio super-resolution model in each distillation round. The number of inference sampling steps in each distillation round can be reduced by half, 2 / 3, 3 / 4, or the like.

[0092] Preferably, in order to improve the audio super-resolution effect of the distillation-optimized Schrödinger bridge audio super-resolution model, so that it can be as consistent as possible with the audio super-resolution effect of the trained Schrödinger bridge audio super-resolution model obtained in step 303, the number of inference sampling steps can be halved in each distillation round. After multiple distillation rounds of model distillation, the final distillation-optimized Schrödinger bridge audio super-resolution model is obtained, which can effectively avoid the problem of model quality degradation caused by excessive reduction of the number of sampling steps in a distillation round.

[0093] For example, if the number of sampling steps of the trained Schrödinger bridge audio super-resolution model is 32 steps, then after one round of distillation, the number of sampling steps can be reduced to 16 steps, and after another round of distillation, the number of sampling steps can be reduced to 8 steps; after another round of distillation, the number of sampling steps can be reduced to 4 steps.

[0094] In the disclosed embodiments, the number of sampling steps of the distillation-optimized Schrödinger bridge audio super-resolution model can be a pre-set value. Depending on the actual application scenario, the number of rounds of distillation optimization of the Schrödinger bridge audio super-resolution model can be different, and the corresponding number of sampling steps of the distillation-optimized Schrödinger bridge audio super-resolution model can be different. For example, in scenarios with high real-time requirements, the number of sampling steps of the distillation-optimized Schrödinger bridge audio super-resolution model can be relatively small, such as 4 steps; in scenarios with moderate real-time requirements, the number of sampling steps of the distillation-optimized Schrödinger bridge audio super-resolution model can be relatively large, such as 8 steps.

[0095] Based on the embodiments of the present disclosure, when training an audio super-resolution model, multiple super-resolution data pairs are first obtained, and the super-resolution data pairs include low-resolution waveform signals and high-resolution waveform signals representing unified audio. For each super-resolution data pair, a Schrödinger bridge model is used to establish a solvable path from the low-resolution waveform signal to the high-resolution waveform signal. Based on the solvable path, the Schrödinger bridge model is trained to obtain a trained Schrödinger bridge audio super-resolution model. Then, the above-mentioned multiple super-resolution data pairs are used to perform at least one round of distillation optimization on the trained Schrödinger bridge audio super-resolution model. The technical solution disclosed in the present invention utilizes super-resolution data to train a Schrödinger bridge audio super-resolution model from a low-resolution waveform signal to a high-resolution waveform signal. The model has a small number of parameters and a lightweight characteristic. Moreover, through multiple rounds of distillation optimization, the sampling path of the Schrödinger bridge audio super-resolution model can be compressed, the reasoning efficiency can be improved, and the lightweight processing of the model can be achieved. This greatly reduces the computational consumption in the audio super-resolution processing process, thereby enhancing the practical feasibility of the Schrödinger bridge audio super-resolution model in low-latency application scenarios, such as edge computing and mobile terminals. It is suitable for a wide range of practical application scenarios. In addition, the Schrödinger bridge audio super-resolution model can directly optimize the low-resolution audio in the audio waveform space to obtain high-resolution audio. While maintaining the low-frequency data, it can restore the audio details of the high-frequency part, which can greatly improve the efficiency of the audio super-resolution processing and the fidelity of the high-frequency details.

[0096] like Figure 4 As shown in the above Figure 3 Based on the illustrated embodiment, step 304 includes the following steps:

[0097] Step 341: Initialize the student model and the teacher model based on the trained Schrödinger bridge audio super-resolution model to obtain an initialized student model and an initialized teacher model.

[0098] In the disclosed embodiment, the student model and the teacher model may be initialized to a trained Schrödinger bridge audio super-resolution model. The initialized student model is identical to the initialized teacher model and is consistent with the trained Schrödinger bridge audio super-resolution model.

[0099] Step 342: Obtain the sampling steps to be trained of the trained Schrödinger bridge audio super-resolution model.

[0100] In some optional implementations, a sampling step can be randomly selected from the sampling steps of the trained Schrödinger bridge audio super-resolution model as the sampling step to be trained. In other optional implementations, a sampling step can be sequentially selected from the sampling steps of the trained Schrödinger bridge audio super-resolution model as the sampling step to be trained. The current training sampling step can be determined by the random selection or sequential selection method described above. After multiple selections, each sampling step of the Schrödinger bridge audio super-resolution model can be used as the sampling step to be trained for subsequent operations.

[0101] Among them, the sampling step to be trained is the sampling step to be optimized for model distillation.

[0102] In step 343, for each super-resolved data pair, the initialized teacher model is used to obtain the predicted audio waveform signal of multiple sampling steps after the sampling step to be trained to obtain the supervised audio waveform signal; and the initialized student model is used to obtain the predicted audio waveform signal of one sampling step after the sampling step to be trained.

[0103] The super-resolution data pair may be the super-resolution data pair obtained in step 301 or a new super-resolution data pair. The super-resolution data pair also includes a low-resolution waveform signal and a high-resolution waveform signal representing the same audio.

[0104] In the embodiment of the present disclosure, based on the above-mentioned forward stochastic differential equation and the waveform state of the randomly obtained sampling step to be trained, the waveform state of the predicted audio waveform signal one sampling step after the sampling step to be trained can be predicted. According to the forward stochastic differential equation and the waveform state of the sampling step one sampling step after the sampling step to be trained, the waveform state of the predicted audio waveform signal second sampling step after the sampling step to be trained can be predicted. By predicting in sequence, the predicted audio waveform signals of multiple sampling steps after the sampling step to be trained can be predicted.

[0105] The predicted audio waveform signal for a number of sampling steps after the training sampling step can be determined based on the reduction in the number of sampling steps in each distillation round. For example, if the number of sampling steps in each distillation round is to be halved, the predicted audio waveform signal for two sampling steps after the training sampling step can be predicted, and the predicted audio waveform signal for two sampling steps after the training sampling step can be used as the supervisory audio waveform signal. If the number of sampling steps in each distillation round is to be reduced to 1 / 4, the predicted audio waveform signal for four sampling steps after the training sampling step can be predicted, and the predicted audio waveform signal for four sampling steps after the training sampling step can be used as the supervisory audio waveform signal.

[0106] For example, if the sampling step to be trained is 20 steps, the predicted audio waveform signal of sampling step - 19 can be predicted, and the predicted audio signal of sampling step - 18 can be predicted based on the predicted audio waveform signal of sampling step - 19, and the predicted audio signal of sampling step - 18 is used as the supervision audio waveform signal.

[0107] In step 344 , a model distillation loss is obtained based on the supervised audio waveform signal and the predicted audio waveform signal at one sampling step after the sampling step to be trained.

[0108] In the disclosed embodiment, the model distillation loss may be determined based on the supervised audio waveform signal and the predicted audio waveform signal one sampling step after the sampling step to be trained, as shown in Formula (4).

[0109] Formula (4)

[0110] In formula (4), is the waveform of one sampling step after the training sampling step predicted by the student model, is the waveform of multiple sampling steps after the sampling step to be trained predicted by the teacher model, is the model distillation loss.

[0111] Step 345: Based on the model distillation loss, the initialized student model is trained to obtain the Schrödinger bridge audio super-resolution model of the current distillation round.

[0112] In the disclosed embodiment, the initialized student model can be trained by minimizing the model distillation loss to obtain the Schrödinger bridge audio super-resolution model of the current distillation round.

[0113] The student model training process is a process of finding an optimal solution, and the process of fitting the model to the optimal solution is primarily iterative, minimizing the error. For a super-resolution data pair, the preset loss function is used to calculate the difference between the output of the teacher model and the output of the student model. This difference is then propagated through the backpropagation algorithm to the connections between each neuron in the student model. The difference signal transmitted to each connection represents the contribution of that connection to the overall error. The student model is then optimized using a gradient descent algorithm, gradually reducing the model distillation loss calculated during the iterative training process.

[0114] By repeatedly executing steps 342 through 345, i.e., iteratively training the student model using multiple super-resolved data pairs (including low-resolution and high-resolution waveform signals of the same audio), when the parameter-adjusted student model meets the training termination criteria, the current student model becomes the Schrödinger bridge audio super-resolved model for the current distillation round. The training termination criteria may include, but are not limited to, at least one of the following: convergence of the loss value of the aforementioned loss function, exceeding a preset training duration, or exceeding a preset number of training cycles.

[0115] Step 346 , determining that the Schrödinger bridge audio super-resolution model obtained in the current distillation round is the trained Schrödinger bridge audio super-resolution model.

[0116] Step 347, iteratively execute the above model distillation process to achieve the next round of model distillation until the preset distillation completion conditions are met, and the distilled and optimized Schrödinger bridge audio super-resolution model is obtained from the trained Schrödinger bridge audio super-resolution model.

[0117] In an embodiment of the present disclosure, after completing a distillation round of model distillation process through the above steps 341 to 345, the Schrödinger bridge audio super-resolution model obtained in the current distillation round can be used as the trained Schrödinger bridge audio super-resolution model, and the above model distillation process is iteratively executed to realize the next round of model distillation until the preset distillation completion conditions are met, thereby obtaining the final distillation-optimized Schrödinger bridge audio super-resolution model.

[0118] Among them, the preset distillation completion conditions may include but are not limited to at least one of the following: the distillation rounds reach a preset number of times, and the sampling steps of the Schrödinger bridge audio super-resolution model after distillation reach a preset number, such as the sampling step is 4 steps.

[0119] Based on the embodiments of the present disclosure, by performing multiple rounds of model distillation on the trained Schrödinger bridge audio super-resolution model, the distilled and optimized Schrödinger bridge audio super-resolution model can maintain the output of high-quality audio signals through reasoning with fewer sampling steps, greatly improving the model reasoning efficiency and application practicality; in addition, by utilizing the teacher-student iterative method, the multi-step trajectory sampling in the trained Schrödinger bridge audio super-resolution model is compressed into a few-step decision process, while retaining the complete information migration path, greatly reducing the computational burden and significantly reducing the delay, making it suitable for real-time scenarios such as voice communication and smart devices, and further expanding the application space of the Schrödinger bridge audio super-resolution model in voice super-resolution tasks.

[0120] In the above Figure 3On the basis of the embodiment shown, in order to further improve the super-resolution effect of the trained Schrödinger bridge audio super-resolution model, in the process of training the Schrödinger bridge audio super-resolution model, an additional multi-scale auxiliary loss can be combined to fine-tune the model to guide the model to adaptively adjust the supervision target according to the time step in the sampling step space, so that the Schrödinger bridge model dynamically focuses on low-frequency or high-frequency modeling on the inference path. The supervision method of the multi-scale auxiliary loss establishes a mapping relationship between the sampling step and the frequency band attention, which improves the spectrum fidelity and enhances the naturalness of the speech. Figure 5 As shown in Figure 2, the process of calculating the multi-scale auxiliary loss includes the following steps:

[0121] Step 501 : Based on the predicted audio waveform signal (ie, the predicted high-frequency waveform signal) of each sampling step, obtain the mean value of the audio waveform of the predicted forward process of each sampling step.

[0122] Among them, the mean value of the audio waveform of the predicted forward process is represented by the sampling step t , the expected value of the audio waveform predicted by the model‌.

[0123] In the embodiment of the present disclosure, the mean value of the audio waveform of the predicted forward process at each sampling step can be determined by formula (5).

[0124] Formula (5)

[0125] In formula (5), is the sampling step t The mean of the audio waveform of the forward process of prediction, is a variable that changes with time, is a function of the sampling step, is the preset value, To predict the audio signal using the Schrödinger bridge model, is the low-resolution sample waveform in the super-resolution data pair, is a variable that changes with time, is a function of the sampling step, It is a preset function about the sampling step, which changes with the sampling step.

[0126] Step 502: Using short-time Fourier transform, based on multiple window functions, convert the mean of the audio waveform of the predicted forward process of each sampling step into multiple frequency domain signals corresponding to the multiple window functions, thereby obtaining multiple predicted mean frequency domain signals of each sampling step; and using short-time Fourier transform, based on multiple window functions, convert the mean of the audio waveform of the actual forward process of each sampling step into multiple frequency domain signals corresponding to the multiple window functions, thereby obtaining multiple real mean frequency domain signals of each sampling step.

[0127] The frequency domain signal is a signal used to describe the frequency characteristics of the audio signal. The horizontal axis represents the frequency and the vertical axis represents the amplitude of the frequency signal. The mean value of the audio waveform of the actual forward process of each sampling step represents the system at the sampling step. t The true average value of the audio waveform of the forward process at each sampling step can be determined by formula (6).

[0128] Formula (6)

[0129] In formula (6), is a variable that changes with time, is a function of the sampling step, is the preset value, To predict the audio signal using the Schrödinger bridge model, is the low-resolution sample waveform in the super-resolution data pair, is a variable that changes with time, is a function of the sampling step, It is a pre-set function about the sampling step, which changes with the sampling step. is the sampling step t The mean of the audio waveform of the true forward process.

[0130] Specifically, when determining the multi-scale auxiliary loss, it can be determined according to formula (7).

[0131] Formula (7)

[0132] In formula (7), is the sampling step t The mean of the audio waveform of the forward process of prediction, is the sampling step t The mean of the audio waveform of the real forward process, It is a multi-scale loss.

[0133] In the embodiment of the present disclosure, the auxiliary loss can be determined independently at each sampling step of the model. and contributes it to the total loss, as above Combined, train the Schrödinger bridge model.

[0134] In the disclosed embodiment, the predicted resolution waveform of each sampling step may be converted into a frequency domain signal through a short-time Fourier transform (STFT), thereby determining the frequency and phase of the frequency domain signal.

[0135] When implementing it, select a time-frequency localized window function, assuming that the analysis window function g(t)It is stationary (pseudo-stationary) within a short time interval. By moving the window function, the power spectrum at different moments is determined, and a frequency domain signal with a resolution that matches the window function is obtained. If you want to change the resolution, you need to reselect the window function.

[0136] In the embodiment of the present disclosure, for the mean value of the audio waveform of the predicted forward process of each sampling step, the short-time Fourier transform can be used to convert the mean value of the audio waveform of the predicted forward process of each sampling step into multiple frequency domain signals of different resolutions based on multiple window functions, thereby obtaining multiple predicted mean frequency domain signals of each sampling step; and, using the short-time Fourier transform, the mean value of the audio waveform of the real forward process of each sampling step can be converted into multiple frequency domain signals corresponding to the multiple window functions based on multiple window functions, thereby obtaining multiple real mean frequency domain signals.

[0137] Furthermore, after obtaining multiple predicted mean frequency domain signals and multiple true mean frequency domain signals of each sampling step, step 503 may be executed to obtain the short-time Fourier transform amplitude loss, and step 504 may be executed to obtain the anti-wrapping phase loss.

[0138] Step 503 : Obtain short-time Fourier transform amplitude loss based on the amplitude information of the predicted mean frequency domain signal of each sampling step and the amplitude information of the true mean frequency domain signal of each sampling step.

[0139] In this example, after using different window functions to obtain predicted mean frequency domain signals of multiple resolutions and true mean frequency domain signals of multiple resolutions, a set of short-time Fourier transform amplitude losses can be calculated based on the predicted mean frequency domain signal and the true mean frequency domain signal of each resolution, and then the short-time Fourier transform amplitude loss of the Schrödinger bridge model can be obtained based on the short-time Fourier transform amplitude loss obtained for each resolution.

[0140] Step 504 : Obtain an anti-wrapping phase loss based on the phase information of the predicted mean frequency domain signal of each sampling step and the phase information of the true mean frequency domain signal of each sampling step.

[0141] In this example, after using different window functions to obtain predicted mean frequency domain signals of multiple resolutions and true mean frequency domain signals of multiple resolutions, a set of anti-wrapping phase losses can be calculated based on the predicted mean frequency domain signals and true mean frequency domain signals of each resolution, and then the anti-wrapping phase loss of the Schrödinger bridge model can be obtained based on the anti-wrapping phase loss obtained for each resolution.

[0142] The anti-wrapping phase loss indicates the phase loss determined by the instantaneous phase, group delay, and angular frequency errors based on the anti-wrapping strategy. Specifically, three sets of short-time Fourier transform (SFT) parameters (i.e., fast Fourier transform (FFT) sizes of 512, 1024, and 2048, respectively) are used to determine the anti-wrapping phase loss based on the phase information of the predicted mean frequency domain signal at each sampling step and the phase information of the true mean frequency domain signal at each sampling step.

[0143] Step 505 : Determine the weighted sum of the short-time Fourier transform amplitude loss and the anti-wrapping phase loss to obtain a multi-scale auxiliary loss.

[0144] Among them, the multi-scale auxiliary loss includes short-time Fourier transform amplitude loss and anti-wrapping phase loss.

[0145] A multi-scale auxiliary loss is obtained by calculating the weighted sum of the short-time Fourier transform amplitude loss and the anti-wrapping phase loss. The weights of the short-time Fourier transform amplitude loss and the anti-wrapping phase loss can be adjusted and optimized during model training to ensure that the obtained multi-scale auxiliary loss can better capture high-frequency features while maintaining consistency in the low-frequency part.

[0146] In the disclosed embodiment, the Schrödinger bridge audio super-resolution model trained using the Schrödinger bridge model loss is fine-tuned by using multi-scale auxiliary losses (short-time Fourier transform amplitude loss and anti-wrapping phase loss), so that the fine-tuned Schrödinger bridge audio super-resolution model can significantly improve the accuracy of spectral details and phase information in the super-resolution audio super-resolution processing process, so that the super-resolution audio can achieve better results in both waveform space and frequency domain space, effectively solving technical problems such as consonant swallowing, high and low frequency mismatch, and poor low-frequency retention in audio super-resolution tasks, and further enhancing the generation of the trained Schrödinger bridge audio super-resolution model. performance, and more efficiently generate high-resolution audio waveforms; in addition, the multi-scale auxiliary loss is determined by using the mean of the audio waveform of the predicted forward process of each sampling step and the mean of the audio waveform of the actual forward process of each sampling step. The model trained based on the auxiliary loss can distinguish the outputs of different sampling steps during the inference process, thereby achieving more emphasis on maintaining the ground structure in the stage close to the prior, and gradually focusing on the generation of high-frequency features and strengthening the construction of high-frequency information in the stage close to the target distribution, effectively enhancing the model's frequency domain resolution and multi-frequency collaborative expression capabilities, and significantly improving the perceptual quality and detail restoration accuracy of audio super-resolution.

[0147] In order to better perceive the difference between low-resolution waveform signals and high-resolution waveform signals and improve the effect of high-frequency detail data generated by the trained Schrödinger bridge audio super-resolution model, in the above Figure 3 and / or Figure 5Based on the embodiment shown, the embodiment of the present disclosure can also perform scaling processing after obtaining the super-resolution data pair, such as Figure 6 As shown in FIG, the audio super-resolution model training process includes the following steps.

[0148] Step 601: Acquire multiple super-resolution data pairs, each super-resolution data pair including a low-resolution waveform signal and a high-resolution waveform signal representing the same audio.

[0149] Step 602 : Amplify the low-resolution waveform signal and the high-resolution waveform signal in each super-resolution data pair using a scaling factor to obtain amplified low-resolution waveform signals and high-resolution waveform signals.

[0150] Among them, the amplified low-resolution waveform signal and high-resolution waveform signal are used to train the Schrödinger bridge model.

[0151] In the present disclosure, the same scaling factor may be used to amplify the low-resolution waveform signal and the high-resolution waveform signal respectively.

[0152] In the present disclosure, the scaling factor can be obtained based on the low-resolution waveform signal and the high-resolution waveform signal, as shown in formula (8):

[0153] Formula (8)

[0154] In formula (8), is a high-resolution waveform signal, For high-resolution waveform signals, s is a scaling factor. By multiplying the low-resolution waveform signal and the high-resolution waveform signal with the above scaling factor respectively, the statistical variance between the low-resolution waveform signal and the high-resolution waveform signal can be amplified to 1, thereby increasing the optimization space and efficiency, and significantly improving the ability of the Schrödinger bridge model to capture high-frequency details.

[0155] Step 603 : For each super-resolved data pair, a solvable path from the amplified low-resolution waveform signal to the amplified high-resolution waveform signal is established using a Schrödinger bridge model.

[0156] Step 604: Based on the solvable path, the Schrödinger bridge model is trained to obtain a trained Schrödinger bridge audio super-resolution model.

[0157] In the embodiment of the present disclosure, after amplifying the low-resolution waveform signal and the high-resolution waveform signal in each super-resolution data pair using a scaling factor to obtain the amplified low-resolution waveform signal and the high-resolution waveform signal, when determining the Schrödinger bridge model loss of the Schrödinger bridge model, the obtained Schrödinger bridge model loss is the Schrödinger bridge model loss after the scaling factor is applied, as shown in Formula (9):

[0158] Formula (9)

[0159] In formula (9), is the predicted high-resolution waveform signal, The sampling step predicted by the Schrödinger bridge model t The waveform, is a low-resolution sample waveform, is a high-resolution sample waveform, s is the scaling factor, is the loss of the Schrödinger bridge model after the scaling factor.

[0160] The implementation of the above steps 603 to 604 can be found in Figure 3 The description of steps 302 to 303 in the illustrated embodiment will not be repeated here in detail.

[0161] Step 605 : Using multiple super-resolution data pairs, perform at least one round of distillation optimization on the trained Schrödinger bridge audio super-resolution model to obtain a distillation-optimized Schrödinger bridge audio super-resolution model.

[0162] The implementation of the above step 605 can be found in Figure 3 The description of step 304 in the illustrated embodiment will not be repeated here in detail.

[0163] In the embodiment of the present disclosure, by amplifying and processing the low-resolution waveform signal and the high-resolution waveform signal in the super-resolution data pair respectively, the Schrödinger bridge model can better perceive the difference between the low-resolution waveform signal and the high-resolution waveform signal, ensuring that the loss function can remain sensitive, which helps to improve the training efficiency and stability of the model.

[0164] Exemplary devices

[0165] Figure 7 This is a structural diagram of an audio super-resolution processing device provided by an exemplary embodiment of the present disclosure. The device can be applied to electronic devices such as terminal devices, such as Figure 7 As shown, the device includes:

[0166] An acquisition module 71 is used to acquire a low-resolution signal to be processed;

[0167] An adjustment module 72 is configured to interpolate the low-resolution signal to be processed using a linear interpolation algorithm to generate prior information;

[0168] The generation module 73 is used to generate a high-resolution target waveform based on the generated prior information and using the distilled and optimized Schrödinger bridge audio super-resolution model. The distilled and optimized Schrödinger bridge audio super-resolution model keeps the low-frequency information in the generated prior information unchanged in each sampling step of generating the high-resolution target waveform.

[0169] Figure 8 This is a structural diagram of an audio super-resolution model training device provided by an exemplary embodiment of the present disclosure. The device can be applied to electronic devices such as terminal devices and servers. The device includes:

[0170] A sample acquisition module 81 is configured to acquire a plurality of super-resolution data pairs, each super-resolution data pair including a low-resolution waveform signal and a high-resolution waveform signal representing the same audio;

[0171] a path generation module 82 for establishing, for each super-resolved data pair, a solvable path from the low-resolution waveform signal to the high-resolution waveform signal using a Schrödinger bridge model;

[0172] The audio super-resolution model training module 83 is used to train the to-be-trained Schrödinger bridge model based on the solvable path to obtain a trained Schrödinger bridge audio super-resolution model;

[0173] The model distillation module 84 is used to use multiple super-resolution data pairs to perform at least one round of distillation optimization on the trained Schrödinger bridge audio super-resolution model to obtain a distilled and optimized Schrödinger bridge audio super-resolution model.

[0174] Figure 9 FIG. 1 is a structural diagram of an audio super-resolution model training device provided by another exemplary embodiment of the present disclosure. Figure 9 As shown, in Figure 8 Based on the illustrated embodiments, in some embodiments, the Schrödinger bridge model uses a stochastic differential equation to generate a solvable path; the stochastic differential equation is defined in the form of an asymmetric noise scheduling strategy.

[0175] In some embodiments, the model distillation module 84 may include:

[0176] An initialization submodule 841 is used to initialize the student model and the teacher model based on the trained Schrödinger bridge audio super-resolution model to obtain an initialized student model and an initialized teacher model;

[0177] The sampling step acquisition submodule 842 is used to obtain the sampling steps to be trained of the trained Schrödinger bridge audio super-resolution model;

[0178] The model sampling submodule 843 is configured to, for each super-resolved data pair, use the initialized teacher model to obtain a predicted audio waveform signal multiple sampling steps after the sampling step to be trained, thereby obtaining a supervisory audio waveform signal; and use the initialized student model to obtain a predicted audio waveform signal one sampling step after the sampling step to be trained;

[0179] a loss determination submodule 844 for obtaining a model distillation loss based on the supervised audio waveform signal and the predicted audio waveform signal one sampling step after the training sampling step;

[0180] A model training submodule 845 is used to train the initialized student model based on the model distillation loss to obtain a Schrödinger bridge audio super-resolution model of the current distillation round;

[0181] A model determination submodule 846 is used to determine whether the Schrödinger bridge audio super-resolution model obtained in the current distillation round is a trained Schrödinger bridge audio super-resolution model;

[0182] The iterative submodule 847 is used to iteratively execute the above-mentioned model distillation process to realize the next round of model distillation until the preset distillation completion conditions are met, and the distilled and optimized Schrödinger bridge audio super-resolution model is obtained from the trained Schrödinger bridge audio super-resolution model.

[0183] In some embodiments, the audio super-resolution model training module 83 is configured to guide the training of the Schrödinger bridge model based on the solvable path using the Schrödinger bridge model loss and the multi-scale auxiliary loss to obtain a trained Schrödinger bridge audio super-resolution model;

[0184] Among them, the Schrödinger bridge model loss is used to characterize the error between the predicted audio waveform signal and the true high-frequency waveform signal of the Schrödinger bridge model at each sampling step, and the multi-scale auxiliary loss is used to characterize the multi-scale auxiliary loss of the Schrödinger bridge model at each sampling step.

[0185] In some embodiments, the multi-scale auxiliary loss includes a short-time Fourier transform amplitude loss and an anti-wrapping phase loss;

[0186] The audio super-resolution model training module 83 includes:

[0187] A signal determination submodule 831 is configured to obtain a mean value of the audio waveform of the forward process of the prediction for each sampling step based on the predicted audio waveform signal for each sampling step;

[0188] The signal conversion submodule 832 is configured to utilize a short-time Fourier transform (SFT) to convert the predicted mean value of the forward audio waveform at each sampling step into multiple frequency domain signals corresponding to the multiple window functions based on multiple window functions, thereby obtaining multiple predicted mean frequency domain signals at each sampling step; and to utilize a short-time Fourier transform (SFT) to convert the actual mean value of the forward audio waveform at each sampling step into multiple frequency domain signals corresponding to the multiple window functions based on multiple window functions, thereby obtaining multiple actual mean frequency domain signals at each sampling step.

[0189] An amplitude loss submodule 833 is configured to obtain a short-time Fourier transform amplitude loss based on the amplitude information of the predicted mean frequency domain signal of each sampling step and the amplitude information of the true mean frequency domain signal of each sampling step;

[0190] A phase loss submodule 834 is configured to obtain an anti-wrapping phase loss based on the phase information of the predicted mean frequency domain signal at each sampling step and the phase information of the true mean frequency domain signal at each sampling step;

[0191] The auxiliary loss submodule 835 is used to determine the weighted sum of the short-time Fourier transform amplitude loss and the anti-wrapping phase loss to obtain a multi-scale auxiliary loss.

[0192] In some embodiments, the audio super-resolution model training module 83 is configured to obtain a Schrödinger bridge model loss based on the error between the predicted audio waveform signal and the true high-frequency waveform signal at each sampling step.

[0193] In some embodiments, the sample acquisition module 81 is configured to perform replication processing on the high-resolution waveform signal using a signal processing filter to obtain a low-resolution waveform signal corresponding to the high resolution.

[0194] In some embodiments, the sample acquisition module 81 is used to use a scaling factor to amplify the low-resolution waveform signal and the high-resolution waveform signal in each super-resolution data pair to obtain an amplified low-resolution waveform signal and a high-resolution waveform signal. The amplified low-resolution waveform signal and the high-resolution waveform signal are used to train the Schrödinger bridge model.

[0195] It should be noted that the modules in the present device can be decomposed and / or reassembled, and such decompositions and / or reassemblies should be regarded as equivalent solutions of the present device.

[0196] The exemplary embodiment of this device corresponds to the exemplary method described above, and the relevant contents can be referenced and cited to each other. The beneficial technical effects corresponding to the exemplary embodiment of this device can be referred to the corresponding beneficial technical effects of the exemplary method described above, and will not be repeated here.

[0197] Exemplary electronic devices

[0198] Figure 10 A structural diagram of an electronic device provided in an embodiment of the present disclosure includes at least one processor 101 and a memory 102.

[0199] The processor 101 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.

[0200] Memory 102 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and processor 101 may execute one or more computer program instructions to implement the vehicle posture detection method and / or other desired functions described in the various embodiments of the present disclosure.

[0201] In one example, the electronic device may further include an input device 103 and an output device 104 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0202] The input device 103 may also include, for example, a keyboard, a mouse, a touch screen, a sound pickup device (such as a microphone array), etc.

[0203] The output device 104 can output various information to the outside, and may include, for example, a display, a speaker, a printer, a communication network and its connected remote output devices, etc.

[0204] Of course, to simplify, Figure 10 Only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device may further include any other appropriate components according to specific application scenarios.

[0205] Exemplary systems, computer program products, and computer-readable storage media

[0206] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the audio super-resolution processing method and the audio super-resolution model training method according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.

[0207] The computer program product may be written in any combination of one or more programming languages ​​to implement the operations of the disclosed embodiments, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0208] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the steps of the audio super-resolution processing method and the audio super-resolution model training method according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.

[0209] Computer-readable storage media can be any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0210] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.

[0211] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments are sufficient. For system embodiments, since they largely correspond to method embodiments, their description is relatively simple. For relevant parts, references to the description of the method embodiments are sufficient.

[0212] The block diagrams of the devices, apparatuses, and equipment involved in this disclosure are intended only as illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, apparatuses, and equipment may be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and may be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and may be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and may be used interchangeably therewith.

[0213] The methods and apparatus of the present disclosure may be implemented in many ways. For example, the methods and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the methods according to the present disclosure. Therefore, the present disclosure also covers recording media that store programs for executing the methods according to the present disclosure.

[0214] It should also be noted that in the apparatus, device, and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.

[0215] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0216] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A method for training an audio super-resolution model, characterized in that: include: Acquire a plurality of super-resolution data pairs, each super-resolution data pair including a low-resolution waveform signal and a high-resolution waveform signal representing the same audio; For each super-resolved data pair, a solvable path from the low-resolution waveform signal to the high-resolution waveform signal is established using a Schrödinger bridge model; Based on the solvable path, the Schrödinger bridge model is trained to obtain a trained Schrödinger bridge audio super-resolution model; Using the multiple super-resolution data pairs, at least one round of distillation optimization is performed on the trained Schrödinger bridge audio super-resolution model to obtain a distillation-optimized Schrödinger bridge audio super-resolution model.

2. The method according to claim 1, characterized in that The method of performing at least one round of distillation optimization on the trained Schrödinger bridge audio super-resolution model using the multiple super-resolution data pairs to obtain the distillation-optimized Schrödinger bridge audio super-resolution model comprises: Initializing the student model and the teacher model based on the trained Schrödinger bridge audio super-resolution model to obtain an initialized student model and an initialized teacher model; Obtaining the sampling steps to be trained of the trained Schrödinger bridge audio super-resolution model; For each super-resolved data pair, using the initialized teacher model, obtain a predicted audio waveform signal for multiple sampling steps after the sampling step to be trained to obtain a supervision audio waveform signal; and using the initialized student model, obtain a predicted audio waveform signal for one sampling step after the sampling step to be trained; Obtaining a model distillation loss based on the supervised audio waveform signal and a predicted audio waveform signal one sampling step after the sampling step to be trained; Based on the model distillation loss, the initialized student model is trained to obtain a Schrödinger bridge audio super-resolution model of the current distillation round; Determining that the Schrödinger bridge audio super-resolution model obtained in the current distillation round is the trained Schrödinger bridge audio super-resolution model; The above-mentioned model distillation process is iteratively executed to realize the next round of model distillation until the preset distillation completion conditions are met, and the distilled and optimized Schrödinger bridge audio super-resolution model is obtained from the trained Schrödinger bridge audio super-resolution model.

3. The method according to claim 1, characterized in that The Schrödinger bridge model uses a stochastic differential equation to generate the solvable path; the stochastic differential equation is defined in the form of an asymmetric noise scheduling strategy.

4. The method according to claim 1, wherein The training of the Schrödinger bridge model based on the solvable path includes: Based on the solvable path, using the Schrödinger bridge model loss and the multi-scale auxiliary loss to guide the training of the Schrödinger bridge model, to obtain a trained Schrödinger bridge audio super-resolution model; Among them, the Schrödinger bridge model loss is used to characterize the error between the predicted audio waveform signal and the true high-frequency waveform signal of the Schrödinger bridge model at each sampling step, and the multi-scale auxiliary loss is used to characterize the multi-scale auxiliary loss of the Schrödinger bridge model at each sampling step. The multi-scale auxiliary loss is obtained based on the mean of the audio waveform of the predicted forward process of each sampling step and the mean of the audio waveform of the true forward process.

5. The method according to claim 4, characterized in that The multi-scale auxiliary loss includes a short-time Fourier transform amplitude loss and an anti-wrapping phase loss; calculating the multi-scale auxiliary loss includes: Based on the predicted audio waveform signal of each sampling step, the mean value of the audio waveform of the predicted forward process of each sampling step is obtained; Using short-time Fourier transform, based on multiple window functions, the mean of the audio waveform of the forward process predicted at each sampling step is converted into multiple frequency domain signals corresponding to the multiple window functions, thereby obtaining multiple predicted mean frequency domain signals at each sampling step; and Using short-time Fourier transform, based on multiple window functions, the mean of the audio waveform of the real forward process of each sampling step is converted into multiple frequency domain signals corresponding to the multiple window functions, thereby obtaining multiple real mean frequency domain signals of each sampling step; Obtaining the short-time Fourier transform amplitude loss based on the amplitude information of the predicted mean frequency domain signal of each sampling step and the amplitude information of the true mean frequency domain signal of each sampling step; Obtaining the anti-wrapping phase loss based on the phase information of the predicted mean frequency domain signal of each sampling step and the phase information of the true mean frequency domain signal of each sampling step; A weighted sum of the short-time Fourier transform amplitude loss and the anti-wrapping phase loss is determined to obtain the multi-scale auxiliary loss.

6. The method according to claim 4, characterized in that Calculating the Schrödinger bridge model loss includes: The Schrödinger bridge model loss is obtained based on the error between the predicted audio waveform signal and the true high-frequency waveform signal at each sampling step.

7. The method according to any one of claims 1 to 6, characterized in that: The obtaining of multiple super-resolution data pairs includes: The high-resolution waveform signal is copied using a signal processing filter to obtain a low-resolution waveform signal corresponding to the high-resolution waveform signal.

8. The method according to any one of claims 1 to 6, characterized in that: The obtaining of multiple super-resolution data pairs includes: The low-resolution waveform signal and the high-resolution waveform signal in each super-resolution data pair are amplified using a scaling factor to obtain an amplified low-resolution waveform signal and a high-resolution waveform signal. The amplified low-resolution waveform signal and the high-resolution waveform signal are used to train the Schrödinger bridge model.

9. An audio super-resolution processing method, characterized in that: include: obtaining a low-resolution signal to be processed; Using a linear interpolation algorithm, interpolating the low-resolution signal to be processed to obtain prior information; Based on the generated prior information, a high-resolution target waveform is generated using the distilled and optimized Schrödinger bridge audio super-resolution model. The distilled and optimized Schrödinger bridge audio super-resolution model keeps the low-frequency information in the generated prior information unchanged in each sampling step of generating the high-resolution target waveform.

10. An audio super-resolution model training device, characterized in that: include: A sample acquisition module is used to acquire multiple super-resolution data pairs, each super-resolution data pair includes a low-resolution waveform signal and a high-resolution waveform signal representing the same audio; a path generation module for establishing, for each super-resolved data pair, a solvable path from the low-resolution waveform signal to the high-resolution waveform signal using a Schrödinger bridge model; An audio super-resolution model training module, configured to train the Schrödinger bridge model based on the solvable path to obtain a trained Schrödinger bridge audio super-resolution model; The model distillation module is used to use the multiple super-resolution data pairs to perform at least one round of distillation optimization on the trained Schrödinger bridge audio super-resolution model to obtain a distilled and optimized Schrödinger bridge audio super-resolution model.

11. An audio super-resolution processing device, characterized in that: include: An acquisition module, used for acquiring a low-resolution signal to be processed; An adjustment module is used to interpolate the low-resolution signal to be processed using a linear interpolation algorithm to generate prior information; A generation module is used to generate a high-resolution target waveform based on the generated prior information and using the distilled and optimized Schrödinger bridge audio super-resolution model, wherein the distilled and optimized Schrödinger bridge audio super-resolution model maintains the low-frequency information in the generated prior information unchanged in each sampling step of generating the high-resolution target waveform.

12. A computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 9.

13. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Image super-resolution method based on knowledge distillation

    CN114359039A

  • Super-resolution audio generation method, computer equipment and storage medium

    CN114400015A