Audio processing method, electronic equipment, storage medium and chip system

By combining the L1 loss function, L2 loss function and signal-to-noise ratio loss function in the audio processing model, the interactive relationship between sound objects is learned, and the problem of sound details loss in the prior art is solved, improving the effect of music source separation and user experience.

CN120472926APending Publication Date: 2025-08-12HONOR DEVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411314200.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the prior art, when music sources are separated, the extracted sound components are prone to lose signals and cannot effectively retain sound details.

Method used

An audio processing model based on the L1 loss function and the L2 loss function combined with the signal-to-noise ratio loss function is adopted. By calculating the combination of each sound object in the sound source, the interactive relationship between different sound objects is learned, which is the training target of the audio processing model.

Benefits of technology

It achieves better preservation of spectrum details, improves the separation effect and user experience of the audio processing model, and enhances the learning ability of small signals and signal vector integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472926A_ABST
    Figure CN120472926A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio processing method, electronic equipment, a storage medium and a chip system, and relates to the technical field of terminals. The method comprises the following steps: considering each sound object in a sound source, calculating an L1 loss function of a combination of one or more sound objects in a time domain and a frequency domain or an L2 loss function of the combination of the one or more sound objects in the time domain and the frequency domain, and calculating a comprehensive loss function as a training target of an audio processing model by combining a signal-to-noise ratio loss function. Therefore, the interaction relationship between different sound objects can be learned by combining the time domain and frequency domain characteristics of the audio data, and an audio processing model which can retain spectrum details and has a better effect is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of terminal technology, and in particular to an audio processing method, an electronic device, a storage medium, and a chip system. Background Art

[0002] Users use electronic devices to listen to music, watch videos, and watch movies. To provide users with a better listening experience, electronic devices can perform music source separation on the sounds in music, videos, and movies. Music source separation can be understood as extracting the sound components from the original mixed music and processing them as separate channels to achieve better sound effects.

[0003] However, in some implementations, when performing music source separation, signals may be lost in the extracted sound components, and the sound details may not be better preserved. Summary of the Invention

[0004] The audio processing method, electronic device, storage medium, and chip system provided in the embodiments of the present application can consider each sound object in the sound source, calculate the L1 loss function of the combination of one or more sound objects in the time domain and frequency domain, or the L2 loss function in the time domain and frequency domain, and combine it with the signal-to-noise ratio loss function to calculate a comprehensive loss function as the training target of the audio processing model. In this way, the interaction between different sound objects can be learned by combining the time domain and frequency domain characteristics of the audio data, resulting in an audio processing model that can retain spectral details and has better results.

[0005] In a first aspect, an embodiment of the present application provides an audio processing method, applied to an electronic device, the method comprising:

[0006] Acquire a first audio signal; extract t sound objects from the first audio signal based on an audio processing model; wherein the audio processing model is trained based on a target loss function, the target loss function is related to the loss function of the ith combination, the ith combination is a combination of any one or more sound objects from the t sound objects, the loss function of the ith combination is related to the p-norm and signal-to-noise ratio loss function, and t is a positive integer; post-process one or more sound objects from the t sound objects to obtain a second audio signal. In this way, the combination of each sound object in the first audio signal can be considered, and the p-norm loss function and the signal-to-noise ratio loss function can be combined as the training target of the audio processing model. In this way, the characteristics of the p-norm and signal-to-noise ratio loss functions can be combined to learn the interactive relationship between each sound object, obtain a more accurate separation result, and thus improve the user experience.

[0007] In one possible implementation, the loss function for the i-th combination is related to a p-norm L2 loss function and an L2-norm signal-to-noise ratio loss function. This combination not only preserves spectral details and produces more complete results, but also enhances the audio processing model's ability to learn small signals and signal vector integrity, improving small signal separation.

[0008] In one possible implementation, the loss function of the i-th combination is related to at least two of the following: an L2 loss function in the frequency domain, a frequency domain signal-to-noise ratio loss function based on the L2 norm, an L2 loss function in the time domain, and a time domain signal-to-noise ratio loss function based on the L2 norm. The embodiment of the present application can comprehensively consider time domain information and frequency domain information, and calculate the L2 loss function and the signal-to-noise ratio loss function on time domain waveform data and frequency domain spectrum data. This can be more conducive to the convergence of the audio processing model, allowing the audio processing model to better learn the characteristics of the time domain and frequency domain.

[0009] In one possible implementation, the loss function of the i-th combination is Loss i Satisfy the following formula:

[0010]

[0011] in, represents the L2 loss function of the i-th combination in the frequency domain, represents the frequency domain signal-to-noise ratio loss function based on the L2 norm of the i-th combination, represents the L2 loss function of the i-th combination in the time domain, represents the time domain signal-to-noise ratio loss function based on the L2 norm of the i-th combination, α1 represents The hyperparameter, α2 represents The hyperparameters, α3 represents The hyperparameters, α4 represents The loss function of the i-th combination considers the L2 loss function and the signal-to-noise ratio loss function based on the L2 norm in the time domain, as well as the L2 loss function and the signal-to-noise ratio loss function based on the L2 norm in the frequency domain. This not only preserves more complete small signals, but also ensures that the separation results of the audio processing model have more accurate envelopes and more complete spectral structures.

[0012] In one possible implementation, the L2 loss function of the i-th combination in the frequency domain and the L2 loss function of the i-th combination in the time domain both satisfy the following L2 loss function formula:

[0013]

[0014] in, represents the L2 loss function of the i-th combination, y i is the estimated sound object of the i-th combination, represents the target sound object of the i-th combination; the frequency domain signal-to-noise ratio loss function based on the L2 norm of the i-th combination and the time domain signal-to-noise ratio loss function based on the L2 norm of the i-th combination both satisfy the following signal-to-noise ratio loss function formula:

[0015]

[0016] in, represents the signal-to-noise ratio loss function of the i-th combination, x represents the first audio signal, (y i ) T represents the transpose of the estimated sound object of the i-th combination, ‖y i ‖ represents the bi-norm of the estimated sound object of the i-th combination, represents the bi-norm of the target sound object of the i-th combination, ρ i and (1-ρ i ) represent the coordination coefficients, respectively. The audio processing model can be trained based on the L2 loss function in the time and frequency domains, as well as the signal-to-noise ratio loss function in the time and frequency domains, thereby updating the parameters in the audio processing model and continuously optimizing the results, so that the estimated sound object gradually approaches the target sound object, thereby more accurately separating the sound objects in the audio.

[0017] In one possible implementation, the loss function for the i-th combination is related to a p-norm L1 loss function and an L1-norm signal-to-noise ratio loss function. In the embodiment of the present application, the use of a p-norm L1 loss function and a signal-to-noise ratio loss function can not only tend to extract cleaner results, but also enhance the audio processing model's ability to learn small signals and signal vector integrity.

[0018] In one possible implementation, the loss function for the i-th combination is related to at least two of the following: an L1 loss function in the frequency domain, a frequency domain signal-to-noise ratio loss function based on the L1 norm, an L1 loss function in the time domain, and a time domain signal-to-noise ratio loss function based on the L1 norm. By comprehensively considering both time and frequency domain information, the L1 loss function and the signal-to-noise ratio loss function are calculated on both the time domain waveform data and the frequency domain spectrum data. This allows the audio processing model to better learn the characteristics of both the time and frequency domains, resulting in cleaner audio separation results.

[0019] In one possible implementation, the loss function of the i-th combination is Loss i Satisfies the following formula:

[0020]

[0021] in, represents the L1 loss function of the i-th combination in the frequency domain, represents the frequency domain signal-to-noise ratio loss function based on the L1 norm of the i-th combination, represents the L1 loss function of the i-th combination in the time domain, represents the time domain signal-to-noise ratio loss function based on the L1 norm of the i-th combination, and α5 represents The hyperparameters, α6 represents The hyperparameters, α7 represents The hyperparameters, α8 represents The loss function of the i-th combination considers the L1 loss function and the signal-to-noise ratio loss function based on the L1 norm in the time domain, as well as the L1 loss function and the signal-to-noise ratio loss function based on the L1 norm in the frequency domain. This can produce cleaner extraction results and achieve better extraction effects when extracting sounds with low loudness, high missing rate, and relatively sparse sounds.

[0022] In one possible implementation, the L1 loss function of the i-th combination in the frequency domain and the L1 loss function of the i-th combination in the time domain both satisfy the following L1 loss function formula:

[0023]

[0024] in, represents the L1 loss function of the i-th combination, y i is the estimated sound object of the i-th combination, represents the target sound object of the i-th combination; the frequency domain signal-to-noise ratio loss function based on the L1 norm of the i-th combination and the time domain signal-to-noise ratio loss function based on the L1 norm of the i-th combination both satisfy the following signal-to-noise ratio loss function formula:

[0025]

[0026] in, represents the signal-to-noise ratio loss function of the i-th combination, x represents the first audio signal, (y i ) T represents the transpose of the estimated sound object of the i-th combination, |y i | represents the absolute value of the estimated sound object of the i-th combination, represents the absolute value of the target sound object of the i-th combination, ρ i and (1-ρ i) represent the coordination coefficients, respectively. The audio processing model can be trained based on the L1 loss function in the time and frequency domains, as well as the signal-to-noise ratio loss function in the time and frequency domains, thereby updating the parameters in the audio processing model and continuously optimizing the results, so that the estimated sound object gradually approaches the target sound object, thereby more accurately separating the sound objects in the audio.

[0027] In one possible implementation, one or more of the t sound objects are post-processed, including performing one or more of the following processing on the one or more sound objects: gain processing, similarity calculation, and sound field expansion. Gain processing can adjust the intensity of the sound object, thereby balancing the volume of different sound objects, making the second audio signal obtained more harmonious and improving the auditory effect. Similarity calculation can separate the center voice and background voice in the human voice, thereby performing different audio processing on the center voice and background voice respectively, thereby improving the user's listening experience. Sound field expansion can improve the overall quality and effect of the audio experience by optimizing the spatial distribution and positioning of the sound, enhancing the listener's sense of immersion, and making the listener feel as if they are at the scene or in a specific environment. Different sound field expansion methods can also be performed on sounds in different frequency bands to make the sound clearer and more layered.

[0028] In a second aspect, an embodiment of the present application provides a method for training an audio processing model, which is applied to an electronic device, and the method includes:

[0029] Inputting audio signal samples into an audio processing model to be trained; adjusting the audio processing model to be trained based on a target loss function until the target loss function converges, thereby obtaining a trained audio processing model; wherein the target loss function is related to the p-norm and the signal-to-noise ratio loss function. The embodiment of the present application makes the target loss function related to the p-norm and the signal-to-noise ratio loss function, and takes into account the combination between each prediction category and the time-frequency domain characteristics of the signal. In this way, the characteristics of the p-norm and the signal-to-noise ratio loss function can be combined to learn the interaction relationship between each sound object, thereby obtaining an audio processing model that can retain spectral details and has a better separation effect.

[0030] In one possible implementation, the audio signal sample includes a first audio signal, which includes t sound objects. The target loss function is related to the loss function of the i-th combination, where the i-th combination is any combination of one or more sound objects from the t sound objects. The loss function of the i-th combination is related to the p-norm and signal-to-noise ratio loss functions. In this way, the final calculated loss function is the sum of the loss functions of multiple combinations. By continuously optimizing the results, more accurate music source separation can be achieved.

[0031] In one possible implementation, the loss function of the i-th combination is related to at least two of the following: an L2 loss function in the frequency domain, a frequency domain signal-to-noise ratio loss function based on the L2 norm, an L2 loss function in the time domain, and a time domain signal-to-noise ratio loss function based on the L2 norm. The embodiment of the present application can comprehensively consider time domain information and frequency domain information, and calculate the L2 loss function and the signal-to-noise ratio loss function on time domain waveform data and frequency domain spectrum data. This can be more conducive to the convergence of the audio processing model, allowing the audio processing model to better learn the characteristics of the time domain and frequency domain.

[0032] In one possible implementation, the loss function of the i-th combination is related to at least two of the following: an L1 loss function in the frequency domain, a frequency domain signal-to-noise ratio loss function based on the L1 norm, an L1 loss function in the time domain, and a time domain signal-to-noise ratio loss function based on the L1 norm. The embodiment of the present application can comprehensively consider time domain information and frequency domain information, and calculate the L1 loss function and the signal-to-noise ratio loss function on the time domain waveform data and the frequency domain spectrum data. In this way, the audio processing model can better learn the characteristics of the time domain and frequency domain, and obtain a cleaner audio separation result.

[0033] In a third aspect, an embodiment of the present application provides an audio processing device, which may be an electronic device or a chip or chip system within an electronic device. The device may include a processing unit. The processing unit is used to implement any processing-related method performed by the electronic device in the first aspect or any possible implementation of the first aspect. When the device is an electronic device, the processing unit may be a processor. The device may also include a storage unit, which may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to enable the electronic device to implement the method described in the first aspect or any possible implementation of the first aspect. When the device is a chip or chip system within an electronic device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit to enable the electronic device to implement the method described in the first aspect or any possible implementation of the first aspect. The storage unit may be a storage unit within the chip (for example, a register, a cache, etc.), or a storage unit within the electronic device located outside the chip (for example, a read-only memory, a random access memory, etc.).

[0034] In a fourth aspect, an embodiment of the present application provides a training device for an audio processing model, which may be an electronic device or a chip or chip system within an electronic device. The device may include a processing unit. The processing unit is used to implement any method related to processing performed by the electronic device in the second aspect or any possible implementation of the second aspect. When the device is an electronic device, the processing unit may be a processor. The device may also include a storage unit, which may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to enable the electronic device to implement the method described in the second aspect or any possible implementation of the second aspect. When the device is a chip or chip system within an electronic device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit to enable the electronic device to implement the method described in the second aspect or any possible implementation of the second aspect. The storage unit may be a storage unit within the chip (for example, a register, a cache, etc.), or it may be a storage unit within the electronic device located outside the chip (for example, a read-only memory, a random access memory, etc.).

[0035] In a fifth aspect, an embodiment of the present application provides an electronic device comprising one or more processors and a memory, the memory being coupled to the one or more processors, the memory being used to store computer program code, the computer program code comprising computer instructions, and the one or more processors being used to call computer instructions to execute the method described in the first aspect or any possible implementation of the first aspect, or to execute the method described in the second aspect or any possible implementation of the second aspect.

[0036] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction is run on a computer, the computer executes the method described in the first aspect or any possible implementation of the first aspect, or executes the method described in the second aspect or any possible implementation of the second aspect.

[0037] In the seventh aspect, an embodiment of the present application provides a computer program product including a computer program. When the computer program is run on a computer, it enables the computer to execute the method described in the first aspect or any possible implementation of the first aspect, or execute the method described in the second aspect or any possible implementation of the second aspect.

[0038] In an eighth aspect, the present application provides a chip or chip system, comprising at least one processor and a communication interface, wherein the communication interface and the at least one processor are interconnected by a line, and the at least one processor is configured to run a computer program or instruction to execute the method described in the first aspect or any possible implementation of the first aspect, or to execute the method described in the second aspect or any possible implementation of the second aspect. The communication interface in the chip may be an input / output interface, a pin, or a circuit, etc.

[0039] In one possible implementation, the chip or chip system described above in this application further includes at least one memory, in which instructions are stored. The memory may be a storage unit within the chip, such as a register, a cache, etc., or a storage unit of the chip (e.g., a read-only memory, a random access memory, etc.).

[0040] It should be understood that the second to eighth aspects of the present application correspond to the technical solutions of the first or second aspect of the present application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation methods are similar and will not be repeated. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0042] Figure 2 A schematic diagram of the software structure of an electronic device provided in an embodiment of the present application;

[0043] Figure 3 A schematic diagram of music separation and extraction provided in an embodiment of the present application;

[0044] Figure 4 A schematic diagram of a sound object combination provided in an embodiment of the present application;

[0045] FIG5( a ) is a comparison diagram of the sound extraction effects of an L1 solution and an L2 solution provided in an embodiment of the present application;

[0046] FIG5( b ) is a comparison diagram of the sound extraction effects of another L1 solution and an L2 solution provided in an embodiment of the present application;

[0047] Figure 6 A comparison chart of the sound extraction effects of an L1 loss function and an L1 solution provided in an embodiment of the present application;

[0048] Figure 7 A schematic diagram of audio processing after extracting sound objects provided in an embodiment of the present application;

[0049] Figure 8A schematic diagram of an audio processing method provided in an embodiment of the present application;

[0050] Figure 9 A schematic diagram of a training method for an audio processing model provided in an embodiment of the present application;

[0051] Figure 10 A schematic diagram of the structure of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] To facilitate a clear description of the technical solutions of the embodiments of the present application, some of the terms and technologies involved in the embodiments of the present application are briefly introduced below:

[0053] 1. Terminology

[0054] In the embodiments of this application, terms such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. For example, the terms "first chip" and "second chip" are used solely to distinguish between different chips and do not define their order. Those skilled in the art will understand that terms such as "first" and "second" do not define the quantity or execution order, and do not necessarily define differences.

[0055] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0056] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, a--c, bc, or abc, where a, b, c can be single or multiple.

[0057] 2. Electronic devices

[0058] The electronic device of the embodiment of the present application may also be a terminal device in any form. For example, the electronic device may include: a mobile phone, a tablet computer, a PDA, a laptop computer, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, a cellular phone, a cordless phone, a session initiation protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA), a handheld device with wireless communication function, a computing device or other processing device connected to a wireless modem, an in-vehicle device, a wearable device, an electronic device in a 5G network or a future evolved public land mobile communication network (PLMN) Mobile network, PLMN) and other electronic devices, and the embodiments of the present application are not limited to this.

[0059] As an example and not a limitation, in the embodiments of the present application, the electronic device may also be a wearable device. Wearable devices may also be referred to as wearable smart devices, which are a general term for wearable devices that are intelligently designed and developed using wearable technology for daily wear, such as glasses, gloves, watches, clothing, and shoes. A wearable device is a portable device that is worn directly on the body or integrated into the user's clothes or accessories. Wearable devices are not only hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are fully functional, large in size, and can achieve complete or partial functions without relying on smartphones, such as smart watches or smart glasses, as well as those that only focus on a certain type of application function and need to be used in conjunction with other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.

[0060] In addition, in the embodiment of the present application, the electronic device can also be an electronic device in the Internet of Things (IoT) system. IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection.

[0061] The electronic devices in the embodiments of the present application may also be referred to as: user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, etc.

[0062] In the embodiments of the present application, the electronic device or each network device includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory (also known as main memory). The operating system can be any one or more computer operating systems that implement business processing through processes, such as a Linux operating system, a Unix operating system, an Android operating system, an iOS operating system, or a Windows operating system. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software.

[0063] For example, Figure 1 A schematic structural diagram of an electronic device is shown.

[0064] The electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0065] It is understood that the structures illustrated in the embodiments of the present application do not constitute specific limitations on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0066] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors. The controller may generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.

[0067] The processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory may store instructions or data that the processor 110 has just used or is reusing. If the processor 110 needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system. For example, in an embodiment of the present application, the processor 110 may be used to process the relevant processes of extracting sound objects based on a loss function of an audio processing model.

[0068] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is only for illustrative purposes and does not constitute a structural limitation on the electronic device. In other embodiments of the present application, the electronic device may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0069] The internal memory 121 can be used to store computer executable program code, and the executable program code includes instructions. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function, etc. The data storage area may store data created during the use of the electronic device, etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device by running instructions stored in the internal memory 121, and / or instructions stored in a memory provided in the processor. For example, in an embodiment of the present application, the internal memory 121 can be used to store extracted sound objects, and can also be used to store relevant codes executed by the audio processing model, etc.

[0070] The electronic device can implement audio functions, such as audio playback or recording, through the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor. For example, in the embodiment of the present application, the electronic device can extract multiple sound objects in the audio, thereby performing different processing on different sounds to obtain an audio signal that meets the listening requirements.

[0071] Figure 2This is a block diagram of the software structure of the electronic device in an embodiment of the present application. The layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers: the application layer, the application framework layer, the Android runtime and system library, the hardware abstraction layer (HAL), and the kernel layer.

[0072] The application layer can also be called the application layer, which can include a series of application packages. Figure 2 As shown, the application package can include applications such as phone, music, calendar, camera, game, memo, video, etc. Applications can include system applications and third-party applications.

[0073] The application framework layer, also known as the Framework layer, provides an application programming interface (API) and programming framework for applications in the application layer. The Framework layer may include some predefined functions.

[0074] like Figure 2 As shown, the Framework layer may include an activity manager, a window manager, a resource manager, a notification manager, a content provider, a view system, etc. For details, please refer to the relevant technologies and will not be described in detail.

[0075] The Android runtime includes the core library and the virtual machine. The Android runtime is responsible for controlling and managing the Android system.

[0076] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.

[0077] The application layer and framework layer run in a virtual machine. The virtual machine executes the Java files in the application and framework layers as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection. For example, in the embodiments of the present application, the virtual machine can be used to execute functions such as extracting sound objects based on a loss function in an audio processing model.

[0078] The system library can also be called the Native layer. The Native layer can include multiple functional modules, such as the media library, function library, and graphics processing library.

[0079] The HAL is an abstract layer between the kernel layer and the Android runtime. The hardware abstraction layer can be a package for hardware drivers, providing a unified interface for upper-layer applications to call.

[0080] The kernel layer is the layer between hardware and software. The kernel layer may include display driver, camera driver, audio driver, battery driver, Bluetooth driver, CPU driver, USB driver, etc.

[0081] It should be noted that the embodiments of the present application are only illustrated using the Android system as an example. In other operating systems (such as Windows system, IOS system, etc.), as long as the functions implemented by each functional module are similar to those in the embodiments of the present application, the solutions of the present application can also be implemented.

[0082] Users can use electronic devices to listen to music, watch videos, and watch movies. In order to provide users with a better listening experience, electronic devices can perform music source separation on the sounds in music, videos, and movies. Music source separation can be understood as extracting the sound components in the music from the mixed original music and processing them as separate channels. In this way, different sound processing methods are adopted for different types of sound components to create better sound effects. Among them, the original music can also be called source music or original audio. For the sake of convenience, the original audio will be used in the following explanations.

[0083] It's understandable that since each sound component can be considered an object, in some scenarios, a sound component can also be referred to as a sound object. A sound object can be understood as a sound category extracted from multiple sound categories in the mixed original audio. For example, the original audio may include multiple sound categories such as vocals, drums, and bass, and a sound object can be any of these multiple sound categories.

[0084] like Figure 3 As shown, for music, the original audio may include, but is not limited to, one or more of the following sound objects: vocals, drums, bass, piano, and guitar. For audio from videos or movies, the original audio may include vocals, background music, sound effects, ambient sounds, noise, and other sounds. Background music may include drums, bass, piano, guitar, and / or other instruments.

[0085] In some implementations, electronic devices can use audio processing models to extract sound objects. The audio processing models can use the p-norm of time domain data or frequency domain data as a training target. In some scenarios, the audio processing models can also be referred to as music source separation models or sound source separation models.

[0086] The p-norm is a standard training objective for regression tasks. It can be used to measure the amplitude or energy of a signal, or to describe the size of a vector or function. The p-norm can include the L1 norm and the L2 norm.

[0087] The L1 norm tends to produce sparser results. In some scenarios, the L1 norm is also referred to as the L1 loss function or L1 loss. For a consistent description, the L1 loss function will be used as an example in the following explanations. As you can understand, since the L1 loss function's gradient is 1 when the loss function value is close to 0, the L1 loss function tends to extract cleaner results, but it is prone to losing subtle details.

[0088] The L2 norm can be used to measure the energy or amplitude of a signal. In some scenarios, the L2 norm is also referred to as the L2 loss function or L2 loss. For a consistent description, the L2 loss function will be used as an example in the following explanations. Because the gradient of the L2 loss function gradually decays as the loss value approaches 0, the L2 loss function tends to preserve spectral details and produce more complete results, but it is prone to residual signals or background noise.

[0089] In other implementations, the audio processing model can use the signal to noise ratio loss (SNR Loss) as a training target. The signal to noise ratio loss function is based on the evaluation index SNR of music source separation. It is understandable that in the music source separation task, the performance of the signal to noise ratio loss function alone is not as good as the L1 loss function or the L2 loss function. Because the signal to noise ratio loss function is not sensitive to the signal size, the ratio of the separation result to the target result is the main factor that determines the size of the loss function. When the large signal has not been well separated, the rules of the small signal will be more difficult to learn. Therefore, using the signal to noise ratio loss function alone is not conducive to training. However, the appropriate use of the signal to noise ratio loss function can enhance the audio processing model's ability to learn small signals and signal vector integrity, which is beneficial to the improvement of small signals. In particular, in an embodiment of the present application, small signals may include signals with a higher missing rate and / or a smaller amplitude.

[0090] Furthermore, most existing loss functions calculate each sound object independently, failing to account for the interrelationships between them. For example, piano and guitar, or drums and bass, often form a chordal relationship and are not independent of each other. Therefore, the loss function calculations for two sound objects should not be independent of each other. Therefore, it is possible to combine individual sound objects and calculate the loss function using the combined prediction result and the target result to learn the interactions between them.

[0091] Signals have different characteristics in the time domain and frequency domain. Therefore, many existing loss functions are constructed by considering both time domain and frequency domain signals.

[0092] In view of this, the audio processing method provided in the embodiments of the present application can consider each sound object in the sound source, calculate the L1 loss function of the combination of one or more sound objects in the time domain and frequency domain, or the L2 loss function in the time domain and frequency domain, and combine it with the signal-to-noise ratio loss function to calculate a comprehensive loss function as the training target of the audio processing model. In this way, the interaction between different sound objects can be learned by combining the time domain and frequency domain characteristics of the audio data, resulting in an audio processing model that can preserve spectral details and has better results.

[0093] The following describes the method of the embodiment of the present application in detail through specific embodiments. The following embodiments can be combined with each other or implemented independently, and the same or similar concepts or processes may not be repeated in some embodiments.

[0094] It is understood that the extraction of different sound objects is not necessarily independent of each other. In order to learn the interactions between different sound objects, in embodiments of the present application, all combinations of the sound objects to be extracted can be considered when calculating the loss function. For example, all combinations of the sound objects to be extracted can be listed, and the sum of all sound objects in the combination can be used as the audio for calculating the loss function.

[0095] like Figure 4 As shown, the embodiment of the present application takes the extraction of 4 sound objects as an example to show a schematic diagram of the sound object combination.

[0096] For example, when each sound object is a combination, there can be 4 combinations, including sound object 1, sound object 2, sound object 3, and sound object 4. It can also be understood that when each sound object is a combination, there can be combinations.

[0097] When every two sound objects form a combination, there can be 6 combinations, including the combination of sound object 1 and sound object 2, the combination of sound object 1 and sound object 3, the combination of sound object 1 and sound object 4, the combination of sound object 2 and sound object 3, the combination of sound object 2 and sound object 4, and the combination of sound object 3 and sound object 4. It can also be understood that when every two sound objects form a combination, there can be combinations.

[0098] When every three sound objects form a combination, there can be four combinations, including the combination of sound object 1, sound object 2 and sound object 3, the combination of sound object 1, sound object 2 and sound object 4, the combination of sound object 1, sound object 3 and sound object 4, and the combination of sound object 2, sound object 3 and sound object 4. It can also be understood that when every three sound objects form a combination, there can be four combinations. combinations.

[0099] Therefore, the total number N of combinations of 4 sound objects can satisfy the following formula:

[0100]

[0101] By analogy, the total number N of combinations of t sound objects can satisfy the following formula:

[0102]

[0103] The formula can be understood as including t sound objects, j changes from 1 to t-1, For example, taking t=4 as an example, the above formula can be converted to

[0104] In this way, considering the combination of sound objects to be extracted and summing the sound objects in the combination as the audio for calculating the loss function can enable the audio processing model to learn the relationship between multiple sound objects.

[0105] For example, the audio intervals of some sound objects overlap, or the characteristics of the sound objects are similar. Take drums and bass as an example. Both drums and bass are in the low-frequency range. If they are extracted separately, their respective characteristics cannot be better learned. However, if they can be considered together, the audio processing model can first learn and analyze the drums and bass, learn the overall characteristics of the combined sound, and then make the signal envelope more accurate when learning the characteristics of each sound. More accurate learning results can be obtained, thereby extracting a more complete sound object. In addition, when two sound objects are often in a chord relationship (such as piano and guitar), the appearance of one sound object is often accompanied by the appearance of the other sound object in a certain pattern. Using a combined loss function can better capture the relationship between the two.

[0106] It can be understood that since the L1 loss function and the L2 loss function are more conducive to network convergence, the signal-to-noise ratio loss function is beneficial to the retention of small signals and the integrity of the spectrum vector, and the time domain loss function can enable the audio processing model to learn a more accurate envelope, and the frequency domain loss function can enable the audio processing model to have a more complete spectrum structure. Therefore, the embodiment of the present application can comprehensively consider the time domain information and the frequency domain information, calculate the L1 loss function and the signal-to-noise ratio loss function based on the L1 norm on the time domain waveform data and the frequency domain spectrum data, or calculate the L2 loss function and the signal-to-noise ratio loss function based on the L2 norm, and select a suitable hyperparameter α for matching. In this way, the audio processing model can better learn the characteristics of the signal in the time domain and frequency domain.

[0107] In one possible implementation, in an embodiment based on the L2 loss function, for each combination i of sound objects to be extracted, the corresponding loss function satisfies the following formula:

[0108]

[0109] Among them, Loss i It represents the loss function of the i-th combination finally calculated. The i-th combination can be understood as any combination of the N combinations of the t sound objects mentioned above. The i-th combination can also be referred to as combination i.

[0110] represents the L2 loss function of the frequency domain spectrum data of combination i, α1 represents the hyperparameter of the L2 loss function of the frequency domain spectrum data, which is used to represent In Los i In some scenarios, hyperparameters can also be called ratio hyperparameters.

[0111] represents the frequency domain signal-to-noise ratio loss function based on the L2 norm of combination i. In some scenarios, It can also be called the L2 version of the signal-to-noise ratio loss function of the frequency domain spectrum data of combination i, that is, the signal-to-noise ratio loss function of the frequency domain spectrum data implemented based on the L2 scheme. α2 represents the hyperparameter of the frequency domain signal-to-noise ratio loss function based on the L2 norm, which is used to represent In Los i The ratio in .

[0112] represents the L2 loss function of the time domain waveform data of combination i, α3 represents the hyperparameter of the L2 loss function of the time domain waveform data, which is used to represent In Los i The ratio in .

[0113] represents the time domain signal-to-noise ratio loss function based on the L2 norm of combination i. In some scenarios, It can also be called the L2 version of the signal-to-noise ratio loss function of the time domain waveform data of combination i, that is, the signal-to-noise ratio loss function of the time domain waveform data implemented based on the L2 scheme. α4 represents the hyperparameter of the time domain signal-to-noise ratio loss function based on the L2 norm, which is used to represent In Los i The ratio in .

[0114] In this embodiment of the application, the values of the hyperparameters α1, α2, α3 and α4 can be adjusted to make the Loss i The values of each data in the are more balanced, so that Loss iIt can consider the influence of various factors in a balanced manner, that is, it has a comprehensive consideration of the L2 loss function in the frequency domain and time domain, as well as the signal-to-noise ratio loss function. It is understandable that the hyperparameters can be parameters set before training the model, and the values of the hyperparameters can be determined through experiments and verification. In a possible implementation, according to laboratory tests, the values of α1, α2, α3 and α4 can be in the range of 0 to 10, for example, α1 can be set to 0.1, α2 can be set to 3, α3 can be set to 0.01, and α4 can be set to 0.5. In the embodiment of the present application, the specific values of α1, α2, α3 and α4 are not limited, and can be set in advance in the electronic device, or can be flexibly adjusted according to actual conditions.

[0115] In possible implementations, the frequency domain and time domain The calculation method is similar, and both satisfy the calculation formula of the L2 loss function:

[0116]

[0117] in, represents the L2 loss function of combination i, Can include frequency domain It can also include time domain y i is the estimated sound object of combination i, The target sound object for group i.

[0118] It can be understood that the estimated sound object is a sound object of any combination of N sound object combinations composed of t sound objects extracted by the audio processing model, and the estimated sound object can be understood as the estimated audio data extracted by the audio processing model. The target sound object is a sound object of any combination of N sound object combinations composed of t sound objects in the original audio, and the target sound object can be understood as the real audio data in the original audio. In some scenarios, the estimated sound object can also be referred to as an estimated signal or an estimated audio object, and the target sound object can also be referred to as a target signal or a target audio object, which is not limited in the embodiments of the present application.

[0119] In the frequency domain When calculating the time domain, the estimated sound object can be an estimated sound object in the frequency domain, and the target sound object can be a target sound object in the frequency domain. When the estimated sound object may be an estimated sound object in the time domain, the target sound object may be a target sound object in the time domain.

[0120] Frequency domain and time domain The calculation method is similar, and both satisfy the calculation formula of the signal-to-noise ratio loss function:

[0121]

[0122] in, represents the signal-to-noise ratio loss function based on the L2 norm of combination i, It can include a frequency domain signal-to-noise ratio loss function based on the L2 norm, or a time domain signal-to-noise ratio loss function based on the L2 norm. x represents the original audio, which can be understood as a mixed signal of multiple sound objects or as the original audio to be extracted. (y i ) T is the transpose of the estimated sound object of combination i. i ‖ represents the bi-norm of the estimated sound object of combination i, Represents the two-norm of the target sound object of combination i.

[0123] It is understood that when calculating the frequency domain signal-to-noise ratio loss function based on the L2 norm, the second norm of the estimated sound object of combination i can be the second norm of the estimated sound object in the frequency domain, and the second norm of the target sound object can be the second norm of the target sound object in the frequency domain. When calculating the time domain signal-to-noise ratio loss function based on the L2 norm, the second norm of the estimated sound object can be the second norm of the estimated sound object in the time domain, and the second norm of the target sound object can be the second norm of the target sound object in the time domain.

[0124] ρ i Satisfy the following formula:

[0125]

[0126] It is understandable that ρ i and (1-ρ i ) can be used to proportionally coordinate the values of the first and second terms, ρ i It can be an adaptive value, that is, according to x and y i The value calculated from the value of .

[0127] because And Loss i in like It is not easy to formulate reasonable hyperparameter ratio values. You can Add 1, so

[0128] Add the loss functions of N sound object combinations to obtain the final loss function that satisfies the following formula:

[0129]

[0130] Where N is the total number of sound object combinations, i varies from 1 to N, and the loss functions of N combinations are accumulated.

[0131] In a possible implementation, the original audio is input into the audio processing model. After the model is calculated, multiple sound objects can be output from the original audio, such as human voice, drum sound, bass, etc. Taking three sound objects as an example, the audio processing model can output There are 6 combinations, namely sound object 1, sound object 2, sound object 3, the combination of sound object 1 and sound object 2, the combination of sound object 1 and sound object 3, and the combination of sound object 2 and sound object 3. The final loss function is the sum of the loss functions of the 6 combinations:

[0132] The estimated sound object may include the estimated sound object y of combination 1 1 , estimated sound object y of combination 2 2 , estimated sound object y of combination 3 3 , estimated sound object y of combination 4 4 , estimated sound object y of combination 5 5 , estimated sound object y of combination 6 6 During the training of the audio processing model, the parameters in the audio processing model can be updated based on the loss function in the L2 scheme, and the results can be continuously optimized so that the estimated sound object gradually approaches the target sound object, thereby more accurately separating the music sources.

[0133] It is understandable that for the signal-to-noise ratio loss function based on the L2 norm In the calculation formula, the first item is the estimated sound object calculated by the audio processing model, and the second item is the original audio minus the estimated result. As can be seen from the formula, the calculation of the first and second items are complementary. Taking 4 sound objects as an example, when the combination is sound object 1, the first item calculates sound object 1, and the second item calculates the combination of sound object 2, sound object 3 and sound object 4; and when the combination is sound object 2, sound object 3 and sound object 4, the first item calculates the combination of sound object 2, sound object 3 and sound object 4, and the second item calculates sound object 1. Therefore, it can be understood that the combination of each sound object is symmetrical when calculating the signal-to-noise ratio loss function, so only half of the combined signal-to-noise ratio loss function can be calculated. In this way, the computing power of electronic equipment can be reduced and the audio processing speed can be improved.

[0134] In another possible implementation, in an embodiment based on the L1 loss function, for each combination i of sound objects to be extracted, the corresponding loss function satisfies the following formula:

[0135]

[0136] Among them, Loss i Denotes the loss function of the finally calculated i-th combination. Similar to the above embodiment, the i-th combination can be understood as any combination of the N combinations of the t sound objects, and the i-th combination can also be referred to as combination i.

[0137] represents the L1 loss function of the frequency domain spectrum data of combination i, α5 represents the hyperparameter of the L1 loss function of the frequency domain spectrum data, which is used to represent In Los i The ratio in .

[0138] represents the frequency domain signal-to-noise ratio loss function based on the L1 norm of combination i. In some scenarios, It can also be understood as the L1 version of the signal-to-noise ratio loss function of the frequency domain spectrum data of combination i, that is, the signal-to-noise ratio loss function of the frequency domain spectrum data implemented based on the L1 scheme. α6 represents the hyperparameter of the frequency domain signal-to-noise ratio loss function based on the L1 norm, which is used to represent In Los i The ratio in .

[0139] represents the L1 loss function of the time domain waveform data of combination i, α7 represents the hyperparameter of the L1 loss function of the time domain waveform data, which is used to represent In Los i The ratio in .

[0140] represents the time domain signal-to-noise ratio loss function based on the L1 norm of combination i. In some scenarios, It can also be understood as the L1 version of the signal-to-noise ratio loss function of the time domain waveform data of combination i, that is, the signal-to-noise ratio loss function of the time domain waveform data implemented based on the L1 scheme. α8 represents the hyperparameter of the time domain signal-to-noise ratio loss function based on the L1 norm, which is used to represent In Los i The ratio in .

[0141] Similarly, in this embodiment of the application, the values of α5, α6, α7 and α8 can be adjusted to make Loss i The values of the various data in the can be more balanced, so that the Loss i The effects of various factors are balanced, i.e., the L1 loss function in the frequency or time domain, as well as the signal-to-noise ratio loss function, are comprehensively considered. There are no specific restrictions on the values of α5, α6, α7, and α8. These can be pre-set in the electronic device or flexibly adjusted based on actual conditions.

[0142] In possible implementations, the frequency domain and time domain The calculation method is similar, and both satisfy the calculation formula of the L1 loss function:

[0143]

[0144] in, represents the L1 loss function of combination i, Can include frequency domain It can also include time domain y i is the estimated sound object of combination i, The target sound object for group i.

[0145] It is understandable that in the frequency domain When calculating the time domain, the estimated sound object can be an estimated sound object in the frequency domain, and the target sound object can be a target sound object in the frequency domain. When the estimated sound object may be an estimated sound object in the time domain, the target sound object may be a target sound object in the time domain.

[0146] Frequency domain and time domain The calculation method is similar, and both satisfy the calculation formula of the signal-to-noise ratio loss function:

[0147]

[0148] in, represents the signal-to-noise ratio loss function based on the L1 norm of combination i, It can include a frequency domain signal-to-noise ratio loss function based on the L1 norm, or a time domain signal-to-noise ratio loss function based on the L1 norm. x is the original audio, which can be understood as a mixed signal of multiple sound objects, or as the original audio to be extracted. (y i ) T is the transpose of the estimated sound object of combination i. |y i | represents the absolute value of the estimated sound object of combination i, Indicates the absolute value of the target sound object for combination i.

[0149] It is understood that when calculating the frequency domain signal-to-noise ratio loss function based on the L1 norm, the absolute value of the estimated sound object of combination i can be the absolute value of the estimated sound object in the frequency domain, and the absolute value of the target sound object can be the absolute value of the target sound object in the frequency domain. When calculating the time domain signal-to-noise ratio loss function based on the L1 norm, the absolute value of the estimated sound object can be the absolute value of the estimated sound object in the time domain, and the absolute value of the target sound object can be the absolute value of the target sound object in the time domain.

[0150] ρ i Satisfies the following formula:

[0151]

[0152] It is understandable that ρ i and (1-ρ i ) can be used to proportionally coordinate the values of the first and second terms, ρ i It can be an adaptive value, that is, according to x and y i The value calculated from the value of .

[0153] because And Loss i in like It is not easy to formulate reasonable hyperparameter ratio values. You can Add 1, so

[0154] Add the loss functions of N sound object combinations to obtain the final loss function that satisfies the following formula:

[0155]

[0156] Where N is the total number of sound object combinations, i varies from 1 to N, and the loss functions of N combinations are accumulated.

[0157] For the sake of convenience of description, the embodiments of the present application refer to the implementation method based on the L1 loss function as the L1 solution, and the implementation method based on the L2 loss function as the L2 solution.

[0158] Figures 5(a) and 5(b) show the sound extraction results obtained using the L1 and L2 schemes in the embodiments of this application. Figure 5(a) shows the sound extraction results for a sound segment, while Figure 5(b) shows the sound extraction results for a silent segment. As can be seen from the figures, the L1 scheme tends to produce cleaner extraction results. The L2 scheme can retain some details of the background noise, resulting in a more complete extraction result.

[0159] It is understood that the specific choice of L1 or L2 solution can be made based on actual needs. For example, if you want to obtain a cleaner extraction result, you can use the L1 solution; if you want to retain some details of the background noise to ensure a more comfortable listening experience, you can use the L2 solution. This embodiment of the application is not limited to this.

[0160] For example, the L1 loss function and the L1 scheme are used to extract the sound objects in the original audio. Figure 6 As shown, the original audio may include multiple sound categories, for example, category 1, category 2, category 3, category 4, category 5, category 6, etc. Figure 6 (a) is the sound object extraction effect diagram when only L1 loss function is used. Figure 6 (b) is the sound object extraction effect diagram when using the L1 scheme.

[0161] It is understandable that in Figure 6 a and Figure 6 In b, the sound objects in the example area have low loudness and a high missing rate, resulting in different extraction results using different schemes. Low loudness can include a low average loudness or a low maximum loudness. A high missing rate can be interpreted as sound objects that occasionally appear in the original audio, such as ringtones and horns. These sound objects account for a small proportion of the original audio and are relatively sparse.

[0162] As can be seen from the zoomed-in image, when using only the L1 loss function, the sound objects in the sample area are sometimes lost. This means that the L1 loss function tends to miss details with low loudness and a high loss rate. Using the L1 solution preserves the details of the sound objects in the sample area, meaning that details with low average loudness and a high loss rate are not lost. Therefore, the L1 solution is more effective for extracting sounds with low loudness, high loss rates, and sparse sound quality.

[0163] Alternatively, you can use both the L2 loss function and the L2 scheme to extract the sound objects from the original audio. Compared to using only the L2 loss function, the L2 scheme can also avoid data loss in categories with low loudness and high missing rates, resulting in better extraction results. This will not be discussed in detail here.

[0164] like Figure 7 As shown, after extracting the sound objects in the original audio, the electronic device can process each sound object. For example, taking the application of human voice and drum sound by the electronic device as an example, the electronic device can perform gain enhancement on the human voice and drum sound. After the gain enhancement of the human voice, the center vocals and background vocals can also be extracted from the human voice by similarity calculation. Among them, the center vocals can include the lead singer or main vocal part of the human voice. The center vocals occupy a central position in the human voice and are the part that the audience pays more attention to. The background vocals can include the part of the human voice that assists the center vocals, which is used to support and enhance the center vocals. For example, the background vocals can include harmonies, repeated phrases, shouts, etc.

[0165] Electronic devices can use the sound field expansion module to expand the sound field of background vocals, gain-enhanced drum sounds, and other sound objects, and mix them with the center vocals to output the processed audio. Sound field expansion can improve the overall quality and effect of the audio experience by optimizing the spatial distribution and positioning of sound. For example, sound field expansion can make listeners feel that sound is coming from all directions, not just the front or left and right sides, enhancing the listener's sense of immersion and making them feel like they are at the scene or in a specific environment. Sound field expansion can also perform different sound field expansion methods on sounds in different frequency bands, making the sound clearer and more layered.

[0166] It is understandable that in the processed audio, the vocal components and drum sounds can be more obvious, for example, the vocals are clearer, the drum sounds are more powerful, etc. Of course, the electronic device can also process other sound objects separately, and this embodiment of the application is not limited thereto.

[0167] The audio processing method of the embodiment of the present application can be applied to various audio separation scenarios, such as the extraction of sound objects in music, videos, and movies mentioned above, and can also be applied to the extraction of sound objects in noise. It is not limited to specific audio scenarios and has good portability and versatility.

[0168] Figure 8 An audio processing method provided in an embodiment of the present application is shown, which is applied to an electronic device. The method includes:

[0169] S801: Acquire a first audio signal.

[0170] In the embodiment of the present application, the first audio signal can be understood as an audio signal to be processed. The first audio signal can include audio signals in scenes such as music, video, and movie, which are not limited in the embodiment of the present application. For example, the first audio signal can include the original audio or original music in the above embodiment.

[0171] Taking original music as an example, the first audio signal may include k sound categories. For example, the k sound categories may include human voice, drums, bass, piano, guitar, etc. The k sound categories can be understood as the categories of all sounds that make up the first audio signal, and k is a positive integer.

[0172] S802. Extract t sound objects from the first audio signal based on an audio processing model; wherein the audio processing model is trained based on the above-mentioned target loss function, the target loss function is related to the loss function of the i-th combination, the i-th combination is a combination of any one or more sound objects among the t sound objects, the loss function of the i-th combination is related to the p-norm and the signal-to-noise ratio loss function, and t is a positive integer.

[0173] In an embodiment of the present application, the audio processing model can be used to extract different types of sound objects from mixed original audio.

[0174] Here, t can be understood as the number of sound objects that the audio processing model needs to extract from the k sound categories of the first audio signal, and t is less than or equal to k. For example, taking the first audio signal as including five sound categories: vocals, drums, bass, piano, and guitar, if the audio processing model needs to extract vocals and drums from the first audio signal, the value of t is 2; if the audio processing model needs to extract vocals, drums, bass, piano, and guitar from the first audio signal, the value of t is 5.

[0175] The target loss function can be understood as the sum of the loss functions of the combination of t sound objects finally calculated in the above embodiment. The target loss function satisfies the following formula:

[0176]

[0177] Wherein, N represents the total number of combinations of any one or more sound objects among t sound objects, i varies from 1 to N, and the target loss function can be obtained by accumulating the loss functions of N combinations, wherein the loss functions of the N combinations can be any one of the aforementioned loss function implementation plans, and the loss function of each combination can be a deformed loss function, for example, the loss function of each combination can have a corresponding coefficient.

[0178] S803: Post-process one or more sound objects among the t sound objects to obtain a second audio signal.

[0179] In the embodiment of the present application, the process of post-processing one or more of the t sound objects can refer to Figure 7 The relevant descriptions in the corresponding embodiments, such as gain processing of vocals and drum sounds, similarity calculation of vocals to separate the center vocals and background vocals, sound field expansion processing of background vocals, drum sounds and other sounds, and final synthesis processing with the center vocals, etc., are not repeated here.

[0180] The second audio signal can be understood as an audio signal that meets actual needs after performing audio processing on one or more of the t sound objects extracted from the first audio signal based on the audio processing model. For example, Figure 7 In the corresponding embodiment, the drum sound is gain-processed separately. Then, in the processed audio, the drum sound can be more obvious and powerful, thereby improving the user's auditory experience.

[0181] The audio processing method provided in the embodiments of this application can consider the combination of various sound objects in the first audio signal and integrate the p-norm loss function and the signal-to-noise ratio loss function as the training target of the audio processing model. In this way, the interactive relationship between various sound objects can be learned by combining the characteristics of the p-norm and signal-to-noise ratio loss functions, resulting in more accurate separation results, thereby improving the user experience.

[0182] Optional, in Figure 8 On the basis of the corresponding embodiment, the loss function of the i-th combination is related to the L2 loss function of the p-norm and the signal-to-noise ratio loss function based on the L2 norm.

[0183] It is understandable that the p-norm L2 loss function easily produces residues of other signals or background noise, and the signal-to-noise ratio loss function has difficulty learning the laws of the signal. In the embodiment of the present application, the use of the p-norm L2 loss function and the signal-to-noise ratio loss function can not only preserve spectral details and obtain more complete results, but also enhance the audio processing model's ability to learn small signals and signal vector integrity, which is conducive to improving the small signal separation effect.

[0184] Optional, in Figure 8 Based on the corresponding embodiment, the loss function of the i-th combination is related to at least two of the following: the L2 loss function in the frequency domain, the frequency domain signal-to-noise ratio loss function based on the L2 norm, the L2 loss function in the time domain, and the time domain signal-to-noise ratio loss function based on the L2 norm.

[0185] It is understandable that since the L2 loss function can better preserve spectral details and obtain more complete results, the signal-to-noise ratio loss function can improve the learning ability of small signals and signal vector integrity, the time domain loss function can make the separation results of the audio processing model have a more accurate envelope, and the frequency domain loss function can make the audio processing model have a more complete spectral structure.

[0186] Therefore, the embodiments of the present application can comprehensively consider time domain information and frequency domain information, and calculate the L2 loss function and signal-to-noise ratio loss function on the time domain waveform data and frequency domain spectrum data. This can be more conducive to the convergence of the audio processing model, allowing the audio processing model to better learn the characteristics of the time domain and frequency domain.

[0187] Optional, in Figure 8 Based on the corresponding embodiment, the loss function Loss of the i-th combination i Satisfies the following formula:

[0188]

[0189] in, represents the L2 loss function of the i-th combination in the frequency domain, represents the frequency domain signal-to-noise ratio loss function based on the L2 norm of the i-th combination, represents the L2 loss function of the i-th combination in the time domain, represents the time domain signal-to-noise ratio loss function of the i-th combination based on the L2 norm, α1 represents The hyperparameter, α2 represents The hyperparameters, α3 represents The hyperparameters, α4 represents hyperparameters.

[0190] In the embodiment of the present application, the loss function of the i-th combination can refer to the relevant description in the above-mentioned embodiment based on the L2 loss function, and will not be repeated here. As can be seen from the above formula, the loss function of the i-th combination takes into account the L2 loss function in the time domain and the signal-to-noise ratio loss function based on the L2 norm, as well as the L2 loss function in the frequency domain and the signal-to-noise ratio loss function based on the L2 norm. This not only retains a more complete small signal, but also makes the separation result of the audio processing model have a more accurate envelope and a more complete spectral structure.

[0191] Optional, in Figure 8 On the basis of the corresponding embodiment, the L2 loss function of the i-th combination in the frequency domain and the L2 loss function of the i-th combination in the time domain both satisfy the following L2 loss function The formula is:

[0192]

[0193] in, represents the L2 loss function of the i-th combination, y i is the estimated sound object of the i-th combination, represents the target sound object of the i-th combination; the frequency domain signal-to-noise ratio loss function based on the L2 norm of the i-th combination and the time domain signal-to-noise ratio loss function based on the L2 norm of the i-th combination both satisfy the following signal-to-noise ratio loss function formula:

[0194]

[0195] in, represents the signal-to-noise ratio loss function of the i-th combination, x represents the first audio signal, (y i ) T represents the transpose of the estimated sound object of the i-th combination, ‖y i ‖ represents the bi-norm of the estimated sound object of the i-th combination, represents the bi-norm of the target sound object of the i-th combination, ρ i and (1-ρ i ) represent the coordination coefficients respectively.

[0196] In the embodiment of the present application, the L2 loss function and the signal-to-noise ratio loss function can refer to the relevant description in the above embodiment based on the L2 loss function, and will not be repeated here. i and (1-ρ i ) can be used to proportionally coordinate the values of the first and second terms.

[0197] It can be understood that when calculating the loss function in the frequency domain, Represents the L2 loss function in the frequency domain, the estimated sound object can be the estimated sound object in the frequency domain, and the target sound object can be the target sound object in the frequency domain. When calculating the loss function in the time domain, represents an L2 loss function in the time domain, the estimated sound object may be an estimated sound object in the time domain, and the target sound object may be a target sound object in the time domain.

[0198] The audio processing model can be trained based on the L2 loss function in the time domain and frequency domain, as well as the signal-to-noise ratio loss function in the time domain and frequency domain, so as to update the parameters in the audio processing model and continuously optimize the results so that the estimated sound object gradually approaches the target sound object, thereby more accurately separating the sound objects in the audio.

[0199] Optional, in Figure 8 On the basis of the corresponding embodiment, the loss function of the i-th combination is related to the L1 loss function of the p-norm and the signal-to-noise ratio loss function based on the L1 norm.

[0200] It is understandable that the p-norm L1 loss function tends to lose detail signals with lower loudness, while the signal-to-noise ratio loss function is beneficial for improving small signal separation. In the embodiments of the present application, the use of the p-norm L1 loss function and the signal-to-noise ratio loss function not only tends to extract cleaner results, but also enhances the audio processing model's ability to learn small signals and signal vector integrity.

[0201] Optional, in Figure 8 Based on the corresponding embodiments, the loss function of the i-th combination is related to at least two of the following: the L1 loss function in the frequency domain, the frequency domain signal-to-noise ratio loss function based on the L1 norm, the L1 loss function in the time domain, and the time domain signal-to-noise ratio loss function based on the L1 norm.

[0202] It can be understood that since the L1 loss function can better extract cleaner results, the signal-to-noise ratio loss function can improve the learning ability of small signals and signal vector integrity, the time domain loss function can make the separation result of the audio processing model have a more accurate envelope, and the frequency domain loss function can make the audio processing model have a more complete spectral structure. Therefore, the embodiment of the present application can comprehensively consider the time domain information and the frequency domain information, and calculate the L1 loss function and the signal-to-noise ratio loss function on the time domain waveform data and the frequency domain spectrum data. In this way, the audio processing model can better learn the characteristics of the time domain and the frequency domain, and obtain a cleaner audio separation result.

[0203] Optional, in Figure 8 Based on the corresponding embodiment, the loss function Loss of the i-th combination i Satisfies the following formula:

[0204]

[0205] in, represents the L1 loss function of the i-th combination in the frequency domain, represents the frequency domain signal-to-noise ratio loss function based on the L1 norm of the i-th combination, represents the L1 loss function of the i-th combination in the time domain, represents the time domain signal-to-noise ratio loss function based on the L1 norm of the i-th combination, and α5 represents The hyperparameters, α6 represents The hyperparameters, α7 represents The hyperparameters, α8 represents hyperparameters.

[0206] In the embodiment of the present application, the loss function of the i-th combination can refer to the relevant description in the above-mentioned embodiment based on the L1 loss function, and will not be repeated here. As can be seen from the above formula, the loss function of the i-th combination takes into account the L1 loss function in the time domain and the signal-to-noise ratio loss function based on the L1 norm, as well as the L1 loss function in the frequency domain and the signal-to-noise ratio loss function based on the L1 norm. This can obtain a cleaner extraction result, and can achieve better extraction effects when extracting sounds with lower loudness, higher missing rates, and sparser sounds.

[0207] Optional, in Figure 8 On the basis of the corresponding embodiment, the L1 loss function of the i-th combination in the frequency domain and the L1 loss function of the i-th combination in the time domain both satisfy the following L1 loss function The formula is:

[0208]

[0209] in, represents the L1 loss function of the i-th combination, y i is the estimated sound object of the i-th combination, represents the target sound object of the i-th combination; the frequency domain signal-to-noise ratio loss function based on the L1 norm of the i-th combination and the time domain signal-to-noise ratio loss function based on the L1 norm of the i-th combination both satisfy the following signal-to-noise ratio loss function formula:

[0210]

[0211] in, represents the signal-to-noise ratio loss function of the i-th combination, x represents the first audio signal, (y i ) T represents the transpose of the estimated sound object of the i-th combination, |y i | represents the absolute value of the estimated sound object of the i-th combination, represents the absolute value of the target sound object of the i-th combination, ρ i and (1-ρ i ) represent the coordination coefficients respectively.

[0212] In the embodiment of the present application, the L1 loss function and the signal-to-noise ratio loss function can refer to the relevant description in the above embodiment based on the L1 loss function, and will not be repeated here. i and (1-ρ i ) can be used to proportionally coordinate the values of the first and second terms.

[0213] It can be understood that when calculating the loss function in the frequency domain, Represents the L1 loss function in the frequency domain, the estimated sound object can be the estimated sound object in the frequency domain, and the target sound object can be the target sound object in the frequency domain. When calculating the loss function in the time domain, represents an L1 loss function in the time domain, the estimated sound object may be an estimated sound object in the time domain, and the target sound object may be a target sound object in the time domain.

[0214] The audio processing model can be trained based on the L1 loss function in the time domain and frequency domain, as well as the signal-to-noise ratio loss function in the time domain and frequency domain, so as to update the parameters in the audio processing model and continuously optimize the results so that the estimated sound object gradually approaches the target sound object, thereby more accurately separating the sound objects in the audio.

[0215] Optional, in Figure 8 On the basis of the corresponding embodiment, post-processing is performed on one or more sound objects among the t sound objects, including performing one or more of the following processing on the one or more sound objects: gain processing, similarity calculation, and sound field expansion.

[0216] In the embodiment of the present application, the process of post-processing one or more sound objects in the t sound objects can refer to Figure 7 The relevant descriptions in the corresponding embodiments are omitted here.

[0217] Gain processing can adjust the intensity of the sound object, thereby balancing the volume of different sound objects, making the second audio signal more harmonious and improving the auditory effect.

[0218] Similarity calculation can separate the center voice and background voice in the human voice, so as to perform different audio processing on the center voice and background voice respectively, thereby improving the user's listening experience.

[0219] Sound field expansion can improve the overall quality and effect of the audio experience by optimizing the spatial distribution and positioning of sound, enhancing the listener's sense of immersion and making the listener feel as if they are at the scene or in a specific environment. It can also perform different sound field expansion methods on sounds in different frequency bands, making the sound clearer and more layered.

[0220] Figure 9 The present invention provides an embodiment of an audio processing model training method, which is applied to electronic devices. The method includes:

[0221] S901: Input audio signal samples into an audio processing model to be trained.

[0222] In the embodiment of the present application, the audio signal sample can be understood as a sample for training the audio processing model. It is understood that the audio signal sample can include one or more audio signals for model training, and any audio signal for model training can include a signal to be extracted and multiple sound categories in the signal to be extracted. For example, the signal to be extracted can include the above Figure 8 For the first audio signal in step S801 of the corresponding embodiment, the multiple sound categories in the signal to be extracted may include t sound objects in the first audio signal in step S802. The first audio signal includes samples during training and a mixed signal input to the model for separation after model training is completed.

[0223] Optionally, the multiple sound categories in the signal to be extracted may further include a combination of multiple objects among the k sound objects constituting the first audio signal, and the combination of the multiple sound objects may be used as one of the t sound objects.

[0224] Audio processing models can be used to extract different types of sound objects from mixed raw audio.

[0225] S902. Adjust the audio processing model to be trained based on the target loss function until the target loss function converges to obtain a trained audio processing model; wherein the target loss function is related to the p-norm and the signal-to-noise ratio loss function.

[0226] In the embodiments of the present application, a target loss function serves as the training target for the audio processing model, quantifying the gap between the predicted values and the true values in the audio processing model. By calculating the value of the target loss function, the performance of the model under the current parameter settings can be evaluated. It is understood that the smaller the value of the target loss function, the closer the audio processing model's prediction results are to the true values, and the better the audio processing model's performance.

[0227] The embodiment of the present application makes the target loss function related to the p-norm and signal-to-noise ratio loss function, and takes into account the combination between each prediction category and the time-frequency domain characteristics of the signal. In this way, the characteristics of the p-norm and signal-to-noise ratio loss function can be combined to learn the interaction relationship between each sound object, and obtain an audio processing model that can retain spectral details and has a better separation effect.

[0228] Optional, in Figure 9 On the basis of the corresponding embodiment, the audio signal sample includes a first audio signal, the first audio signal includes t sound objects, the target loss function is related to the loss function of the i-th combination, the i-th combination is a combination of any one or more sound objects among the t sound objects, and the loss function of the i-th combination is related to the p-norm and signal-to-noise ratio loss function.

[0229] In the embodiment of the present application, the first audio signal can be understood as any one of the audio signals to be extracted from the one or more audio signals used for model training included in the audio signal sample. When the audio signal used for model training is used in actual applications, the first audio signal can also be Figure 8 This corresponds to the first audio signal in the embodiment.

[0230] t can be understood as the sound object that the audio processing model wants to extract from the first audio signal. It can be understood that the first audio signal can also include more than t sound objects. For example, the first audio signal can include k sound objects, and t is less than or equal to k.

[0231] The target loss function can be understood as the sum of the loss functions of the combination of t sound objects finally calculated in the above embodiment. The target loss function satisfies the following formula:

[0232]

[0233] Wherein, N represents the total number of combinations of any one or more sound objects among t sound objects, i varies from 1 to N, and the target loss function can be obtained by accumulating the loss functions of N combinations, wherein the loss functions of the N combinations can be any one of the aforementioned loss function implementation plans, and the loss function of each combination can be a deformed loss function, for example, the loss function of each combination can have a corresponding coefficient.

[0234] Optional, in Figure 9 Based on the corresponding embodiment, the loss function of the i-th combination is related to at least two of the following: the L2 loss function in the frequency domain, the frequency domain signal-to-noise ratio loss function based on the L2 norm, the L2 loss function in the time domain, and the time domain signal-to-noise ratio loss function based on the L2 norm.

[0235] It can be understood that the L2 loss function can better preserve the spectral details and obtain more complete results. The signal-to-noise ratio loss function can improve the learning ability of small signals and signal vector integrity. The time domain loss function can make the separation results of the audio processing model have a more accurate envelope. The frequency domain loss function can make the audio processing model have a more complete spectral structure.

[0236] Therefore, the embodiments of the present application can comprehensively consider time domain information and frequency domain information, and calculate the L2 loss function and signal-to-noise ratio loss function on the time domain waveform data and frequency domain spectrum data. This can be more conducive to the convergence of the audio processing model, allowing the audio processing model to better learn the characteristics of the time domain and frequency domain.

[0237] Optional, in Figure 9 Based on the corresponding embodiments, the loss function of the i-th combination is related to at least two of the following: the L1 loss function in the frequency domain, the frequency domain signal-to-noise ratio loss function based on the L1 norm, the L1 loss function in the time domain, and the time domain signal-to-noise ratio loss function based on the L1 norm.

[0238] It is understandable that the L1 loss function can better extract cleaner results, the signal-to-noise ratio loss function can improve the learning ability of small signals and signal vector integrity, the time domain loss function can make the separation results of the audio processing model have a more accurate envelope, and the frequency domain loss function can make the audio processing model have a more complete spectral structure.

[0239] Therefore, the embodiments of the present application can comprehensively consider time domain information and frequency domain information, and calculate the L1 loss function and signal-to-noise ratio loss function on the time domain waveform data and frequency domain spectrum data. In this way, the audio processing model can better learn the characteristics of the time domain and frequency domain, and obtain cleaner audio separation results.

[0240] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0241] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of method. In order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily appreciate that, in combination with the method steps of each example described in the embodiment disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0242] The embodiment of the present application can divide the functional modules of the device implementing the method according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. In actual implementation, there may be other division methods.

[0243] like Figure 10 FIG. 1 is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip 1000 includes one or more (including two) processors 1001 , a communication line 1002 , a communication interface 1003 , and a memory 1004 .

[0244] In some embodiments, the memory 1004 stores the following elements: executable modules or data structures, or a subset thereof, or an extended set thereof.

[0245] The method described in the above embodiment of the present application can be applied to the processor 1001, or implemented by the processor 1001. The processor 1001 may be an integrated circuit chip with signal processing capabilities. During the implementation process, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor 1001 or an instruction in the form of software. The above-mentioned processor 1001 can be a general-purpose processor (for example, a microprocessor or a conventional processor), a digital signal processor (digital signal processing, DSP), an application specific integrated circuit (application specific integrated circuit, ASIC), a field-programmable gate array (field-programmable gate array, FPGA) or other programmable logic devices, discrete gates, transistor logic devices or discrete hardware components. The processor 1001 can implement or execute the methods, steps and logic block diagrams related to each processing disclosed in the embodiment of the present application.

[0246] The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. Among them, the software module can be located in a storage medium mature in the art such as a random access memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable read only memory (EEPROM). The storage medium is located in the memory 1004, and the processor 1001 reads the information in the memory 1004 and completes the steps of the above method in combination with its hardware.

[0247] The processor 1001 , the memory 1004 , and the communication interface 1003 can communicate with each other via the communication line 1002 .

[0248] In the above embodiment, the instructions stored in the memory for execution by the processor may be implemented in the form of a computer program product, wherein the computer program product may be pre-written in the memory or downloaded and installed in the memory in the form of software.

[0249] The present application also provides a computer program product including one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or data center that includes one or more available media integrated. For example, the available medium can include magnetic media (e.g., floppy disk, hard disk or tape), optical media (e.g., digital versatile disc (DVD)), or semiconductor media (e.g., solid state disk (SSD)).

[0250] The present application also provides a computer-readable storage medium. The methods described in the above embodiments can be implemented in whole or in part via software, hardware, firmware, or any combination thereof. Computer-readable media can include computer storage media and communication media, and can also include any medium that can transfer a computer program from one location to another. The storage medium can be any target medium that can be accessed by a computer.

[0251] As one possible design, computer-readable media may include compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM or other optical disc storage; computer-readable media may include magnetic disk storage or other magnetic disk storage devices. Moreover, any connecting line may also be appropriately referred to as a computer-readable medium. For example, if the software is transmitted from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, DSL or wireless technologies such as infrared, radio and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technologies such as infrared, radio and microwave are included in the definition of medium. Disk and disc as used herein include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks usually reproduce data magnetically, while discs reproduce data optically using lasers.

[0252] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processing unit of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

Claims

1. An audio processing method, applied to electronic equipment, characterized in that: The method comprises: Acquire a first audio signal; extracting t sound objects from the first audio signal based on an audio processing model; wherein the audio processing model is trained based on a target loss function, the target loss function is related to a loss function of an i-th combination, the i-th combination being a combination of any one or more sound objects among the t sound objects, the loss function of the i-th combination being related to a p-norm and a signal-to-noise ratio loss function, and t being a positive integer; Post-process one or more sound objects among the t sound objects to obtain a second audio signal.

2. The method according to claim 1, characterized in that The loss function of the i-th combination is related to the L2 loss function of the p-norm and the signal-to-noise ratio loss function based on the L2 norm.

3. The method according to claim 2, characterized in that The loss function of the i-th combination is related to at least two of the following: an L2 loss function in the frequency domain, a frequency domain signal-to-noise ratio loss function based on an L2 norm, an L2 loss function in the time domain, and a time domain signal-to-noise ratio loss function based on an L2 norm.

4. The method according to claim 2 or 3, characterized in that The loss function Loss of the i-th combination i Satisfy the following formula: in, represents the L2 loss function of the i-th combination in the frequency domain, represents the frequency domain signal-to-noise ratio loss function based on the L2 norm of the i-th combination, represents the L2 loss function of the i-th combination in the time domain, represents the time domain signal-to-noise ratio loss function based on the L2 norm of the i-th combination, α1 represents The hyperparameter, α2 represents The hyperparameters, α3 represents The hyperparameters, α4 represents hyperparameters.

5. The method according to claim 3 or 4, characterized in that The L2 loss function of the i-th combination in the frequency domain and the L2 loss function of the i-th combination in the time domain both satisfy the following L2 loss function formula: in, represents the L2 loss function of the i-th combination, y i is the estimated sound object of the i-th combination, represents the target sound object of the i-th combination; The frequency domain signal-to-noise ratio loss function based on the L2 norm of the i-th combination and the time domain signal-to-noise ratio loss function based on the L2 norm of the i-th combination both satisfy the following signal-to-noise ratio loss function formula: in, represents the signal-to-noise ratio loss function of the i-th combination, x represents the first audio signal, (y i ) T represents the transpose of the estimated sound object of the i-th combination, ‖y i ‖ represents the bi-norm of the estimated sound object of the i-th combination, represents the bi-norm of the target sound object of the i-th combination, ρ i and (1-ρ i ) represent the coordination coefficients respectively.

6. The method according to claim 1, characterized in that The loss function of the i-th combination is related to the L1 loss function of the p-norm and the signal-to-noise ratio loss function based on the L1 norm.

7. The method according to claim 6, characterized in that The loss function of the i-th combination is related to at least two of the following: an L1 loss function in the frequency domain, a frequency domain signal-to-noise ratio loss function based on an L1 norm, an L1 loss function in the time domain, and a time domain signal-to-noise ratio loss function based on an L1 norm.

8. The method according to claim 6 or 7, characterized in that The loss function Loss of the i-th combination i Satisfy the following formula: in, represents the L1 loss function of the i-th combination in the frequency domain, represents the frequency domain signal-to-noise ratio loss function based on the L1 norm of the i-th combination, represents the L1 loss function of the i-th combination in the time domain, represents the time domain signal-to-noise ratio loss function based on the L1 norm of the i-th combination, α5 represents The hyperparameters, α6 represents The hyperparameters, α7 represents The hyperparameters, α8 represents hyperparameters.

9. The method according to claim 7 or 8, characterized in that The L1 loss function of the i-th combination in the frequency domain and the L1 loss function of the i-th combination in the time domain both satisfy the following L1 loss function formula: in, Represents the L1 loss function of the i-th combination, y i is the estimated sound object of the i-th combination, represents the target sound object of the i-th combination; The frequency domain signal-to-noise ratio loss function based on the L1 norm of the i-th combination and the time domain signal-to-noise ratio loss function based on the L1 norm of the i-th combination both satisfy the following signal-to-noise ratio loss function formula: in, represents the signal-to-noise ratio loss function of the i-th combination, x represents the first audio signal, (y i ) T represents the transpose of the estimated sound object of the i-th combination, |y i | represents the absolute value of the estimated sound object of the i-th combination, represents the absolute value of the target sound object of the i-th combination, ρ i and (1-ρ i ) represent the coordination coefficients respectively.

10. The method according to any one of claims 1 to 9, characterized in that The post-processing of one or more sound objects among the t sound objects includes performing one or more of the following processing on the one or more sound objects: gain processing, similarity calculation, and sound field expansion.

11. A training method for an audio processing model, applied to electronic equipment, characterized in that: The method comprises: Input the audio signal sample into the audio processing model to be trained; The audio processing model to be trained is adjusted based on a target loss function until the target loss function converges, thereby obtaining a trained audio processing model; wherein the target loss function is related to a p-norm and a signal-to-noise ratio loss function.

12. The method according to claim 11, characterized in that The audio signal sample includes a first audio signal, which includes t sound objects. The target loss function is related to the loss function of the i-th combination, and the i-th combination is a combination of any one or more sound objects among the t sound objects. The loss function of the i-th combination is related to the p-norm and signal-to-noise ratio loss function.

13. The method according to claim 12, characterized in that The loss function of the i-th combination is related to at least two of the following: an L2 loss function in the frequency domain, a frequency domain signal-to-noise ratio loss function based on an L2 norm, an L2 loss function in the time domain, and a time domain signal-to-noise ratio loss function based on an L2 norm.

14. The method according to claim 12, characterized in that The loss function of the i-th combination is related to at least two of the following: an L1 loss function in the frequency domain, a frequency domain signal-to-noise ratio loss function based on an L1 norm, an L1 loss function in the time domain, and a time domain signal-to-noise ratio loss function based on an L1 norm.

15. An electronic device, characterized in that: The electronic device includes: one or more processors and memory; The memory is coupled to the one or more processors, and the memory is used to store computer program code, where the computer program code includes computer instructions. The one or more processors call the computer instructions to enable the electronic device to execute the method according to any one of claims 1 to 10, or the method according to any one of claims 11 to 14.

16. A chip system, characterized in that: The chip system is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions to enable the electronic device to execute the method according to any one of claims 1-10, or the method according to any one of claims 11-14.

17. A computer-readable storage medium, characterized in that The computer-readable storage medium includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the method according to any one of claims 1 to 10 or the method according to any one of claims 11 to 14.

18. A computer program product, characterized in that The computer program product includes a computer program code, and when the computer program code runs on an electronic device, the electronic device executes the method according to any one of claims 1 to 10, or the method according to any one of claims 11 to 14.

Citation Information

Patent Citations

  • Device voice noise reduction, electronic device and storage medium

    CN114121031A

  • Speech enhancement model training method, speech enhancement model recognition method, electronic equipment and storage medium

    CN114283795A

  • Wireless communication data processing method and lane early warning method

    CN116052452A

  • Voiceprint recognition method, model training method, and server

    US20210050020A1

  • Speech separation model training method and apparatus, storage medium and computer device

    US20220172708A1