Methods, devices and electronic equipment for music library audio quality restoration

By performing high-low frequency separation, semantic encoding, and decoding on low-quality recorded songs, and fusing low-frequency and high-frequency vectors, the problem of low-efficiency sound quality restoration of songs recorded by low-quality recording equipment is solved, achieving efficient and reliable sound quality improvement.

CN120708636BActive Publication Date: 2025-10-28CHENGDU XIAOCHANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511198517.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-10-28
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

In existing technologies, the sound quality restoration efficiency of songs recorded by low-quality recording equipment is low, and it relies on manual editing, which is also inefficient.

Method used

By performing high and low frequency separation processing on the recorded songs to be repaired in the target music library, low-frequency time-domain data and high-frequency time-domain data of the songs are formed. Frequency domain conversion is then performed, and combined with semantic encoding and decoding, the low-frequency vector and high-frequency vector of the songs are fused to form a song fusion vector, ultimately achieving sound quality restoration.

Benefits of technology

It achieves efficient audio quality restoration of recorded songs, improves restoration efficiency, avoids local information loss, and enhances the reliability and audio quality of restoration results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708636B_ABST
    Figure CN120708636B_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, and electronic device for restoring the audio quality of a music library, relating to the field of data processing technology. In this application, firstly, high and low frequency separation processing is performed to form low-frequency time-domain data and high-frequency time-domain data of the song, followed by frequency domain conversion to form high-frequency frequency-domain data. Secondly, semantic encoding in the time domain or frequency domain is performed on the recorded song to be restored, forming a global vector for the song. Then, semantic encoding in the time domain is performed on the low-frequency time-domain data to form a low-frequency vector. Next, semantic encoding in the frequency domain is performed on the high-frequency frequency-domain data to form a high-frequency vector. Further, the low-frequency vector and the high-frequency vector are fused into the global vector to form a fused vector. Finally, semantic decoding is performed on the fused vector to form the target recorded song to be restored. Based on the above, the relatively low efficiency of audio quality restoration of recorded songs in existing technologies can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically, to a method, apparatus, and electronic device for restoring the audio quality of a music library. Background Technology

[0002] High-quality recording equipment employs advanced audio processing technologies, typically supporting high-precision sampling rates, such as 44.1kHz or higher. These technologies help capture a wider spectral range, reduce distortion and noise, and provide a cleaner audio signal. Low-quality recording equipment has limited audio processing capabilities and a lower sampling rate, which may lead to degraded sound quality (e.g., the inability to capture high-frequency parts of drums or vocals) and unclear sound. Therefore, choosing the right recording equipment is crucial for achieving high-quality recordings, especially in music production and recording studios. However, for songs in a music library recorded with low-quality equipment, the focus is on audio quality restoration. Currently, audio quality restoration is generally achieved through manual editing, which relies on high levels of expertise and is relatively inefficient. Therefore, there is an urgent need for a solution that can efficiently restore the audio quality of recorded songs. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a method, apparatus and electronic device for restoring the audio quality of a music library, so as to improve the problem of relatively low efficiency in restoring the audio quality of songs in the prior art.

[0004] To achieve the above objectives, this application adopts the following technical solution:

[0005] A method for restoring the audio quality of a music library, characterized by comprising:

[0006] The songs to be repaired in the target music library are processed by high-low frequency separation in the time domain to form low-frequency time domain data and high-frequency time domain data of the songs. The high-frequency time domain data of the songs is then converted into frequency domain data to form high-frequency frequency domain data of the songs.

[0007] The song to be repaired is semantically encoded in the time domain or semantically encoded in the frequency domain of the frequency domain transformation result of the song to be repaired to form a global vector of the song.

[0008] The low-frequency time-domain data of the song is semantically encoded in the time domain to form a low-frequency vector of the song;

[0009] The high-frequency domain data of the song is semantically encoded in the frequency domain to form a high-frequency vector of the song;

[0010] The low-frequency vector and the high-frequency vector of the song are fused into the global vector of the song to form a fused vector of the song.

[0011] The fusion vector of the song is semantically decoded to form the target repaired recorded song.

[0012] In a preferred embodiment of this application, in the aforementioned method for restoring the audio quality of a music library, the step of performing temporal semantic encoding on the low-frequency temporal data of the song to form a low-frequency vector of the song includes:

[0013] The low-frequency time-domain data of the song is convolved to form a low-frequency time-domain convolution vector;

[0014] Long-term average characteristic parameters are extracted from the low-frequency time-domain data of the song to obtain the long-term average characteristic distribution corresponding to the low-frequency time-domain data of the song, and the long-term average characteristic distribution is convolved to form a long-term average characteristic vector.

[0015] Based on the long-term average characteristic vector, the low-frequency temporal convolution vector is semantically enhanced to form the song's low-frequency vector.

[0016] In a preferred embodiment of this application, in the aforementioned method for restoring the audio quality of a music library, the step of semantically enhancing the low-frequency temporal convolution vector based on the long-term average characteristic vector to form a low-frequency vector for the song includes:

[0017] The low-frequency temporal convolution vector is upsampled by A depths to form a low-frequency temporal depth vector with A depths.

[0018] For the a-th depth among the A depths, semantic enhancement processing is performed based on the low-frequency time-domain depth vector of the a-th depth and the enhancement auxiliary vector of the a-th depth to form the low-frequency enhancement vector of the song at the a-th depth, where a is a positive integer less than or equal to A. When a=1, the enhancement auxiliary vector of the a-th depth is the long-term average characteristic vector, and when a>1, the enhancement auxiliary vector of the a-th depth is the low-frequency enhancement vector of the song at the (a-1)-th depth.

[0019] Based on the A-depth low-frequency enhancement vectors and the A-depth low-frequency time-domain depth vectors, the low-frequency vectors of the song are determined.

[0020] In a preferred embodiment of this application, in the aforementioned method for restoring the audio quality of a music library, the step of performing semantic enhancement processing on the a-th depth among the A depths, based on the low-frequency temporal depth vector of the a-th depth and the enhancement auxiliary vector of the a-th depth, to form the low-frequency enhancement vector of the song at the a-th depth, includes:

[0021] Using the b-th semantic enhancement unit among B semantic enhancement units, the input vector of the b-th semantic enhancement unit is processed to form the semantic enhancement output vector of the b-th semantic enhancement unit, where b is a positive integer less than or equal to B. When b=1, the input vector of the b-th semantic enhancement unit includes the low-frequency time-domain depth vector of the a-th depth and the enhancement auxiliary vector of the a-th depth. When b>1, the input vector of the b-th semantic enhancement unit includes the low-frequency time-domain depth vector of the a-th depth and the semantic enhancement output vector of the (b-1)-th semantic enhancement unit.

[0022] The semantic enhancement output vector of the Bth semantic enhancement unit is used as the low-frequency enhancement vector of the song at the ath depth.

[0023] In a preferred embodiment of this application, in the aforementioned method for restoring the audio quality of a music library, the step of processing the input vector of the b-th semantic enhancement unit among B semantic enhancement units to form the semantic enhancement output vector of the b-th semantic enhancement unit includes:

[0024] Using the b-th semantic enhancement unit among the B semantic enhancement units, the low-frequency time-domain depth vector of the a-th depth is mapped in two different ways to form a first time-domain depth vector and a second time-domain depth vector.

[0025] The semantic enhancement output vector included in the input vector of the b-th semantic enhancement unit is offset and removed to form a semantic enhancement offset vector. The offset removal includes calculating the difference between each vector parameter in the semantic enhancement output vector and the mean of each vector parameter in the semantic enhancement output vector.

[0026] Perform a bitwise multiplication operation on the first temporal depth vector and the semantic enhancement offset vector to form a third temporal depth vector;

[0027] The standard deviations of each vector parameter in the third temporal depth vector and each vector parameter in the semantic enhancement output vector are respectively calculated to form the fourth temporal depth vector;

[0028] The fourth temporal depth vector and the second temporal depth vector are added bitwise to form the semantic enhancement output vector of the b-th semantic enhancement unit.

[0029] In a preferred embodiment of this application, the step of performing frequency domain semantic encoding on the high-frequency data of the song to form a high-frequency vector in the above-mentioned music library audio quality restoration method includes:

[0030] The high-frequency domain data of the song is convolved to form a high-frequency domain convolution vector;

[0031] Two different pooling methods are applied to the high-frequency domain convolution vector to form a first frequency domain pooling vector and a second frequency domain pooling vector, wherein the vector size of the first frequency domain pooling vector and the vector size of the second frequency domain pooling vector are the same.

[0032] The first frequency domain pooling vector is mapped to form a gated parameter distribution, wherein the vector size of the gated parameter distribution is the same as the vector size of the first frequency domain pooling vector, and each vector parameter in the gated parameter distribution is greater than or equal to 0 and less than or equal to 1.

[0033] The gating parameter distribution and the second frequency domain pooling vector are multiplied bitwise to form the high-frequency vector of the song.

[0034] In a preferred embodiment of this application, in the aforementioned method for restoring the audio quality of a music library, the step of fusing the low-frequency vector and the high-frequency vector of the song into the global vector of the song to form a fused song vector includes:

[0035] When the global vector of the song belongs to the result of semantic encoding in the time domain, the low-frequency vector of the song is used as the first vector to be fused, and the high-frequency vector of the song is used as the second vector to be fused.

[0036] When the global vector of the song belongs to the semantic coding result of the frequency domain, the low-frequency vector of the song is used as the second vector to be fused, and the high-frequency vector of the song is used as the first vector to be fused.

[0037] The first vector to be fused is fused into the global vector of the song based on a gating mechanism, and the second vector to be fused is fused into the global vector of the song based on an attention mechanism, so as to form a fused vector of the song.

[0038] In a preferred embodiment of this application, in the aforementioned music library audio quality restoration method, the steps of fusing the first vector to be fused into the global vector of the song based on a gating mechanism, and fusing the second vector to be fused into the global vector of the song based on an attention mechanism to form a song fusion vector, include:

[0039] The first vector to be fused is mapped to form a gating mapping parameter, and the gating mapping parameter and the global vector of the song are multiplied bitwise to form a first fused vector;

[0040] The global vector of the song is connected to the first fusion vector to form a first connection vector;

[0041] Determine the attention parameter distribution between the second vector to be fused and the first connection vector, and perform a weighted summation operation on the first connection vector based on the attention parameter distribution to form the second fused vector;

[0042] The first connection vector is connected to the second fusion vector to form the second connection vector;

[0043] Based on the second connection vector, an iterative fusion process is performed to form a song fusion vector. The iterative fusion process includes fusing the first vector to be fused into the second connection vector based on a gating mechanism, and fusing the second vector to be fused into the second connection vector based on an attention mechanism to form a new second connection vector. The last newly formed second connection vector is used as the song fusion vector.

[0044] This application also provides a music library audio quality restoration device, including:

[0045] The recording song conversion module is used to perform high and low frequency separation processing on the recorded songs to be repaired in the target music library in the time domain to form low frequency time domain data and high frequency time domain data of the songs, and to perform frequency domain conversion on the high frequency time domain data of the songs to form high frequency frequency domain data of the songs.

[0046] The first encoding module is used to perform temporal semantic encoding on the recorded song to be repaired or to perform frequency domain semantic encoding on the frequency domain conversion result of the recorded song to be repaired, so as to form a global vector of the song.

[0047] The second encoding module is used to perform temporal semantic encoding on the low-frequency time-domain data of the song to form a low-frequency vector of the song.

[0048] The third encoding module is used to perform semantic encoding of the high-frequency domain data of the song in the frequency domain to form a high-frequency vector of the song;

[0049] The vector fusion module is used to fuse the low-frequency vector and the high-frequency vector of the song into the global vector of the song to form a fused vector of the song.

[0050] The semantic decoding module is used to perform semantic decoding on the song fusion vector to form the target repair recorded song.

[0051] Based on the above, this application also provides an electronic device, including:

[0052] Memory, used to store computer programs;

[0053] A processor connected to the memory is used to execute the computer program stored in the memory to implement the above-described method for restoring the audio quality of the music library.

[0054] The music library audio quality restoration method, apparatus, and electronic device provided in this application firstly separate the high and low frequencies of the recorded songs to be restored in the target music library to form low-frequency time-domain data and high-frequency time-domain data, and then perform frequency domain conversion to form high-frequency frequency-domain data. Secondly, the recorded songs to be restored are semantically encoded in the time domain or frequency domain to form a global vector. Then, the low-frequency time-domain data is semantically encoded in the time domain to form a low-frequency vector. After that, the high-frequency frequency-domain data is semantically encoded in the frequency domain to form a high-frequency vector. Further, the low-frequency vector and the high-frequency vector are fused into the global vector to form a fused vector. Finally, the fused vector is semantically decoded to form the target restored recorded song. Based on the above, on the one hand, because the recorded songs to be restored can be semantically encoded and decoded, the audio quality restoration of the recorded songs can be intelligently realized, which is more efficient than conventional restoration schemes based on manual editing, thus improving the problem of relatively low efficiency in the restoration of recorded song audio quality in the prior art. On the other hand, since low-frequency signals change slowly over time, meaning that signal values ​​change little between adjacent time points, the slow-changing characteristic of low-frequency signals (i.e., good stationarity) and the latent semantic relationship between low-frequency and high-frequency signals can be utilized to represent the lost high-frequency signals. This involves mining the low-frequency time-domain vector of the song's low-frequency data in the time domain, thus ensuring that the target restored recording obtained through semantic decoding contains the lost high-frequency signals. Furthermore, it should be noted that the song fusion vector from semantic decoding includes not only the low-frequency vector but also the high-frequency vector and the global vector, resulting in a more comprehensive semantic representation. Therefore, it effectively avoids the problem of local information loss in the semantic decoding results. Based on this, the reliability of music library audio quality restoration can be improved. Attached Figure Description

[0055] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings.

[0056] Figure 1 This is a structural block diagram of an electronic device provided in an embodiment of this application.

[0057] Figure 2 This is a flowchart illustrating the music library audio quality restoration method provided in this application embodiment.

[0058] Figure 3 This is a schematic diagram illustrating the implementation process of the music library audio quality restoration method provided in the embodiments of this application.

[0059] Figure 4 This is a schematic diagram of semantic enhancement processing provided in an embodiment of this application.

[0060] Figure 5 This is a schematic diagram of vector fusion provided in an embodiment of this application.

[0061] Figure 6 A block diagram illustrating the music library audio quality restoration device provided in this application embodiment. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0063] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0064] like Figure 1 As shown in the illustration, this application provides an electronic device. The electronic device may include a memory, a processor, and a music library audio quality restoration device.

[0065] In detail, the memory and the processor are electrically connected directly or indirectly to enable data transmission or interaction. For example, the memory and the processor can be electrically connected via one or more communication buses or signal lines. The music library audio quality restoration device includes at least one software functional module stored in the memory in the form of software or firmware. The processor is used to execute executable computer programs stored in the memory, such as the software functional modules and computer programs included in the music library audio quality restoration device, to implement the music library audio quality restoration method provided in this application embodiment.

[0066] Optionally, the memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0067] Furthermore, the processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a system on chip (SoC), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0068] I understand. Figure 1 The structure shown is for illustrative purposes only; the electronic device may also include components that are more advanced than those shown. Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown may include, for example, a communication unit for exchanging information with other devices.

[0069] Combination Figure 2 This application also provides a method for restoring the audio quality of a music library, applicable to the aforementioned electronic device. The method steps defined in the process of the music library audio quality restoration method can be implemented by the electronic device. The following will describe... Figure 2 The specific process shown will be explained in detail.

[0070] Step S110: Perform high-low frequency separation processing on the recorded songs to be repaired in the target music library in the time domain to form low-frequency time domain data and high-frequency time domain data of the songs; and perform frequency domain conversion on the high-frequency time domain data of the songs to form high-frequency frequency domain data of the songs.

[0071] In this embodiment, the electronic device can perform high-low frequency separation processing on the recorded songs to be repaired in the target music library in the time domain, forming low-frequency time domain data and high-frequency time domain data of the songs (combined with...). Figure 3This is achieved through a low-pass filter and a high-pass filter. Specifically, the low-pass filter is used to filter the recorded song to be repaired, obtaining the low-frequency time-domain data of the song, and the high-pass filter is used to filter the recorded song to be repaired, obtaining the high-frequency time-domain data of the song. In addition, the high-frequency time-domain data of the song is transformed into frequency-domain data to form the high-frequency frequency-domain data of the song. For example, the high-frequency time-domain data of the song can be Fourier transformed to obtain a spectrum, which is the high-frequency frequency-domain data of the song.

[0072] Step S120: Perform temporal semantic encoding on the recorded song to be repaired or perform frequency domain semantic encoding on the frequency domain conversion result of the recorded song to be repaired to form a global vector of the song.

[0073] In this embodiment, the electronic device can perform temporal semantic encoding on the recorded song to be repaired or perform frequency domain semantic encoding on the frequency domain transformation result of the recorded song to be repaired to form a global song vector. That is, it can either directly perform semantic encoding on the recorded song to be repaired in the temporal domain to obtain the global song vector, or first perform a Fourier transform on the recorded song to be repaired in the temporal domain to obtain the corresponding spectrogram, and then perform semantic encoding on the spectrogram to obtain the global song vector. For example, the aforementioned semantic encoding can include convolution and self-attention processing. For instance, convolution processing can be performed on the recorded song to be repaired in the temporal domain, and then self-attention processing can be performed on the resulting convolution vector to obtain the global song vector, thus representing the global semantic information of the recorded song to be repaired. Another example is that convolution processing can be performed on the spectrogram, and then self-attention processing can be performed on the resulting convolution vector to obtain the global song vector. Alternatively, any existing semantic encoding method can be used.

[0074] Step S130: Perform temporal semantic encoding on the low-frequency time-domain data of the song to form a low-frequency vector of the song.

[0075] In this embodiment, after forming the low-frequency time-domain data of the song, the electronic device can perform time-domain semantic encoding on the low-frequency time-domain data of the song to form a low-frequency vector of the song. That is, since the low-frequency signal changes slowly over time, it means that the signal value changes little between adjacent time points. Therefore, the slow change characteristic of the low-frequency signal (i.e., it has good stability) and the potential semantic relationship between the low-frequency signal and the high-frequency signal can be used to represent the lost high-frequency signal, that is, to mine the low-frequency vector of the song in the time domain. The potential semantic relationship between the low-frequency signal and the high-frequency signal can actually be represented by the parameters of the neural network model that performs semantic encoding (that is, during the learning process, the neural network model will learn the sample of the song to be repaired (low quality, missing some high-frequency information) and the label of the target recorded song to be repaired (high quality, no loss of high-frequency information), thereby capturing the potential semantic relationship between the low-frequency information and the lost high-frequency information in the sample of the song to be repaired).

[0076] Step S140: Perform semantic encoding of the high-frequency domain data of the song in the frequency domain to form a high-frequency vector of the song.

[0077] In this embodiment, after forming the high-frequency time-domain data of the song, the electronic device can perform semantic encoding on the high-frequency frequency-domain data of the song to form a high-frequency vector of the song. Since high-frequency information can characterize detailed information, it is also necessary to perform semantic encoding on the high-frequency frequency-domain data of the song separately, so as to obtain a high-frequency vector of the song that can fully characterize the detailed information.

[0078] Step S150: The low-frequency vector and the high-frequency vector of the song are fused into the global vector of the song to form a fused vector of the song.

[0079] In this embodiment, after forming the song's high-frequency vector, low-frequency vector, and high-frequency vector, the electronic device can fuse the low-frequency vector and high-frequency vector into the song's global vector to form a fused song vector. That is, based on the global semantic information represented by the song's global vector, the semantic information represented by the low-frequency and high-frequency vectors can be further fused, resulting in a fused song vector that represents semantic information with both high richness and high accuracy.

[0080] Step S160: Semantic decoding is performed on the song fusion vector to form the target repaired recorded song.

[0081] In this embodiment, after forming the song fusion vector, the electronic device can perform semantic decoding on the song fusion vector to form the target restored recording song. It is understood that the semantic decoding can directly generate the time-domain target restored recording song, or a corresponding spectrogram can be generated first, and then the time-domain target restored recording song can be obtained by performing an inverse Fourier transform on the spectrogram. For example, generating the corresponding spectrogram can be implemented using a decoder, where the decoder's network architecture is as follows:

[0082] Fully connected layer: Outputs a 256×16×16 feature map, with ReLU or LeakyReLU activation function;

[0083] Deconvolutional layer 1: Input size is 256×16×16, output size is 128×32×32, kernel size is 3x3, stride is 2, padding is 1, activation function is LeakyReLU;

[0084] Deconvolutional layer 2: Input size is 128×32×32, output size is 64×64×64, kernel size is 3x3, stride is 2, padding is 1, activation function is LeakyReLU;

[0085] Deconvolutional layer 3: Input size is 64×64×64, output size is 1×128×128, kernel size is 3x3, stride is 2, padding is 1, activation function is Tanh or Sigmoid;

[0086] Based on this, the size of the obtained spectrogram is 128×128.

[0087] Based on the above, on the one hand, by semantically encoding and decoding the recorded songs to be restored, intelligent sound quality restoration can be achieved, resulting in higher efficiency compared to conventional restoration methods based on manual editing. This addresses the relatively low efficiency of existing technologies for restoring the sound quality of recorded songs. On the other hand, since low-frequency signals change slowly over time, meaning that signal values ​​do not change significantly between adjacent time points, the slow-changing characteristics (i.e., good stability) of low-frequency signals and the latent semantic relationship between low-frequency and high-frequency signals can be utilized to represent the lost high-frequency signals. This involves mining the low-frequency vector of the song in the time domain, ensuring that the target recorded song for restoration obtained through semantic decoding contains the lost high-frequency signals. Furthermore, it should be noted that the song fusion vector from semantic decoding includes not only the low-frequency vector but also the high-frequency vector and the global vector, resulting in a more comprehensive semantic representation. Therefore, it effectively avoids the problem of local information loss in the semantic decoding results. Based on this, the reliability of music library sound quality restoration can be improved.

[0088] Firstly, regarding step S130, it should be noted that the specific method of performing temporal semantic encoding on the low-frequency temporal data of the song is not limited and can be selected according to actual needs.

[0089] For example, in an alternative implementation, the low-frequency time-domain data of the song can be convolved and self-attention processed to obtain the low-frequency vector of the song.

[0090] For example, in another alternative implementation, in order to improve the semantic representation capability of the formed low-frequency vector of the song, the above step S130 may further include steps S131, S132 and S133, the specific contents of each step are as follows.

[0091] Step S131: Perform convolution processing on the low-frequency time-domain data of the song to form a low-frequency time-domain convolution vector.

[0092] In this embodiment, the low-frequency time-domain data of the song can be convolved to form a low-frequency time-domain convolution vector. That is, primary semantic information can be extracted from the low-frequency time-domain data of the song through convolution processing, thus obtaining the low-frequency time-domain convolution vector.

[0093] Step S132: Extract long-term average characteristic parameters from the low-frequency time-domain data of the song to obtain the long-term average characteristic distribution corresponding to the low-frequency time-domain data of the song, and perform convolution processing on the long-term average characteristic distribution to form a long-term average characteristic vector.

[0094] In this embodiment, long-term average characteristic parameters can be extracted from the low-frequency time-domain data of the song to obtain the long-term average characteristic distribution corresponding to the low-frequency time-domain data. The long-term average characteristic distribution can then be convolved to form a long-term average characteristic vector. It should be noted that the average characteristics of high-frequency signals over a long period, such as the average value and autocorrelation, may be related to the characteristics of low-frequency signals. Due to their stationarity, low-frequency signals can be used to approximate the behavior of high-frequency signals under long-term averaging. This characteristic is particularly important in signal processing, especially when high-frequency signals are missing, as low-frequency signals can provide important reference information. Therefore, the long-term average characteristic distribution can be extracted from the low-frequency time-domain data of the song. For example, a sliding window process can be applied to the low-frequency time-domain data of the song, and the mean of the time-domain signal within each window can be calculated. This yields the corresponding long-term average characteristic distribution. Then, the long-term average characteristic distribution can be convolved to obtain the long-term average characteristic vector.

[0095] Step S133: Based on the long-term average characteristic vector, semantic enhancement is performed on the low-frequency temporal convolution vector to form the song's low-frequency vector.

[0096] In this embodiment, after forming the long-term average characteristic vector and the low-frequency temporal convolution vector, semantic enhancement can be performed on the low-frequency temporal convolution vector based on the long-term average characteristic vector to form a song low-frequency vector. That is, since the low-frequency temporal convolution vector can represent low-frequency semantic information holistically, but since the long-term average characteristic vector can reflect some characteristics of high-frequency signals, the semantic information in the long-term average characteristic vector can be fused into the low-frequency temporal convolution vector to obtain the song low-frequency vector.

[0097] It is understood that the specific method of semantic enhancement of the low-frequency temporal convolution vector in step S133 above is not limited. For example, in an alternative implementation, in order to fully fuse the long-term average feature vector and the low-frequency temporal convolution vector to ensure the reliability of semantic enhancement, step S133 above may further include steps S133a, S133b and S133c, the specific contents of each step are as follows.

[0098] Step S133a: The low-frequency time-domain convolution vector is upsampled by A depths to form a low-frequency time-domain depth vector with A depths.

[0099] In this embodiment, the low-frequency temporal convolution vector can be upsampled by A depths to form a low-frequency temporal depth vector of A depths, where A can be a positive integer greater than or equal to 2. This means that multiple semantic vectors with different depths can be obtained, thus fully representing the semantic information of the low-frequency temporal domain at different depths. For example, a first upsampling network upsamples the low-frequency temporal convolution vector, outputting a 128*128 low-frequency temporal depth vector. A second upsampling network upsamples the low-frequency temporal depth vector output by the first upsampling network, obtaining a 256*256 low-frequency temporal depth vector. A third upsampling network upsamples the low-frequency temporal depth vector output by the second upsampling network, obtaining a 512*512 low-frequency temporal depth vector. A fourth upsampling network upsamples the low-frequency temporal depth vector output by the third upsampling network, obtaining a 1024*1024 low-frequency temporal depth vector.

[0100] Step S133b: For the a-th depth among the A depths, perform semantic enhancement processing based on the low-frequency time-domain depth vector of the a-th depth and the enhancement auxiliary vector of the a-th depth to form the low-frequency enhancement vector of the song at the a-th depth.

[0101] In this embodiment, after obtaining A low-frequency temporal depth vectors of different depths, semantic enhancement processing can be performed on the a-th depth among the A depths, based on the low-frequency temporal depth vector of the a-th depth and the enhancement auxiliary vector of the a-th depth, to form the song low-frequency enhancement vector of the a-th depth. Here, a is a positive integer less than or equal to A. When a=1, the enhancement auxiliary vector of the a-th depth is the long-term average characteristic vector; when a>1, the enhancement auxiliary vector of the a-th depth is the song low-frequency enhancement vector of the (a-1)-th depth. For example, semantic enhancement processing can be performed on the low-frequency temporal depth vector of the first depth and the long-term average characteristic vector to form the song low-frequency enhancement vector of the first depth. And, semantic enhancement processing can be performed on the low-frequency temporal depth vector of the second depth and the song low-frequency enhancement vector of the first depth to form the song low-frequency enhancement vector of the second depth.

[0102] Step S133c: Determine the low-frequency vector of the song based on the A-depth low-frequency enhancement vectors and the A-depth low-frequency time-domain depth vectors.

[0103] In this embodiment, after forming the A-depth low-frequency enhancement vectors of the song, the low-frequency vector of the song can be determined based on the A-depth low-frequency enhancement vectors and the A-depth low-frequency temporal depth vectors. For example, the A-depth low-frequency enhancement vectors of the song and the A-depth low-frequency temporal depth vectors can be concatenated (either by first expanding and then concatenating, or by upsampling or downsampling to make the dimensions the same before concatenation) to form a concatenated vector. Then, convolution and self-attention processing can be performed on the concatenated vector to obtain the corresponding low-frequency vector of the song. This avoids the semantic distortion problem caused by a larger depth of semantic enhancement processing, thereby ensuring the accuracy of the semantic representation of the determined low-frequency vector of the song.

[0104] It is understood that in step S133b above, the specific method of semantic enhancement processing based on the low-frequency time-domain depth vector of the a-th depth and the enhancement auxiliary vector of the a-th depth is not limited. For example, in an alternative implementation, in order to further improve the effect of fusing different vectors in the process of semantic enhancement processing, that is, to achieve full fusion between the semantic information of different vectors, step S133b above may further include steps b1 and b2, the specific contents of each step are as follows.

[0105] Step b1: Using the bth semantic enhancement unit among the B semantic enhancement units, process the input vector of the bth semantic enhancement unit to form the semantic enhancement output vector of the bth semantic enhancement unit.

[0106] In this embodiment, the input vector of the b-th semantic enhancement unit among B semantic enhancement units can be processed to form the semantic enhancement output vector of the b-th semantic enhancement unit. Here, b is a positive integer less than or equal to B (B can be a positive integer greater than or equal to 2). When b=1, the input vector of the b-th semantic enhancement unit includes the low-frequency time-domain depth vector of the a-th depth and the enhancement auxiliary vector of the a-th depth. When b>1, the input vector of the b-th semantic enhancement unit includes the low-frequency time-domain depth vector of the a-th depth and the semantic enhancement output vector of the (b-1)-th semantic enhancement unit. For example, in the first semantic enhancement unit, the low-frequency time-domain depth vector of the a-th depth and the enhancement auxiliary vector of the a-th depth can be processed to obtain the semantic enhancement output vector of the first semantic enhancement unit. In the second semantic enhancement unit, the low-frequency time-domain depth vector of the a-th depth and the semantic enhancement output vector of the first semantic enhancement unit can be processed to obtain the semantic enhancement output vector of the second semantic enhancement unit.

[0107] Step b2: Use the semantic enhancement output vector of the Bth semantic enhancement unit as the low-frequency enhancement vector of the song at the ath depth.

[0108] In this embodiment of the application, after obtaining the semantic enhancement output vector of the Bth semantic enhancement unit, that is, after obtaining the semantic enhancement output vector of the last semantic enhancement unit, the semantic enhancement output vector of the last semantic enhancement unit can be used as the low-frequency enhancement vector of the song at the ath depth.

[0109] It is understood that in step b1 above, the specific method of processing the input vector of the b-th semantic enhancement unit is not limited. For example, in an alternative implementation, in order to achieve full fusion between the two semantic vectors, step b1 above may further include the following (in conjunction with...). Figure 4 ):

[0110] First, the b-th semantic enhancement unit among the B semantic enhancement units can be used to perform two different mappings on the low-frequency temporal depth vector at the a-th depth, forming a first temporal depth vector and a second temporal depth vector. For example, the mappings can both be linear mappings, such as y1=a1x+b1 and y2=a2x+b2, where y1 and y2 are the first and second temporal depth vectors, respectively, x is the low-frequency temporal depth vector at the a-th depth, a1 and a2 are both weight parameters, and b1 and b2 are both bias parameters, all formed during the training process of the corresponding neural network model. In this way, the representation capability of the semantic vector can be further optimized by using two linear mappings.

[0111] Secondly, the semantic enhancement output vector included in the input vector of the b-th semantic enhancement unit (when b=1, the semantic enhancement output vector is the enhancement auxiliary vector of the a-th depth; when b is greater than 1, the semantic enhancement output vector is the semantic enhancement output vector of the (b-1)-th semantic enhancement unit) can be offset to form a semantic enhancement offset vector. The offset removal includes calculating the difference between each vector parameter in the semantic enhancement output vector and the mean of each vector parameter in the semantic enhancement output vector. Based on this, the operation of subtracting each vector parameter from its mean can achieve decentralization. This means that the mean of the input data is removed, making the mean of each vector parameter become zero. In this way, the decentralized data can eliminate the offset, making the data more evenly distributed. Moreover, after removing the mean, the local differences of each feature can be better highlighted, and it can help the network better capture the differences and relationships between different elements.

[0112] Then, the first temporal depth vector and the semantic enhancement offset vector can be multiplied bitwise to form a third temporal depth vector. Based on this, since the semantic enhancement offset vector is decentralized, different parameters can represent different levels of importance. Thus, through bitwise multiplication, the first temporal depth vector can be weighted based on the semantic enhancement offset vector, i.e., gating and filtering can be achieved, thereby completing the full fusion of the two semantic vectors.

[0113] Furthermore, the standard deviations of each vector parameter in the third temporal depth vector and each vector parameter in the semantic enhancement output vector can be used to calculate the quotient, forming a fourth temporal depth vector. Based on this, by quoting each vector parameter with the standard deviation of the semantic enhancement output vector, the scale of each local feature is actually adjusted to eliminate the differences between different dimensions. Specifically, the standard deviation represents the dispersion of features. Dividing by the standard deviation allows features to be compared and fused at the same scale, which helps to eliminate problems that may be caused by different feature scales and ensures that the data processed by the network has a uniform scale.

[0114] Finally, the fourth time-domain depth vector and the second time-domain depth vector can be added bit by bit to form the semantic enhancement output vector of the b-th semantic enhancement unit. That is, since the fourth time-domain depth vector is actually the result of fusing the low-frequency time-domain depth vector and the semantic enhancement output vector, semantic distortion may occur as the fusion deepens. Therefore, by adding the second time-domain depth vector, which can also represent the low-frequency time-domain depth vector, to the fourth time-domain depth vector bit by bit, the resulting semantic enhancement output vector can be guaranteed to have high semantic representation accuracy.

[0115] Secondly, regarding step S140, it should be noted that the specific method of semantic encoding of the high-frequency domain data of the song is not limited and can be selected according to actual needs.

[0116] For example, in an alternative implementation, the high-frequency domain data of the song can be convolved and self-attention processed to obtain the high-frequency vector of the song.

[0117] For example, in another alternative implementation, in order to fully mine the semantic information in the high-frequency domain data of the song while avoiding the problem of excessive computation caused by attention processing, step S140 above may further include:

[0118] First, the high-frequency domain data of the song can be convolved to form a high-frequency domain convolution vector;

[0119] Secondly, two different pooling methods can be applied to the high-frequency domain convolution vector to form a first frequency domain pooling vector (such as the result of max pooling) and a second frequency domain pooling vector (such as the result of mean pooling), wherein the vector size of the first frequency domain pooling vector and the vector size of the second frequency domain pooling vector are the same.

[0120] Then, the first frequency domain pooling vector can be mapped (e.g., through nonlinear mapping using the sigmoid function) to form a gated parameter distribution, wherein the vector size of the gated parameter distribution is the same as the vector size of the first frequency domain pooling vector, and each vector parameter in the gated parameter distribution is greater than or equal to 0 and less than or equal to 1.

[0121] Finally, the gating parameter distribution and the second frequency domain pooling vector can be multiplied bitwise to form the high-frequency vector of the song.

[0122] Thirdly, regarding step S150, it should be noted that the specific method of fusing the low-frequency vector and the high-frequency vector of the song into the global vector of the song is not limited and can be selected according to actual needs.

[0123] For example, in an alternative implementation, the low-frequency vector, the high-frequency vector, and the global vector of the song can be added together to achieve fusion, that is, to obtain a song fusion vector that can represent the semantic information in the three semantic vectors.

[0124] For example, in another alternative implementation, in order to improve the fusion effect and avoid semantic distortion during the fusion process, the above step S150 may further include steps S151, S152 and S153, the specific contents of each step are as follows.

[0125] Step S151: When the global vector of the song belongs to the result of semantic encoding in the time domain, the low-frequency vector of the song is used as the first vector to be fused, and the high-frequency vector of the song is used as the second vector to be fused.

[0126] In this embodiment, when the song's global vector belongs to the result of temporal semantic coding, the song's low-frequency vector can be used as the first vector to be fused, and the song's high-frequency vector can be used as the second vector to be fused. That is, the first vector to be fused and the song's global vector are vectors of the same domain, while the second vector to be fused and the song's global vector are vectors of different domains.

[0127] Step S152: When the global vector of the song belongs to the result of semantic encoding in the frequency domain, the low-frequency vector of the song is used as the second vector to be fused, and the high-frequency vector of the song is used as the first vector to be fused.

[0128] In this embodiment, when the song's global vector belongs to the result of semantic coding in the frequency domain, the song's low-frequency vector can be used as the second vector to be fused, and the song's high-frequency vector can be used as the first vector to be fused. That is, the first vector to be fused and the song's global vector are vectors in the same domain, while the second vector to be fused and the song's global vector are vectors in different domains.

[0129] Step S153: Based on the gating mechanism, the first vector to be fused is fused into the global vector of the song, and based on the attention mechanism, the second vector to be fused is fused into the global vector of the song to form a fused song vector.

[0130] In this embodiment, after determining the first and second vectors to be fused based on steps S151 and S152, the first vector to be fused can be fused into the global song vector using a gating mechanism, and the second vector to be fused can be fused into the global song vector using an attention mechanism to form a fused song vector. That is, since the first vector to be fused and the global song vector are vectors in the same domain, they can be directly fused using a gating mechanism. This avoids semantic distortion while reducing computation and improving processing efficiency. Furthermore, since the second vector to be fused and the global song vector are vectors in different domains, to avoid semantic distortion, an attention mechanism can be used to fuse the second vector to be fused into the global song vector. This fully utilizes the cross-modal processing capability of the attention mechanism, resulting in better fusion and ensuring the accuracy of the semantic representation of the obtained fused song vector.

[0131] It is understood that the specific method of forming the song fusion vector in step S153 above is not limited. For example, in an alternative implementation, in order to ensure that both the first vector to be fused and the second vector to be fused can be fully fused into the global song vector, step S153 above may further include the following (in conjunction with...). Figure 5 ):

[0132] The first step is to map the first vector to be fused (refer to the relevant description above) to form a gating mapping parameter, and then perform a bitwise multiplication operation on the gating mapping parameter and the global vector of the song to form the first fused vector;

[0133] The second step is to connect the global vector of the song to the first fusion vector to form a first connection vector. For example, the global vector of the song and the first fusion vector can be added together. This can avoid the problem that the quality of the first fusion vector formed by the fusion is not high when the quality of the first vector to be fused is not high, thereby avoiding the quality problem of gated fusion.

[0134] The third step is to determine the attention parameter distribution between the second vector to be fused and the first connection vector (for example, the transpose of the query vector corresponding to the second vector to be fused and the key vector corresponding to the first connection vector can be multiplied to obtain the attention parameter distribution), and a weighted summation operation is performed on the first connection vector based on the attention parameter distribution (such as a weighted summation operation is performed on the value vector corresponding to the first connection vector) to form the second fused vector;

[0135] Fourth, the first connection vector can be connected to the second fusion vector to form the second connection vector; for example, the first connection vector and the second fusion vector can be added together to avoid the semantic distortion problem that occurs during attention processing.

[0136] The fifth step involves iterative fusion processing based on the second connection vector to form a song fusion vector. Each iteration of this fusion process includes fusing the first vector to be fused into the second connection vector using a gating mechanism, and then fusing the second vector to be fused into the second connection vector using an attention mechanism to form a new second connection vector. The final new second connection vector is then used as the song fusion vector. In other words, steps one, two, three, and four can be executed in reverse order. For example, in the first iteration, the global song vector can be replaced with the first second connection vector to form a second second connection vector. In the second iteration, the global song vector can be replaced with the second second connection vector to form a third second connection vector. In the third iteration, the global song vector can be replaced with the third second connection vector to form a fourth second connection vector.

[0137] In detail, during the first iteration:

[0138] First, the first vector to be fused can be mapped to form a gating mapping parameter, and the gating mapping parameter and the first second connection vector can be multiplied bitwise to form a second first fusion vector;

[0139] Secondly, the first second connection vector can be connected to the second first fusion vector to form a second first connection vector;

[0140] Then, the attention parameter distribution between the second vector to be fused and the second first connection vector can be determined, and a weighted summation operation can be performed on the second first connection vector based on the attention parameter distribution to form the second second fusion vector;

[0141] Finally, the second first connection vector can be connected to the second second fusion vector to form the second second connection vector.

[0142] Combination Figure 6 This application also provides a music library audio quality restoration device applicable to the aforementioned electronic devices. The music library audio quality restoration device includes a recorded song conversion module, a first encoding module, a second encoding module, a third encoding module, a vector fusion module, and a semantic decoding module.

[0143] The recorded song conversion module is used to perform high-low frequency separation processing on the recorded songs to be repaired in the target music library in the time domain, forming low-frequency time domain data and high-frequency time domain data of the songs; and to perform frequency domain conversion on the high-frequency time domain data of the songs, forming high-frequency frequency domain data of the songs. In this embodiment of the application, the recorded song conversion module can be used to perform... Figure 2 For details regarding the recorded song conversion module shown in step S110, please refer to the previous description of step S110.

[0144] The first encoding module is used to perform temporal semantic encoding on the recorded song to be repaired or to perform frequency domain semantic encoding on the frequency domain transformation result of the recorded song to be repaired, forming a global vector for the song. In this embodiment, the first encoding module can be used to perform... Figure 2 The relevant content regarding the first encoding module in step S120 shown can be found in the previous description of step S120.

[0145] The second encoding module is used to perform temporal semantic encoding on the low-frequency time-domain data of the song to form a low-frequency vector of the song. In this embodiment, the second encoding module can be used to perform... Figure 2 For details regarding step S130 shown, please refer to the previous description of step S130 for information about the second encoding module.

[0146] The third encoding module is used to perform semantic encoding of the high-frequency domain data of the song in the frequency domain to form a high-frequency vector of the song. In this embodiment, the third encoding module can be used to perform... Figure 2 The relevant content regarding the third encoding module in step S140 shown can be found in the preceding description of step S140.

[0147] The vector fusion module is used to fuse the low-frequency vector and the high-frequency vector of the song into the global vector of the song, forming a fused song vector. In this embodiment, the vector fusion module can be used to perform... Figure 2 The details of step S150 shown above, and the relevant content regarding the vector fusion module, can be found in the preceding description of step S150.

[0148] The semantic decoding module is used to perform semantic decoding on the song fusion vector to form the target repaired recorded song. In this embodiment, the semantic decoding module can be used to perform... Figure 2 The relevant content regarding the semantic decoding module in step S160 shown can be found in the preceding description of step S160.

[0149] In this embodiment of the application, corresponding to the above-described method for restoring the sound quality of a music library applied to the electronic device, a computer-readable storage medium is also provided, which stores a computer program that executes the various steps of the music library sound quality restoration method when the computer program is run.

[0150] The steps executed by the aforementioned computer program during runtime will not be described in detail here, but can be found in the explanation of the music library audio quality restoration method above.

[0151] In summary, the music library audio quality restoration method, apparatus, and electronic device provided in this application firstly separates the high and low frequencies of the recorded songs to be restored in the target music library to form low-frequency time-domain data and high-frequency time-domain data, and then performs frequency domain conversion to form high-frequency frequency-domain data. Secondly, the recorded songs to be restored are semantically encoded in the time domain or the frequency domain to form a global vector. Then, the low-frequency time-domain data is semantically encoded in the time domain to form a low-frequency vector. Afterward, the high-frequency frequency-domain data is semantically encoded in the frequency domain to form a high-frequency vector. Further, the low-frequency vector and the high-frequency vector are fused into the global vector to form a fused vector. Finally, the fused vector is semantically decoded to form the target restored recorded song. Based on the above, on the one hand, because the recorded songs to be restored can be semantically encoded and decoded, the audio quality restoration of the recorded songs can be intelligently realized, which is more efficient than conventional restoration schemes based on manual editing, thus improving the problem of relatively low efficiency in the restoration of recorded song audio quality in the prior art. On the other hand, since low-frequency signals change slowly over time, meaning that signal values ​​change little between adjacent time points, the slow-changing characteristic of low-frequency signals (i.e., good stationarity) and the latent semantic relationship between low-frequency and high-frequency signals can be utilized to represent the lost high-frequency signals. This involves mining the low-frequency time-domain vector of the song's low-frequency data in the time domain, thus ensuring that the target restored recording obtained through semantic decoding contains the lost high-frequency signals. Furthermore, it should be noted that the song fusion vector from semantic decoding includes not only the low-frequency vector but also the high-frequency vector and the global vector, resulting in a more comprehensive semantic representation. Therefore, it effectively avoids the problem of local information loss in the semantic decoding results. Based on this, the reliability of music library audio quality restoration can be improved.

[0152] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0153] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0154] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, electronic device, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. In the absence of further restrictions, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0155] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for restoring the audio quality of a music library, characterized in that, include: The songs to be repaired in the target music library are processed by high-low frequency separation in the time domain to form low-frequency time domain data and high-frequency time domain data of the songs. The high-frequency time domain data of the songs is then converted into frequency domain data to form high-frequency frequency domain data of the songs. The song to be repaired is semantically encoded in the time domain or semantically encoded in the frequency domain of the frequency domain transformation result of the song to be repaired to form a global vector of the song. The low-frequency time-domain data of the song is semantically encoded in the time domain to form a low-frequency vector of the song; The high-frequency domain data of the song is semantically encoded in the frequency domain to form a high-frequency vector of the song; The low-frequency vector and the high-frequency vector of the song are fused into the global vector of the song to form a fused vector of the song. The fusion vector of the song is semantically decoded to form the target repaired recorded song.

2. The method for restoring the audio quality of a music library according to claim 1, characterized in that, The step of performing temporal semantic encoding on the low-frequency time-domain data of the song to form a low-frequency vector of the song includes: The low-frequency time-domain data of the song is convolved to form a low-frequency time-domain convolution vector; Long-term average characteristic parameters are extracted from the low-frequency time-domain data of the song to obtain the long-term average characteristic distribution corresponding to the low-frequency time-domain data of the song, and the long-term average characteristic distribution is convolved to form a long-term average characteristic vector. Based on the long-term average characteristic vector, the low-frequency temporal convolution vector is semantically enhanced to form the song's low-frequency vector.

3. The method for restoring the sound quality of a music library according to claim 2, characterized in that, The step of semantically enhancing the low-frequency temporal convolution vector based on the long-term average feature vector to form the song's low-frequency vector includes: The low-frequency temporal convolution vector is upsampled by A depths to form a low-frequency temporal depth vector with A depths. For the a-th depth among the A depths, semantic enhancement processing is performed based on the low-frequency time-domain depth vector of the a-th depth and the enhancement auxiliary vector of the a-th depth to form the low-frequency enhancement vector of the song at the a-th depth, where a is a positive integer less than or equal to A. When a=1, the enhancement auxiliary vector of the a-th depth is the long-term average characteristic vector, and when a>1, the enhancement auxiliary vector of the a-th depth is the low-frequency enhancement vector of the song at the (a-1)-th depth. Based on the A-depth low-frequency enhancement vectors and the A-depth low-frequency time-domain depth vectors, the low-frequency vectors of the song are determined.

4. The method for restoring the audio quality of a music library according to claim 3, characterized in that, The step of performing semantic enhancement processing on the a-th depth among the A depths, based on the low-frequency time-domain depth vector of the a-th depth and the enhancement auxiliary vector of the a-th depth, to form the song low-frequency enhancement vector of the a-th depth, includes: Using the b-th semantic enhancement unit among B semantic enhancement units, the input vector of the b-th semantic enhancement unit is processed to form the semantic enhancement output vector of the b-th semantic enhancement unit, where b is a positive integer less than or equal to B. When b=1, the input vector of the b-th semantic enhancement unit includes the low-frequency time-domain depth vector of the a-th depth and the enhancement auxiliary vector of the a-th depth. When b>1, the input vector of the b-th semantic enhancement unit includes the low-frequency time-domain depth vector of the a-th depth and the semantic enhancement output vector of the (b-1)-th semantic enhancement unit. The semantic enhancement output vector of the Bth semantic enhancement unit is used as the low-frequency enhancement vector of the song at the ath depth.

5. The method for restoring the audio quality of a music library according to claim 4, characterized in that, The step of processing the input vector of the b-th semantic enhancement unit among B semantic enhancement units to form the semantic enhancement output vector of the b-th semantic enhancement unit includes: Using the b-th semantic enhancement unit among the B semantic enhancement units, the low-frequency time-domain depth vector of the a-th depth is mapped in two different ways to form a first time-domain depth vector and a second time-domain depth vector. The semantic enhancement output vector included in the input vector of the b-th semantic enhancement unit is offset and removed to form a semantic enhancement offset vector. The offset removal includes calculating the difference between each vector parameter in the semantic enhancement output vector and the mean of each vector parameter in the semantic enhancement output vector. Perform a bitwise multiplication operation on the first temporal depth vector and the semantic enhancement offset vector to form a third temporal depth vector; The standard deviations of each vector parameter in the third temporal depth vector and each vector parameter in the semantic enhancement output vector are respectively calculated to form the fourth temporal depth vector; The fourth temporal depth vector and the second temporal depth vector are added bitwise to form the semantic enhancement output vector of the b-th semantic enhancement unit.

6. The method for restoring the audio quality of a music library according to claim 1, characterized in that, The step of performing semantic encoding of the high-frequency domain data of the song in the frequency domain to form a high-frequency vector of the song includes: The high-frequency domain data of the song is convolved to form a high-frequency domain convolution vector; Two different pooling methods are applied to the high-frequency domain convolution vector to form a first frequency domain pooling vector and a second frequency domain pooling vector, wherein the vector size of the first frequency domain pooling vector and the vector size of the second frequency domain pooling vector are the same. The first frequency domain pooling vector is mapped to form a gated parameter distribution, wherein the vector size of the gated parameter distribution is the same as the vector size of the first frequency domain pooling vector, and each vector parameter in the gated parameter distribution is greater than or equal to 0 and less than or equal to 1. The gating parameter distribution and the second frequency domain pooling vector are multiplied bitwise to form the high-frequency vector of the song.

7. The method for restoring the audio quality of a music library according to any one of claims 1-6, characterized in that, The step of fusing the low-frequency vector and the high-frequency vector of the song into the global vector of the song to form a fused vector includes: When the global vector of the song belongs to the result of semantic encoding in the time domain, the low-frequency vector of the song is used as the first vector to be fused, and the high-frequency vector of the song is used as the second vector to be fused. When the global vector of the song belongs to the semantic coding result of the frequency domain, the low-frequency vector of the song is used as the second vector to be fused, and the high-frequency vector of the song is used as the first vector to be fused. The first vector to be fused is fused into the global vector of the song based on a gating mechanism, and the second vector to be fused is fused into the global vector of the song based on an attention mechanism, so as to form a fused vector of the song.

8. The method for restoring the audio quality of a music library according to claim 7, characterized in that, The steps of fusing the first vector to be fused into the global vector of the song based on a gating mechanism, and fusing the second vector to be fused into the global vector of the song based on an attention mechanism to form a fused song vector, include: The first vector to be fused is mapped to form a gating mapping parameter, and the gating mapping parameter and the global vector of the song are multiplied bitwise to form a first fused vector; The global vector of the song is connected to the first fusion vector to form a first connection vector; Determine the attention parameter distribution between the second vector to be fused and the first connection vector, and perform a weighted summation operation on the first connection vector based on the attention parameter distribution to form the second fused vector; The first connection vector is connected to the second fusion vector to form the second connection vector; Based on the second connection vector, an iterative fusion process is performed to form a song fusion vector. The iterative fusion process includes fusing the first vector to be fused into the second connection vector based on a gating mechanism, and fusing the second vector to be fused into the second connection vector based on an attention mechanism to form a new second connection vector. The last newly formed second connection vector is used as the song fusion vector.

9. A music library audio quality restoration device, characterized in that, include: The recording song conversion module is used to perform high and low frequency separation processing on the recorded songs to be repaired in the target music library in the time domain to form low frequency time domain data and high frequency time domain data of the songs, and to perform frequency domain conversion on the high frequency time domain data of the songs to form high frequency frequency domain data of the songs. The first encoding module is used to perform temporal semantic encoding on the recorded song to be repaired or to perform frequency domain semantic encoding on the frequency domain conversion result of the recorded song to be repaired, so as to form a global vector of the song. The second encoding module is used to perform temporal semantic encoding on the low-frequency time-domain data of the song to form a low-frequency vector of the song. The third encoding module is used to perform semantic encoding of the high-frequency domain data of the song in the frequency domain to form a high-frequency vector of the song; The vector fusion module is used to fuse the low-frequency vector and the high-frequency vector of the song into the global vector of the song to form a fused vector of the song. The semantic decoding module is used to perform semantic decoding on the song fusion vector to form the target repair recorded song.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor connected to the memory is used to execute the computer program stored in the memory to implement the music library audio quality restoration method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Encoding method and device for audio signals

    CN103035248A

  • High-frequency optimization method and device for audios and medium

    CN112562703A