Music library sound quality restoration method and device and electronic equipment

By separating the high and low frequencies and semantically encoding songs recorded with low-quality recording equipment, utilizing the smoothness and potential semantic relationship of low-frequency signals, and fusing the low-frequency vectors and high-frequency vectors of songs, the problem of low efficiency in sound quality restoration of songs recorded with low-quality recording equipment is solved, and efficient and intelligent sound quality restoration is achieved.

CN120708636AActive Publication Date: 2025-09-26CHENGDU XIAOCHANG TECH CO LTD

Patent Information

Application Number
CN202511198517.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-09-26
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

In the existing technology, the sound quality restoration of songs recorded with low-quality recording equipment is inefficient and relies on manual editing, which is not very efficient.

Method used

By separating the high and low frequencies of songs in the target music library, combining semantic encoding and decoding, and utilizing the smoothness and potential semantic relationship of low-frequency signals, the song's low-frequency vector, song's high-frequency vector and song's global vector are fused to form a song fusion vector, ultimately achieving sound quality restoration.

Benefits of technology

It improves the efficiency and reliability of song sound quality restoration, avoids local information loss in semantic decoding results, and realizes efficient and intelligent sound quality restoration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708636A_ABST
    Figure CN120708636A_ABST
Patent Text Reader

Abstract

The invention provides a music library tone quality restoration method and device and electronic equipment, and relates to the technical field of data processing. In the application, firstly, high-low frequency separation processing is performed to form song low-frequency time domain data and song high-frequency time domain data, and frequency domain conversion is performed to form song high-frequency frequency domain data; secondly, performing time-domain semantic coding or frequency-domain semantic coding on the recorded song to be repaired to form a song global vector; secondly, performing time-domain semantic coding on the low-frequency time-domain data of the song to form a low-frequency vector of the song; thirdly, performing frequency domain semantic coding on the high-frequency frequency domain data of the song to form a song high-frequency vector; further, fusing the song low-frequency vector and the song high-frequency vector into the song global vector to form a song fusion vector; and finally, performing semantic decoding on the song fusion vector to form a target repair recorded song. On the basis of the content, the problem that in the prior art, the sound quality repairing efficiency of recorded songs is relatively low can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method, device and electronic device for repairing the sound quality of a music library. Background Art

[0002] High-quality recording equipment uses advanced audio processing technology and usually supports high-precision sampling rates, such as 44.1kHz or higher. These technologies help capture a wider frequency spectrum, reduce distortion and noise, and provide a purer audio signal. Low-quality recording equipment has limited audio processing capabilities and low sampling rates, which may result in degraded sound quality (for example, the high-frequency parts of drums and vocals cannot be captured) and the sound is not clear enough. Therefore, choosing the right recording equipment is crucial to achieving high-quality recording effects, especially in the fields of music production and recording studio use. However, for songs in the music library that have been recorded using low-quality recording equipment, the focus is on restoring the sound quality of low-quality recorded songs. However, in existing technologies, sound quality restoration is generally achieved through manual editing, which relies on high professional skills and is relatively inefficient. Therefore, there is an urgent need for a solution that can efficiently repair the sound quality of recorded songs. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a method, device and electronic equipment for repairing the sound quality of a music library, so as to improve the problem of relatively low efficiency of song sound quality repair in the prior art.

[0004] To achieve the above objectives, this application adopts the following technical solutions: A method for repairing the sound quality of a music library, comprising: Performing high- and low-frequency separation processing on the recorded songs to be repaired in the target music library in the time domain to form song low-frequency time domain data and song high-frequency time domain data, and performing frequency domain conversion on the song high-frequency time domain data to form song high-frequency frequency domain data; Performing semantic coding of the recorded song to be repaired in the time domain or performing semantic coding of the frequency domain conversion result of the recorded song to be repaired in the frequency domain to form a global song vector; Performing time-domain semantic encoding on the low-frequency time-domain data of the song to form a low-frequency vector of the song; Performing frequency domain semantic encoding on the high-frequency frequency domain data of the song to form a high-frequency vector of the song; Fusing the song low-frequency vector and the song high-frequency vector into the song global vector to form a song fusion vector; The song fusion vector is semantically decoded to form a target repaired recorded song.

[0005] In a preferred embodiment of the present application, in the above-mentioned music library sound quality restoration method, the step of performing semantic encoding of the low-frequency time domain data of the song in the time domain to form a low-frequency vector of the song includes: Performing convolution processing on the low-frequency time-domain data of the song to form a low-frequency time-domain convolution vector; Extracting long-term average characteristic parameters from the low-frequency time domain data of the song to obtain a long-term average characteristic distribution corresponding to the low-frequency time domain data of the song, and performing convolution processing on the long-term average characteristic distribution to form a long-term average characteristic vector; Based on the long-term average characteristic vector, the low-frequency time-domain convolution vector is semantically enhanced to form a low-frequency vector of the song.

[0006] In a preferred embodiment of the present application, in the above-mentioned music library sound quality restoration method, the step of semantically enhancing the low-frequency time domain convolution vector based on the long-time average characteristic vector to form a low-frequency vector of the song includes: Performing an upsampling process on the low-frequency time-domain convolution vector by A depths to form a low-frequency time-domain depth vector of A depths; For the a-th depth among the A depths, semantic enhancement processing is performed based on the low-frequency time domain depth vector of the a-th depth and the enhanced auxiliary vector of the a-th depth to form a low-frequency enhanced vector of the song at the a-th depth, wherein a is a positive integer less than or equal to A, when a=1, the enhanced auxiliary vector of the a-th depth is the long-term average characteristic vector, and when a>1, the enhanced auxiliary vector of the a-th depth is the low-frequency enhanced vector of the song at the a-th depth; Based on the song low-frequency enhancement vectors of the A depths and the low-frequency time-domain depth vectors of the A depths, a song low-frequency vector is determined.

[0007] In a preferred embodiment of the present application, in the above-mentioned music library sound quality restoration method, the step of performing semantic enhancement processing on the a-th depth among the A depths based on the low-frequency time domain depth vector of the a-th depth and the enhanced auxiliary vector of the a-th depth to form the low-frequency enhanced vector of the song at the a-th depth includes: Using the bth semantic enhancement unit among B semantic enhancement units, the input vector of the bth semantic enhancement unit is processed to form a semantic enhancement output vector of the bth semantic enhancement unit, wherein b is a positive integer less than or equal to B, when b=1, the input vector of the bth semantic enhancement unit includes the low-frequency time domain depth vector of the ath depth and the enhancement auxiliary vector of the ath depth, and when b>1, the input vector of the bth semantic enhancement unit includes the low-frequency time domain depth vector of the ath depth and the semantic enhancement output vector of the b-1th semantic enhancement unit; The semantic enhancement output vector of the Bth semantic enhancement unit is used as the low-frequency enhancement vector of the song at the ath depth.

[0008] In a preferred embodiment of the present application, in the above-mentioned music library sound quality restoration method, the step of using the bth semantic enhancement unit among the B semantic enhancement units to process the input vector of the bth semantic enhancement unit to form the semantic enhancement output vector of the bth semantic enhancement unit includes: Using the bth semantic enhancement unit among the B semantic enhancement units, performing two different mappings on the low-frequency time-domain depth vector of the ath depth to form a first time-domain depth vector and a second time-domain depth vector; performing offset removal on the semantic enhancement output vector included in the input vector of the b-th semantic enhancement unit to form a semantic enhancement offset vector, wherein the offset removal comprises performing difference calculation on each vector parameter in the semantic enhancement output vector and a mean value of each vector parameter of the semantic enhancement output vector; Performing a bitwise multiplication operation on the first temporal depth vector and the semantic enhancement offset vector to form a third temporal depth vector; performing quotient calculations on the standard deviations of the vector parameters in the third time-domain depth vector and the vector parameters of the semantic enhancement output vector to form a fourth time-domain depth vector; A bitwise addition operation is performed on the fourth time-domain depth vector and the second time-domain depth vector to form a semantic enhancement output vector of the b-th semantic enhancement unit.

[0009] In a preferred embodiment of the present application, in the above-mentioned music library sound quality restoration method, the step of performing frequency domain semantic encoding on the high-frequency frequency domain data of the song to form a high-frequency vector of the song includes: Performing convolution processing on the high-frequency frequency domain data of the song to form a high-frequency frequency domain convolution vector; Performing two different pooling operations on the high-frequency frequency domain convolution vector to form a first frequency domain pooling vector and a second frequency domain pooling vector, wherein the vector size of the first frequency domain pooling vector is the same as the vector size of the second frequency domain pooling vector; Mapping the first frequency-domain pooling vector to form a gating parameter distribution, wherein a vector size of the gating parameter distribution is the same as a vector size of the first frequency-domain pooling vector, and each vector parameter in the gating parameter distribution is greater than or equal to 0 and less than or equal to 1; A bitwise multiplication operation is performed on the gating parameter distribution and the second frequency domain pooling vector to form a high-frequency vector of the song.

[0010] In a preferred embodiment of the present application, in the above-mentioned music library sound quality restoration method, the step of fusing the song low-frequency vector and the song high-frequency vector into the song global vector to form a song fusion vector includes: When the song global vector is a result of semantic encoding in the time domain, the song low-frequency vector is used as the first vector to be fused, and the song high-frequency vector is used as the second vector to be fused; When the song global vector is a result of semantic encoding in the frequency domain, the song low-frequency vector is used as the second vector to be fused, and the song high-frequency vector is used as the first vector to be fused; The first vector to be fused is fused into the song global vector based on a gating mechanism, and the second vector to be fused is fused into the song global vector based on an attention mechanism to form a song fusion vector.

[0011] In a preferred embodiment of the present application, in the above-mentioned music library sound quality restoration method, the step of fusing the first vector to be fused into the song global vector based on the gating mechanism, and fusing the second vector to be fused into the song global vector based on the attention mechanism to form a song fusion vector includes: Mapping the first vector to be fused to form a gate mapping parameter, and performing a bitwise multiplication operation on the gate mapping parameter and the song global vector to form a first fused vector; Connecting the song global vector to the first fusion vector to form a first connection vector; determining an attention parameter distribution between the second to-be-fused vector and the first connection vector, and performing a weighted sum operation on the first connection vector based on the attention parameter distribution to form a second fused vector; Connecting the first connection vector to the second fusion vector to form a second connection vector; An iterative fusion process is performed based on the second connection vector to form a song fusion vector, wherein one fusion process in the iterative fusion process includes fusing the first vector to be fused into the second connection vector based on a gating mechanism, and fusing the second vector to be fused into the second connection vector based on an attention mechanism to form a new second connection vector, and the new second connection vector formed for the last time is used as the song fusion vector.

[0012] This application also provides a music library sound quality repair device, comprising: The recorded song conversion module is used to perform high- and low-frequency separation processing on the recorded songs to be repaired in the target music library in the time domain to form low-frequency time domain data and high-frequency time domain data of the songs, and to perform frequency domain conversion on the high-frequency time domain data of the songs to form high-frequency frequency domain data of the songs; A first encoding module is used to perform semantic encoding of the recorded song to be repaired in the time domain or perform semantic encoding of the frequency domain conversion result of the recorded song to be repaired in the frequency domain to form a global song vector; A second encoding module is used to perform time-domain semantic encoding on the low-frequency time-domain data of the song to form a low-frequency vector of the song; A third encoding module is used to perform frequency domain semantic encoding on the high-frequency frequency domain data of the song to form a high-frequency vector of the song; A vector fusion module, configured to fuse the low-frequency vector of the song and the high-frequency vector of the song into the global vector of the song to form a song fusion vector; The semantic decoding module is used to perform semantic decoding on the song fusion vector to form a target repaired recorded song.

[0013] Based on the above, the present application further provides an electronic device, including: memory for storing computer programs; The processor connected to the memory is used to execute the computer program stored in the memory to implement the above-mentioned music library sound quality repair method.

[0014] The present application provides a method, device and electronic device for repairing the sound quality of a music library. First, the recorded songs to be repaired in the target music library are subjected to high and low frequency separation processing to form low-frequency time domain data and high-frequency time domain data of the songs, and frequency domain conversion is performed to form high-frequency frequency domain data of the songs; secondly, the recorded songs to be repaired are subjected to semantic encoding in the time domain or semantic encoding in the frequency domain to form a global vector of the song; then, the low-frequency time domain data of the songs are subjected to semantic encoding in the time domain to form a low-frequency vector of the songs; thereafter, the high-frequency frequency domain data of the songs are subjected to semantic encoding in the frequency domain to form a high-frequency vector of the songs; further, the low-frequency vector of the songs and the high-frequency vector of the songs are fused into the global vector of the songs to form a fusion vector of the songs; finally, the fusion vector of the songs is semantically decoded to form the target repaired recorded songs. Based on the above content, on the one hand, because the sound quality repair of the recorded songs can be intelligently realized by semantically encoding and decoding the recorded songs to be repaired, it can have higher efficiency than the conventional repair scheme based on manual editing, thereby improving the problem of relatively low efficiency of the sound quality repair of recorded songs in the prior art. On the other hand, since low-frequency signals vary slowly over time, meaning that signal values ​​between adjacent time points do not vary much, the slow-changing nature of low-frequency signals (i.e., their good stability) and the potential semantic relationship between low-frequency and high-frequency signals can be exploited to represent the lost high-frequency signals. This means mining the song's low-frequency vector in the time domain, which represents the song's low-frequency time-domain data. This ensures that the target restored recording, obtained through semantic decoding, retains the lost high-frequency signal. Furthermore, it should be noted that the semantically decoded song fusion vector includes not only the song's low-frequency vector, but also its high-frequency vector and global vector, making the semantic representation more comprehensive. This effectively avoids the problem of local information loss in the semantic decoding results. This improves the reliability of sound quality restoration in music libraries. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings.

[0016] Figure 1 This is a structural block diagram of an electronic device provided in an embodiment of the present application.

[0017] Figure 2 A flowchart of the method for repairing the sound quality of a music library provided in an embodiment of the present application.

[0018] Figure 3 Schematic diagram of the implementation process of the music library sound quality repair method provided in the embodiment of the present application.

[0019] Figure 4 A schematic diagram of the semantic enhancement processing provided in an embodiment of the present application.

[0020] Figure 5 A schematic diagram of vector fusion provided in an embodiment of the present application.

[0021] Figure 6 A block diagram of the device for repairing the sound quality of a music library provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.

[0023] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for protection, but merely represents selected embodiments of the present application. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments in the present application without creative work are within the scope of protection of the present application.

[0024] like Figure 1 As shown, an embodiment of the present application provides an electronic device, wherein the electronic device may include a memory, a processor, and a music library sound quality restoration device.

[0025] In detail, the memory and the processor are electrically connected directly or indirectly to achieve data transmission or interaction. For example, the memory and the processor can be electrically connected through one or more communication buses or signal lines. The music library sound quality repair device includes at least one software function module stored in the memory in the form of software or firmware. The processor is used to execute the executable computer program stored in the memory, for example, the software function module and computer program included in the music library sound quality repair device, so as to implement the music library sound quality repair method provided in the embodiment of the present application.

[0026] Optionally, the memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0027] Furthermore, the processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a system on chip (SoC), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0028] I understand. Figure 1 The structure shown is only for illustration, and the electronic device may also include Figure 1 More or fewer components than shown, or with Figure 1 The different configurations shown may, for example, further include a communication unit for exchanging information with other devices.

[0029] Combine Figure 2 , the embodiment of the present application also provides a method for repairing the sound quality of a music library that can be applied to the above electronic device. Among them, the method steps defined in the process related to the method for repairing the sound quality of the music library can be implemented by the electronic device. Figure 2 The specific process shown is explained in detail.

[0030] Step S110, performing high and low frequency separation processing on the recorded songs to be repaired in the target music library in the time domain to form low-frequency time domain data and high-frequency time domain data of the songs, and performing frequency domain conversion on the high-frequency time domain data of the songs to form high-frequency frequency domain data of the songs.

[0031] In the embodiment of the present application, the electronic device can perform high and low frequency separation processing on the recorded songs to be repaired in the target music library in the time domain to form song low frequency time domain data and song high frequency time domain data (combined with Figure 3, achieved through a low-pass filter and a high-pass filter, that is, filtering the recorded song to be repaired through a low-pass filter to obtain low-frequency time domain data of the song, and filtering the recorded song to be repaired through a high-pass filter to obtain high-frequency time domain data of the song), and performing frequency domain conversion on the high-frequency time domain data of the song to form high-frequency frequency domain data of the song. For example, the high-frequency time domain data of the song can be Fourier transformed to obtain a spectrum diagram, that is, the high-frequency frequency domain data of the song.

[0032] Step S120 , performing semantic coding in the time domain on the recorded song to be repaired or performing semantic coding in the frequency domain on the frequency domain conversion result of the recorded song to be repaired, to form a global song vector.

[0033] In an embodiment of the present application, the electronic device can perform semantic encoding on the recorded song to be repaired in the time domain or perform semantic encoding on the frequency domain conversion result of the recorded song to be repaired in the frequency domain to form a global vector of the song. That is to say, the recorded song to be repaired in the time domain can be directly semantically encoded to obtain a global vector of the song, or the recorded song to be repaired in the time domain can be Fourier transformed to obtain a corresponding spectrum graph, and then the spectrum graph can be semantically encoded to obtain a global vector of the song. Exemplarily, the aforementioned semantic encoding can include convolution and self-attention processing. For example, the recorded song to be repaired in the time domain can be convolved, and then the obtained convolution vector can be self-attention processed to obtain a global vector of the song, that is, to characterize the global semantic information of the recorded song to be repaired. For another example, the spectrum graph can be convolved, and then the obtained convolution vector can be self-attention processed to obtain a global vector of the song. Alternatively, any existing semantic encoding method can be adopted.

[0034] Step S130: performing time-domain semantic encoding on the low-frequency time-domain data of the song to form a low-frequency vector of the song.

[0035] In an embodiment of the present application, after forming the low-frequency time domain data of the song, the electronic device can perform time domain semantic encoding on the low-frequency time domain data of the song to form a low-frequency vector of the song. That is, since the low-frequency signal changes slowly in time, it means that the signal value between adjacent time points does not change much. Therefore, the slow-changing characteristics of the low-frequency signal (i.e., it has good stability) and the potential semantic relationship between the low-frequency signal and the high-frequency signal can be used to achieve the characterization of the lost high-frequency signal, that is, to mine the low-frequency vector of the song in the time domain of the low-frequency time domain data of the song. Among them, the potential semantic relationship between the low-frequency signal and the high-frequency signal can actually be represented by the parameters of the neural network model that performs semantic encoding (i.e., during the learning process, the neural network model will learn the recorded song sample to be repaired (low quality, missing some high-frequency information) and the target repaired recorded song label (high quality, no loss of high-frequency information), thereby capturing the potential semantic relationship between the low-frequency information in the recorded song sample to be repaired and the lost high-frequency information).

[0036] Step S140: performing frequency domain semantic encoding on the high-frequency frequency domain data of the song to form a high-frequency vector of the song.

[0037] In an embodiment of the present application, after forming the high-frequency time-domain data of the song, the electronic device may perform frequency-domain semantic encoding on the high-frequency frequency-domain data of the song to form a high-frequency vector of the song. Since high-frequency information can characterize detailed information, it is also necessary to separately perform semantic encoding on the high-frequency frequency-domain data of the song to form a high-frequency vector of the song that fully characterizes the detailed information.

[0038] Step S150: Fusing the song low-frequency vector and the song high-frequency vector into the song global vector to form a song fusion vector.

[0039] In an embodiment of the present application, after forming the song's high-frequency vector, the song's low-frequency vector, and the song's high-frequency vector, the electronic device may fuse the song's low-frequency vector and the song's high-frequency vector into the song's global vector to form a song fusion vector. That is, based on the global semantic information represented by the song's global vector, the semantic information represented by the song's low-frequency vector and the song's high-frequency vector may be further fused, so that the semantic information represented by the formed song fusion vector has both high richness and high precision.

[0040] Step S160: semantically decode the song fusion vector to form a target repaired recorded song.

[0041] In an embodiment of the present application, after forming the song fusion vector, the electronic device can perform semantic decoding on the song fusion vector to form a target repaired recorded song. It is understandable that the target repaired recorded song in the time domain can be directly generated through the semantic decoding, or the corresponding spectrogram can be generated first, and then the spectrogram is inverse Fourier transformed to obtain the target repaired recorded song in the time domain. For example, the generation of the corresponding spectrogram can be achieved by a decoder, wherein the network architecture of the decoder is as follows: Fully connected layer: outputs a 256×16×16 feature map, with activation function of ReLU or LeakyReLU; Deconvolution layer 1: Input size is 256×16×16, output size is 128×32×32, convolution kernel size is 3x3, stride 2, padding 1, activation function is LeakyReLU; Deconvolution layer 2: input size is 128×32×32, output size is 64×64×64, convolution kernel size is 3x3, stride 2, padding 1, activation function is LeakyReLU; Deconvolution layer 3: Input size is 64×64×64, output size is 1×128×128, convolution kernel size is 3x3, stride 2, padding 1, activation function is Tanh or Sigmoid; Based on this, the size of the obtained spectrogram is 128×128. Based on the above, intelligent song restoration can be achieved by semantically encoding and decoding the song to be restored. This approach offers greater efficiency than conventional restoration methods based on manual editing, thereby addressing the relatively low efficiency of song restoration in existing techniques. Furthermore, since low-frequency signals vary slowly over time, meaning that signal values ​​between adjacent time points do not vary significantly, the slow-changing nature of low-frequency signals (i.e., their good stationarity) and the potential semantic relationship between low-frequency and high-frequency signals can be leveraged to represent the missing high-frequency signals. This involves mining the song's low-frequency vector in the time domain, thereby ensuring that the target song, obtained through semantic decoding, retains the missing high-frequency signals. Furthermore, the semantically decoded song fusion vector includes not only the song's low-frequency vector, but also its high-frequency vector and global vector, providing a more comprehensive semantic representation. This effectively avoids the problem of local information loss in the semantic decoding results. Consequently, the reliability of song restoration can be improved.

[0042] First, it should be noted that for step S130, the specific method of performing time-domain semantic encoding on the low-frequency time-domain data of the song is not limited and can be selected according to actual needs.

[0043] For example, in an alternative embodiment, convolution and self-attention processing can be performed on the low-frequency time domain data of the song to obtain a low-frequency vector of the song.

[0044] For example, in another alternative embodiment, in order to improve the semantic representation ability of the formed low-frequency vector of the song, the above-mentioned step S130 can further include step S131, step S132 and step S133, and the specific content of each step is described as follows.

[0045] Step S131, performing convolution processing on the low-frequency time domain data of the song to form a low-frequency time domain convolution vector.

[0046] In an embodiment of the present application, the low-frequency time domain data of the song may be subjected to convolution processing to form a low-frequency time domain convolution vector. In other words, primary semantic information may be extracted from the low-frequency time domain data of the song through convolution processing, thereby obtaining the low-frequency time domain convolution vector.

[0047] Step S132, extracting the long-time average characteristic parameters of the low-frequency time domain data of the song to obtain the long-time average characteristic distribution corresponding to the low-frequency time domain data of the song, and performing convolution processing on the long-time average characteristic distribution to form a long-time average characteristic vector.

[0048] In an embodiment of the present application, the long-term average characteristic parameters of the low-frequency time domain data of the song can be extracted to obtain the long-term average characteristic distribution corresponding to the low-frequency time domain data of the song, and the long-term average characteristic distribution can be convolved to form a long-term average characteristic vector. It should be noted that the average characteristics of the high-frequency signal over a long period of time, such as the average value, autocorrelation, etc., may be related to the characteristics of the low-frequency signal. Due to its stationary nature, the low-frequency signal can be used to approximate the behavior of the high-frequency signal under long-term averaging. This characteristic is particularly important in signal processing, especially when the high-frequency signal is missing, the low-frequency signal can provide important reference information. Therefore, the long-term average characteristic distribution can be first extracted from the low-frequency time domain data of the song. For example, the low-frequency time domain data of the song can be subjected to sliding window processing, and the mean of the time domain signal in each window can be calculated respectively. In this way, the corresponding long-term average characteristic distribution can be obtained. Then, the long-term average characteristic distribution can be convolved to obtain the long-term average characteristic vector.

[0049] Step S133: Based on the long-term average characteristic vector, semantic enhancement is performed on the low-frequency time-domain convolution vector to form a low-frequency vector of the song.

[0050] In an embodiment of the present application, after forming the long-term average characteristic vector and the low-frequency time-domain convolution vector, the low-frequency time-domain convolution vector can be semantically enhanced based on the long-term average characteristic vector to form a low-frequency vector of the song. That is to say, since the low-frequency time-domain convolution vector can represent the low-frequency semantic information as a whole, but since the long-term average characteristic vector can reflect some characteristics of the high-frequency signal, the semantic information in the long-term average characteristic vector can be fused into the low-frequency time-domain convolution vector to obtain the low-frequency vector of the song.

[0051] It can be understood that in the above-mentioned step S133, the specific method of semantically enhancing the low-frequency time domain convolution vector is not limited. For example, in an alternative embodiment, in order to fully fuse the long-term average feature vector and the low-frequency time domain convolution vector, thereby ensuring the reliability of semantic enhancement, the above-mentioned step S133 can further include step S133a, step S133b and step S133c, and the specific content of each step is described below.

[0052] Step S133a: performing upsampling processing on the low-frequency time-domain convolution vector by A depths to form a low-frequency time-domain depth vector of A depths.

[0053] In an embodiment of the present application, the low-frequency time domain convolution vector can be upsampled by A depths to form a low-frequency time domain depth vector of A depths, where A can be a positive integer greater than or equal to 2. That is, a plurality of semantic vectors with different depths can be obtained, thereby fully representing the semantic information of the low-frequency time domain at different depths. Exemplarily, the low-frequency time domain convolution vector is upsampled by the first upsampling network to output a 128*128 low-frequency time domain depth vector. The low-frequency time domain depth vector output by the first upsampling network is upsampled by the second upsampling network to obtain a 256*256 low-frequency time domain depth vector. The low-frequency time domain depth vector output by the second upsampling network is upsampled by the third upsampling network to obtain a 512*512 low-frequency time domain depth vector. The low-frequency time domain depth vector output by the third upsampling network is upsampled by the fourth upsampling network to obtain a 1024*1024 low-frequency time domain depth vector. Step S133b: For the a-th depth among the A depths, semantic enhancement processing is performed based on the low-frequency time domain depth vector of the a-th depth and the enhanced auxiliary vector of the a-th depth to form a low-frequency enhancement vector of the song at the a-th depth.

[0054] In an embodiment of the present application, after obtaining the low-frequency time domain depth vector of A depths, for the a-th depth among the A depths, semantic enhancement processing can be performed based on the low-frequency time domain depth vector of the a-th depth and the enhancement auxiliary vector of the a-th depth to form the low-frequency enhancement vector of the song of the a-th depth. Wherein, a is a positive integer less than or equal to A. When a=1, the enhancement auxiliary vector of the a-th depth is the long-term average characteristic vector. When a>1, the enhancement auxiliary vector of the a-th depth is the low-frequency enhancement vector of the song of the a-1-th depth. Exemplarily, semantic enhancement processing can be performed based on the low-frequency time domain depth vector of the first depth and the long-term average characteristic vector to form the low-frequency enhancement vector of the song of the first depth. And, semantic enhancement processing can be performed based on the low-frequency time domain depth vector of the second depth and the low-frequency enhancement vector of the song of the first depth to form the low-frequency enhancement vector of the song of the second depth.

[0055] Step S133c: Determine the song low-frequency vector based on the song low-frequency enhancement vectors of the A depths and the low-frequency time-domain depth vectors of the A depths.

[0056] In an embodiment of the present application, after forming the song low-frequency enhancement vector of the A depth, the song low-frequency vector can be determined based on the song low-frequency enhancement vector of the A depth and the low-frequency time-domain depth vector of the A depth. Exemplarily, the song low-frequency enhancement vector of the A depth and the low-frequency time-domain depth vector of the A depth can be spliced ​​(they can be expanded first and then spliced, or they can be upsampled or downsampled to the same size and then spliced) to form a spliced ​​vector. Then, the spliced ​​vector can be convolved and self-attention processed to obtain the corresponding song low-frequency vector. In this way, the problem of semantic distortion caused by the greater depth of semantic enhancement processing can be avoided, thereby ensuring the accuracy of the semantic representation of the determined song low-frequency vector.

[0057] It can be understood that, in the above-mentioned step S133b, the specific method of performing semantic enhancement processing based on the low-frequency time domain depth vector of the a-th depth and the enhanced auxiliary vector of the a-th depth is not limited. For example, in an alternative embodiment, in order to further improve the effect of fusing different vectors in the process of semantic enhancement processing, that is, to achieve full fusion of semantic information between different vectors, the above-mentioned step S133b can further include step b1 and step b2, and the specific content of each step is described as follows.

[0058] Step b1: using the bth semantic reinforcement unit among the B semantic reinforcement units, processing the input vector of the bth semantic reinforcement unit to form a semantic reinforcement output vector of the bth semantic reinforcement unit.

[0059] In an embodiment of the present application, the bth semantic reinforcement unit among B semantic reinforcement units can be used to process the input vector of the bth semantic reinforcement unit to form the semantic reinforcement output vector of the bth semantic reinforcement unit. Wherein, b is a positive integer less than or equal to B (B can be a positive integer greater than or equal to 2), when b=1, the input vector of the bth semantic reinforcement unit includes the low-frequency time domain depth vector of the ath depth and the reinforcement auxiliary vector of the ath depth, and when b>1, the input vector of the bth semantic reinforcement unit includes the low-frequency time domain depth vector of the ath depth and the semantic reinforcement output vector of the b-1th semantic reinforcement unit. Exemplarily, in the first semantic reinforcement unit, the low-frequency time domain depth vector of the ath depth and the reinforcement auxiliary vector of the ath depth can be processed to obtain the semantic reinforcement output vector of the first semantic reinforcement unit. In the second semantic enhancement unit, the low-frequency time-domain depth vector of the a-th depth and the semantic enhancement output vector of the first semantic enhancement unit can be processed to obtain the semantic enhancement output vector of the second semantic enhancement unit.

[0060] Step b2: Use the semantic enhancement output vector of the Bth semantic enhancement unit as the low-frequency enhancement vector of the song at the ath depth.

[0061] In an embodiment of the present application, after obtaining the semantic reinforcement output vector of the Bth semantic reinforcement unit, that is, after obtaining the semantic reinforcement output vector of the last semantic reinforcement unit, the semantic reinforcement output vector of the last semantic reinforcement unit can be used as the low-frequency reinforcement vector of the song at the ath depth.

[0062] It is understandable that in the above step b1, the specific manner of processing the input vector of the bth semantic enhancement unit is not limited. For example, in an alternative embodiment, in order to achieve full fusion between the two semantic vectors, the above step b1 may further include the following contents (combined with Figure 4 ): First, the bth semantic enhancement unit among the B semantic enhancement units can be used to perform two different mappings on the low-frequency time-domain depth vector of the ath depth to form a first time-domain depth vector and a second time-domain depth vector; illustratively, the mappings can both be linear mappings, such as y1=a1x+b1, y2=a2x+b2, y1 and y2 are the first time-domain depth vector and the second time-domain depth vector, x is the low-frequency time-domain depth vector of the ath depth, a1 and a2 are both weight parameters, b1 and b2 are both bias parameters, and are all formed during the training process of the corresponding neural network model; in this way, using two linear mappings, the representation ability of the semantic vector can be further optimized; Secondly, the semantic enhancement output vector included in the input vector of the b-th semantic enhancement unit (when b=1, the semantic enhancement output vector is the enhancement auxiliary vector of the a-th depth, and when b is greater than 1, the semantic enhancement output vector is the semantic enhancement output vector of the b-1-th semantic enhancement unit) can be offset removed to form a semantic enhancement offset vector, wherein the offset removal includes respectively calculating the difference between each vector parameter in the semantic enhancement output vector and the mean of each vector parameter of the semantic enhancement output vector; based on this, the operation of subtracting each vector parameter from its mean can achieve decentralization, which means removing the mean of the input data so that the mean of each vector parameter becomes zero. In this way, the decentralized data can eliminate the offset and make the data more evenly distributed. After removing the mean, the local differences of each feature can be better highlighted, and it can help the network better capture the differences and relationships between different elements. Then, the first time-domain depth vector and the semantic enhancement offset vector can be bitwise multiplied to form a third time-domain depth vector. Based on this, since the semantic enhancement offset vector is decentralized, different parameters can represent different levels of importance. In this way, through the bitwise multiplication operation, the first time-domain depth vector can be weighted based on the semantic enhancement offset vector, that is, gated screening can be achieved, thereby completing the full fusion of the two semantic vectors. Furthermore, the standard deviations of the vector parameters in the third time-domain depth vector and the vector parameters of the semantic enhancement output vector can be quotiented to form a fourth time-domain depth vector. Based on this, by quotienting each vector parameter with the standard deviation of the semantic enhancement output vector, each local feature is actually scaled to eliminate possible differences between different dimensions. Specifically, the standard deviation represents the degree of dispersion of the feature. Dividing by the standard deviation allows the features to be compared and fused at the same scale, which helps to eliminate problems that may arise from different feature scales and ensures that the data processed by the network has a uniform scale. Finally, the fourth time-domain depth vector and the second time-domain depth vector can be bitwise added to form the semantic enhancement output vector of the b-th semantic enhancement unit; that is, since the fourth time-domain depth vector is actually the fusion result of the two semantic vectors, the low-frequency time-domain depth vector and the semantic enhancement output vector, and as the fusion continues to deepen, semantic distortion may occur. Therefore, by performing a bitwise addition operation on the second time-domain depth vector that can also represent the low-frequency time-domain depth vector and the fourth time-domain depth vector, it can be ensured that the formed semantic enhancement output vector has a higher semantic representation accuracy.

[0063] Secondly, it should be noted that for step S140, the specific method of performing frequency domain semantic encoding on the high-frequency frequency domain data of the song is not limited and can be selected according to actual needs.

[0064] For example, in an alternative embodiment, the high-frequency frequency domain data of the song can be convolved and self-attention processed to obtain a high-frequency vector of the song.

[0065] For example, in another alternative embodiment, in order to fully mine the semantic information in the high-frequency frequency domain data of the song while avoiding the problem of excessive computational complexity caused by attention processing, the above-mentioned step S140 may further include: First, convolution processing can be performed on the high-frequency frequency domain data of the song to form a high-frequency frequency domain convolution vector; Secondly, two different pooling operations can be performed on the high-frequency frequency domain convolution vector to form a first frequency domain pooling vector (such as the result of maximum pooling) and a second frequency domain pooling vector (such as the result of mean pooling), wherein the vector size of the first frequency domain pooling vector and the vector size of the second frequency domain pooling vector are the same; Then, the first frequency domain pooling vector can be mapped (e.g., nonlinearly mapped by a sigmoid function) to form a gating parameter distribution, wherein the vector size of the gating parameter distribution is the same as the vector size of the first frequency domain pooling vector, and each vector parameter in the gating parameter distribution is greater than or equal to 0 and less than or equal to 1; Finally, the gating parameter distribution and the second frequency domain pooling vector may be bitwise multiplied to form a high-frequency vector of the song.

[0066] Thirdly, it should be noted that for step S150, the specific method of fusing the song low-frequency vector and the song high-frequency vector into the song global vector is not limited and can be selected according to actual needs.

[0067] For example, in an alternative embodiment, the song low-frequency vector, the song high-frequency vector and the song global vector can be added together to achieve fusion, that is, to obtain a song fusion vector that can represent the semantic information in the three semantic vectors.

[0068] For example, in another alternative implementation, in order to improve the fusion effect and avoid the problem of semantic distortion during the fusion process, the above-mentioned step S150 can further include step S151, step S152 and step S153, and the specific content of each step is as follows.

[0069] Step S151, when the global vector of the song is the result of semantic encoding in the time domain, the low-frequency vector of the song is used as the first vector to be fused, and the high-frequency vector of the song is used as the second vector to be fused.

[0070] In an embodiment of the present application, when the song global vector is the result of semantic encoding in the time domain, the song low-frequency vector can be used as the first vector to be fused, and the song high-frequency vector can be used as the second vector to be fused. In other words, the first vector to be fused and the song global vector are vectors in the same domain, while the second vector to be fused and the song global vector are vectors in different domains.

[0071] Step S152, when the global vector of the song is the result of semantic encoding in the frequency domain, the low-frequency vector of the song is used as the second vector to be fused, and the high-frequency vector of the song is used as the first vector to be fused.

[0072] In an embodiment of the present application, when the song global vector is the result of semantic encoding in the frequency domain, the song low-frequency vector can be used as the second vector to be fused, and the song high-frequency vector can be used as the first vector to be fused. In other words, the first vector to be fused and the song global vector are vectors in the same domain, while the second vector to be fused and the song global vector are vectors in different domains.

[0073] Step S153: The first vector to be fused is fused into the global song vector based on a gating mechanism, and the second vector to be fused is fused into the global song vector based on an attention mechanism to form a song fusion vector.

[0074] In an embodiment of the present application, after determining the first vector to be fused and the second vector to be fused based on step S151 and step S152, the first vector to be fused can be fused into the song global vector based on the gating mechanism, and the second vector to be fused can be fused into the song global vector based on the attention mechanism to form a song fusion vector. That is to say, since the first vector to be fused and the song global vector are vectors of the same domain, they can be directly fused based on the gating mechanism. In this way, while avoiding the problem of semantic distortion, the amount of calculation can be reduced and the processing efficiency can be improved. Moreover, since the second vector to be fused and the song global vector are vectors of different domains, in order to avoid the problem of semantic distortion, the attention mechanism can be used to fuse the second vector to be fused into the song global vector, that is, the cross-modal processing capability of the attention mechanism is fully utilized to make the fusion effect better and ensure the accuracy of the semantic representation of the obtained song fusion vector.

[0075] It is understandable that in the above step S153, the specific method of forming the song fusion vector is not limited. For example, in an alternative embodiment, in order to enable the first vector to be fused and the second vector to be fused to be fully fused into the song global vector, the above step S153 may further include the following contents (combined with Figure 5 ): In the first step, the first vector to be fused may be mapped (refer to the relevant description above) to form a gate mapping parameter, and the gate mapping parameter and the song global vector are bitwise multiplied to form a first fused vector; In a second step, the song global vector may be connected to the first fusion vector to form a first connection vector. For example, the song global vector and the first fusion vector may be added together. This can avoid the problem of low quality of the first fusion vector formed by fusion when the quality of the first vector to be fused is not high, thereby avoiding the problem of quality of gated fusion. In a third step, an attention parameter distribution between the second vector to be fused and the first connection vector may be determined (for example, the attention parameter distribution may be obtained by multiplying the query vector corresponding to the second vector to be fused and the transposed result of the key vector corresponding to the first connection vector), and a weighted sum operation is performed on the first connection vector based on the attention parameter distribution (for example, a weighted sum operation is performed on the value vector corresponding to the first connection vector) to form a second fused vector. In a fourth step, the first connection vector may be connected to the second fusion vector to form a second connection vector. For example, the first connection vector and the second fusion vector may be added together to avoid semantic distortion during attention processing. In the fifth step, an iterative fusion process can be performed based on the second connection vector to form a song fusion vector, wherein one fusion process in the iterative fusion process includes fusing the first vector to be fused into the second connection vector based on a gating mechanism, and fusing the second vector to be fused into the second connection vector based on an attention mechanism to form a new second connection vector, and the new second connection vector formed for the last time is used as the song fusion vector. In other words, the first, second, third, and fourth steps above can be executed in reverse. For example, during the first iteration, the song global vector can be replaced with the first second connection vector to form the second second connection vector. During the second iteration, the song global vector can be replaced with the second second connection vector to form the third second connection vector. During the third iteration, the song global vector can be replaced with the third second connection vector to form the fourth second connection vector.

[0076] In detail, during the first iteration: First, the first vector to be fused may be mapped to form a gate mapping parameter, and the gate mapping parameter and the first second connection vector may be bitwise multiplied to form a second first fused vector; Secondly, the first second connection vector may be connected to the second first fusion vector to form a second first connection vector; Then, an attention parameter distribution between the second vector to be fused and the second first connection vector may be determined, and a weighted sum operation may be performed on the second first connection vector based on the attention parameter distribution to form a second second fused vector; Finally, the second first connection vector may be connected to the second second fusion vector to form a second second connection vector.

[0077] Combine Figure 6 The present application also provides a device for restoring the sound quality of a music library applicable to the aforementioned electronic device. The device comprises a recorded song conversion module, a first encoding module, a second encoding module, a third encoding module, a vector fusion module, and a semantic decoding module.

[0078] The recorded song conversion module is used to perform high and low frequency separation processing on the recorded songs to be repaired in the target music library in the time domain to form song low-frequency time domain data and song high-frequency time domain data, and to perform frequency domain conversion on the song high-frequency time domain data to form song high-frequency frequency domain data. In this embodiment of the application, the recorded song conversion module can be used to perform Figure 2 As shown in step S110, for the relevant content of the recorded song conversion module, please refer to the description of step S110 above.

[0079] The first encoding module is used to perform semantic encoding of the time domain of the recorded song to be repaired or to perform semantic encoding of the frequency domain conversion result of the recorded song to be repaired in the frequency domain to form a global song vector. In the embodiment of the present application, the first encoding module can be used to perform Figure 2 As shown in step S120, for the relevant content of the first encoding module, please refer to the above description of step S120.

[0080] The second encoding module is used to perform semantic encoding of the low-frequency time domain data of the song in the time domain to form a low-frequency vector of the song. In the embodiment of the present application, the second encoding module can be used to perform Figure 2 As shown in step S130, for the relevant content of the second encoding module, please refer to the above description of step S130.

[0081] The third encoding module is used to perform frequency domain semantic encoding on the high frequency frequency domain data of the song to form a high frequency vector of the song. In the embodiment of the present application, the third encoding module can be used to perform Figure 2 As shown in step S140, for the relevant content of the third encoding module, please refer to the above description of step S140.

[0082] The vector fusion module is used to fuse the low-frequency vector of the song and the high-frequency vector of the song into the global vector of the song to form a song fusion vector. In this embodiment of the application, the vector fusion module can be used to perform Figure 2 As shown in step S150, for the relevant content of the vector fusion module, reference may be made to the above description of step S150.

[0083] The semantic decoding module is used to perform semantic decoding on the song fusion vector to form a target repair recording song. In the embodiment of the present application, the semantic decoding module can be used to perform Figure 2 As shown in step S160, for the relevant content of the semantic decoding module, reference can be made to the above description of step S160.

[0084] In an embodiment of the present application, corresponding to the above-mentioned method for repairing the sound quality of a music library applied to the electronic device, a computer-readable storage medium is also provided, in which a computer program is stored. When the computer program is run, the various steps of the method for repairing the sound quality of a music library are executed.

[0085] Among them, the steps executed when the aforementioned computer program is running will not be described here one by one, and reference can be made to the above explanation of the method for repairing the sound quality of the music library.

[0086] In summary, the music library sound quality repair method, device and electronic device provided by the present application first perform high and low frequency separation processing on the recorded songs to be repaired in the target music library to form song low-frequency time domain data and song high-frequency time domain data, and perform frequency domain conversion to form song high-frequency frequency domain data; secondly, perform time domain semantic encoding or frequency domain semantic encoding on the recorded songs to be repaired to form a song global vector; then, perform time domain semantic encoding on the song low-frequency time domain data to form a song low-frequency vector; after that, perform frequency domain semantic encoding on the song high-frequency frequency domain data to form a song high-frequency vector; further, fuse the song low-frequency vector and the song high-frequency vector into the song global vector to form a song fusion vector; finally, perform semantic decoding on the song fusion vector to form the target repaired recorded song. Based on the above content, on the one hand, because the recorded songs to be repaired can be semantically encoded and decoded, the sound quality repair of the recorded songs can be intelligently realized, which makes it more efficient than the conventional repair scheme based on manual editing, thereby improving the problem of relatively low efficiency of recorded song sound quality repair in the existing technology. On the other hand, since low-frequency signals vary slowly over time, meaning that signal values ​​between adjacent time points do not vary much, the slow-changing nature of low-frequency signals (i.e., their good stability) and the potential semantic relationship between low-frequency and high-frequency signals can be exploited to represent the lost high-frequency signals. This means mining the song's low-frequency vector in the time domain, which represents the song's low-frequency time-domain data. This ensures that the target restored recording, obtained through semantic decoding, retains the lost high-frequency signal. Furthermore, it should be noted that the semantically decoded song fusion vector includes not only the song's low-frequency vector, but also its high-frequency vector and global vector, making the semantic representation more comprehensive. This effectively avoids the problem of local information loss in the semantic decoding results. This improves the reliability of sound quality restoration in music libraries.

[0087] In the several embodiments provided in the embodiments of the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0088] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0089] If the functions are implemented in the form of software modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, electronic device, or network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. It should be noted that, in this document, the terms "comprise," "include," or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. Without further constraints, an element defined by the phrase "comprises a..." does not preclude the existence of additional identical elements in the process, method, article or apparatus that includes the element.

[0090] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A method for repairing the sound quality of a music library, characterized in that: include: Performing high- and low-frequency separation processing on the recorded songs to be repaired in the target music library in the time domain to form song low-frequency time domain data and song high-frequency time domain data, and performing frequency domain conversion on the song high-frequency time domain data to form song high-frequency frequency domain data; Performing semantic coding of the recorded song to be repaired in the time domain or performing semantic coding of the frequency domain conversion result of the recorded song to be repaired in the frequency domain to form a global song vector; Performing time-domain semantic encoding on the low-frequency time-domain data of the song to form a low-frequency vector of the song; Performing frequency domain semantic encoding on the high-frequency frequency domain data of the song to form a high-frequency vector of the song; Fusing the song low-frequency vector and the song high-frequency vector into the song global vector to form a song fusion vector; The song fusion vector is semantically decoded to form a target repaired recorded song.

2. The method for repairing the sound quality of a music library according to claim 1, characterized in that: The step of performing time-domain semantic encoding on the low-frequency time-domain data of the song to form a low-frequency vector of the song includes: Performing convolution processing on the low-frequency time-domain data of the song to form a low-frequency time-domain convolution vector; Extracting long-term average characteristic parameters from the low-frequency time domain data of the song to obtain a long-term average characteristic distribution corresponding to the low-frequency time domain data of the song, and performing convolution processing on the long-term average characteristic distribution to form a long-term average characteristic vector; Based on the long-term average characteristic vector, the low-frequency time-domain convolution vector is semantically enhanced to form a low-frequency vector of the song.

3. The method for repairing the sound quality of a music library according to claim 2, characterized in that: The step of semantically enhancing the low-frequency time-domain convolution vector based on the long-term average characteristic vector to form a low-frequency vector of the song includes: Performing an upsampling process on the low-frequency time-domain convolution vector by A depths to form a low-frequency time-domain depth vector of A depths; For the a-th depth among the A depths, semantic enhancement processing is performed based on the low-frequency time domain depth vector of the a-th depth and the enhanced auxiliary vector of the a-th depth to form a low-frequency enhanced vector of the song at the a-th depth, wherein a is a positive integer less than or equal to A, when a=1, the enhanced auxiliary vector of the a-th depth is the long-term average characteristic vector, and when a>1, the enhanced auxiliary vector of the a-th depth is the low-frequency enhanced vector of the song at the a-th depth; Based on the song low-frequency enhancement vectors of the A depths and the low-frequency time-domain depth vectors of the A depths, a song low-frequency vector is determined.

4. The method for repairing the sound quality of a music library according to claim 3, wherein: The step of performing semantic enhancement processing on the a-th depth among the A depths based on the low-frequency time-domain depth vector of the a-th depth and the enhanced auxiliary vector of the a-th depth to form the low-frequency enhanced vector of the song at the a-th depth includes: Using the bth semantic enhancement unit among B semantic enhancement units, the input vector of the bth semantic enhancement unit is processed to form a semantic enhancement output vector of the bth semantic enhancement unit, wherein b is a positive integer less than or equal to B, when b=1, the input vector of the bth semantic enhancement unit includes the low-frequency time domain depth vector of the ath depth and the enhancement auxiliary vector of the ath depth, and when b>1, the input vector of the bth semantic enhancement unit includes the low-frequency time domain depth vector of the ath depth and the semantic enhancement output vector of the b-1th semantic enhancement unit; The semantic enhancement output vector of the Bth semantic enhancement unit is used as the low-frequency enhancement vector of the song at the ath depth.

5. The method for repairing the sound quality of a music library according to claim 4, characterized in that: The step of using the bth semantic enhancement unit among the B semantic enhancement units to process the input vector of the bth semantic enhancement unit to form the semantic enhancement output vector of the bth semantic enhancement unit includes: Using the bth semantic enhancement unit among the B semantic enhancement units, performing two different mappings on the low-frequency time-domain depth vector of the ath depth to form a first time-domain depth vector and a second time-domain depth vector; performing offset removal on the semantic enhancement output vector included in the input vector of the b-th semantic enhancement unit to form a semantic enhancement offset vector, wherein the offset removal comprises performing difference calculation on each vector parameter in the semantic enhancement output vector and a mean value of each vector parameter of the semantic enhancement output vector; Performing a bitwise multiplication operation on the first temporal depth vector and the semantic enhancement offset vector to form a third temporal depth vector; performing quotient calculations on the standard deviations of the vector parameters in the third time-domain depth vector and the vector parameters of the semantic enhancement output vector to form a fourth time-domain depth vector; A bitwise addition operation is performed on the fourth time-domain depth vector and the second time-domain depth vector to form a semantic enhancement output vector of the b-th semantic enhancement unit.

6. The method for repairing the sound quality of a music library according to claim 1, characterized in that: The step of performing frequency domain semantic encoding on the high-frequency frequency domain data of the song to form a high-frequency vector of the song includes: Performing convolution processing on the high-frequency frequency domain data of the song to form a high-frequency frequency domain convolution vector; Performing two different pooling operations on the high-frequency frequency domain convolution vector to form a first frequency domain pooling vector and a second frequency domain pooling vector, wherein the vector size of the first frequency domain pooling vector is the same as the vector size of the second frequency domain pooling vector; Mapping the first frequency-domain pooling vector to form a gating parameter distribution, wherein a vector size of the gating parameter distribution is the same as a vector size of the first frequency-domain pooling vector, and each vector parameter in the gating parameter distribution is greater than or equal to 0 and less than or equal to 1; A bitwise multiplication operation is performed on the gating parameter distribution and the second frequency domain pooling vector to form a high-frequency vector of the song.

7. The method for repairing the sound quality of a music library according to any one of claims 1 to 6, characterized in that: The step of fusing the song low-frequency vector and the song high-frequency vector into the song global vector to form a song fusion vector includes: When the song global vector is a result of semantic encoding in the time domain, the song low-frequency vector is used as the first vector to be fused, and the song high-frequency vector is used as the second vector to be fused; When the song global vector is a result of semantic encoding in the frequency domain, the song low-frequency vector is used as the second vector to be fused, and the song high-frequency vector is used as the first vector to be fused; The first vector to be fused is fused into the song global vector based on a gating mechanism, and the second vector to be fused is fused into the song global vector based on an attention mechanism to form a song fusion vector.

8. The method for repairing the sound quality of a music library according to claim 7, characterized in that: The step of fusing the first vector to be fused into the global song vector based on a gating mechanism, and fusing the second vector to be fused into the global song vector based on an attention mechanism to form a song fusion vector includes: Mapping the first vector to be fused to form a gate mapping parameter, and performing a bitwise multiplication operation on the gate mapping parameter and the song global vector to form a first fused vector; Connecting the song global vector to the first fusion vector to form a first connection vector; determining an attention parameter distribution between the second to-be-fused vector and the first connection vector, and performing a weighted sum operation on the first connection vector based on the attention parameter distribution to form a second fused vector; Connecting the first connection vector to the second fusion vector to form a second connection vector; An iterative fusion process is performed based on the second connection vector to form a song fusion vector, wherein one fusion process in the iterative fusion process includes fusing the first vector to be fused into the second connection vector based on a gating mechanism, and fusing the second vector to be fused into the second connection vector based on an attention mechanism to form a new second connection vector, and the new second connection vector formed for the last time is used as the song fusion vector.

9. A device for repairing the sound quality of a music library, characterized in that: include: The recorded song conversion module is used to perform high- and low-frequency separation processing on the recorded songs to be repaired in the target music library in the time domain to form low-frequency time domain data and high-frequency time domain data of the songs, and to perform frequency domain conversion on the high-frequency time domain data of the songs to form high-frequency frequency domain data of the songs; A first encoding module is used to perform semantic encoding of the recorded song to be repaired in the time domain or perform semantic encoding of the frequency domain conversion result of the recorded song to be repaired in the frequency domain to form a global song vector; A second encoding module is used to perform time-domain semantic encoding on the low-frequency time-domain data of the song to form a low-frequency vector of the song; A third encoding module is used to perform frequency domain semantic encoding on the high-frequency frequency domain data of the song to form a high-frequency vector of the song; A vector fusion module, configured to fuse the low-frequency vector of the song and the high-frequency vector of the song into the global vector of the song to form a song fusion vector; The semantic decoding module is used to perform semantic decoding on the song fusion vector to form a target repaired recorded song.

10. An electronic device, characterized in that: include: memory for storing computer programs; A processor connected to the memory is used to execute the computer program stored in the memory to implement the music library sound quality repair method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Encoding method and device for audio signals

    CN103035248A

  • High-frequency optimization method and device for audios and medium

    CN112562703A

  • Audio restoration method and device, storage medium and electronic equipment

    CN117594059A

  • Audio processing method, electronic equipment and storage medium

    CN118248157A

  • Human voice quality optimization method and device, equipment and medium

    CN120375838A

Cited By

  • Risk event identification method and system based on monitoring video

    CN121214354A