A method, device and medium for generating low-frequency songs

CN117912429BActive Publication Date: 2026-08-14TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

因此这类人群在收听音乐时,通常只能听到音乐中的低频声音分量,音乐收听效果较差

Benefits of technology

[0050]可见,本申请通过按照预设配器轨种类将原始歌曲分离为歌声轨和若干个配器轨,并分别对每一所述配器轨进行能量分析,以根据能量分析结果确定出所述原始歌曲中存在的目标配器轨;确定所述目标配器轨的能量信号在歌曲总时段中出现的目标时段,并在所述目标时段上计算所述目标配器轨在预设低频段的能量集中度;若基于所述能量集中度确定所述目标配器轨的能量信号未集中在所述预设低频段,则对所述目标配器轨进行降调处理,得到处理后配器轨;获取对所述歌声轨进行垫音处理后得到的处理后歌声轨,并对所述处理后歌声轨和所述处理后配器轨进行合成,以生成所述原始歌曲对应的低频歌曲。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117912429B_ABST
    Figure CN117912429B_ABST
Patent Text Reader

Abstract

This application discloses a method, device, and medium for generating low-frequency songs, relating to the field of audio processing technology. The method includes: separating an original song into a vocal track and several instrumentation tracks, and performing energy analysis on each instrumentation track to determine the target instrumentation track present in the original song; determining the target time period in which the energy signal of the target instrumentation track appears within the total song duration, and calculating the energy concentration of the target instrumentation track in a preset low-frequency band during the target time period; if, based on the energy concentration, the energy signal of the target instrumentation track is not concentrated in the preset low-frequency band, then performing pitch reduction processing on the target instrumentation track to obtain a processed instrumentation track; obtaining the processed vocal track after applying backing tone processing to the vocal track, and synthesizing the processed vocal track and the processed instrumentation track to generate a low-frequency song. This application, by performing energy analysis and pitch reduction processing on the instrumentation tracks, and then synthesizing them with the vocal track after backing tone processing, can generate low-frequency songs specifically for people with high-frequency hearing loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and in particular to a method, device and medium for generating low-frequency songs. Background Technology

[0002] High-frequency hearing loss is one of the most common types of hearing loss. People with high-frequency hearing loss are those who experience symptoms of high-frequency hearing loss. They are typically not sensitive enough to high-frequency sound components and have difficulty hearing higher-pitched sounds. Their ability to perceive mid-to-high frequencies, especially audio frequencies above 4kHz, declines sharply. Therefore, when listening to music, these individuals usually only hear the low-frequency sound components, resulting in a poor listening experience.

[0003] Existing technologies address these issues by manually adjusting the frequency, but this approach incurs extremely high labor and time costs and is inefficient.

[0004] In summary, how to automatically generate corresponding low-frequency songs for people with high-frequency hearing loss, so that they can also perceive the dynamics and color of the songs to the greatest extent, is a problem that needs to be solved. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a method, device, and medium for generating low-frequency songs, which can automatically generate corresponding low-frequency songs for people with high-frequency hearing loss, thereby enabling them to perceive the dynamics and color of songs to the greatest extent possible. The specific solution is as follows:

[0006] In a first aspect, this application discloses a method for generating low-frequency songs, including:

[0007] The original song is separated into vocal tracks and several instrumentation tracks according to the preset instrumentation track types, and energy analysis is performed on each of the instrumentation tracks to determine the target instrumentation tracks in the original song based on the energy analysis results.

[0008] Determine the target time period in which the energy signal of the target orchestration track appears in the total duration of the song, and calculate the energy concentration of the target orchestration track in a preset low-frequency band during the target time period;

[0009] If, based on the energy concentration, it is determined that the energy signal of the target instrument track is not concentrated in the preset low-frequency band, then the target instrument track is down-modulated to obtain the processed instrument track.

[0010] The processed vocal track is obtained by adding backing to the vocal track, and the processed vocal track and the processed instrumentation track are synthesized to generate a low-frequency song corresponding to the original song.

[0011] Optionally, the process of performing energy analysis on each of the instrumentation tracks to determine the target instrumentation tracks present in the original song based on the energy analysis results further includes:

[0012] Based on the preset frame shift and preset frame length, the audio corresponding to the current tuner track to be analyzed is divided into several frame signals, and a short-time Fourier transform is performed on each frame signal to obtain the power spectrum of each frame signal.

[0013] The corresponding power value is calculated based on the power spectrum of each frame signal, and the total power value of the current tuner track to be analyzed is obtained based on the power value of each frame signal.

[0014] The total loudness value of the instrumentation track is determined based on the total power value. If the total loudness value is greater than a preset loudness threshold, the current instrumentation track to be analyzed is determined to be the target instrumentation track in the original song.

[0015] Optionally, determining the target time period in which the energy signal of the target orchestration track appears within the total duration of the song includes:

[0016] The power spectrum of each frame signal of the target tuner track is smoothed using the target smoothing kernel function to obtain the smoothed power value of each frame signal.

[0017] If the smoothed power value of any frame signal in each frame signal is greater than the preset loudness threshold, then the time period corresponding to any frame signal is determined as the target time period in which the energy signal of the target orchestration track appears in the total time period of the song.

[0018] If the smoothed power value of any frame signal is not greater than the preset loudness threshold, then it is determined that there is no energy signal of the target orchestration track in the time period corresponding to any frame signal.

[0019] Optionally, the low-frequency song generation method further includes:

[0020] Determine the length of the time window used to smooth the power spectrum, and determine the number of signal frames for one smoothing process based on the length of the time window and the preset frame shift;

[0021] Increment the signal frame number by one to obtain the updated signal frame number, and construct an initial smoothing kernel function based on the updated signal frame number;

[0022] The initial smoothing kernel function is normalized to obtain the target smoothing kernel function.

[0023] Optionally, calculating the energy concentration of the target transmitter track in a preset low-frequency band during the target time period includes:

[0024] For each frame of signal corresponding to the target time period, obtain the first signal power and value of the corresponding frequency point within the preset low frequency band, and obtain the second signal power and value of the total frequency point;

[0025] Calculate the ratio of the sum of the first signal power to the sum of the second signal power, and smooth the ratio using the target smoothing kernel function to obtain the smoothed ratio;

[0026] The energy concentration of the target tuner track in the preset low-frequency band is obtained based on the smoothed ratio of all frame signals in the target time period.

[0027] Optionally, the low-frequency song generation method further includes:

[0028] If the energy concentration is greater than the preset low-frequency proportion threshold, it is determined that the energy signal of the target instrument rail is concentrated in the preset low-frequency band.

[0029] If the energy concentration is not greater than the preset low-frequency proportion threshold, it is determined that the energy signal of the target instrument track is not concentrated in the preset low-frequency band.

[0030] Optionally, after calculating the energy concentration of the target transmitter track in the preset low-frequency band during the target time period, the method further includes:

[0031] If the energy signal of the target orchestration track is determined to be concentrated in the preset low-frequency band based on the energy concentration, then the energy signal of the target orchestration track is subjected to full-pass filtering to obtain the processed orchestration track.

[0032] Optionally, the step of down-tuning the target instrument track to obtain the processed instrument track includes:

[0033] A preset frequency divider is used to classify the audio in the target tuner track located in the preset low-frequency band as low-frequency audio, and the remaining audio besides the low-frequency audio as high-frequency audio; wherein, the preset frequency divider is constructed based on an LR filter;

[0034] The high-frequency audio is down-pitched based on a preset octave down-pitch strategy to obtain down-pitch audio, and the low-frequency audio is subjected to full-pass filtering to obtain filtered audio.

[0035] The down-pitched audio and the filtered audio are mixed to obtain the processed instrumentation track.

[0036] Optionally, obtaining the processed vocal track after applying backing vocals to the vocal track includes:

[0037] The fundamental frequency information of the vocal track is detected using a fundamental frequency extraction tool and a CREPE model, and the fundamental frequency concentration of the fundamental frequency information within a preset frequency band is calculated.

[0038] If the fundamental frequency concentration is greater than a preset threshold, then the vocal track is subjected to full-pass filtering to obtain the processed vocal track.

[0039] If the fundamental frequency concentration is not greater than the preset threshold, the vocal track is down-tuned based on the octave down-tuning strategy to obtain the down-tuned vocals, and the down-tuned vocals and the first preset weighting parameter are used to perform backing sound processing on the vocal track to obtain the processed vocal track.

[0040] Optionally, obtaining the processed vocal track after applying backing vocals to the vocal track includes:

[0041] The fundamental frequency information of the vocal track is detected using a fundamental frequency extraction tool and a CREPE model, and the target fundamental frequency that is not located in the preset frequency band is determined.

[0042] The target fundamental frequency is down-tuned based on a preset octave down-tuning strategy to obtain the down-tuned fundamental frequency. The down-tuned audio and the second preset weighting parameter are then used to perform backing sound processing on the time period of the target fundamental frequency to obtain the processed vocal track.

[0043] Optionally, after synthesizing the processed vocal track and the processed instrumentation track to generate a low-frequency song corresponding to the original song, the method further includes:

[0044] The degree of hearing loss of the target user is determined, and the loudness adjustment parameters corresponding to the degree of hearing loss are determined based on the generalized dynamic range compression technology;

[0045] The loudness of the low-frequency song is adjusted using the loudness adjustment parameters, and the adjusted low-frequency song is then played.

[0046] Secondly, this application discloses an electronic device, comprising:

[0047] Memory, used to store computer programs;

[0048] A processor is configured to execute the computer program to implement the steps of the aforementioned disclosed low-frequency song generation method.

[0049] Thirdly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the aforementioned disclosed low-frequency song generation method.

[0050] As can be seen, this application separates the original song into a vocal track and several instrumentation tracks according to preset instrumentation track types, and performs energy analysis on each instrumentation track to determine the target instrumentation track in the original song based on the energy analysis results; determines the target time period in which the energy signal of the target instrumentation track appears in the total time period of the song, and calculates the energy concentration of the target instrumentation track in a preset low-frequency band during the target time period; if it is determined based on the energy concentration that the energy signal of the target instrumentation track is not concentrated in the preset low-frequency band, then the target instrumentation track is down-tuned to obtain a processed instrumentation track; obtains the processed vocal track obtained after adding backing to the vocal track, and synthesizes the processed vocal track and the processed instrumentation track to generate a low-frequency song corresponding to the original song.

[0051] Therefore, this application first separates the original song into vocal tracks and several instrumentation tracks according to preset instrumentation track types, and performs energy analysis on each instrumentation track to determine the target instrumentation tracks present in the original song based on the energy analysis results. In other words, this embodiment first determines which instrumentation tracks exist in the original song. Further, this application determines which time periods in the total song duration the energy signal of the target instrumentation track appears in, thus obtaining the target time period. The energy concentration of the target instrumentation track in a preset low-frequency band is calculated within the target time period to determine whether the energy signal of the target instrumentation track is concentrated in the preset low-frequency band. If the energy signal of the target instrumentation track is not concentrated in the preset low-frequency band, it indicates that there are high-frequency components on the target instrumentation track that are difficult for hearing-impaired individuals to perceive. Therefore, the target instrumentation track needs to be down-tuned to obtain a processed instrumentation track. Furthermore, this application also needs to perform backing sound processing on the separated vocal tracks to obtain a processed vocal track. The processed vocal track and the processed instrumentation track are then synthesized to generate a low-frequency song corresponding to the original song. In this way, by performing energy analysis and pitch reduction processing on the separated instrument tracks one after another, and then synthesizing them with the vocal track after the backing tone processing, this application can automatically generate low-frequency songs for people with high-frequency hearing loss. This allows people with high-frequency hearing loss to perceive the dynamics and color of the songs to the greatest extent. Moreover, this application does not require human intervention, saving a lot of manpower and time costs. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0053] Figure 1A schematic diagram of the system framework applicable to the low-frequency song generation scheme disclosed in this application;

[0054] Figure 2 This application discloses a flowchart of a method for generating low-frequency songs.

[0055] Figure 3 This application discloses a specific method for generating low-frequency songs.

[0056] Figure 4 This is a schematic diagram of a process for generating high-frequency hearing-loss songs disclosed in this application;

[0057] Figure 5 This application discloses a flowchart of a track down-tuning process;

[0058] Figure 6 This is a schematic diagram of a DRC curve for adjusting loudness disclosed in this application;

[0059] Figure 7 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0060] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0061] People with high-frequency hearing loss typically lack sensitivity to high-frequency sound components, making it difficult for them to hear higher-pitched sounds. Their ability to perceive mid-to-high frequencies, especially audio above 4kHz, declines sharply. Therefore, when listening to music, these individuals usually only hear the low-frequency components, resulting in a poor listening experience. Current technology addresses this by manually adjusting frequencies, but this method incurs extremely high labor and time costs and is inefficient.

[0062] Therefore, this application discloses a method, device and medium for generating low-frequency songs, which can automatically generate corresponding low-frequency songs for people with high-frequency hearing loss, so that people with high-frequency hearing loss can also perceive the dynamics and colors of the songs to the greatest extent.

[0063] The system framework used in the low-frequency song generation scheme of this application can be found in [reference needed]. Figure 1 As shown, it may specifically include an original song database 11, a low-frequency song generation platform 12, and may also include a song playback platform 13.

[0064] The original song database 11 stores the original songs in the music library and sends them to the low-frequency song generation platform 12. After receiving the original songs, the low-frequency song generation platform 12 processes them based on the low-frequency song generation method provided in this application to generate corresponding low-frequency songs. Additionally, it can upload the low-frequency songs to the song playback platform 13 for playback upon user-triggered or automatic triggering. After receiving the low-frequency songs sent by the low-frequency song generation platform 12, the song playback platform 13 can play the low-frequency songs through the user's local audio player in response to song listening requests initiated by users with high-frequency hearing loss.

[0065] The original song database 11 and the low-frequency song generation platform 12 can be located on the same server, and the low-frequency song generation platform 12 can also be located on the same server as the song playback platform 13, or the original song database 11, the low-frequency song generation platform 12, and the song playback platform 13 can be located on different servers. This embodiment does not limit the positional relationship between the original song database 11, the low-frequency song generation platform 12, and the song playback platform 13.

[0066] See Figure 2 As shown in the figure, this application discloses a method for generating low-frequency songs, the method comprising:

[0067] Step S11: Separate the original song into vocal tracks and several instrumentation tracks according to the preset instrumentation track types, and perform energy analysis on each of the instrumentation tracks to determine the target instrumentation tracks in the original song based on the energy analysis results.

[0068] In this embodiment, the original song is first separated into a vocal track and several instrumentation tracks according to preset instrumentation track types. The specific technology used in this process is the music separation technology in TME Studio, which is based on deep learning technology to separate any song into vocals and various instruments (i.e., musical instruments). In this embodiment, the instrumentation track types specifically include keyboard, guitar, bass, drums, etc., and other instruments are categorized as "other types." The sum of all instrumentation tracks is the accompaniment track. Therefore, in one specific implementation, the above process can also be understood as first using vocal-accompaniment separation technology to separate the original song into a vocal track and an accompaniment track, and then dividing the accompaniment track into several instrumentation tracks according to the instrumentation track types.

[0069] However, not every song contains all of the above instrumentation tracks. Therefore, it is necessary to perform energy analysis on each instrumentation track separately to determine the target instrumentation tracks present in the original song based on the energy analysis results. That is, among the several instrumentation tracks initially separated, some instrumentation tracks may have zero energy, indicating that this instrumentation does not exist in the original song. Therefore, energy analysis can determine which specific types of instruments are present in the original song.

[0070] In a specific implementation, the process of performing energy analysis on each of the aforementioned instrumentation tracks to determine the target instrumentation track in the original song based on the energy analysis results further includes: dividing the audio corresponding to the current instrumentation track to be analyzed into several frame signals based on a preset frame shift and a preset frame length, and performing a short-time Fourier transform on each frame signal to obtain the power spectrum of each frame signal; calculating the corresponding power value based on the power spectrum of each frame signal, and obtaining the total power value of the current instrumentation track to be analyzed based on the power value of each frame signal; determining the total loudness value of the instrumentation track based on the total power value, and if the total loudness value is greater than a preset loudness threshold, then the current instrumentation track to be analyzed is determined to be the target instrumentation track in the original song. In this embodiment, the discrete signal corresponding to the track to be analyzed is defined as x(i), where i = 0, 1, 2, ... represents the sample index. First, based on a preset frame shift and a preset frame length, signal x(i) is divided into several frames. In this embodiment, the preset frame shift is specifically set to 10ms, and the preset frame length is specifically set to 30ms. Therefore, 0ms-30ms is the first frame, 10ms-40ms is the second frame, 20ms-50ms is the third frame, and so on. Further, this embodiment requires performing a short-time Fourier transform on each frame signal to obtain the power spectrum of each frame signal. The main idea of ​​the short-time Fourier transform is to window the signal and then perform a Fourier transform on the windowed signal.

[0071] In this implementation, a Hanning window is added to each frame signal to obtain the segmented windowed frame signal sequence: x w (n,i)=x(L·n+i)·w hann (i), where i represents the i-th sample point, L represents the frame shift (10ms in this case), and n represents the frame position index. The Hanning window is defined as: Here, N represents the window length, corresponding to a duration of 30ms. Then, a Fourier transform is performed on the windowed frame signal sequence to obtain: Where N1 represents the number of points in the Fourier transform. Based on this, the power, i.e., the squared amplitude, is calculated: P(n,k)=||X(n,k)|| 2Let n = 0, 1, ..., N²-1, where N² represents the total number of frames after the current signal undergoes Fourier transform. P(n,k) represents the power spectrum of the k-th frequency point in the n-th frame.

[0072] After obtaining the power spectrum of each frame, the corresponding power value is calculated based on the power spectrum: N2 represents the total number of frequency points of the nth frame signal, and the total power value of the current tuner track to be analyzed is obtained based on the power value of each frame signal. Finally, the total loudness value of the orchestration track is determined based on the total power value, that is, the power is quantified in decibels (dB). In addition, this embodiment also includes a preset loudness threshold. If the total loudness value loudness greater than the preset loudness threshold If the current instrumentation track to be analyzed is determined to be the target instrumentation track that exists in the original song, then it means that the instrumentation track does not exist in the original song.

[0073] Step S12: Determine the target time period in which the energy signal of the target instrumentation track appears in the total time period of the song, and calculate the energy concentration of the target instrumentation track in the preset low frequency band during the target time period.

[0074] In this embodiment, the energy signals of the target instrumentation track appearing in which time periods within the total song duration are determined, thus obtaining the target time period. The energy concentration of the target instrumentation track in a preset low-frequency band is then calculated within the target time period. Specifically, the preset low-frequency band is below 4kHz; that is, for the instrumentation tracks present in the original song, it is analyzed whether their main energy is concentrated below 4kHz.

[0075] In a specific implementation, determining the target time period in which the energy signal of the target orchestration track appears within the total song duration includes: smoothing the power spectrum of each frame signal of the target orchestration track using a target smoothing kernel function to obtain the smoothed power value corresponding to each frame signal; if the smoothed power value of any frame signal is greater than the preset loudness threshold, then the time period corresponding to that frame signal is determined as the target time period in which the energy signal of the target orchestration track appears within the total song duration; if the smoothed power value of any frame signal is not greater than the preset loudness threshold, then it is determined that the energy signal of the target orchestration track does not exist in the time period corresponding to that frame signal. In other words, this embodiment smooths the power spectrum to determine the effective time period for the appearance of the target orchestration track. Specifically, it uses a target smoothing kernel function to smooth the power spectrum of each frame signal of the target orchestration track to obtain the smoothed power value corresponding to each frame signal. The effective time period for the appearance of the target orchestration track is determined by judging the relative magnitude of the smoothed power value and the preset loudness threshold. Specifically, if the smoothed power value of any frame signal is greater than the preset loudness threshold, it means that the energy signal of the target orchestration track exists in the time period corresponding to that frame signal, and this time period is the target time period in which the energy signal of the target orchestration track appears in the total time period of the song; conversely, if the smoothed power value of any frame signal is not greater than the preset loudness threshold, it means that the energy signal of the target orchestration track does not exist in the time period corresponding to that frame signal.

[0076] The construction process of the target smoothing kernel function is as follows: determine the length of the time window used to smooth the power spectrum, and determine the number of signal frames for one smoothing process based on the time window length and the preset frame shift; increment the number of signal frames to obtain the updated number of signal frames, and construct an initial smoothing kernel function based on the updated number of signal frames; normalize the initial smoothing kernel function to obtain the target smoothing kernel function.

[0077] In this embodiment, the time window length for smoothing the power spectrum is set to 0.4s. This means that a 0.4s frame signal is used for smoothing via convolution. The number of signal frames for one smoothing operation is then determined based on the time window length and a preset frame shift. Since the frame shift is 0.01s, the corresponding number of signal frames is calculated as: M = 0.4 / 0.01 = 20; that is, the number of frames M = 40. Then... in This indicates that rounding up yields an initial smoothing kernel function of length M+1 points:

[0078]

[0079] Here, B is half the length of the time window. The reason for introducing the value of B is that this embodiment uses "spline" smoothing. The smoothing window is a triangular window, and B is equivalent to the horizontal coordinate position of the fixed point of this triangular window. However, at this time, M = 40, which is an even number, while the triangular window has only one vertex, which means that its length must be an odd number. If B = M / 2 = 20, then there need to be 20 points on the left and 20 points on the right. In this way, the window length will be updated to 41 (i.e., M+1) points: that is, the horizontal coordinate index of the left rising curve is 0 to 19 (a total of 20 points), the fixed point index of the triangular window is 20 (1 point), and the index of the right falling curve is 21 to 40 (a total of 20 points), which adds up to 41 (M+1) points.

[0080] Furthermore, the initial smoothing kernel function needs to be normalized to obtain the target smoothing kernel function:

[0081]

[0082] Therefore, the power spectrum can be smoothed using the target smoothing kernel function, and the smoothed power value is:

[0083]

[0084] By judgment and The magnitude of the signal can determine the target time period in the total duration of the song where the energy signal of the target orchestration track appears.

[0085] In a specific implementation, the above-mentioned calculation of the energy concentration of the target tuner track in the preset low-frequency band during the target time period includes: obtaining the first signal power sum value of the corresponding frequency points within the preset low-frequency band and obtaining the second signal power sum value of the total frequency points for each frame signal corresponding to the target time period; calculating the ratio of the first signal power sum value to the second signal power sum value, and smoothing the ratio using the target smoothing kernel function to obtain a smoothed ratio value; and obtaining the energy concentration of the target tuner track in the preset low-frequency band based on the smoothed ratio value corresponding to all frame signals during the target time period.

[0086] In this embodiment, the preset low-frequency band is specifically below 4kHz. That is, after determining the target time period where the orchestration track energy signal exists, the relative proportion of energy concentrated in the 4kHz frequency band in each frame signal is further calculated. Specifically, for each frame signal corresponding to the target time period, the sum of signal power at frequency points below 4kHz is obtained to obtain the first signal power sum value. And obtain the sum of signal power at all frequency points to get the second signal power sum value. Where K represents the highest frequency band, K 4kThe frequency band is 4kHz. Then, the ratio of the sum of the first signal power to the sum of the second signal power is calculated: This ratio can be used to measure the degree of concentration of low frequencies.

[0087] To reduce anomalies in individual frame processing results, the trend of low-frequency energy concentration during the effective time period is tracked, and r is also adjusted accordingly. 4kHz (n) Perform frame-level smoothing to obtain the smoothed ratio, specifically: Finally, based on the smoothed ratio of all frame signals in the target time period, the energy concentration of the target instrumentation track in the preset low-frequency band is obtained, which is also the low-frequency concentration of the target instrumentation track in the entire song: Ratio 4kHz This refers to audio frames marked as "voiced". The mean, because It can be viewed as a function with the independent variable "frame index" n. Some frames are silent, while others have sound. In this embodiment, the subscript "voiced" is used to mark the frames with sound.

[0088] Taking the song "Sunbathing with Grandma in Winter" as an example, after the above steps, the loudness of each instrument track (i.e., bass, drums, guitar, keyboard, and other types) is obtained. and low-frequency concentration ratio 4kHz The distribution is shown in Table 1:

[0089] Table 1

[0090] Bass drum Guitar keyboard other Loudness -62 -55 -36 -74 -38 Low frequency concentration 0.998 0.452 0.996 0.998 0.998

[0091] Step S13: If it is determined based on the energy concentration that the energy signal of the target instrument track is not concentrated in the preset low frequency band, then the target instrument track is down-modulated to obtain the processed instrument track.

[0092] In one specific implementation, if the energy signal of the target instrumentation track is not concentrated in the preset low frequency band based on the energy concentration, it indicates that there are high-frequency components on the target instrumentation track that are difficult for people with high hearing loss to perceive. Therefore, it is necessary to down-modulate the target instrumentation track to obtain the processed instrumentation track.

[0093] In another specific implementation, if the energy signal of the target orchestration track is determined to be concentrated in a preset low-frequency band based on the energy concentration, it means that almost all the energy of the orchestration track is concentrated in the low frequency. Therefore, the energy signal of the target orchestration track is subjected to full-pass filtering to obtain the processed orchestration track. In other words, it can also be understood that no processing is required on the target orchestration track.

[0094] Specifically, this embodiment determines whether the energy signal of the target instrument track is concentrated in the preset low-frequency band by comparing the energy concentration with a preset low-frequency proportion threshold. In a specific implementation, the preset low-frequency proportion threshold can be set to 0.9. Therefore, if the energy concentration is greater than the preset low-frequency proportion threshold of 0.9, it is determined that the energy signal of the target instrument track is concentrated in the preset low-frequency band; if the energy concentration is not greater than the preset low-frequency proportion threshold of 0.9, it is determined that the energy signal of the target instrument track is not concentrated in the preset low-frequency band.

[0095] Step S14: Obtain the processed vocal track after adding backing to the vocal track, and synthesize the processed vocal track and the processed instrumentation track to generate a low-frequency song corresponding to the original song.

[0096] In this embodiment, the separated vocal track needs to be padded with background music to obtain a processed vocal track. The processed vocal track and the processed instrumentation track are then synthesized to generate a low-frequency song corresponding to the original song. In this way, by sequentially performing energy analysis and pitch reduction processing on each separated instrumentation track, and then synthesizing it with the padded vocal track, this application can automatically generate low-frequency songs for people with high-frequency hearing loss. This allows people with high-frequency hearing loss to perceive the dynamics and color of the song to the greatest extent possible. Furthermore, this application requires no manual intervention, saving significant manpower and time costs.

[0097] As can be seen, this application separates the original song into a vocal track and several instrumentation tracks according to preset instrumentation track types, and performs energy analysis on each instrumentation track to determine the target instrumentation track in the original song based on the energy analysis results; determines the target time period in which the energy signal of the target instrumentation track appears in the total time period of the song, and calculates the energy concentration of the target instrumentation track in a preset low-frequency band during the target time period; if it is determined based on the energy concentration that the energy signal of the target instrumentation track is not concentrated in the preset low-frequency band, then the target instrumentation track is down-tuned to obtain a processed instrumentation track; obtains the processed vocal track obtained after adding backing to the vocal track, and synthesizes the processed vocal track and the processed instrumentation track to generate a low-frequency song corresponding to the original song.

[0098] Therefore, this application first separates the original song into vocal tracks and several instrumentation tracks according to preset instrumentation track types, and performs energy analysis on each instrumentation track to determine the target instrumentation tracks present in the original song based on the energy analysis results. In other words, this embodiment first determines which instrumentation tracks exist in the original song. Further, this application determines which time periods in the total song duration the energy signal of the target instrumentation track appears in, thus obtaining the target time period. The energy concentration of the target instrumentation track in a preset low-frequency band is calculated within the target time period to determine whether the energy signal of the target instrumentation track is concentrated in the preset low-frequency band. If the energy signal of the target instrumentation track is not concentrated in the preset low-frequency band, it indicates that there are high-frequency components on the target instrumentation track that are difficult for hearing-impaired individuals to perceive. Therefore, the target instrumentation track needs to be down-tuned to obtain a processed instrumentation track. Furthermore, this application also needs to perform backing sound processing on the separated vocal tracks to obtain a processed vocal track. The processed vocal track and the processed instrumentation track are then synthesized to generate a low-frequency song corresponding to the original song. In this way, by performing energy analysis and pitch reduction processing on the separated instrument tracks one after another, and then synthesizing them with the vocal track after the backing tone processing, this application can automatically generate low-frequency songs for people with high-frequency hearing loss. This allows people with high-frequency hearing loss to perceive the dynamics and color of the songs to the greatest extent. Moreover, this application does not require human intervention, saving a lot of manpower and time costs.

[0099] See Figure 3 and Figure 4 As shown in the illustration, this application discloses a specific method for generating low-frequency songs. Compared to the previous embodiment, this embodiment further explains and optimizes the technical solution. Specifically, it includes:

[0100] Step S21: Separate the original song into vocal tracks and several instrumentation tracks according to the preset instrumentation track types, and perform energy analysis on each instrumentation track to determine the target instrumentation track in the original song based on the energy analysis results.

[0101] Step S22: Determine the target time period in which the energy signal of the target instrumentation track appears in the total time period of the song, and calculate the energy concentration of the target instrumentation track in the preset low frequency band during the target time period.

[0102] Step S23: If it is determined based on the energy concentration that the energy signal of the target instrumentation track is not concentrated in the preset low frequency band, then the audio in the target instrumentation track located in the preset low frequency band is used as low frequency audio and the remaining audio other than the low frequency audio is used as high frequency audio; wherein, the preset frequency divider is constructed based on the LR filter.

[0103] In this embodiment, when the energy signal of the target instrument track is not concentrated in the preset low-frequency band, it is necessary to down-modulate the target instrument track. For example... Figure 5 As shown, a preset crossover is first used to classify the audio in the target accompaniment track located in a preset low-frequency band as low-frequency audio, and the remaining audio is classified as high-frequency audio. The preset crossover is based on an LR (Linkwitz-Riley) filter. This patent sets the cutoff frequency of the preset crossover to 4kHz. The preset crossover is used to perform a high-pass filter on the signal to obtain the high-frequency audio, and a low-pass filter to obtain the low-frequency audio. The LR filter is specifically composed of two Butterworth filters cascaded together. The Butterworth filter is specifically represented as: [b,a]=butter(n,wn); where n=4, and the cutoff frequency is f. c =4kHz, the corresponding normalized cutoff frequency wn = fc / (fs / 2), where fs represents the sampling rate.

[0104] Step S24: Based on a preset octave down-pitch strategy, the high-frequency audio is down-pitch processed to obtain down-pitch audio, and the low-frequency audio is subjected to full-pass filtering to obtain filtered audio. The down-pitch audio and the filtered audio are mixed to obtain the processed orchestration track.

[0105] In this embodiment, to maintain the original texture structure of the accompaniment, a preset octave down-pitch strategy is used to down-pitch the target instrumentation track. Simultaneously, to maximize the consistency of low-frequency audio perception, only the high-frequency components of the target instrumentation track are down-pitched. Figure 5 As shown, a full-pass filter is applied to the low-frequency components to obtain the filtered audio. Full-pass filtering can also be understood as not processing the low-frequency audio at all. Finally, the down-pitch audio and the filtered audio are mixed to obtain the processed orchestration track. Down-pitching of the song can be achieved using the open-source MATLAB tool TSM (Time Scale Modification), or by referring to the Rubberband system solution.

[0106] Step S25: Obtain the processed vocal track after adding backing to the vocal track, and synthesize the processed vocal track and the processed instrumentation track to generate a low-frequency song corresponding to the original song.

[0107] In one specific implementation, obtaining the processed vocal track after applying backing vocals to the vocal track includes: detecting the fundamental frequency information of the vocal track using a fundamental frequency extraction tool and a CREPE model, and calculating the fundamental frequency concentration of the fundamental frequency information within a preset frequency band; if the fundamental frequency concentration is greater than a preset threshold, then performing full-pass filtering on the vocal track to obtain the processed vocal track; if the fundamental frequency concentration is not greater than the preset threshold, then performing pitch reduction processing on the vocal track based on an octave pitch reduction strategy to obtain a pitch-reduced vocal, and then applying backing vocals to the vocal track using the pitch-reduced vocal and a first preset weight parameter to obtain the processed vocal track. In this embodiment, the efficient fundamental frequency extraction tool pyin is specifically used in conjunction with the CREPE model to detect the fundamental frequency information of the vocal track and calculate the fundamental frequency concentration of the fundamental frequency information within a preset frequency band of 1kHz. It should be noted that users with severe high-frequency hearing loss cannot hear vocals with a fundamental frequency higher than 1kHz. Furthermore, if the fundamental frequency concentration is greater than a preset threshold, the vocal track undergoes full-pass filtering to obtain a processed vocal track. If the fundamental frequency concentration is not greater than the preset threshold, the vocal track is down-pitched based on an octave reduction strategy to obtain a down-pitched vocal. The down-pitched vocal and a first preset weight parameter are then used to add backing vocals to the vocal track, resulting in a processed vocal track. In this implementation, the specific values ​​of the preset threshold and the first preset weight parameter are not limited; for example, the preset threshold can be set to 0.5, and the first preset weight parameter can be set to 0.4. That is, by calculating whether the proportion of fundamental frequencies below 1kHz in the entire song is greater than 0.5, if it is greater than 0.5, no down-pitch backing vocal is added; if it is less than 0.5, the entire vocal track is down-pitched by an octave, and the down-pitched vocal is then superimposed onto the original vocal track using the first preset weight parameter of 0.4 as backing vocals. The octave reduction can also be achieved using the TSM toolkit, with the relevant parameters for protecting the tone enabled. Using backing tracks throughout the song can maximize the harmony of its timbre and scale.

[0108] In another specific embodiment, obtaining the processed vocal track after applying backing vocals to the vocal track includes: using a fundamental frequency extraction tool and a CREPE model to detect the fundamental frequency information of the vocal track, and determining the target fundamental frequency that is not located within a preset frequency band; performing pitch reduction processing on the target fundamental frequency based on a preset octave down-pitch strategy to obtain a down-pitch fundamental frequency, and using the down-pitch audio and a second preset weight parameter to apply backing vocals to the time period where the target fundamental frequency is located to obtain the processed vocal track. In this embodiment, the efficient fundamental frequency extraction tool pyin is also used in conjunction with the CREPE model to detect the fundamental frequency information of the vocal track, determine the target fundamental frequency that is not located within a preset frequency band of 1kHz, then perform octave down-pitch processing on the target fundamental frequency to obtain a down-pitch fundamental frequency, and finally use the down-pitch audio and the second preset weight parameter to apply backing vocals to the time period where the target fundamental frequency is located to obtain the processed vocal track. In other words, this method only applies padding to the periods below 1kHz where there is no fundamental frequency, while leaving the periods below 1kHz where there is no fundamental frequency unaffected. This approach is more flexible and ensures that hearing-impaired users can hear the fundamental frequency. It is understandable that in certain special cases, only a very few keys in a song may be too high (i.e., no fundamental frequency below 1kHz). Therefore, this embodiment only applies lowering padding to these few keys.

[0109] Finally, the processed vocal track and the processed instrumentation track are combined to generate a low-frequency version of the original song.

[0110] Step S26: Determine the degree of hearing loss of the target user, and determine the loudness adjustment parameter corresponding to the degree of hearing loss based on the generalized dynamic range compression technology. Then, adjust the loudness of the low-frequency song using the loudness adjustment parameter, and play the adjusted low-frequency song.

[0111] In this embodiment, after obtaining the aforementioned low-frequency song, loudness adjustment processing with different parameters is required based on the target user's degree of hearing loss. That is, the user can upload their hearing loss information. This embodiment categorizes hearing loss into mild, moderate, moderately severe, severe, and profound hearing loss. Therefore, based on the different degrees of hearing loss for each user, corresponding loudness adjustment parameters are determined using Wide Dynamic Range Compression (WDRC) technology. These parameters are then used to dynamically adjust the loudness of the low-frequency song, resulting in the adjusted low-frequency song, which is also the high-frequency hearing-impaired song. Finally, the high-frequency hearing-impaired song is played. The DRC (Dynamic Range Control) curve used for loudness adjustment in this embodiment is shown below. Figure 6 As shown.

[0112] For more detailed processing procedures of steps S21 and S22, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.

[0113] As can be seen, in this embodiment, when down-tuning the target instrumentation track, a preset crossover is first used to divide the entire target instrumentation track into low-frequency and high-frequency audio. To ensure the consistency of the perceived low-frequency audio, only the high-frequency part is down-tuned by an octave, resulting in the down-tuned audio. The low-frequency part can be left unprocessed or filtered using a full-pass filter, resulting in the filtered audio. Finally, the two are mixed to obtain the processed instrumentation track. For the vocal track, full backing can be used to maintain the harmony of the song's timbre / scale. Alternatively, backing can be applied only to the sections below 1kHz where there is no fundamental frequency, while leaving the sections below 1kHz where there is no fundamental frequency untreated. This approach is more flexible and ensures that hearing-impaired users can hear the fundamental frequency. Finally, loudness adjustments are made according to the degree of hearing impairment of the target user, and the adjusted low-frequency song is played. In this way, the above scheme can maximize the display of the song's dynamics and color to hearing-impaired individuals.

[0114] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Specifically, it may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the low-frequency song generation method performed by the electronic device disclosed in any of the foregoing embodiments.

[0115] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0116] The processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 21 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0117] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored on it include operating system 221, computer program 222 and data 223, etc., and the storage method can be temporary storage or permanent storage.

[0118] The operating system 221 manages and controls the various hardware devices and computer programs 222 on the electronic device 20 to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. The operating system 221 can be Windows, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of performing the low-frequency song generation method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the electronic device from external devices, as well as data collected by its own input / output interface 25.

[0119] Furthermore, embodiments of this application also disclose a computer-readable storage medium storing a computer program, which, when loaded and executed by a processor, implements the low-frequency song generation method steps disclosed in any of the foregoing embodiments.

[0120] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0121] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0122] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, compact disc read-only memory (CD-ROM), or any other form of storage medium known in the art.

[0123] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0124] The above provides a detailed description of a low-frequency song generation method, device, and storage medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for generating low-frequency songs, characterized in that, include: The original song is separated into vocal tracks and several instrumentation tracks according to the preset instrumentation track types, and energy analysis is performed on each of the instrumentation tracks to determine the target instrumentation tracks in the original song based on the energy analysis results. Determine the target time period in which the energy signal of the target orchestration track appears in the total duration of the song, and calculate the energy concentration of the target orchestration track in a preset low-frequency band during the target time period; If, based on the energy concentration, it is determined that the energy signal of the target instrument track is not concentrated in the preset low-frequency band, then the target instrument track is down-modulated to obtain the processed instrument track. The processed vocal track is obtained by adding backing to the vocal track, and the processed vocal track and the processed instrumentation track are synthesized to generate a low-frequency song corresponding to the original song.

2. The low-frequency song generation method according to claim 1, characterized in that, The process of performing energy analysis on each of the instrumentation tracks to determine the target instrumentation tracks present in the original song based on the energy analysis results also includes: Based on the preset frame shift and preset frame length, the audio corresponding to the current tuner track to be analyzed is divided into several frame signals, and a short-time Fourier transform is performed on each frame signal to obtain the power spectrum of each frame signal. The corresponding power value is calculated based on the power spectrum of each frame signal, and the total power value of the current tuner track to be analyzed is obtained based on the power value of each frame signal. The total loudness value of the instrumentation track is determined based on the total power value. If the total loudness value is greater than a preset loudness threshold, the current instrumentation track to be analyzed is determined to be the target instrumentation track in the original song.

3. The low-frequency song generation method according to claim 2, characterized in that, The determination of the target time period in which the energy signal of the target orchestration track appears within the total duration of the song includes: The power spectrum of each frame signal of the target tuner track is smoothed using the target smoothing kernel function to obtain the smoothed power value of each frame signal. If the smoothed power value of any frame signal in each frame signal is greater than the preset loudness threshold, then the time period corresponding to any frame signal is determined as the target time period in which the energy signal of the target orchestration track appears in the total time period of the song. If the smoothed power value of any frame signal is not greater than the preset loudness threshold, then it is determined that there is no energy signal of the target orchestration track in the time period corresponding to any frame signal.

4. The low-frequency song generation method according to claim 3, characterized in that, Also includes: Determine the length of the time window used to smooth the power spectrum, and determine the number of signal frames for one smoothing process based on the length of the time window and the preset frame shift; Increment the signal frame number by one to obtain the updated signal frame number, and construct an initial smoothing kernel function based on the updated signal frame number; The initial smoothing kernel function is normalized to obtain the target smoothing kernel function.

5. The low-frequency song generation method according to claim 3, characterized in that, The calculation of the energy concentration of the target instrument track in the preset low-frequency band during the target time period includes: For each frame of signal corresponding to the target time period, obtain the first signal power and value of the corresponding frequency point within the preset low frequency band, and obtain the second signal power and value of the total frequency point; Calculate the ratio of the sum of the first signal power to the sum of the second signal power, and smooth the ratio using the target smoothing kernel function to obtain the smoothed ratio; The energy concentration of the target tuner track in the preset low-frequency band is obtained based on the smoothed ratio of all frame signals in the target time period.

6. The low-frequency song generation method according to claim 1, characterized in that, Also includes: If the energy concentration is greater than the preset low-frequency proportion threshold, it is determined that the energy signal of the target instrument rail is concentrated in the preset low-frequency band. If the energy concentration is not greater than the preset low-frequency proportion threshold, it is determined that the energy signal of the target instrument track is not concentrated in the preset low-frequency band.

7. The low-frequency song generation method according to claim 1, characterized in that, After calculating the energy concentration of the target tuner track in the preset low-frequency band during the target time period, the method further includes: If the energy signal of the target orchestration track is determined to be concentrated in the preset low-frequency band based on the energy concentration, then the energy signal of the target orchestration track is subjected to full-pass filtering to obtain the processed orchestration track.

8. The low-frequency song generation method according to claim 1, characterized in that, The process of down-tuning the target orchestration track to obtain the processed orchestration track includes: A preset frequency divider is used to classify the audio in the target tuner track located in the preset low-frequency band as low-frequency audio, and the remaining audio besides the low-frequency audio as high-frequency audio; wherein, the preset frequency divider is constructed based on an LR filter; The high-frequency audio is down-pitched based on a preset octave down-pitch strategy to obtain down-pitch audio, and the low-frequency audio is subjected to full-pass filtering to obtain filtered audio. The down-pitched audio and the filtered audio are mixed to obtain the processed instrumentation track.

9. The low-frequency song generation method according to claim 1, characterized in that, The process of obtaining the processed vocal track after adding backing vocals to the vocal track includes: The fundamental frequency information of the vocal track is detected using a fundamental frequency extraction tool and a CREPE model, and the fundamental frequency concentration of the fundamental frequency information within a preset frequency band is calculated. If the fundamental frequency concentration is greater than a preset threshold, then the vocal track is subjected to full-pass filtering to obtain the processed vocal track. If the fundamental frequency concentration is not greater than the preset threshold, the vocal track is down-tuned based on the octave down-tuning strategy to obtain the down-tuned vocals, and the down-tuned vocals and the first preset weighting parameter are used to perform backing sound processing on the vocal track to obtain the processed vocal track.

10. The low-frequency song generation method according to claim 1, characterized in that, The process of obtaining the processed vocal track after adding backing vocals to the vocal track includes: The fundamental frequency information of the vocal track is detected using a fundamental frequency extraction tool and a CREPE model, and the target fundamental frequency that is not located in the preset frequency band is determined. The target fundamental frequency is down-tuned based on a preset octave down-tuning strategy to obtain a down-tuned fundamental frequency. The down-tuned fundamental frequency and a second preset weighting parameter are then used to perform backing vocal processing on the time period of the target fundamental frequency to obtain the processed vocal track.

11. The low-frequency song generation method according to any one of claims 1 to 10, characterized in that, After synthesizing the processed vocal track and the processed instrumentation track to generate the low-frequency song corresponding to the original song, the method further includes: The degree of hearing loss of the target user is determined, and the loudness adjustment parameters corresponding to the degree of hearing loss are determined based on the generalized dynamic range compression technology; The loudness of the low-frequency song is adjusted using the loudness adjustment parameters, and the adjusted low-frequency song is then played.

12. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the low-frequency song generation method as described in any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the low-frequency song generation method as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Audio processing method and device, electronic device and storage medium

    CN109524016A

  • Singing intonation scoring method and device, equipment, medium and product

    CN116110431A