Audio signal processing method and system
Through the audio signal processing method of sampling, framed and moving the real-time audio stream, the problem that existing sound editing software cannot handle human voice electronic music in real time is solved, and the electronic music effect adjustment of high sound quality and low distortion is achieved.
Patent Information
- Application Number
- CN202310094072.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-03
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-02-03
AI Technical Summary
Existing voice revision software cannot realize real-time adjustable voice change processing for human voice, and the vocal distortion is serious.
By acquiring the real-time audio stream, sampling and frame processing, determining the fundamental frequency and converting it to a scale, the fundamental frequency is moved by using the improved real-time PSOLA algorithm, and adjusting the electronic music strength and strength in combination with predetermined parameters to realize real-time processing of the audio signal.
Provides excellent audio and audio effects for vocals in high real-time scenarios, reducing voice output distortion, improving sound quality, simplifying user operations, and providing free adjustment of electronic music effects.
Smart Images

Figure CN116092457B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technologies, and in particular, to an audio signal processing method, system, electronic device, and computer-readable storage medium. Background Art
[0002] Electronic music generally refers to audio produced using electronic musical instruments and electronic music technologies. In the field of vocal tuning, it generally refers to using means such as pitch shifting to transform the frequency of a human voice to a specific scale and obtain a special listening experience. In application scenarios such as audio-visual editing and live performances, users will use voice tuning software to achieve the vocal electronic music effect, thereby achieving purposes such as beautifying the voice, adding fun, and concealing identity. Currently, the vocal electronic music on the market mainly uses voice tuning software such as Autotune and Melodyne, and uses signal processing algorithms such as pitch shifting to automatically adjust the human voice to the nearest scale. However, most current voice tuning software has problems such as a high threshold, inability to adjust the sound effect using parameters, inability to process in real time, and serious vocal distortion. Summary of the Invention
[0003] The main purpose of this application is to propose an audio signal processing method, system, electronic device, and computer-readable storage medium, aiming to solve the problem of how to perform real-time adjustable vocal electronic music voice conversion processing and ensure less distortion of the output sound.
[0004] To achieve the above purpose, an embodiment of this application provides an audio signal processing method, and the method includes:
[0005] Obtain a real-time input audio stream;
[0006] Sample the audio stream at a predetermined sampling rate to obtain a series of discrete sampling points, and perform frame division processing on the sampling points to obtain a plurality of input speech frames;
[0007] Determine the fundamental frequency of each input speech frame;
[0008] Convert the fundamental frequency to a scale and discretize it;
[0009] Convert the discretized scale back to a second frequency value, and obtain a fundamental frequency shift ratio according to the second frequency value and the fundamental frequency;
[0010] According to the fundamental frequency shift ratio, shift the fundamental frequency of each input speech frame;
[0011] Overlay and splice each shifted input speech frame to obtain a processed audio stream.
[0012] Optionally, after obtaining the real-time input audio stream, the method further includes:
[0013] Obtain a predetermined parameter input by a user, where the predetermined parameter represents the strength degree of electronic music; and
[0014] When discretizing the scale, it further includes discretizing according to the predetermined parameter.
[0015] Optionally, the determining the fundamental frequency of each input speech frame includes:
[0016] Determine the fundamental frequency of the input speech frame through a time-domain autocorrelation algorithm or a frequency-domain harmonic-to-noise ratio algorithm.
[0017] Optionally, the discretizing according to the predetermined parameter includes:
[0018] Use the predetermined parameter as the granularity of the discretization, round the scale, and obtain an integer value.
[0019] Optionally, the fundamental frequency shift multiple is the ratio of the second frequency value to the fundamental frequency.
[0020] Optionally, the shifting the fundamental frequency of each input speech frame includes:
[0021] Shift the fundamental frequency of each input speech frame through a real-time pitch synchronous overlap algorithm.
[0022] Optionally, the shifting the fundamental frequency of each input speech frame through a real-time pitch synchronous overlap algorithm includes:
[0023] Obtain the fundamental frequency of the input speech frame;
[0024] Decompose the input speech frame according to the period corresponding to the fundamental frequency to obtain a plurality of pitch period segments;
[0025] Use a window function to extract speech segments from each pitch period segment;
[0026] Copy and shift the speech segments according to the fundamental frequency shift multiple, and obtain the shifted result of the input speech frame after superposition.
[0027] Optionally, the decomposing the input speech frame according to the period corresponding to the fundamental frequency to obtain a plurality of pitch period segments includes:
[0028] According to the period corresponding to the fundamental frequency, identify the maximum value in each period, and use the current maximum value as the center and the two adjacent maximum values before and after as the boundaries to form a pitch period segment.
[0029] In addition, to achieve the above object, an embodiment of the present application further provides an audio signal processing system, and the system includes:
[0030] An acquisition module, configured to acquire a real-time input audio stream;
[0031] A sampling module, configured to sample the audio stream at a predetermined sampling rate to obtain a series of discrete sampling points, and perform frame division processing on the sampling points to obtain a plurality of input speech frames;
[0032] A determination module, configured to determine the fundamental frequency of each of the input speech frames;
[0033] A conversion module, configured to convert the fundamental frequency into a musical scale and perform discretization;
[0034] A calculation module, configured to convert the discretized musical scale back to a second frequency value, and obtain a fundamental frequency shift ratio according to the second frequency value and the fundamental frequency;
[0035] A movement module, configured to move the fundamental frequency of each of the input speech frames according to the fundamental frequency shift ratio;
[0036] A splicing module, configured to superimpose and splice each of the moved input speech frames to obtain a processed audio stream.
[0037] To achieve the above object, an embodiment of the present application further provides an electronic device, where the electronic device includes: a memory, a processor, and an audio signal processing program stored on the memory and executable on the processor, and when the audio signal processing program is executed by the processor, the audio signal processing method as described above is implemented.
[0038] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium, where an audio signal processing program is stored on the computer-readable storage medium, and when the audio signal processing program is executed by a processor, the audio signal processing method as described above is implemented.
[0039] The audio signal processing method, system, electronic device, and computer-readable storage medium proposed in the embodiments of the present application can provide a human voice electro-acoustic tuning service in a high-real-time scenario. By analyzing the signal characteristics of the input audio signal stream uploaded by the user frame by frame through a real-time algorithm, the fundamental frequency of the voice is transformed into a specific musical scale to change the pitch of the voice, ensuring that the output of the tuned voice has less distortion and higher sound quality, and improving the user's auditory experience. At the same time, the strength of the electro-acoustic effect can be adjusted using predetermined parameters, giving the user the freedom to adjust the human voice electro-acoustic effect using the parameters while simplifying the user operation and reducing the usage threshold. Description of the Drawings
[0040] Figure 1 An application environment architecture diagram for implementing various embodiments of the present application;
[0041] Figure 2Flowchart of an audio signal processing method proposed in the first embodiment of this application;
[0042] Figure 3 For Figure 2 Schematic diagram of the detailed process of step S210 in
[0043] Figure 4 Schematic diagram of the pitch period segment in the first embodiment of this application;
[0044] Figure 5 According to Figure 4 Extracted speech segment and fundamental frequency shift schematic diagram;
[0045] Figure 6 Flowchart of an audio signal processing method proposed in the second embodiment of this application;
[0046] Figure 7 Schematic diagram of the hardware architecture of an electronic device proposed in the third embodiment of this application;
[0047] Figure 8 Schematic diagram of the modules of an audio signal processing system proposed in the fourth embodiment of this application. Detailed implementation manners
[0048] In order to make the objectives, technical solutions and advantages of this application clearer, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0049] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of this application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by this application.
[0050] Please refer to Figure 1 , Figure 1 is an application environment architecture diagram for implementing various embodiments of this application. This application can be applied to an application environment including, but not limited to, a client 2, a server 4, and a network 6.
[0051] Among them, the client 2 is used to receive the audio stream input by the user in real time, receive the predetermined parameters, and output the voice conversion result to the user, etc. The client 2 can be a terminal device such as a PC (Personal Computer), mobile phone, tablet computer, portable computer, wearable device, etc., or a dedicated audio processing device, such as a voice changer, etc.
[0052] The server 4 is used to provide technical support for the client 2, including obtaining the audio stream and the predetermined parameters from the client 2, performing sampling and frame segmentation processing on the audio stream, determining the fundamental frequency of each input voice frame, converting the fundamental frequency into a musical scale, and performing discretization according to the predetermined parameters, calculating the fundamental frequency shift magnification, moving the fundamental frequency of each input voice frame through a real-time algorithm, splicing each moved input voice frame, and outputting the voice conversion result, etc. The server 4 can be a computing device such as a rack server, blade server, tower server, or cabinet server, and can be an independent server or a server cluster composed of multiple servers.
[0053] The network 6 can be a wireless or wired network such as an enterprise intranet (Intranet), Internet, Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, etc. The server 4 and one or more clients 2 are communicatively connected through the network 6 for data transmission and interaction.
[0054] Embodiment 1
[0055] As Figure 2 shown, it is a flowchart of an audio signal processing method proposed in the first embodiment of the present application. It can be understood that the flowchart in the embodiment of this method is not used to limit the order of execution steps. According to needs, some steps in this flowchart can also be added or deleted. The audio signal processing method can be applied to the client or the server. The following takes the server as the execution subject to illustrate this method.
[0056] This method includes the following steps:
[0057] S200, obtain the audio stream input in real time.
[0058] Currently, existing voice tuning software generally cannot perform real-time voice conversion on audio, while this embodiment can perform real-time processing. After the user inputs the audio stream to be processed in real time on the client side, the server obtains the audio stream for real-time processing, and then returns the voice conversion result to the client for real-time output to the user.
[0059] S202, sample the audio stream at a predetermined sampling rate to obtain a series of discrete sampling points, and perform frame segmentation on the sampling points to obtain multiple input speech frames.
[0060] After obtaining the audio stream, first, sample the speech signal in the audio stream. The sampling rate refers to how many times the analog signal is sampled per unit time. The higher the sampling frequency, the more real and natural the waveform of the mechanical wave is. In this embodiment, the audio stream can be sampled at any feasible sampling rate to obtain a series of discrete sampling points. Generally speaking, the commonly used sampling rates are mainly 44100 Hz (hertz) and 48000 Hz. Then, perform frame segmentation on these sampling points, that is, take a predetermined number of sampling points as an input speech frame to obtain multiple input speech frames. For example, take 1024 sampling points as an input speech frame.
[0061] S204, determine the fundamental frequency of each input speech frame.
[0062] The fundamental frequency refers to the most basic vibration frequency in the speech signal, which represents the degree to which the human ear can distinguish the pitch of a sound in terms of auditory perception. The methods for measuring the fundamental frequency of a speech frame can be roughly divided into the time domain method and the frequency domain method. Among them, the time domain method takes the waveform of the sound as the input, and its basic principle is to find the smallest positive period of the waveform. Of course, the periodicity of the actual signal can only be approximate. The frequency domain method will first perform a Fourier transform on the signal to obtain the frequency spectrum (only take the magnitude spectrum and discard the phase spectrum). There will be spikes at integer multiples of the fundamental frequency on the frequency spectrum, and the basic principle of the frequency domain method is to find the greatest common divisor of these spike frequencies.
[0063] In this embodiment, after obtaining multiple input speech frames, for each input speech frame, the fundamental frequency of the input speech frame can be determined by means of the time domain autocorrelation algorithm or the frequency domain harmonic-to-noise ratio algorithm, etc. In other embodiments, any other feasible existing method can also be used to determine the fundamental frequency, which is not limited here.
[0064] S206, convert the fundamental frequency to a musical scale and discretize it.
[0065] The scale refers to a series of tones with frequencies increasing from low to high, increasing according to an exponential relationship. In this embodiment, the scale can be represented by MIDI (Musical Instrument Digital Interface) values. Under the commonly used twelve-tone system in the current music field, there is an exponential function relationship between the scale and the corresponding frequency. Specifically, for every twelve scales increased, the pitch rises by one octave, and the corresponding frequency becomes twice the original, that is, the ratio of the frequencies corresponding to two adjacent scales is 2^(1 / 12) = 1.05946. At the same time, in music theory, it is stipulated that the scale number corresponding to middle C (C4) is 60, and the corresponding frequency is 261.6 Hz. Therefore, the conversion formula from the frequency value to the scale value is: MIDI = 60 + 12 * log2(F / 261.6), where F represents the frequency value, that is, the fundamental frequency of each of the input voice frames, in Hz. According to the above formula, the fundamental frequency of each input voice frame can be converted into a MIDI value, that is, converted into a scale.
[0066] In the traditional tuning scheme, vocal electronic music is generally achieved by converting the frequency of the input voice signal to the nearest integer MIDI value. For example, assume that the fundamental frequency of the input voice frame is 250 Hz, and the corresponding MIDI value is 59.215, which can be converted to the nearest integer MIDI value 59, that is, at 246.9 Hz. In this embodiment, such a scheme can also be adopted to discretize the scale. That is to say, after converting the fundamental frequency of each input voice frame into a scale, the scale is then converted to the nearest integer MIDI value.
[0067] S208, convert the discretized scale back to the second frequency value, and obtain the fundamental frequency shift ratio according to the second frequency value and the fundamental frequency.
[0068] After discretizing the scale, that is, converting it into an integer MIDI value, according to the above conversion formula, it can be calculated in reverse, and the scale value (integer MIDI value) can be converted back to the corresponding frequency value, which is called the second frequency value. Then, taking the ratio of the second frequency value and the fundamental frequency of the input voice frame, the corresponding fundamental frequency shift ratio can be obtained. That is to say, the fundamental frequency shift ratio of each input voice frame = second frequency value / fundamental frequency.
[0069] S210, shift the fundamental frequency of each input voice frame according to the fundamental frequency shift ratio.
[0070] In the existing Pitch Synchronous Overlap and Add (PSOLA) algorithm, generally the fundamental frequency of the entire audio segment is shifted, so real-time processing cannot be achieved. In this embodiment, through the improved real-time PSOLA algorithm, the fundamental frequency period segments can be decomposed based on each of the input speech frames, and then the fundamental frequency of each segment is shifted, rather than shifting the entire audio segment, thereby achieving real-time processing of the audio signal.
[0071] In this embodiment, the basic idea of the real-time PSOLA algorithm is to decompose the speech signal into segments corresponding to the fundamental frequency period, and then perform operations such as copying and translation on these segments. Since the short-term correlation in the signal is not changed, that is, the signal within the segment does not change, the spectral envelope of the signal is maintained, and only the fundamental frequency of the signal is changed, and the quality of the speech signal is retained to the greatest extent.
[0072] Specifically, further refer to Figure 3 , which is a schematic diagram of the refined process of the above step S210. It can be understood that this flowchart is not used to limit the order of executing steps. According to needs, some steps in this flowchart can also be added or deleted. In this embodiment, the step S210 specifically includes:
[0073] S2100, obtain the fundamental frequency of the input speech frame.
[0074] S2102, decompose the input speech frame according to the period corresponding to the fundamental frequency to obtain a plurality of fundamental frequency period segments.
[0075] In this embodiment, in order to perform real-time voice conversion processing, after the audio stream is frame-divided, each of the input speech frames is segmented according to the fundamental frequency period, and subsequent movement is performed on each segment. According to the period corresponding to the fundamental frequency, the maximum value in each period can be identified. Taking the current maximum value as the center and the two adjacent maximum values before and after as the boundaries, a fundamental frequency period segment can be formed.
[0076] As Figure 4 shown, it is a schematic diagram of the fundamental frequency period segment in this embodiment. In Figure 4 , the maximum value identified in each period is represented by the x point, and then taking each x point as the center and the adjacent x points on the left and right as the boundaries, it is a fundamental frequency period segment.
[0077] S2104, use a window function to extract speech segments from each of the fundamental frequency period segments.
[0078] By applying the window function to the fundamental frequency period segment, the speech segment corresponding to each of the fundamental frequency period segments can be extracted. The window length of the window function is the length of the fundamental frequency period segment, that is, taking the two adjacent maximum values before and after as the boundaries. AsFigure 5 As shown, it is Figure 4 the extracted voice segment and the fundamental frequency shift schematic diagram. In Figure 5 , from the input voice signal at the top, several voice segments can be extracted, such as Figure 5 shown in the middle. Each of the pitch period segments corresponds to one of the voice segments.
[0079] S2106, copy and shift the voice segment according to the fundamental frequency shift ratio, and after superposition, obtain the shift result of the input voice frame.
[0080] According to the fundamental frequency shift ratio calculated previously, copy and perform fundamental frequency shift on each of the voice segments, and then the shift result of each of the voice segments can be obtained, such as Figure 5 shown in the middle. The fundamental frequency shift means changing the frequency of the fundamental frequency part in the signal through an algorithm without changing the frequencies of other harmonic parts. Then, by superposing the shift results of each of the voice segments, the shift result of the input voice frame can be obtained, such as Figure 5 shown in the lower part.
[0081] Return to Figure 2 , S212, superpose and splice each of the shifted input voice frames to obtain the processed audio stream.
[0082] In this embodiment, each of the input voice frames is processed separately. When the shift result of each of the input voice frames is obtained, they are re-spliced frame by frame and output, and then the voice conversion result of the audio stream can be obtained.
[0083] The audio signal processing method proposed in this embodiment can analyze the signal characteristics of the input audio signal stream uploaded by the user frame by frame through the real-time PSOLA algorithm, transform the fundamental frequency of the voice to a specific scale, change the pitch of the voice, ensure that the voice output after tuning has less distortion and higher sound quality, realize providing a better auditory experience of human voice electroacoustic sound effects in high-real-time scenarios, and improve the user's auditory experience. Moreover, the human voice electroacoustic sound effects are realized with less computational amount, the feedback speed is fast, the interaction is simple, and the user experience during the tuning process is improved.
[0084] Embodiment 2
[0085] Such as Figure 6As shown in the figure, it is a flowchart of an audio signal processing method proposed in the second embodiment of the present application. In the second embodiment, based on the above first embodiment, the audio signal processing method adds the setting of predetermined parameters, and discretizes the musical scale according to the predetermined parameters, so as to adjust the strength of the electronic voice of voice conversion. It can be understood that the flowchart in the embodiment of the present method is not used to limit the order of execution steps. According to needs, some steps in this flowchart can also be added or deleted.
[0086] The method includes the following steps:
[0087] S300, Obtain the real-time input audio stream and predetermined parameters.
[0088] Generally, the current existing pitch correction software generally cannot perform real-time voice conversion on audio, while this embodiment can perform real-time processing. After the user inputs the audio stream to be processed in real time on the client side, the server obtains the audio stream for real-time processing, and then returns the voice conversion result to the client side for real-time output to the user. In addition, this embodiment also allows the user to customize the electronic voice effect through predetermined parameters, and the user can input the value of the predetermined parameters on the client side. The predetermined parameter represents the strength of the electronic voice, which can be represented by a, and is used for subsequent discretization of the converted musical scale.
[0089] S302, Sample the audio stream according to a predetermined sampling rate to obtain a series of discrete sampling points, and perform frame division processing on the sampling points to obtain multiple input speech frames.
[0090] After obtaining the audio stream, the voice signal in the audio stream needs to be sampled first. In this embodiment, the audio stream can be sampled according to any feasible sampling rate to obtain a series of discrete sampling points. Generally speaking, the commonly used sampling rates are mainly 44100Hz and 48000Hz. Then, frame division processing is performed on these sampling points to obtain multiple input speech frames. For example, 1024 sampling points are used as an input speech frame.
[0091] S304, Determine the fundamental frequency of each input speech frame.
[0092] In this embodiment, after obtaining multiple input speech frames, for each input speech frame, the fundamental frequency of the input speech frame can be determined by means of time-domain autocorrelation algorithm or frequency-domain harmonic-to-noise ratio algorithm, etc. In other embodiments, any other feasible existing method can also be used to determine the fundamental frequency, which is not limited here.
[0093] S306, Convert the fundamental frequency to a musical scale, and discretize it according to the predetermined parameters.
[0094] In this embodiment, the musical scale can be represented by MIDI values. The conversion formula from frequency values to musical scale values is: MIDI = 60 + 12 * log2(F / 261.6), where F represents the frequency value, that is, the fundamental frequency of each input speech frame, and the unit is Hz. According to the above formula, the fundamental frequency of each input speech frame can be converted into a MIDI value, that is, converted into a musical scale.
[0095] After converting the fundamental frequency into a musical scale, discretize the musical scale according to the predetermined parameter input by the user. In this embodiment, the user can use a custom predetermined parameter a to represent the intensity of the electronic sound. The predetermined parameter a represents the granularity of discretization when the fundamental frequency is converted into a MIDI value. At this time, after converting the fundamental frequency into a musical scale MIDI value, the discretization formula is a * round(MIDI / a), where the round function represents rounding, that is, taking an integer. For example, when a = 1, it represents converting the frequency of the human voice speech signal to the nearest integer MIDI value in the traditional tuning scheme. Assume that the fundamental frequency of the input speech frame is 250 Hz, then the corresponding MIDI value is 59.215. Discretize it with a = 1 as the granularity and convert it to the nearest integer MIDI value, which is 59, that is, at 246.9 Hz. When a = 2.5, discretize it with 2.5 as the granularity, and the converted integer MIDI value is 2.5 * round(59.215 / 2.5) = 60, that is, at 261.6 Hz. The larger the value of the predetermined parameter a, the stronger the electronic sound effect in the obtained voice-changing result.
[0096] S308, convert the discretized musical scale back to the second frequency value, and obtain the fundamental frequency shift ratio according to the second frequency value and the fundamental frequency.
[0097] After discretizing the musical scale with the predetermined parameter a as the granularity, according to the above conversion formula, it can be calculated in reverse, and the musical scale value can be converted back to the corresponding frequency value, which is called the second frequency value. Then, take the ratio of the second frequency value and the fundamental frequency of the input speech frame to obtain the corresponding fundamental frequency shift ratio. That is, the fundamental frequency shift ratio of each input speech frame = second frequency value / fundamental frequency.
[0098] S310, move the fundamental frequency of each input speech frame according to the fundamental frequency shift ratio.
[0099] In the existing PSOLA algorithm, generally, the fundamental frequency of the entire audio segment is shifted, so real-time processing cannot be achieved. In this embodiment, through the improved real-time PSOLA algorithm, the segments of the fundamental frequency period can be decomposed based on each input speech frame, and then the fundamental frequency of each segment is shifted instead of shifting the entire audio segment, thereby realizing the real-time processing of the audio signal.
[0100] In this embodiment, the basic idea of the real-time PSOLA algorithm is to decompose the speech signal into segments corresponding to the fundamental frequency period, and then perform operations such as copying and translation on these segments. Since the short-term correlation in the signal is not changed, that is, the signal within the segment does not change, the spectral envelope of the signal is maintained, and only the fundamental frequency of the signal is changed, so the sound quality of the speech signal is retained to the greatest extent.
[0101] Specifically, first, obtain the fundamental frequency of the input speech frame, and decompose the input speech frame according to the period corresponding to the fundamental frequency to obtain multiple pitch period segments. Then, use a window function to extract speech segments from each pitch period segment. Then, copy and move the speech segments according to the fundamental frequency shift ratio, and finally, after superposition, obtain the movement result of the input speech frame.
[0102] S312, superimpose and splice each of the moved input speech frames to obtain the processed audio stream.
[0103] In this embodiment, each input speech frame is processed separately. When the movement result of each input speech frame is obtained, they are re-spliced frame by frame and output, and then the voice-changing result of the audio stream can be obtained.
[0104] The audio signal processing method proposed in this embodiment deeply explores the user's needs, provides a vocal electro-tuning service in a high-real-time scenario. Through the real-time PSOLA algorithm, it analyzes the signal characteristics of the input audio signal stream uploaded by the user frame by frame, transforms the fundamental frequency of the speech to a specific scale, changes the pitch of the speech, ensures that the output speech after tuning has less distortion and higher sound quality, and improves the user's auditory experience. At the same time, the strength of the electro-music effect can be adjusted using predetermined parameters. On the premise of simplifying the user operation and reducing the usage threshold, it gives the user the freedom to adjust the vocal electro-music effect using parameters. Moreover, the vocal electro-music sound effect is realized with less computational effort, has a fast feedback speed, and simple interaction, improving the user's experience during the tuning process.
[0105] Embodiment III
[0106] Such as Figure 7As shown in the figure, the following is a schematic diagram of the hardware architecture of an electronic device 20 proposed in the third embodiment of the present application. In this embodiment, the electronic device 20 may include, but is not limited to, a memory 21, a processor 22, and a network interface 23 that are communicatively connected to each other via a system bus. It should be noted that Figure 7 Only the electronic device 20 with components 21-23 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. In this embodiment, the electronic device 20 can be the client or the server.
[0107] The memory 21 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 21 can be an internal storage unit of the electronic device 20, such as the hard disk or memory of the electronic device 20. In other embodiments, the memory 21 can also be an external storage device of the electronic device 20, such as a plug-in hard disk equipped on the electronic device 20, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Of course, the memory 21 can also include both the internal storage unit and the external storage device of the electronic device 20. In this embodiment, the memory 21 is generally used to store the operating system and various application software installed on the electronic device 20, such as the program code of the audio signal processing system 60, etc. In addition, the memory 21 can also be used to temporarily store various data that have been output or will be output.
[0108] In some embodiments, the processor 22 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 22 is generally used to control the overall operation of the electronic device 20. In this embodiment, the processor 22 is used to run the program code stored in the memory 21 or process data, such as running the audio signal processing system 60, etc.
[0109] The network interface 23 can include a wireless network interface or a wired network interface, and this network interface 23 is generally used to establish a communication connection between the electronic device 20 and other electronic devices.
[0110] Embodiment Four
[0111] As shown Figure 8 in the figure, a schematic diagram of modules of an audio signal processing system 60 is proposed in the fourth embodiment of the present application. The audio signal processing system 60 can be divided into one or more program modules, and one or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment.
[0112] In this embodiment, the audio signal processing system 60 includes:
[0113] An acquisition module 600, configured to acquire a real-time input audio stream.
[0114] Generally, existing voice tuning software generally cannot perform real-time voice conversion on audio, while this embodiment can perform real-time processing. After a user inputs an audio stream to be processed in real time at a client, the acquisition module 600 acquires the audio stream for real-time processing, and then returns the voice conversion result to the client for real-time output to the user.
[0115] A sampling module 602, configured to sample the audio stream at a predetermined sampling rate to obtain a series of discrete sampling points, and perform frame division processing on the sampling points to obtain a plurality of input speech frames.
[0116] After the audio stream is acquired, the voice signal in the audio stream needs to be sampled first. In this embodiment, the audio stream can be sampled at any feasible sampling rate to obtain a series of discrete sampling points. Generally speaking, the commonly used sampling rates are mainly 44100 Hz and 48000 Hz. Then, frame division processing is performed on these sampling points to obtain a plurality of input speech frames. For example, 1024 sampling points are used as an input speech frame.
[0117] A determination module 604, configured to determine the fundamental frequency of each of the input speech frames.
[0118] In this embodiment, after a plurality of the input speech frames are obtained, for each of the input speech frames, the fundamental frequency of the input speech frame can be determined by means of a time-domain autocorrelation algorithm or a frequency-domain harmonic-to-noise ratio algorithm, etc. In other embodiments, any other feasible existing method can also be used to determine the fundamental frequency, which is not limited herein.
[0119] A conversion module 606, configured to convert the fundamental frequency into a musical scale and perform discretization.
[0120] In this embodiment, the musical scale can be represented by MIDI values. The conversion formula from frequency values to musical scale values is: MIDI = 60 + 12 * log2(F / 261.6), where F represents the frequency value, that is, the fundamental frequency of each input speech frame, and the unit is Hz. According to the above formula, the fundamental frequency of each input speech frame can be converted into a MIDI value, that is, converted into a musical scale.
[0121] In traditional tuning schemes, vocal electro - music is generally achieved by converting the frequency of the input speech signal to the nearest integer MIDI value. For example, assume that the fundamental frequency of the input speech frame is 250 Hz, and the corresponding MIDI value is 59.215, which can be converted to the nearest integer MIDI value 59, that is, at 246.9 Hz. In this embodiment, such a scheme can also be used to discretize the musical scale. That is to say, after converting the fundamental frequency of each input speech frame into a musical scale, the musical scale is then converted to the nearest integer MIDI value.
[0122] The calculation module 608 is used to convert the discretized musical scale back to the second frequency value, and obtain the fundamental frequency shift ratio according to the second frequency value and the fundamental frequency.
[0123] After discretizing the musical scale, according to the above conversion formula, calculating in reverse, the musical scale value can be converted back to the corresponding frequency value, which is called the second frequency value. Then, taking the ratio of the second frequency value and the fundamental frequency of the input speech frame, the corresponding fundamental frequency shift ratio can be obtained. That is to say, the fundamental frequency shift ratio of each input speech frame = second frequency value / fundamental frequency.
[0124] The movement module 610 is used to move the fundamental frequency of each input speech frame according to the fundamental frequency shift ratio.
[0125] In the existing PSOLA algorithm, generally, the fundamental frequency of the entire audio segment is moved, so it cannot be processed in real - time. In this embodiment, through the improved real - time PSOLA algorithm, the fundamental - frequency period segments can be decomposed based on each input speech frame, and then the fundamental frequency of each segment is moved, rather than moving the entire audio segment, so as to achieve real - time processing of the audio signal.
[0126] In this embodiment, the basic idea of the real - time PSOLA algorithm is to decompose the speech signal into segments corresponding to the fundamental - frequency period, and then perform operations such as copying and translation on these segments. Because the short - term correlation in the signal is not changed, that is, the signal within the segment does not change, the spectral envelope of the signal is maintained, and only the fundamental frequency of the signal is changed, so the sound quality of the speech signal is retained to the greatest extent.
[0127] Specifically, first, obtain the fundamental frequency of the input speech frame, and decompose the input speech frame according to the period corresponding to the fundamental frequency to obtain a plurality of pitch period segments. Then, use a window function to extract speech segments from each of the pitch period segments. Next, copy and move the speech segments according to the fundamental frequency shift ratio, and finally, obtain the movement result of the input speech frame after superposition.
[0128] The splicing module 612 is configured to stack and splice each of the moved input speech frames to obtain a processed audio stream.
[0129] In this embodiment, each of the input speech frames is processed separately. After obtaining the movement result of each input speech frame, they are spliced frame by frame and output, and then the pitch-shifting result of the audio stream can be obtained.
[0130] Preferably, the obtaining module 600 is further configured to obtain a predetermined parameter input by the user.
[0131] This embodiment can also allow the user to customize the electroacoustic effect through the predetermined parameter, and the user can input the value of the predetermined parameter on the client. The predetermined parameter represents the strength of the electroacoustic effect and can be represented by a, which is used for subsequent discretization of the converted musical scale.
[0132] The conversion module 606 is further configured to discretize the converted musical scale according to the predetermined parameter.
[0133] In this embodiment, the user can customize the predetermined parameter a to represent the strength of the electroacoustic effect. The predetermined parameter a represents the granularity of discretization when the fundamental frequency is converted to the MIDI value. At this time, after converting the fundamental frequency to the musical scale MIDI value, the discretization formula is a*round(MIDI / a), where the round function represents rounding, that is, taking an integer. For example, when a = 1, it represents the traditional tuning scheme of converting the frequency of the human voice speech signal to the nearest integer MIDI value. Assume that the fundamental frequency of the input speech frame is 250 Hz, then the corresponding MIDI value is 59.215. Discretizing it with a = 1 as the granularity and converting it to the nearest integer MIDI value is 59, that is, at 246.9 Hz. When a = 2.5, discretizing it with 2.5 as the granularity and converting it to the integer MIDI value is 2.5*round(59.215 / 2.5) = 60, that is, at 261.6 Hz. The larger the value of the predetermined parameter a, the stronger the electroacoustic effect in the obtained pitch-shifting result.
[0134] The audio signal processing system proposed in this embodiment can provide human voice pitch tuning services in high-real-time scenarios. Through the real-time PSOLA algorithm, it analyzes the signal characteristics of the input audio signal stream uploaded by the user frame by frame, transforms the fundamental frequency of the voice to a specific scale, changes the pitch of the voice, ensures that the voice output after tuning has less distortion and higher sound quality, and improves the user's auditory experience. At the same time, it can adjust the strength of the electronic music effect using predetermined parameters, giving users the freedom to adjust the human voice electronic music effect using parameters on the premise of simplifying user operations and reducing the usage threshold.
[0135] Embodiment 5
[0136] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing an audio signal processing program, and the audio signal processing program can be executed by at least one processor, so that the at least one processor executes the steps of the audio signal processing method as described above.
[0137] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0138] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0139] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0140] The above are only the preferred embodiments of the embodiments of the present application, and do not limit the patent scope of the embodiments of the present application. Any equivalent structure or equivalent process transformation made by using the description and drawings of the embodiments of the present application, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the embodiments of the present application.
Claims
1. An audio signal processing method, characterized in that The method includes: Obtaining a real-time input audio stream; Sampling the audio stream at a predetermined sampling rate to obtain a series of discrete sampling points, and performing frame processing on the sampling points to obtain a plurality of input speech frames; Determining the fundamental frequency of each of the input speech frames; Based on a preset conversion formula, converting the fundamental frequency into a musical scale and discretizing it; Converting the discretized musical scale back to a second frequency value, and obtaining a fundamental frequency shift ratio according to the second frequency value and the fundamental frequency; According to the fundamental frequency shift ratio, shifting the fundamental frequency of each of the input speech frames; Superposing and splicing each of the shifted input speech frames to obtain a processed audio stream; Wherein, the second frequency value is calculated according to the preset conversion formula and the discretized musical scale.
2. The audio signal processing method according to claim 1, wherein After obtaining the real-time input audio stream, the method further includes: Obtaining a predetermined parameter input by a user, where the predetermined parameter represents the strength of electronic music; and When discretizing the musical scale, discretization is also performed according to the predetermined parameter.
3. The audio signal processing method according to claim 1 or 2, characterized in that The determining the fundamental frequency of each of the input speech frames includes: Determining the fundamental frequency of the input speech frame through a time-domain autocorrelation algorithm or a frequency-domain harmonic-to-noise ratio algorithm.
4. The audio signal processing method according to claim 2, wherein The discretizing according to the predetermined parameter includes: Using the predetermined parameter as the granularity of the discretization, rounding the musical scale to obtain an integer value.
5. The audio signal processing method according to claim 1 or 2, characterized in that, The fundamental frequency shift ratio is the ratio of the second frequency value to the fundamental frequency.
6. The audio signal processing method according to claim 1 or 2, characterized in that, The shifting the fundamental frequency of each of the input speech frames includes: Shifting the fundamental frequency of each of the input speech frames through a real-time pitch synchronous overlap-add algorithm.
7. The audio signal processing method according to claim 6, wherein The shifting the fundamental frequency of each of the input speech frames through a real-time pitch synchronous overlap-add algorithm includes: Obtaining the fundamental frequency of the input speech frame; Decomposing the input speech frame according to the period corresponding to the fundamental frequency to obtain a plurality of pitch period segments; Using a window function to extract speech segments from each of the pitch period segments; Copying and shifting the speech segments according to the fundamental frequency shift ratio, and obtaining the shifted result of the input speech frame after superposition.
8. The audio signal processing method according to claim 7, wherein The decomposing the input speech frame according to the period corresponding to the fundamental frequency to obtain a plurality of pitch period segments includes: According to the period corresponding to the fundamental frequency, identifying the maximum value in each period, and using the current maximum value as the center and the two adjacent maximum values before and after as the boundaries to form a pitch period segment.
9. An audio signal processing system, characterized in that, The system includes: An obtaining module, configured to obtain a real-time input audio stream; A sampling module, configured to sample the audio stream at a predetermined sampling rate to obtain a series of discrete sampling points, and perform frame processing on the sampling points to obtain a plurality of input speech frames; A determining module, configured to determine the fundamental frequency of each of the input speech frames; A conversion module, configured to convert the fundamental frequency into a musical scale based on a preset conversion formula and discretize it; A calculation module, configured to convert the discretized musical scale back to a second frequency value, and obtain a fundamental frequency shift ratio according to the second frequency value and the fundamental frequency; A shifting module, configured to shift the fundamental frequency of each of the input speech frames according to the fundamental frequency shift ratio; A splicing module, configured to splice each of the shifted input voice frames to obtain a processed audio stream after superposition splicing; Wherein, the second frequency value is calculated according to the preset conversion formula and the discretized musical scale.
10. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and an audio signal processing program stored on the memory and executable on the processor. When the audio signal processing program is executed by the processor, the audio signal processing method according to any one of claims 1 to 8 is implemented.
11. A computer-readable storage medium, characterized in that, An audio signal processing program is stored on the computer-readable storage medium. When the audio signal processing program is executed by a processor, the audio signal processing method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Sound processing method and device, electronic equipment and storage medium
CN112365868A
Speech analysis method and device, electronic equipment and computer readable storage medium
CN113763930A