Audio processing method and device for converting voice into microphone in real time and audio equipment
By performing octave reduction and spectral analysis on real-time human voice audio, extracting throat sounds and overtone features, and combining formant and harmonic processing, the real-time performance and naturalness of converting ordinary human voices into throat singing sounds are solved, achieving a low-threshold throat singing simulation effect.
Patent Information
- Application Number
- CN202511084929.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-04
Smart Images

Figure CN120895044A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to an audio processing method, apparatus and audio equipment for real-time human voice to throat singing. Background Technology
[0002] Khoomei, also known as throat singing, is a traditional singing style originating from the Mongolian and Tuvan regions. It possesses a strong ethnic musical character and is widely used in traditional singing and modern crossover music. A key feature of Khoomei is the simultaneous production of a fundamental tone and one or more overtones within a single voice, creating a unique harmonic effect.
[0003] In existing technologies, human voice signal processing mainly focuses on traditional effects such as reverberation, equalization, pitch shifting, pitch transposition, and voice enhancement, while simulation processing of complex vocal mechanisms such as throat singing is relatively scarce.
[0004] A few musicians may attempt to simulate the effect of throat singing through multi-track recording, synthesizer modeling, or sample splicing, but these methods generally suffer from poor real-time performance, insufficient naturalness, or high requirements for the user's singing style, making them difficult to apply to live performances or ordinary users.
[0005] Therefore, there are currently no mature consumer-grade or performance-grade products on the market that can convert ordinary human voices into sounds with throat singing characteristics in real time. Summary of the Invention
[0006] Based on this, it is necessary to provide a real-time audio processing method, device, and audio equipment for converting human voice into throat singing in order to address the aforementioned technical problems. This method can convert ordinary human voices into sounds with throat singing characteristics in real time, achieving a high degree of naturalness, high real-time performance, and low barrier to entry for throat singing simulation. This allows ordinary users to obtain throat singing effects with throat sounds and overtones simply by speaking or singing without needing to master special vocal techniques, thus meeting the diverse needs of music creation and performance.
[0007] A real-time audio processing method for converting human voice to throat singing includes: The real-time human voice audio is acquired, the single-cycle waveform is down-octave, and the continuity of the sound is removed to obtain the throat sounds corresponding to the real-time human voice audio. Real-time human voice audio is acquired, and spectrum analysis, pitch analysis, and consonant and vowel analysis are performed simultaneously to obtain the current analysis results. Based on the current analysis results, the energy of the fundamental multiples of the real-time human voice audio is calculated to perform formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the whistle tone corresponding to the real-time human voice audio. Based on the throat sounds and the whistle sounds, the throat singing audio corresponding to the real-time human voice audio is obtained.
[0008] In one embodiment, based on the current analysis results, energy calculation is performed on the energy of the fundamental multiples of the frequency points of the real-time human voice audio to perform formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the whistle tone corresponding to the real-time human voice audio, including: Based on the current analysis results, the energy of the fundamental multiples of the frequency points of the real-time human voice audio is calculated to perform formant analysis on the real-time human voice audio, track the overtone with the largest energy or the specified overtone, and obtain the overtone result; Based on the previous analysis results, the real-time human voice audio was subjected to harmonic processing to obtain harmonic results; Based on the overtone and harmonic results, the whistle tone corresponding to the real-time human voice audio is obtained.
[0009] In one embodiment, based on the current analysis results, energy calculation is performed on the energy of the fundamental multiples of the frequency points of the real-time human voice audio to perform formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the whistle tone corresponding to the real-time human voice audio, including: Based on the current analysis results, the energy of the fundamental multiples of the frequency points of the real-time human voice audio is calculated to perform formant analysis on the real-time human voice audio, track the overtone with the largest energy or the specified overtone, and obtain the overtone result; Based on the previous analysis results, the real-time human voice audio was subjected to harmonic processing to obtain harmonic results; Based on the current analysis results, the energy of the fundamental multiple frequency points of the real-time human voice audio is calculated to perform formant analysis iteration on the real-time human voice audio, track the second largest or specified overtone, and obtain the superposition result; Based on the overtone results, harmonic results, and superposition results, multiple whistle tones corresponding to real-time human voice audio are obtained.
[0010] In one embodiment, the audio processing method for real-time human voice to throat singing according to claim 1 is characterized in that, based on the current analysis results, energy calculation is performed on the energy of the fundamental multiples of the frequency points of the real-time human voice audio to perform formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the whistle tone corresponding to the real-time human voice audio, including: Based on the current analysis results, the energy of the fundamental multiples of the frequency points of the real-time human voice audio is calculated to perform formant analysis on the real-time human voice audio, track the overtone with the largest energy or the specified overtone, and obtain the overtone result; Based on the previous analysis results, the real-time human voice audio was subjected to harmonic processing to obtain harmonic results; Based on the current analysis results, the energy of the fundamental multiple frequency points of the real-time human voice audio is calculated to perform formant analysis iteration on the real-time human voice audio, track the second largest or specified overtone, and obtain the superposition result; Volume characteristics are added to the overtone results, harmonic results, and superposition results to obtain the volume result; Based on the overtone results, harmonic results, superposition results, and volume results, controllable multiple whistle tones corresponding to real-time human voice audio are obtained.
[0011] In one embodiment, based on the current analysis results, energy calculation is performed on the energy of the fundamental multiples of the frequency points of the real-time human voice audio to perform formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the overtone results, including: Based on the current analysis results, the following judgment should be made: If the current analysis result is a vowel, the energy of the fundamental multiple frequency point of the real-time human voice audio near the spectrum is calculated to perform formant analysis on the real-time human voice audio, track the overtone with the largest energy or the specified overtone, and perform bandpass filter enhancement around the overtone with the largest energy or the specified overtone to obtain the overtone result. If the current analysis result is not a vowel, the real-time human voice audio is passed through. In one embodiment, based on the previous analysis results, harmonic processing is performed on the real-time human voice audio to obtain harmonic results, including: Based on the results of the previous analysis: If the result of the previous analysis is a vowel, a Schmidt trigger is performed to obtain the harmonic result in order to maintain the continuous and stable whistle tone; If the result of the previous analysis is not a vowel, perform a direct pass-through.
[0012] In one embodiment, Schmitt triggering is performed to obtain harmonic results, including: If the current highest-energy harmonic exceeds the previous highest-energy harmonic by a preset multiple, harmonic conversion is performed; otherwise, the previous harmonic is maintained, and the harmonic result is obtained.
[0013] In one embodiment, obtaining the throat singing audio corresponding to the real-time human voice audio based on the throat sounds and the whistle sounds includes: Based on the guttural sounds and the whistle sounds, waveforms are superimposed to combine the characteristics of the guttural sounds and the whistle sounds, thereby obtaining the throat singing audio corresponding to the real-time human voice audio.
[0014] An audio processing device for real-time human voice to throat singing conversion includes: The first module is used to acquire real-time human voice audio, perform octave down-pitch on the single-cycle waveform, and remove the continuity of the sound to obtain the throat sounds corresponding to the real-time human voice audio. The second module is used to acquire real-time human voice audio, and simultaneously perform spectrum analysis, pitch analysis, and consonant and vowel analysis to obtain the current analysis results; based on the current analysis results, the energy of the fundamental multiple frequency points of the real-time human voice audio is calculated to perform formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the whistle tone corresponding to the real-time human voice audio. The third module is used to obtain the throat singing audio corresponding to the real-time human voice audio based on the throat sounds and the whistle sounds.
[0015] An audio device includes, in sequence: an audio input port, an AD converter, a main chip, a DA converter, and an audio output port; The audio input port is used to send the acquired input audio to the AD converter; The AD converter is used to perform analog-to-digital conversion on the input audio to obtain real-time human voice audio; The main chip includes a real-time human voice to vocalization audio processing device, used to convert real-time human voice audio into vocalization audio. The DA converter is used to perform digital-to-analog conversion on the throat singing audio to obtain the output audio. The audio output port is used to output audio to external devices.
[0016] The aforementioned real-time human voice to throat singing audio processing method, device, and audio equipment can convert ordinary human voices into throat singing characteristics in real time, achieving high naturalness, high real-time performance, and low barrier to entry for throat singing simulation. This fills a market gap and possesses the following outstanding substantive features and significant advancements: 1) High real-time performance: The speech feature extraction and synthesis algorithm has been optimized and low-latency processing is supported, making it suitable for live performances or real-time interaction; 2) High degree of naturalness: The generated throat singing effect has significant throat resonance and overtone harmonic characteristics, and the listening experience is close to the traditional throat singing style; 3) High ease of use: Users do not need to master special vocal techniques; they can generate expressive throat singing sounds simply by speaking or singing naturally. 4) High adaptability: It can select more than one overtone, so that the number of overtones and the volume ratio of guttural and whistling sounds can be flexibly adjusted when waveforms are superimposed, to adapt to different languages, genders and singing styles; 5) Easy to implement in hardware: The algorithm has low resource consumption and is easy to integrate into various effects processors, sound cards, mobile devices and other embedded platforms, so as to achieve portability and productization; In conclusion, this application provides a modern digital expression method for traditional throat singing, a national intangible cultural heritage, making it easy for ordinary users, teenagers, and international musicians to access and perform throat singing. This effectively lowers the learning threshold of traditional skills and promotes the inheritance and global dissemination of Chinese folk music culture. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating an audio processing method for real-time human voice to call microphone conversion in one embodiment. Figure 2 This is a structural block diagram of an audio processing device for real-time human voice to call microphone conversion in one embodiment. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0019] Furthermore, the use of terms such as "first" and "second" in this application is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of those features. In the description of this application, "multiple sets" means at least two sets, such as two sets, three sets, etc., unless otherwise explicitly specified.
[0020] In this application, unless otherwise expressly specified and limited, the terms "connection," "fixed," etc., should be interpreted broadly. For example, "fixed" can mean a fixed connection, a detachable connection, or an integral part; it can mean a mechanical connection, an electrical connection, a physical connection, or a wireless communication connection; it can mean a direct connection or an indirect connection through an intermediate medium; it can mean the internal communication of two elements or the interaction between two elements, unless otherwise expressly limited. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0021] Furthermore, the technical solutions of the various embodiments of this application can be combined with each other, but only if they are based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by this application.
[0022] This application provides an audio processing method for real-time human voice to throat singing, such as... Figure 1 The flowchart shown, in one embodiment, includes: Step 101: Acquire real-time human voice audio, perform octave reduction on the single-cycle waveform, and remove the continuity of the sound to obtain the throat sound corresponding to the real-time human voice audio.
[0023] Specifically: Real-time human voice audio is acquired, and the Pitch Superposition Algorithm (PSOLA) is used to maintain the single-cycle waveform by octave down-pitch to obtain the initial audio. For the initial audio, a high-pass filter of about 100Hz is used to remove the continuity of the sound, forming a clear and clean throat tone with strong graininess, thus obtaining the throat tone corresponding to the real-time human voice audio.
[0024] In this step, the Pitch Superposition Algorithm (PSOLA) and the high-pass filter are existing technologies and will not be described in detail here.
[0025] Step 102: Acquire real-time human voice audio and simultaneously perform spectrum analysis, pitch analysis, and consonant and vowel analysis to obtain the current analysis results; based on the current analysis results, calculate the energy of the fundamental multiples of the real-time human voice audio frequency points to perform formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the whistle tone corresponding to the real-time human voice audio.
[0026] Specifically: Acquire real-time human voice audio and simultaneously perform spectrum analysis, pitch analysis, and consonant and vowel analysis to obtain the current analysis results; Based on the current analysis results, the energy of the fundamental multiples of the frequency points of the real-time human voice audio is calculated in order to perform formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the overtone results. Based on the previous analysis results, harmonic processing was performed on the real-time human voice audio to obtain the harmonic results; Based on the overtone and harmonic results, the whistle tone corresponding to the real-time human voice audio is obtained.
[0027] Preferably: Acquire real-time human voice audio and simultaneously perform spectrum analysis, pitch analysis, and consonant and vowel analysis to obtain the current analysis results; Based on the current analysis results, the energy of the fundamental multiples of the frequency points of the real-time human voice audio is calculated in order to perform formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the overtone results. Based on the previous analysis results, harmonic processing was performed on the real-time human voice audio to obtain the harmonic results; Based on the current analysis results, the energy of the fundamental multiples of the real-time human voice audio is calculated to perform formant analysis iteration on the real-time human voice audio, track the second largest or specified overtone, and obtain the superposition result; Based on the overtone results, harmonic results, and superposition results, multiple whistle tones corresponding to real-time human voice audio are obtained.
[0028] Further preferred: Acquire real-time human voice audio and simultaneously perform spectrum analysis, pitch analysis, and consonant and vowel analysis to obtain the current analysis results; Based on the current analysis results, the energy of the fundamental multiples of the frequency points of the real-time human voice audio is calculated in order to perform formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the overtone results. Based on the previous analysis results, harmonic processing was performed on the real-time human voice audio to obtain the harmonic results; Based on the current analysis results, the energy of the fundamental multiples of the real-time human voice audio is calculated to perform formant analysis iteration on the real-time human voice audio, track the second largest or specified overtone, and obtain the superposition result; Volume characteristics are added to the overtone results, harmonic results, and superposition results to obtain the volume result; Based on the overtone results, harmonic results, superposition results, and volume results, controllable multiple whistle tones corresponding to real-time human voice audio are obtained.
[0029] More specifically: Real-time human voice audio is acquired, and real-time spectrum analysis is performed using the Fast Fourier Transform method. Simultaneously, real-time pitch analysis is performed using the autocorrelation method, cepstral method, or neural network method. Additionally, real-time consonant and vowel analysis is performed using the zero-crossing rate method, MFCC method, LPC method, or neural network method to determine whether the pitch is normal, thus obtaining the current analysis results. Based on the current analysis results, the following judgments are made: If the current analysis result is a vowel, the energy of the frequency points near the fundamental multiples of the real-time human voice audio in the spectrum is calculated (harmonic energy refers to the energy of the nth harmonic component of the original signal near the frequency f*n, which can be calculated using RMS, Fast Fourier Transform, filtering, etc.). This is used to perform formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and enhance it with a bandpass filter around the highest energy or specified overtone to obtain the overtone result; if the current analysis result is not a vowel, a pass-through process is performed. Based on the results of the previous analysis: if the result of the previous analysis is a vowel, perform Schmidt triggering to obtain the harmonic result in order to maintain the continuous and stable whistle tone; if the result of the previous analysis is not a vowel, perform pass-through processing. Based on the current analysis results, the energy of the fundamental multiples of the frequency points of the real-time human voice audio is calculated to perform formant analysis and iterative analysis on the real-time human voice audio, track the second largest, third largest and smaller overtones (or specified overtones) to obtain the superposition result; Volume characteristics are added to the overtone results, harmonic results, and superposition results to obtain the volume result; Based on the overtone results, harmonic results, superposition results, and volume results, controllable multiple whistle tones corresponding to real-time human voice audio are obtained.
[0030] This process involves calculating the energy of frequencies at multiples of the fundamental frequency in real-time vocal audio to perform formant analysis and track the highest-energy overtones. This includes calculating the energy of harmonics (overtones) in the real-time vocal audio to perform formant analysis and track the highest-energy overtones within key throat singing harmonics. This ensures that the generated throat singing audio's whistle pitch remains within the traditional Eastern tuning system, avoiding pitch deviation, reducing the gap with the actual vocal audio, and improving the realism and fidelity of the throat singing audio. Key throat singing harmonics are harmonics within a frequency range of 50 cents above and below the fundamental frequency, or key throat singing harmonics are... The harmonics of 4, 5, 6, 8, 9, 10, 12, 16, 18, 20, 24, 27, and 32 correspond to 24, 28, 31, 36, 38, 40, 43, 48, 50, 52, 55, 57, and 60 semitones respectively (the number of semitones corresponding to n harmonics is log(n) / log(2)*12), which are pentatonic scale notes of different higher octaves. For example, for the overtone of 6 times the frequency, the overtone corresponds to 31 semitones above the fundamental frequency, which is 24+7 semitones, that is, two octaves plus a fifth. In other words, when the fundamental frequency of the singing is do, this overtone is the high note so, and so on.
[0031] The process of performing Schmitt triggering to obtain harmonic results includes: if the current highest energy harmonic exceeds the previous highest energy harmonic by a preset multiple (how to determine the preset multiple is the existing technology, such as 5 times), harmonic conversion is performed; otherwise, the previous harmonic is maintained, and the harmonic result is obtained.
[0032] It should be noted that the Fast Fourier Transform, Autocorrelation, Cepstral, Neural Network, Zero-Crossing Rate, MFCC, LPC, RMS, and Filtering methods are all existing technologies and will not be elaborated here. Similarly, how to perform pass-through processing, how to incorporate volume characteristics, and how to perform harmonic conversion are all existing technologies and will not be elaborated here.
[0033] In this step, formant analysis is used to track the highest-energy or specified overtone, simulating a very narrow, sharp, and controllable formant formed by the human voice controlling the mouth, tongue, and throat. This formant highlights a specific harmonic overtone, creating the auditory effect of a higher pitch with an additional tone, thus producing a whistle. Simultaneously, harmonic processing is performed to ensure the whistle is continuous and stable. Furthermore, the user can control the whistle pitch by imitating the mouth shape of throat singing, creating a controllable throat singing effect. Further, energy tracking iterations are repeated to track the second-highest, third-highest, and lower-energy overtones (or specified overtones) to obtain multiple whistles. Even further, volume features are added to obtain controllable multiple whistles.
[0034] In addition to automatically following the formant as described above (finding the most suitable harmonic number for frequency band enhancement), it can also be made to manually adjust the harmonics or formant (open parameters for users to manually adjust to track the specified overtones, that is, to allow users to control the most suitable harmonic number for enhancement, and in actual operation, any harmonic can be arbitrarily selected as the most suitable harmonic). It is only necessary to manually select the harmonic multiple or the center frequency of the formant to achieve a manually controlled whistle.
[0035] Step 103: Obtain the throat singing audio corresponding to the real-time human voice audio based on the throat sound and the whistle sound.
[0036] Specifically: Based on the guttural sounds and the whistle sounds, waveforms are superimposed to combine the characteristics of the guttural sounds and the whistle sounds, thereby obtaining the throat singing audio corresponding to the real-time human voice audio.
[0037] In this step, waveform superposition is an existing technique and will not be described in detail here.
[0038] In this embodiment, the input real-time human voice audio is processed to obtain the corresponding throat sounds and the corresponding controllable multiple whistles, and the waveforms are superimposed to obtain a complete, realistic, and controllable throat singing timbre that can be freely controlled by the throat sounds and multiple whistles.
[0039] The aforementioned real-time human voice to throat singing audio processing method can convert ordinary human voices into throat singing characteristics in real time, achieving high naturalness, high real-time performance, and low barrier to entry for throat singing simulation. It fills a market gap and has the following outstanding substantive features and significant advancements: 1) High real-time performance: The speech feature extraction and synthesis algorithm has been optimized and low-latency processing is supported, making it suitable for live performances or real-time interaction; 2) High degree of naturalness: The generated throat singing effect has significant throat resonance and overtone harmonic characteristics, and the listening experience is close to the traditional throat singing style; 3) High ease of use: Users do not need to master special vocal techniques; they can generate expressive throat singing sounds simply by speaking or singing naturally. 4) High adaptability: It can select more than one overtone, so that the number of overtones and the volume ratio of guttural and whistling sounds can be flexibly adjusted when waveforms are superimposed, to adapt to different languages, genders and singing styles; 5) Easy to implement in hardware: The algorithm has low resource consumption and is easy to integrate into various effects processors, sound cards, mobile devices and other embedded platforms, so as to achieve portability and productization; In conclusion, this application provides a modern digital expression method for traditional throat singing, a national intangible cultural heritage, making it easy for ordinary users, teenagers, and international musicians to access and perform throat singing. This effectively lowers the learning threshold of traditional skills and promotes the inheritance and global dissemination of Chinese folk music culture.
[0040] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0041] This application also provides an audio processing device for real-time human voice to throat singing, such as... Figure 2 As shown, in one embodiment, it includes: a first module 201, a second module 202, and a third module 203, wherein: The first module is used to acquire real-time human voice audio, perform octave down-pitch on the single-cycle waveform, and remove the continuity of the sound to obtain the throat sounds corresponding to the real-time human voice audio. The second module is used to acquire real-time human voice audio, and simultaneously perform spectrum analysis, pitch analysis, and consonant and vowel analysis to obtain the current analysis results; based on the current analysis results, the energy of the fundamental multiple frequency points of the real-time human voice audio is calculated to perform formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the whistle tone corresponding to the real-time human voice audio. The third module is used to obtain the throat singing audio corresponding to the real-time human voice audio based on the throat sounds and the whistle sounds.
[0042] For specific limitations regarding the audio processing device for real-time human voice to throat singing, please refer to the limitations of the audio processing method for real-time human voice to throat singing mentioned above, which will not be repeated here. Each module in the above device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0043] This application also provides an audio device (e.g., an effects processor), including: an audio input port, an AD converter, a main chip, a DA converter, and an audio output port; The audio input ports are connected to external devices and AD converters respectively, and are used for audio acquisition, that is, sending the input audio obtained from external devices (e.g., microphones) to AD converters; The AD converter is connected to the audio input port and the main chip respectively. It is used to perform analog-to-digital conversion on the input audio sent from the audio input port to obtain real-time human voice audio, and then send it to the main chip. The main chip is connected to both an AD converter and a DA converter, and includes a real-time human voice to shouting audio processing device, which is used to perform shouting audio conversion on the real-time human voice audio sent by the AD converter to obtain shouting audio, and then send it to the DA converter. The DA converter is connected to the main chip and the audio output port respectively. It is used to perform digital-to-analog conversion on the throat singing audio sent by the main chip to obtain the output audio and send it to the audio output port. The audio output port is connected to both the DA converter and an external device, and is used to output the audio sent by the DA converter to the external device (e.g., a speaker).
[0044] In one embodiment, a computer device is provided, which may be a terminal. The computer device includes a processor, memory, a network interface, a display screen, and an input device connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a real-time human voice-to-teleprompting audio processing method. The display screen may be a liquid crystal display (LCD) or an e-ink display. The input device may be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.
[0045] Those skilled in the art will understand that the above structure is merely a description of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. A specific computer device may include more or fewer components than those shown in the figures, or may combine certain components, or may have different component arrangements.
[0046] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the method described above.
[0047] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0048] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0049] The contents not described in detail in this specification are existing technologies known to those skilled in the art.
[0050] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0051] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended application documents.
Claims
1. A real-time audio processing method for converting human voice to throat singing, characterized in that, include: The real-time human voice audio is acquired, the single-cycle waveform is down-octave, and the continuity of the sound is removed to obtain the throat sounds corresponding to the real-time human voice audio. Real-time human voice audio is acquired, and spectrum analysis, pitch analysis, and consonant and vowel analysis are performed simultaneously to obtain the current analysis results. Based on the current analysis results, the energy of the fundamental multiples of the real-time human voice audio is calculated to perform formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the whistle tone corresponding to the real-time human voice audio. Based on the throat sounds and the whistle sounds, the throat singing audio corresponding to the real-time human voice audio is obtained.
2. The audio processing method for real-time human voice to throat singing according to claim 1, characterized in that, Based on the current analysis results, energy calculation is performed on the energy of the fundamental multiples of the real-time human voice audio to conduct formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the whistle tone corresponding to the real-time human voice audio, including: Based on the current analysis results, the energy of the fundamental multiples of the frequency points of the real-time human voice audio is calculated to perform formant analysis on the real-time human voice audio, track the overtone with the largest energy or the specified overtone, and obtain the overtone result; Based on the previous analysis results, the real-time human voice audio was subjected to harmonic processing to obtain harmonic results; Based on the overtone and harmonic results, the whistle tone corresponding to the real-time human voice audio is obtained.
3. The audio processing method for real-time human voice to throat singing according to claim 1, characterized in that, Based on the current analysis results, energy calculation is performed on the energy of the fundamental multiples of the real-time human voice audio to conduct formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the whistle tone corresponding to the real-time human voice audio, including: Based on the current analysis results, the energy of the fundamental multiples of the frequency points of the real-time human voice audio is calculated to perform formant analysis on the real-time human voice audio, track the overtone with the largest energy or the specified overtone, and obtain the overtone result; Based on the previous analysis results, the real-time human voice audio was subjected to harmonic processing to obtain harmonic results; Based on the current analysis results, the energy of the fundamental multiple frequency points of the real-time human voice audio is calculated to perform formant analysis iteration on the real-time human voice audio, track the second largest or specified overtone, and obtain the superposition result; Based on the overtone results, harmonic results, and superposition results, multiple whistle tones corresponding to real-time human voice audio are obtained.
4. The audio processing method for real-time human voice to throat singing according to claim 1, characterized in that, Based on the current analysis results, energy calculation is performed on the energy of the fundamental multiples of the real-time human voice audio to conduct formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the whistle tone corresponding to the real-time human voice audio, including: Based on the current analysis results, the energy of the fundamental multiples of the frequency points of the real-time human voice audio is calculated to perform formant analysis on the real-time human voice audio, track the overtone with the largest energy or the specified overtone, and obtain the overtone result; Based on the previous analysis results, the real-time human voice audio was subjected to harmonic processing to obtain harmonic results; Based on the current analysis results, the energy of the fundamental multiple frequency points of the real-time human voice audio is calculated to perform formant analysis iteration on the real-time human voice audio, track the second largest or specified overtone, and obtain the superposition result; Volume characteristics are added to the overtone results, harmonic results, and superposition results to obtain the volume result; Based on the overtone results, harmonic results, superposition results, and volume results, controllable multiple whistle tones corresponding to real-time human voice audio are obtained.
5. The real-time human voice to throat singing audio processing method according to any one of claims 2 to 4, characterized in that, Based on the current analysis results, energy calculation is performed on the energy of the fundamental multiples of the real-time human voice audio to conduct formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the overtone results, including: Based on the current analysis results, the following judgment should be made: If the current analysis result is a vowel, the energy of the fundamental multiple frequency point of the real-time human voice audio near the spectrum is calculated to perform formant analysis on the real-time human voice audio, track the overtone with the largest energy or the specified overtone, and perform bandpass filter enhancement around the overtone with the largest energy or the specified overtone to obtain the overtone result. If the current analysis result is not a vowel, the real-time human voice audio is passed through.
6. The real-time human voice to throat singing audio processing method according to any one of claims 2 to 4, characterized in that, Based on the previous analysis, harmonic processing is performed on the real-time human voice audio to obtain harmonic results, including: Based on the results of the previous analysis: If the result of the previous analysis is a vowel, a Schmidt trigger is performed to obtain the harmonic result in order to maintain the continuous and stable whistle tone; If the result of the previous analysis is not a vowel, perform a direct pass-through.
7. The real-time human voice to throat singing audio processing method according to claim 6, characterized in that, Schmitt triggering is performed to obtain harmonic results, including: If the current highest-energy harmonic exceeds the previous highest-energy harmonic by a preset multiple, harmonic conversion is performed; otherwise, the previous harmonic is maintained, and the harmonic result is obtained.
8. The audio processing method for real-time human voice to throat singing according to any one of claims 1 to 4, characterized in that, Based on the guttural sounds and the whistling sounds, the throat singing audio corresponding to the real-time human voice audio is obtained, including: Based on the guttural sounds and the whistle sounds, waveforms are superimposed to combine the characteristics of the guttural sounds and the whistle sounds, thereby obtaining the throat singing audio corresponding to the real-time human voice audio.
9. An audio processing device for real-time human voice to throat singing, characterized in that, include: The first module is used to acquire real-time human voice audio, perform octave down-pitch on the single-cycle waveform, and remove the continuity of the sound to obtain the throat sounds corresponding to the real-time human voice audio. The second module is used to acquire real-time human voice audio, and simultaneously perform spectrum analysis, pitch analysis, and consonant and vowel analysis to obtain the current analysis results; based on the current analysis results, the energy of the fundamental multiple frequency points of the real-time human voice audio is calculated to perform formant analysis on the real-time human voice audio, track the highest energy or specified overtone, and obtain the whistle tone corresponding to the real-time human voice audio. The third module is used to obtain the throat singing audio corresponding to the real-time human voice audio based on the throat sounds and the whistle sounds.
10. An audio device, characterized in that, It includes, in sequence: audio input port, AD converter, main chip, DA converter and audio output port; The audio input port is used to send the acquired input audio to the AD converter; The AD converter is used to perform analog-to-digital conversion on the input audio to obtain real-time human voice audio; The main chip includes a real-time human voice to vocalization audio processing device, used to convert real-time human voice audio into vocalization audio. The DA converter is used to perform digital-to-analog conversion on the throat singing audio to obtain the output audio. The audio output port is used to output audio to external devices.
Citation Information
Patent Citations
Method for extracting and utilizing audio features for repairing Chinese national folk music audios
CN102842310A
Method and device for beautifying sound
CN109410971A
Method and device for correcting pitch of audio
CN115331682A
Sound modification method and device, electronic equipment and storage medium
CN116504257A
Intelligent speech synthesis system based on electronic artificial throat audio input signal
CN120412534A