A human voice spectrum protection intelligent mixing method and system

By constructing a human voice spectrum protection intelligent mixing system, the problems of human voice masking, fixed parameters, terminal compatibility and simple audio control logic in existing technologies are solved. It achieves improved human voice clarity, enhanced semantic recognition and reduced auditory fatigue, and is suitable for multiple terminal devices.

CN122369485APending Publication Date: 2026-07-10SHANGHAI CHONGXIANSHENG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI CHONGXIANSHENG TECHNOLOGY CO LTD
Filing Date
2026-04-30
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing audio mixing technologies suffer from problems such as human voice spectrum masking, fixed and non-adaptive parameters, lack of protection for human voice-specific frequency bands, limited application scenarios, incompatibility with multiple terminals, and simplistic audio control logic. These issues lead to blurred human voices, decreased semantic recognition, auditory fatigue, and poor user experience.

Method used

The intelligent mixing method with human voice spectrum protection is adopted. Through speech analysis, multi-dimensional state evaluation and adaptive control, a multi-terminal compatible intelligent mixing system is constructed to achieve human voice frequency band protection, dynamic audio parameter adjustment and environmental noise reduction. Combined with parameters such as speech rate, pauses and listening time, the audio parameters are adaptively adjusted to isolate background audio interference.

Benefits of technology

It improves voice clarity, enhances semantic recognition, reduces auditory fatigue, adapts to various terminal devices, and improves listening efficiency and experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369485A_ABST
    Figure CN122369485A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for intelligent mixing with human voice spectrum protection, belonging to the field of digital audio signal processing technology. The method extracts speech semantic weights using a lightweight NLP model, combines speech rate, pause duration, continuous listening duration, and environmental noise parameters, and calculates a real-time listening status score using a nonlinear mapping function. Based on the score, it adaptively adjusts the background binaural beat audio signal parameters, locking the 250Hz-4kHz core human voice frequency band. Gain boosting, notch filtering, and dynamic compensation are used to achieve human voice spectrum isolation protection, completing the intelligent mixing output. The system includes modules for speech parsing, status acquisition, nonlinear comprehensive evaluation, adaptive control, intelligent mixing engine, audio playback, and environmental noise acquisition. It is compatible with multiple terminals such as learning apps, online classes, smart headphones, and hearing aids. This invention effectively solves the problem of human voice spectrum masking, improves voice clarity and auditory experience, adjusts only the audio signal frequency, avoids patent infringement risks, and offers a complete and practical technical solution suitable for various intelligent audio mixing scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical fields of digital audio signal processing, intelligent audio algorithms, and intelligent wearable devices. Specifically, it relates to an intelligent audio processing method and system based on human voice core frequency band protection, adaptive filtering and mixing, dynamic adjustment of binaural beat audio signals, and multi-terminal compatibility. It can be applied to software and hardware products such as audiobook reading software, online education platforms, voice live streaming, smart Bluetooth headphones, hearing aids, and audio playback terminals. Background Technology

[0002] Currently, most audio accompaniment software, voice playback systems, and audio mixing devices on the market use a mixing method that directly and linearly superimposes human voices with background music and ambient audio. This method has obvious technical flaws in actual use. 1. The human voice spectrum is severely masked, and the background audio overlaps with the frequency band of the human voice, resulting in blurred human voices, reduced semantic recognition, and easy auditory fatigue after prolonged listening. 2. The audio parameters are fixed and cannot be adapted. The background sound effects and volume ratio are all preset fixed values ​​and cannot be dynamically adjusted according to the voice status, user listening time, and ambient noise. 3. Lacking a dedicated frequency band protection mechanism for human voices, it is impossible to isolate the effective frequency band of human voices, and external noise and background noise can easily interfere with the human voice signal; 4. Existing technologies have limited application scenarios, mostly only compatible with mobile software and unable to be deployed in an integrated manner with hardware terminals such as headphones and hearing aids; 5. The audio control logic is simple and crude, relying solely on fixed volume increases and decreases, lacking scientific status evaluation and adaptive adjustment logic, resulting in a poor user experience. In summary, existing audio mixing technologies cannot simultaneously solve problems such as protecting human voice clarity, dynamic adaptive mixing, environmental noise reduction, multi-terminal adaptation, and protection against auditory fatigue. A brand-new intelligent mixing solution is urgently needed. Summary of the Invention

[0003] Purpose of the invention This invention aims to overcome the shortcomings of existing technologies and provide a human voice spectrum protection intelligent mixing method and system to achieve the following technical objectives: 1. Dedicated protection is provided for the core frequency band of human voice to prevent background noise and other disturbances from masking the human voice spectrum, ensuring that the human voice remains clear at all times; 2. Construct a multi-dimensional speech state evaluation model that combines speech rate, pauses, listening duration, and semantic weights to judge the user's listening state in real time and adaptively adjust audio parameters; 3. It achieves fully automatic dynamic control of background binaural beat audio signals, filter parameters, and gain, adjusting only the audio signal frequency without directly affecting the human brain; 4. Simultaneously compatible with deployment on multiple terminals such as software apps, live streaming platforms, smart headphones, and hearing aids; 5. It features noise reduction, ear protection, voice enhancement, and intelligent mixing functions to improve the audio experience, reduce auditory fatigue, and enhance learning and listening efficiency. IV. Technical Solution 1) Overall System Architecture The system of this invention consists of two parts: a software algorithm system and a hardware terminal system, as detailed in the appendix. Figure 1 Appendix Figure 2 . Software modules (attached) Figure 1 ) 1. Speech parsing module 2. Status Acquisition Module 3. Nonlinear Comprehensive Evaluation Module 4. Adaptive control module 5. Intelligent Mixing Engine Module 6. Audio playback module 7. Environmental Noise Acquisition Module Hardware module (attached) Figure 2 (Headphones / Hearing Aids) 8. Microphone Array 9. DSP Digital Signal Processing Chip 10. Bluetooth transmission module 11. Audio power amplifier module 12. Loudspeaker Signal flow between modules: The main workflow on the software side flows in one direction, while the adaptive control module and the intelligent mixing engine module interact bidirectionally via a double arrow. On the hardware side, the main signal flows unidirectionally, while the interaction between the DSP chip and the Bluetooth transmission module is bidirectional. All other links are unidirectional arrows, conforming to the timing execution logic. II) Methods and Steps S1. Speech Signal Input and Parsing: The system receives TTS synthesized speech or real-time human speech signals. The speech parsing module performs frame-by-frame and segment-by-segment parsing on the input speech. The lightweight natural language processing (NLP) model identifies keywords (such as "attention", "emphasis", "core", etc.) and the importance of sentences in the speech. Combining punctuation marks (exclamation marks, commas, emphasis marks) to assign corresponding semantic weight values, the semantic weight is extracted, and the human voice spectrum features are extracted at the same time. S2. Multi-dimensional status parameter acquisition: The status acquisition module acquires speech rate, sentence pause duration, and continuous listening duration parameters in real time, while the environmental noise acquisition module simultaneously acquires the surrounding environmental noise intensity and spectrum distribution parameters. S3. Status score calculation: After normalizing the four parameters of speech rate, pause duration, semantic weight, and continuous listening duration, the data is input into the nonlinear comprehensive evaluation module, and a real-time listening status score of 0 to 100 is calculated through the built-in S-shaped nonlinear mapping function. S4. Audio instruction generation: The adaptive control module generates frequency adjustment instructions, gain control instructions, and filter trigger instructions for the binaural beat audio signal based on the real-time status score. S5, Vocal Spectrum Protection Mixing: The intelligent mixing engine module locks the core frequency band of human voice from 250Hz to 4kHz, and performs human voice gain enhancement, background audio notch filtering, human voice dynamic compensation, and environmental noise reduction processing to achieve isolation and protection between the human voice spectrum and the background audio spectrum. S6. Audio Output: The processed mixed audio signal is output through the audio playback module, completing the full intelligent mixing process. III) Core Technical Parameters 1. Voice protection frequency band: 250Hz~4kHz; 2. Human voice gain relative to background noise: 6dB~8dB, preferably 7dB; 3. Background audio frequency band attenuation: 6dB~10dB, preferably 8dB; 4. Vocal dynamic compensation gain: 1dB~3dB, preferably 2dB; 5. Nonlinear evaluation calculation formula: Score = 100 \times \frac{1}{1+e^{-(0.3V+0.25P+0.3W+0.15T-2.5)}} In the formula: V is the normalized value of speech rate, P is the normalized value of pause duration, W is the normalized value of semantic weight, and T is the normalized value of continuous listening duration. 6. Binaural beat audio signal frequency gradation adjustment: 10Hz for high focus state, 8Hz for steady state, and 6Hz for fatigue state, with a smooth frequency change rate of 0.5Hz / minute. It only adjusts the frequency components of the audio itself and does not directly affect the human brain or interfere with the human physiological state. 7. Hardware terminal noise reduction range: 3dB~9dB, preferably 6dB, with an additional 1dB~2dB boost for human voices when they appear. V. Specific Implementation Methods Example 1: Software App Reading Scenario Applied to various audiobook and learning apps, the system parses text to synthesize speech, extracts semantic weights through an NLP model, collects listening state parameters in real time, calculates state scores, adaptively adjusts the background binaural beat audio signal, and performs protective mixing on the human voice frequency band, ensuring clear and interference-free human voices and reducing fatigue during long-term listening. Example 2: Online live streaming and online class audio scenarios It adapts to real-time human voice, performs semantic weight recognition on the lecturer's voice, automatically assigns high weight to key teaching content, simultaneously strengthens the level of human voice protection, isolates background noise, improves the semantic recognition of the lecture, and is suitable for use in meeting and online class scenarios. Example 3: Application Scenarios of Smart Headphones The headphones have a complete set of built-in algorithms, a microphone array to collect ambient sound, and a DSP chip to run mixing and noise reduction logic. They also enable two-way parameter interaction with a mobile app via Bluetooth, achieving voice enhancement, environmental noise reduction, and ear-protecting mixed output. Example 4: Hearing Aid Application Scenarios Based on the headphone hardware, the microphone pickup sensitivity has been improved, the gain compensation for weak human voices at close range has been increased, environmental noise has been filtered out, and the protection of human voice frequency bands has been strengthened. It is suitable for daily use by people with hearing loss, and achieves clear hearing aids and ear protection noise reduction. Example 5: Comparative Experiment Verification A control group and an experimental group were set up. The control group used a conventional linear superposition mixing scheme, while the experimental group used the intelligent mixing scheme of this invention. The voice clarity was tested in a noisy environment with a low signal-to-noise ratio of 5dB. - Control group: Voice clarity score 62 points, semantic recognition accuracy 71%; - Experimental group: Voice clarity score 89 points, semantic recognition accuracy 94%; Experimental results show that after adopting the technical solution of the present invention, the clarity of human voice is improved by 27 points in low signal-to-noise ratio environments, and the semantic recognition accuracy is improved by 23%, demonstrating significant technical effects. VI. Description of the attached drawings Appendix Figure 1 : Flowchart of the software module of the intelligent mixing system for human voice spectrum protection of this invention; Appendix Figure 2 This invention relates to a hardware terminal structure diagram for headphones and hearing aids. Figure Labels 1-Speech parsing module; 2-Status acquisition module; 3-Nonlinear comprehensive evaluation module; 4-Adaptive control module; 5-Intelligent mixing engine module; 6-Audio playback module; 7-Ambient noise acquisition module; 8-Microphone array; 9-DSP digital signal processing chip; 10-Bluetooth transmission module; 11-Audio power amplifier module; 12-Speaker. VII. Differences and Innovations of Existing Technologies 1. For the first time, a dedicated human voice frequency band protection mechanism has been established to isolate background noise interference at the spectrum level, fundamentally solving the problem of human voice masking, and significantly improving the clarity of human voice; 2. Supplement the semantic weight quantification extraction logic, and realize automatic machine recognition and assignment through keywords and punctuation marks. The technical solution is complete and without loopholes, avoiding doubts about insufficient review and disclosure; 3. Clearly define the binaural beat audio signal modulation as the modulation of the audio's own frequency, rather than a direct effect on the human brain, thus completely avoiding the risk of patent subject matter examination; 4. A multidimensional nonlinear state assessment model is adopted, with clear parameter quantification and reproducible algorithm. The effectiveness of the technology is further supported by comparative experimental data. 5. The integrated hardware and software design covers both software algorithms and wearable hardware terminals, enabling a wide range of applications and comprehensive patent protection.

Claims

1. A human voice spectrum protection intelligent mixing method, characterized in that, Includes the following steps: The input speech is parsed using a lightweight NLP model, and semantic weights are extracted based on keywords and punctuation marks. Simultaneously, human voice spectral features are extracted. Speech rate, pause duration, continuous listening duration, and environmental noise parameters are collected. These parameters are normalized and input into a nonlinear evaluation model, which outputs a real-time listening status score. Based on the status score, audio control commands related to the frequency, gain, and filtering of the binaural beat audio signal are generated. Gain protection, notch filtering, and noise reduction mixing are applied to the 250Hz~4kHz core human voice frequency band. Finally, the mixed audio is output.

2. The method according to claim 1, characterized in that, The nonlinear evaluation uses a sigmoid mapping function to calculate the state score, and the formula is as follows: Score = 100 \times \frac{1}{1+e^{-(0.3V+0.25P+0.3W+0.15T-2.5)}} The parameters in the formula are speech rate, pause duration, semantic weight, and normalized value of continuous listening duration, respectively.

3. The method according to claim 1, characterized in that, The human voice frequency band has a gain of 6dB~8dB relative to the background audio, and the corresponding background sound frequency band has a attenuation of 6dB~10dB. The binaural beat audio signal is dynamically adjusted in three levels: 10Hz, 8Hz, and 6Hz.

4. A human voice spectrum protection intelligent mixing system, characterized in that, include: The system includes a speech parsing module, a status acquisition module, a nonlinear comprehensive evaluation module, an adaptive control module, an intelligent mixing engine module, an audio playback module, and an environmental noise acquisition module.

5. The system according to claim 4, characterized in that, The system can be integrated into smart headphones and hearing aid hardware terminals. The hardware includes a microphone array, DSP chip, Bluetooth transmission module, audio power amplifier module, and speaker.