Voice enhancement method and device based on noise perception, equipment and medium

By combining ambient audio and multimodal sensor data in speech enhancement technology, generating personalized enhanced audio signals, and performing playback feedback adjustments, the real-time and personalized timbre issues of speech enhancement in dynamic noise environments are solved, the environmental adaptability and personalized fidelity of speech signals are improved, and the user experience is improved.

CN120673773APending Publication Date: 2025-09-19PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510826685.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing speech enhancement technologies cannot simultaneously take into account the real-time performance, personalized timbre fidelity, and environmental adaptability of speech enhancement in dynamic noise environments. Especially in the fields of financial technology and healthcare, there are problems such as unclear voice quality, misjudgment, and poor interactive experience.

Method used

By acquiring the ambient audio data and multimodal sensor data of the target audio and its environment, the noise perception network is used to generate environmental feature information. Combined with the audio enhancement model and speaker feature extraction module, personalized enhanced audio signals are generated. Playback feedback data is collected for parameter adjustment to achieve adaptive optimization in the time and frequency domains.

Benefits of technology

It improves the adaptability and personalized fidelity of voice signals in complex environments, improves voice clarity, naturalness and style matching, and improves user experience and interaction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673773A_ABST
    Figure CN120673773A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice enhancement method, device and equipment based on noise perception and a medium. Environment feature information is extracted and input into an audio enhancement model to generate an enhanced audio signal; obtaining a reference audio sample, extracting a personalized feature vector, and carrying out personalized processing on the enhanced audio signal; and collecting playing feedback data, determining a playing time domain adjustment parameter and a playing frequency domain adjustment parameter, adjusting the personalized enhanced audio signal, and generating an optimized audio signal. According to the method, dynamic adjustment is realized in combination with the feedback parameters in the playing process by fusing the environmental perception information and the personalized speaker characteristics, clear and natural optimized audio output with personalized styles can be generated in a complex environment, and the voice interaction quality and adaptability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing technology, and in particular to a speech enhancement method, apparatus, device and storage medium based on noise perception. Background Art

[0002] Against the backdrop of the continuous development of speech enhancement technology, existing technologies still have significant limitations when dealing with speech quality issues in dynamic noisy environments, especially in multi-source heterogeneous environment modeling, personalized speech style fusion, and output adaptive optimization.

[0003] In the fintech sector, voice interaction systems are widely used in scenarios such as telephone customer service and voice contract confirmation. However, current voice enhancement solutions typically model only the audio signal itself, lacking integration of other perceptual information in the target audio acquisition environment (such as the ambient visual state or device motion state). This results in an inability to accurately model the environmental context in complex transaction scenarios (such as business halls or multi-terminal concurrent call scenarios), affecting the enhancement model's accuracy in determining noise characteristics, leading to issues such as unclear voice information and misjudgment.

[0004] In the healthcare sector, voice applications are widely used in highly sensitive scenarios such as remote consultations and medical record recording, placing extremely high demands on speech clarity and individual recognition accuracy. Current speaker adaptation methods rely on implicit modeling of timbre features using deep neural networks, but fail to effectively incorporate the spectral structure or resonance features of the reference audio, which are inherently valuable. This results in the personalized voice style of patients or doctors being easily erased during the enhancement process, reducing the consistency and intelligibility of speech expression.

[0005] Furthermore, current mainstream speech enhancement methods often overlook the dynamic role of user feedback during playback. They lack integrated mechanisms for sensing and providing feedback on time and frequency domain deviations during playback, making it difficult to dynamically adjust parameters based on the user's actual hearing experience and environmental conditions. In financial phone interactions or medical speech playback scenarios, systems lacking playback feedback are prone to problems such as abnormal speech speed, unnatural pauses, and frequency imbalance, impacting the interactive experience.

[0006] At the same time, existing models are usually trained and inferred based on static data input, lacking a mechanism for collaborative perception from multiple dimensions such as ambient audio, multimodal sensors, reference voice, and playback feedback, and are unable to provide end-to-end consistency optimization. Especially on mobile devices or in resource-constrained scenarios, there is a lack of a structurally complete but computationally lightweight multi-input fusion processing path, making it difficult to strike a balance between efficiency and quality.

[0007] Therefore, current technologies have shortcomings in the comprehensiveness of dynamic environment modeling, the method of injecting personalized speaker information, and the ability to adaptively adjust the playback quality after enhancement. They are unable to meet the needs of fields such as finance and medical care for high-quality, robust, and consistent speech processing. Summary of the Invention

[0008] The main purpose of the present invention is to provide a speech enhancement method, device, equipment and storage medium based on noise perception, aiming to solve the technical problem that the existing technology cannot simultaneously take into account the real-time performance of speech enhancement, personalized timbre fidelity and environmental adaptability in dynamic noise environments.

[0009] To achieve the above object, the present invention provides a speech enhancement method based on noise perception, comprising:

[0010] Acquire target audio to be processed, and collect ambient audio data and multimodal sensor data of an environment in which the target audio is located;

[0011] Inputting the environmental audio data and the multimodal sensor data into a preset noise perception network to generate environmental feature information;

[0012] Inputting the environmental feature information and the target audio into an audio enhancement model to generate an enhanced audio signal;

[0013] Obtaining a reference audio sample, and processing the reference audio sample through a speaker feature extraction module to generate a personalized feature vector;

[0014] Processing the enhanced audio signal based on the personalized feature vector to generate a personalized enhanced audio signal;

[0015] Collecting feedback data when the personalized enhanced audio signal is played, and determining a playback time domain adjustment parameter and a playback frequency domain adjustment parameter according to the feedback data;

[0016] The time domain characteristics and frequency domain characteristics of the personalized enhanced audio signal are adjusted according to the playback time domain adjustment parameter and the playback frequency domain adjustment parameter to generate an optimized audio signal.

[0017] Furthermore, to achieve the above-mentioned object, the present invention provides a speech enhancement device based on noise perception, comprising:

[0018] A target audio acquisition module is used to acquire target audio to be processed and collect ambient audio data and multimodal sensor data of the environment in which the target audio is located;

[0019] a noise perception module, configured to input the ambient audio data and the multimodal sensor data into a preset noise perception network to generate environmental feature information;

[0020] an audio enhancement module, configured to input the environmental feature information and the target audio into an audio enhancement model to generate an enhanced audio signal;

[0021] A speaker feature extraction module is used to obtain a reference audio sample and process the reference audio sample through the speaker feature extraction module to generate a personalized feature vector;

[0022] a personalized modulation module, configured to process the enhanced audio signal based on the personalized feature vector to generate a personalized enhanced audio signal;

[0023] a playback feedback analysis module, configured to collect feedback data when the personalized enhanced audio signal is played, and determine playback time domain adjustment parameters and playback frequency domain adjustment parameters based on the feedback data;

[0024] The signal optimization module is used to adjust the time domain characteristics and frequency domain characteristics of the personalized enhanced audio signal according to the playback time domain adjustment parameter and the playback frequency domain adjustment parameter to generate an optimized audio signal.

[0025] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and a noise-awareness-based speech enhancement program stored in the memory and executable on the processor. When the noise-awareness-based speech enhancement program is executed by the processor, the steps of the noise-awareness-based speech enhancement method described above are implemented.

[0026] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a noise-perception-based speech enhancement program is stored. When the noise-perception-based speech enhancement program is executed by a processor, the steps of the noise-perception-based speech enhancement method described above are implemented.

[0027] Beneficial effects: The present invention relates to the field of speech processing technology and can be applied to business scenarios such as financial technology and medical health. It discloses a speech enhancement method, device, equipment and medium based on noise perception, including: obtaining environmental audio data and multimodal sensor data of the target audio to be processed and its environment; generating environmental feature information based on a noise perception network; inputting the environmental feature information and the target audio into an audio enhancement model to generate an enhanced audio signal; generating a personalized feature vector by processing a reference audio sample, and using it to enhance the personalized processing of the audio signal to generate a personalized enhanced audio signal; collecting feedback data during the playback of the personalized enhanced audio signal, and determining the playback time domain adjustment parameters and playback frequency domain adjustment parameters based on the feedback data; adjusting the time domain characteristics and frequency domain characteristics of the personalized enhanced audio signal according to the parameters to generate an optimized audio signal. The present invention combines multimodal environmental perception with personalized feature modeling to construct an audio enhancement mechanism for dynamic changes in the environment and individual expression characteristics, and further combines playback feedback to achieve adaptive adjustment of playback parameters, thereby improving the environmental adaptability and individual fidelity of the final audio signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0029] Figure 1 Schematic diagram of an application environment of a noise perception-based speech enhancement method according to an embodiment of the present invention;

[0030] Figure 2 This is a flow chart of an embodiment of a method for speech enhancement based on noise perception according to the present invention;

[0031] Figure 3 Schematic diagram of functional modules of a preferred embodiment of a speech enhancement device based on noise perception according to the present invention;

[0032] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0033] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0034] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0035] The noise perception-based speech enhancement method provided by the embodiment of the present invention can be applied in Figure 1In an application environment, the user end communicates with the server end through a network. The server end can obtain the target audio to be processed and the environmental audio data and multimodal sensor data of its environment through the user end; generate environmental feature information based on the noise perception network; input the environmental feature information and the target audio into the audio enhancement model to generate an enhanced audio signal; generate a personalized feature vector by processing the reference audio sample, and use it to enhance the personalized processing of the audio signal to generate a personalized enhanced audio signal; collect feedback data when the personalized enhanced audio signal is played, and determine the playback time domain adjustment parameters and playback frequency domain adjustment parameters based on the feedback data; adjust the time domain features and frequency domain features of the personalized enhanced audio signal according to the parameters to generate an optimized audio signal. The present invention combines multimodal environmental perception with personalized feature modeling to construct an audio enhancement mechanism oriented to dynamic changes in the environment and individual expression characteristics, and further combines playback feedback to achieve adaptive adjustment of playback parameters, thereby improving the environmental adaptability and personal fidelity of the final audio signal. The user end can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server end can be implemented using an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0036] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of a noise-aware speech enhancement method provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0037] like Figure 2 As shown, the noise perception-based speech enhancement method proposed in the present invention includes the following steps:

[0038] S10, obtaining target audio to be processed, and collecting ambient audio data and multimodal sensor data of an environment in which the target audio is located;

[0039] In this embodiment, obtaining the target audio to be processed generally refers to collecting the data source audio for enhanced processing under the trigger condition that the target user initiates voice interaction or voice input. The target audio can come from the voice signal actively input by the user, or it can be the voice data sent by the system end, such as the voice clip to be played generated in the speech synthesis task. The input end in this process may include a single-channel microphone, an array microphone, a voice chip or a built-in acquisition device, and its triggering mechanism can be based on the voice interaction instructions of the application layer, or it can determine the activation window by the sudden change of sound intensity. During the audio acquisition process, the integrity and high-fidelity characteristics of the original signal need to be maintained. Therefore, it is recommended to set an adaptive sampling rate, such as 16kHz or 48kHz, and a high-order signal-to-noise ratio acquisition module to ensure that the time domain resolution and frequency domain coverage meet the subsequent processing requirements.

[0040] When collecting ambient audio data of the target audio environment, the focus is on obtaining non-speech components in the spatial background sound, such as steady-state background noise, sudden interference, reverberation echoes, etc. The collection of ambient audio data can be completed by microphone channels independent of the main channel, such as edge microphone nodes in the microphone array. At the same time, the spatial layout should consider the microphone spacing and pickup direction to support subsequent spatial sound source separation and noise modeling. Ambient audio data should be synchronized and complete. It is recommended to combine a timestamp management mechanism to ensure consistency with the target audio in terms of collection starting point, frequency, and duration.

[0041] The acquisition of multimodal sensor data addresses the incomplete environmental perception caused by relying solely on audio signals. In complex scenarios, identifying sound sources, predicting noise interference, and assisting in determining user behavior require the integration of multiple non-audio data sources. Multimodal sensors may include an inertial measurement unit (IMU), image sensors, temperature and humidity sensors, and distance detection modules. The IMU uses accelerometers and gyroscopes to capture the device's three-dimensional motion state, helping to determine whether the user is walking, driving, or stationary. Image sensors can be used to identify multiple speakers, strong light, or obstructions in a scene, as well as analyze mouth shape and facial movement trends to support subsequent speech generation and adaptation. This multimodal data must be centrally managed through a fusion time synchronization module to ensure that data from different sensor sources is aligned within the same time reference system. This supports real-time perception of spatial dynamics in subsequent environmental modeling tasks.

[0042] In practical implementation, specific data preprocessing paths must be established for multimodal sensor data. For example, IMU data requires drift removal, low-pass filtering, and attitude resolution; image sensor data requires frame sampling, image denoising, and optical flow tracking to ensure input data quality and consistency with subsequent computational models. Furthermore, all signals should be structured and packaged in a unified format, such as tensors, multidimensional matrices, and time series containers, to support the parallel input requirements of model modules.

[0043] After acquiring audio and environmental information, these signals should be input into subsequent perception networks or enhancement models based on control logic. At this point, format adaptation specifications, such as input dimensions, channel order, and sampling step size, must be defined through the input interface to prevent data loss or dimensional misalignment during transmission. Furthermore, to improve the stability of system responses, it is recommended to set an input buffering strategy, such as using a sliding window to locally cache input data to ensure input integrity in extreme scenarios.

[0044] By integrating a MEMS microphone array in a mobile device, it is possible to simultaneously capture target and ambient audio, and combine it with an IMU module and a front-facing camera to capture motion data and visual scene information. In-vehicle terminals can also have dual microphone channels, one facing the user for voice input and the other facing the window for ambient sound collection. Driving status detection sensors can also be integrated to obtain driving speed and steering angle data. Smart wearable devices can also incorporate a bone conduction microphone to capture clear speech, while integrating it with an infrared depth camera to obtain visual parameters such as the user's facial distance and mouth opening, supplementing the high-frequency sound energy that may be missing from the speech signal.

[0045] Example description: In the field of medical health, it can be used for doctors wearing smart head-mounted devices during diagnosis and treatment. When there are noise sources such as suction machines, electric beds, and monitors in the background, the doctor's voice and ambient sound are collected in real time. Combined with motion sensing and visual images, it is determined whether the doctor is checking on the patient or entering information. It supports the generation of voice records that accurately restore the doctor's voice style, improving the accuracy and efficiency of remote diagnosis and treatment and automatic archiving of medical record systems.

[0046] In the financial field, it can be applied to intelligent customer service scenarios at the counter. When a user submits a voice service request in an open hall, the system separates the user's voice from the surrounding conversation noise, cooperates with the camera to obtain the user's current facial posture and emotional state, and performs style enhancement and environmental adaptation processing on the customer's input voice, thereby generating clear voice data with personalized recognition features for subsequent customer identity authentication and voice file storage, thereby improving user experience and security.

[0047] This embodiment achieves comprehensive modeling of the acoustic environment and user status in the current voice interaction scenario through the synchronous collection of target audio, ambient audio and multimodal data. This not only improves the adaptability of subsequent enhancement modules to real noise scenes, but also provides an input basis for personalized modeling and dynamic adjustment, thereby improving the performance of generated speech in terms of clarity, naturalness and style matching.

[0048] S20, inputting the environmental audio data and the multimodal sensor data into a preset noise perception network to generate environmental feature information;

[0049] In this embodiment, ambient audio data refers to background sound components other than the target voice, recorded simultaneously at the target audio collection site. These typically include continuous or intermittent sound sources such as air conditioning, traffic, crowd noise, and electrical interference. These audio signals often have a spectral structure that is a mixture of steady-state and bursty noises, and they change dynamically over time and spatial position, making them difficult to directly model using traditional filtering methods. To enhance the accuracy of noise perception, in addition to the ambient audio itself, multimodal sensor data must be collected as auxiliary input to reflect the dynamic information and non-acoustic characteristics of the spatial environment.

[0050] Multimodal sensor data typically includes, but is not limited to, the device's inertial motion state, visual image features, and scene illumination metrics. Inertial sensors (such as triaxial accelerometers and gyroscopes) can provide clues about the device's posture and usage, indicating whether the user is moving, tilted, or stationary. Visual data sources (such as cameras and infrared modules) can identify the presence of faces, mouth movements, flash interference, or ambient occlusion in the image, providing support for determining the spatial source of noise and the interference mechanism.

[0051] Ambient audio data and multimodal sensor data are fed into a pre-defined noise perception network capable of integrating spatiotemporal features. Typically, this network consists of two branches: an audio branch extracts features from audio waveforms or spectrograms, extracting features such as frequency distribution, amplitude variation, and energy density through convolutional layers, attention mechanisms, or transformation modules; and a sensor branch processes image frame sequences and IMU data to extract features such as the trajectory of dynamic objects, the direction of noise sources, and device vibration trends.

[0052] After the outputs of the two branches are aligned, a joint representation is constructed through feature fusion mechanisms, such as a cross-attention mechanism or a temporal gated fusion module, to achieve unified modeling of sound and scene information. After fusion, the network outputs a multidimensional environmental representation tensor, which contains a variety of structured metrics such as noise type, spectral distribution, signal-to-noise ratio estimation, sound source direction estimation, and background disturbance level. This output, representing environmental feature information, can serve as auxiliary input for subsequent audio enhancement or speech generation modules.

[0053] This network is typically trained using labeled environmental datasets, collecting audio and sensor signals in various noise environments and labeling the interference types and intensities. The network parameters are then optimized using supervised or self-supervised methods. To enhance generalization, data augmentation techniques such as spectral masking, motion blur simulation, and random perturbation injection can be incorporated into training.

[0054] A multimodal perception network with a dual-branch structure can be used, in which the audio branch uses a spectral encoder based on a convolutional neural network to process the mel-spectrogram, the sensor branch uses a hybrid of a graph neural network and a recurrent structure to model device motion sequences and image frame sequences, and the fusion structure uses a temporal attention mechanism to dynamically weight the response time differences between different modalities. A Transformer structure can also be used to uniformly model the two types of inputs, performing dimensional normalization and embedding mapping on the audio and sensor data before input, and then merging them into a multi-head attention module to generate a temporal contextual representation. A graph structure representation can also be established through a graph convolution module, mapping sensor sampling points and audio feature points to heterogeneous graph nodes, and generating a fusion vector through a message passing mechanism.

[0055] Example: In the healthcare field, patient monitoring devices often emit continuous sounds in hospital rooms. When medical staff wear the devices to record, the device alarm sounds are mixed with the voice. By introducing a camera to capture the medical staff's spatial position relative to the alarm device and combining it with the audio signal to determine whether there is high-intensity periodic interference, the spectral characteristics of the interference source can be accurately identified. The generated environmental characteristics can be used to support automatic cleaning of voice recordings or transcription error correction.

[0056] In the financial sector, bank tellers can encounter situations where multiple customers are talking simultaneously or the background PA system emits notifications. A microphone, combined with a video feed of the teller's actions, determines whether the current task is focused on the teller's account and detects any loud, distant noise interference. Based on this information, the noise perception network generates environmental features including the direction of the sound source, the interference frequency band, and the SNR trend. This information is then used to adjust speech enhancement strategies and assess the quality of customer voice data.

[0057] This embodiment uses the coordinated input of audio and multimodal data to enable the noise perception network to not only identify interference types based on spectral features, but also dynamically estimate the location of interference sources and signal-to-noise ratio fluctuations by combining user actions and environmental images, helping subsequent processing modules to accurately and adaptively match audio enhancement strategies.

[0058] S30, inputting the environmental feature information and the target audio into an audio enhancement model to generate an enhanced audio signal;

[0059] In this embodiment, environmental feature information refers to a multi-dimensional representation of environmental interference generated by sensing and modeling the noise state in the target audio environment. This information typically includes parameters such as noise type identification, sound source direction, spectrum offset, and signal-to-noise ratio evaluation. This information is generated by the fusion perception network in the previous stage and has the ability to characterize the time-varying and non-stationary structure of noise. The target audio refers to the collected raw audio input containing the user's voice, which may be contaminated by environmental interference and contains a mixture of the desired voice signal and background noise.

[0060] The audio enhancement model is a neural network structure used to improve speech quality and clarity. It is designed to extract speech components from noisy audio and suppress interfering signals. During this processing phase, the audio enhancement model receives two types of input: the original target audio signal and the environmental characteristics output by the previous stage. This input structure enables the model to dynamically adjust its enhancement strategy based on different noise scenarios, thereby enhancing its adaptability and generalization capabilities.

[0061] The model structure can include multiple switchable processing paths, each designed for different noise characteristics. For example, the steady-state noise path is used to handle continuous background sounds, while the burst noise path is used to handle unstable interference such as coughs, knocks, and whistles. By analyzing the environmental characteristics, the corresponding path can be activated and related parameters such as noise reduction weight, frequency band priority, and gain control range can be set.

[0062] The model's primary workflow involves performing time-frequency analysis on the target audio to generate a spectrogram or feature frame sequence, fusing this with environmental characteristics at the feature level, and performing frame-by-frame enhancement processing via a deep neural network module. During this fusion process, environmental characteristics influence the model's determination of noise weights for each frequency band through an attention mechanism, enabling differentiated processing of different noise types. The output waveform, after time-domain reconstruction, becomes the enhanced audio signal.

[0063] The final output enhanced audio signal has a higher signal-to-noise ratio, spectral continuity and speech structure integrity, and can be used for subsequent personalized processing or directly for tasks such as playback and recognition.

[0064] This can be done through a multi-branch neural network, in which each branch corresponds to a specific type of noise suppression path, and an environmental feature parsing module automatically selects the branch based on the noise type identification; it can also be done by building a unified shared backbone network, dynamically injecting environmental feature information at different time steps, and adjusting the processing path through a gating mechanism; it can also be done through conditional normalization techniques (such as conditional instance normalization and conditional layer normalization) to embed environmental features into the internal intermediate layers of the model to achieve differentiated weight updates on frequency components or channels.

[0065] Alternatively, an end-to-end time-domain enhancement architecture, such as a TCN or WaveNet model, can be employed. This architecture directly inputs the target audio as a waveform and uses environmental features to control the receptive field of the attention convolution kernel, thereby achieving stable modeling of time-varying noise. Furthermore, joint modeling in the time-frequency domain can be employed, using dual-channel or multi-scale paths to process high- and low-frequency regions separately, achieving separation and enhancement of plosives from background noise.

[0066] Example: In healthcare scenarios, the constant noise of medical equipment and sudden alarms in ICU wards can easily obscure the doctor's voice recordings. By inputting environmental feature information generated by device status and sound sensor data, along with the doctor's voice, into the audio enhancement model, the frequency band characteristics of the alarm sound can be identified and specifically processed during the enhancement process, resulting in a clear output of the doctor's spoken content.

[0067] In financial scenarios, customer service calls are often accompanied by multiple noise sources, such as hall broadcasts and keyboard tapping. By combining device vibration data, sound field layout, and the current audio signal during a call and inputting them into a model, the model can determine whether the current noise is far-field broadcast interference. By adjusting the energy ratio between speech and background frequency bands, the model enhances the customer's voice content and automatically suppresses background broadcasts, thereby improving speech recognition accuracy and customer experience.

[0068] This embodiment simultaneously inputs the target audio and environmental feature information into the audio enhancement model, and can adjust the model's structure and parameters based on the environmental state, so that the output enhanced audio signal still maintains voice integrity, clarity, and intelligibility in complex noise environments, solving the problems of traditional enhancement methods being prone to distortion, slow response, and unstable enhancement effects in non-steady-state scenarios.

[0069] S40, obtaining a reference audio sample, and processing the reference audio sample through a speaker feature extraction module to generate a personalized feature vector;

[0070] In this embodiment, reference audio samples refer to historical speech data that represents the voice characteristics of a specific speaker. This speech data should be semantically complete, have distinguishable sound quality, and have timbre continuity. It can come from a variety of sources, such as past user recordings, voice commands, and identity recognition sentences. The essential function of these samples is to provide a basic sample set for identifying and modeling a speaker's voice style.

[0071] Processing reference audio samples through the speaker feature extraction module essentially involves building a robust and generalizable mapping function that extracts a stable representation of speaker attributes from the audio signal. This module typically includes multiple substructures, including a spectrum analyzer, a speech encoder, a feature normalization network, and a personality modeling layer.

[0072] During initial processing, the reference audio samples undergo a time-frequency transformation (such as a short-time Fourier transform or a Mel-spectrum transform) to obtain speech representations across multiple dimensions, including fundamental frequency distribution, formant distribution, duration structure, vocal tract characteristics, and resonance structure. These representations constitute important indicators for characterizing a speaker's pronunciation. In particular, fundamental frequency distribution is used to capture pitch style, and formant distribution is used to restore timbre expressiveness. These dimensions play a core role in personalized audio synthesis and style transfer tasks.

[0073] A style-adaptive normalization mechanism is introduced within the deep network architecture to extract speaker characteristics while maintaining the stability of speech content. This mechanism dynamically adapts the normalization scaling factors of different styles through a parameter adjustment strategy, enabling timbre differentiation and modeling. On this basis, incremental learning methods can be applied to further refine the feature representation, preventing the model from overfitting to a single speaker and ensuring that the feature representation is transferable and stable across scenarios.

[0074] The final output personalized feature vector should be a high-dimensional dense code, containing multiple layers of information such as timbre, speaking speed rhythm, oral structure influence, and intonation style, which will be used in the subsequent personalized enhancement processing stage to achieve the preservation and integration of the speaker's style.

[0075] A speaker feature extraction network can be constructed based on a dual-tower structure, where one side inputs a spectrogram to extract acoustic features, and the other side inputs the vocal tract model simulation results to model oral features. The two sides of the network are fused and encoded in subsequent layers. Alternatively, audio sequence modeling can be performed based on a pre-trained Transformer model, introducing positional encoding to capture speech rhythm and word order.

[0076] The normalization layer can be designed using conditional instance normalization, dynamically estimating sample-level conditional parameters from reference samples, or using a channel attention mechanism to guide the model to focus on discriminative regions in different frequency bands. Contrastive learning training can also be used to introduce negative sample constraints during sample encoding to enhance the ability of personalized feature vectors to distinguish between different speakers.

[0077] To adapt to edge device deployment environments, a depth-separable convolutional structure or a distilled lightweight version can be used to compress the original model into a low-latency structure while maintaining feature extraction accuracy.

[0078] Example: In healthcare scenarios, clinical voice recordings require unified organization and archiving of recordings from different doctors. Using the doctor's reference voice data to generate personalized feature vectors can then be used to enhance noisy frequencies into clear speech that matches the doctor's unique voice style, improving the accuracy and readability of dictation systems and speech recognition models.

[0079] In financial scenarios, telephone customer service systems can generate user-personalized feature vectors based on reference voice samples extracted from historical user calls, and use them to style-match and enhance user voices in subsequent voice interactions. This ensures identity consistency while improving voice quality and natural interaction, helping to reduce voice misrecognition rates and enhance user trust.

[0080] This embodiment constructs a stable and robust speaker feature extraction mechanism, which can accurately extract individual feature codes from limited speech samples, effectively avoiding the problem of loss of style consistency in current speech enhancement models in multi-speaker scenarios, so that the subsequently generated audio signals can simultaneously meet the requirements of speech clarity and speaker consistency.

[0081] S50, processing the enhanced audio signal based on the personalized feature vector to generate a personalized enhanced audio signal;

[0082] In this embodiment, the personalized feature vector is a multidimensional numerical code used to express the target speaker's vocal style. It includes features such as timbre structure, speaking rate and rhythm, pitch contour, and oral pronunciation habits. These feature vectors provide the foundation for style transfer and voice customization, ensuring that the enhanced audio signal not only has good clarity but also a consistent expressive style with the reference speaker.

[0083] The process of processing enhanced audio signals can be divided into two stages: parameter mapping and feature modulation. During the parameter mapping stage, modulation functions, such as fundamental frequency and formant modulation, are constructed based on the timbre-related parameters contained in the personalized feature vector. These modulation functions are used to guide the audio signal toward the target speaker's style while preserving its semantic content.

[0084] In the feature modulation stage, the spectral features of the enhanced audio signal are first extracted. Spectral features represent the energy distribution of the signal across different frequency components and are an important basis for expressing the timbre of speech. The fundamental frequency and formant components of these spectral features are then manipulated separately. The fundamental frequency modulation function adjusts the primary harmonic components in the spectrum to match the pitch contour of the personalized feature vector; the formant modulation function reconstructs the vocal tract simulation filter to ensure that the resonance of the enhanced speech closely matches the reference timbre.

[0085] Furthermore, to enhance the fusion effect, a frequency-domain attention mechanism can be introduced, allowing the model to focus on frequency ranges with significant individual differences, thereby improving the accuracy of individual feature transfer. Ultimately, the modulated multi-level spectral features are reconstructed into a time-domain audio signal, outputting a personalized enhanced audio signal.

[0086] A modulation network can be designed based on a two-branch structure. One branch receives the personalized feature vector as a conditional input, while the other receives the spectral features of the enhanced audio signal. These branches generate modulation factors and audio activation maps, respectively. The modulation factors are applied to the spectrogram via channel multiplication or feature concatenation to achieve style adjustment.

[0087] Alternatively, a personalized encoding can be introduced as a control vector based on the diffusion-based generative structure, guiding the noise removal path to converge toward the target style during each diffusion iteration. Furthermore, a style discrimination module can be introduced to perform adversarial discrimination training on the generated results based on the target personalized vector to enhance style preservation.

[0088] In lightweight implementation, the computational complexity of the spectrum modulation module can be reduced through low-rank tensor decomposition, or a sparse attention structure can be used to accelerate the feature fusion process to adapt to the real-time response requirements of mobile devices.

[0089] This embodiment modulates the enhanced audio signal through personalized feature vectors, so that the output speech has environmental adaptability and clarity while restoring the timbre and expression style of the target speaker, thereby achieving dual consistency of speech content and speech style, meeting the needs of personality preservation and recognition consistency in multi-user interaction.

[0090] S60, collecting feedback data when the personalized enhanced audio signal is played, and determining a playback time domain adjustment parameter and a playback frequency domain adjustment parameter according to the feedback data;

[0091] In this embodiment, the feedback data generated during the playback of the personalized enhanced audio signal is external sensory data used to perceive the user's actual response to the audio experience at the receiving end. This data includes information such as the state of the sound environment, physiological reaction characteristics, and facial expressions. This feedback data can be collected through multiple sensory channels, typically including audio and image acquisition channels.

[0092] The audio acquisition channel records the superimposed audio signal in the playback environment, including the personalized enhanced audio signal itself, background noise, reverberation, and possible user vocal responses, forming auditory feedback data. The image acquisition channel captures a real-time facial image sequence of the user to identify facial expression features and lip movement trajectory, forming visual feedback data. This image data is valuable in assisting in inferring whether the user has heard correctly, whether they are experiencing any distress, or whether they need to adjust their speaking speed and tone.

[0093] In the feedback data processing stage, background noise separation and sound pressure envelope extraction are first performed on the auditory feedback data to extract the background noise intensity characteristics and instantaneous spectrum offset information; then, temporal image encoding and facial feature point tracking are performed on the visual feedback data to extract facial expression recognition labels and lip opening and closing rhythm characteristics.

[0094] Playback time-domain adjustment parameters control speech speed, tempo fluctuation, and pause rhythm. They are derived primarily from data such as lip opening and closing frequency and blinking frequency, which indicate the user's attention and comprehension rhythm. Playback frequency-domain adjustment parameters control spectral equalization, loudness compensation, and high-frequency enhancement. They are quantified based on the dominant frequency band, spectral energy density, and peak perturbation of the ambient background noise.

[0095] By building a feedback processing module that connects to a high-sensitivity microphone array and image sensor, multi-source feedback data can be collected simultaneously while playing the personalized enhanced audio signal. A Fourier transform-based adaptive frequency band noise analyzer is introduced into the audio channel to extract non-target energy components in the signal and estimate the main interference frequency bands in the environment. A convolutional neural network-based facial expression recognition module is introduced into the image channel to extract expression labels indicating states such as confusion, attention, and response.

[0096] The feedback analysis module constructs a parameter mapping model based on the frequency domain fluctuation range and facial dynamic stability metrics extracted from the feedback data. For example, if the main energy band of the background noise is detected to be frequently shifting, and the facial expression in the visual feedback is stable, the proportion of the high-frequency component can be reduced and the playback frequency domain adjustment parameters can be output. If the facial lip rhythm frequency exceeds the upper limit of the average rhythm and is judged to be insufficiently responsive based on the speech rate, the playback time domain adjustment parameters are generated to slow the playback rhythm.

[0097] The parameter generation model can be obtained based on machine learning training, or it can be constructed based on rule-based mapping logic to adapt to implementation scenarios under different computing resource conditions.

[0098] Example: In a healthcare scenario, patients may receive instructions based on personalized voice feedback during rehabilitation training. Through cameras and sound sensors, their facial tension and reaction delays are collected, and the voice playback speed and tone clarity are adjusted in real time to improve the accuracy of instruction transmission and reduce the patient's hearing burden.

[0099] In the FinTech scenario, the automatic voice notification system plays compliance instructions or transaction receipts to customers, uses a microphone to collect the noise level of the customer's listening environment, and combines it with a camera to identify whether there are expressions of confusion such as frowning or tilting the head. When potential troubles are identified, the playback speed and frequency response structure are adjusted to improve the success rate of conveying key information.

[0100] This embodiment collects user multimodal feedback data during the playback of personalized enhanced audio signals and generates time domain and frequency domain adjustment parameters based on this data. This achieves dynamic adaptive control of the audio playback rhythm and spectrum, enhances the voice output system's perception and response capabilities to user state changes and environmental interference, and further improves the user's auditory experience and interactive comfort in real playback scenarios.

[0101] S70 , adjusting the time domain characteristics and the frequency domain characteristics of the personalized enhanced audio signal according to the playback time domain adjustment parameter and the playback frequency domain adjustment parameter to generate an optimized audio signal.

[0102] In this embodiment, the personalized enhanced audio signal is the output signal generated by fusing the personalized characteristics of the target speaker with the original audio enhancement model. The playback time domain adjustment parameters and playback frequency domain adjustment parameters are dynamically generated control factors based on user feedback during actual playback. The time domain adjustment parameters control the temporal characteristics of the audio signal, such as playback speed, speech tempo, and pause intervals; the frequency domain adjustment parameters control the spectral structure, such as energy distribution, frequency band loudness balance, and high and low frequency enhancement.

[0103] Temporal feature adjustments are performed on the personalized enhanced audio signal, typically involving scaling and shifting its waveform along the time axis. This can be achieved through dynamic time warping or a tempo-based compression algorithm. This can be achieved by adjusting the tempo compression ratio based on playback tempo parameters, controlling speech rate, or resetting sentence boundaries to automatically adapt the playback speed. These adjustments must ensure semantic integrity and rhythmic naturalness, avoiding semantic distortion caused by rhythmic disturbances.

[0104] Frequency domain feature adjustments are performed on the personalized enhanced audio signal, primarily including band gain compensation, spectrum reshaping, and high-frequency fidelity enhancement. Spectrograms are extracted using a short-time Fourier transform (SFT). Frequency domain adjustment parameters are then used to set the gain factor for each frequency band, achieving dynamic equalization in the frequency dimension. For high-frequency distortion areas, a linear predictive coding residual compensation mechanism is introduced to further restore high-frequency speech details such as plosives and sibilants.

[0105] Finally, the audio signal adjusted in the time domain is inversely transformed and fused with the spectral structure output by the frequency domain enhancement module to reconstruct and optimize the waveform of the audio signal and realize the joint correction of multi-dimensional features.

[0106] A joint control structure consisting of a time-domain adjustment submodule and a frequency-domain adjustment submodule can be constructed. The time-domain adjustment submodule receives the playback time-domain adjustment parameters and performs resampling, alignment, and synchronization based on the compression ratio set to adjust the timing structure of the personalized enhanced audio signal. This module employs a duration reconfiguration algorithm based on dynamic rhythm mapping to nonlinearly stretch the original beat structure and maintain the consistency of rhythmic fluctuations. The frequency-domain adjustment submodule receives the playback frequency-domain adjustment parameters and converts the audio signal into a spectrogram based on a short-time Fourier transform. It then performs dynamic gain compensation based on the target frequency band weights set by the parameters. This module incorporates a bandpass filter structure and a signal-to-noise ratio perception model to perform high-pass or low-pass enhancement on specific frequency bands to improve speech intelligibility. The optimized audio signal is obtained by inverse transforming and synchronously fusing the outputs of the two adjustment modules. If necessary, phase compensation and harmonic reconstruction can be performed by a terminal waveform reconstructor to ensure signal waveform continuity and spectral integrity.

[0107] Example description: In a healthcare scenario, when playing training speech for hearing rehabilitation patients, the playback speed can be automatically reduced and the high-frequency energy can be enhanced based on the patient's actual response during reception, making the originally difficult-to-recognize consonant structure clearer, thereby improving the effectiveness and efficiency of rehabilitation training.

[0108] In the fintech scenario, when the automatic customer service system plays contract terms or risk warnings to customers, it automatically slows down the speech speed and enhances the main semantic components in the mid-frequency band based on the background noise and visual fatigue of the customer's current environment, making the key terms easier to understand and remember, thereby reducing the risk of communication misunderstandings.

[0109] This embodiment jointly adjusts the personalized enhanced audio signal through playback time domain adjustment parameters and playback frequency domain adjustment parameters, achieving automatic adaptation of speech rate, frequency response structure, and detail clarity under different user states and playback environments, effectively improving the adaptive playback capabilities of the voice output system and enhancing the terminal user's auditory clarity, information reception integrity, and interactive experience consistency.

[0110] The present invention relates to the field of speech processing technology and can be applied to business scenarios such as financial technology and medical health. A speech enhancement method, device, equipment and medium based on noise perception are disclosed, including: obtaining environmental audio data and multimodal sensor data of the target audio to be processed and its environment; generating environmental feature information based on a noise perception network; inputting the environmental feature information and the target audio into an audio enhancement model to generate an enhanced audio signal; generating a personalized feature vector by processing a reference audio sample, and using it to enhance the personalized processing of the audio signal to generate a personalized enhanced audio signal; collecting feedback data when the personalized enhanced audio signal is played, and determining the playback time domain adjustment parameters and playback frequency domain adjustment parameters based on the feedback data; adjusting the time domain characteristics and frequency domain characteristics of the personalized enhanced audio signal according to the parameters to generate an optimized audio signal. The present invention combines multimodal environmental perception with personalized feature modeling to construct an audio enhancement mechanism oriented to dynamic changes in the environment and individual expression characteristics, and further combines playback feedback to achieve adaptive adjustment of playback parameters, thereby improving the environmental adaptability and individual fidelity of the final audio signal.

[0111] In one embodiment, the above step S10 includes:

[0112] S101, obtaining a target audio signal to be processed through an audio input interface;

[0113] S102, collecting original ambient audio data through a multi-channel microphone array;

[0114] S103, obtaining device motion state data including three-dimensional acceleration and angular velocity information through an inertial measurement unit;

[0115] S104, capturing environmental visual data including scene illumination intensity and motion trajectory of dynamic objects through a camera module;

[0116] S105 , aligning the original environmental audio data, the device motion state data, and the environmental visual data with synchronous timestamps to generate multimodal sensor data.

[0117] In this embodiment, the target audio signal refers to the core voice data that the user or system needs to enhance and optimize, and may include conversation voice, speech synthesis content, or preset broadcast content. The source of this audio signal must have a programmable input path, such as an audio input interface, which may include a wired audio port, a Bluetooth audio receiver module, or the system's internal text-to-speech (TTS) voice output channel. This input channel must have real-time access capabilities and a timing maintenance mechanism to ensure the integrity of subsequent processing.

[0118] The acquisition of raw ambient audio data is based on a multi-channel microphone array. This array consists of multiple spatially distributed audio acquisition units, arranged in linear, circular, or directional configurations. These units are used to detect sound source location, ambient reverberation, and multipath interference. This microphone array not only captures background noise but also provides spatial sound information, enabling spatial resolution for subsequent noise direction estimation and sound source enhancement.

[0119] Device motion data is acquired via an inertial measurement unit (IMU), which includes an accelerometer and gyroscope modules for measuring linear acceleration and angular velocity in three-dimensional space, respectively. These parameters reflect the device's dynamic state, such as hand-held shaking, head rotation, or vehicle vibration. By capturing motion trends, it can help determine whether the noise source is synchronized with the device's resonance, and whether there is frequency drift or directional disturbance.

[0120] The collection of environmental visual data relies on camera modules, capturing visual information including but not limited to scene lighting intensity and the motion trajectory of dynamic objects. Scene lighting influences microphone occlusion detection and echo interference detection, while the trajectory of dynamic objects correlates with the positional changes of sound or noise sources. Visual information is particularly valuable for building multimodal noise models in scenarios involving pedestrians, traffic, or conference interactions.

[0121] The aforementioned audio, motion, and visual information must be aligned in time, using a synchronized timestamp mechanism to align the multi-source data. This alignment process is typically coordinated through a global system time or hardware-level time base, ensuring that each data segment represents the same perceptual state at the same moment within the same frame. This ensures true semantic consistency in the environmental context and generates multimodal sensor data, which serves as the fundamental input for subsequent noise modeling and enhancement strategy selection.

[0122] The audio system's built-in microphone array and an external USB microphone array can be combined to collect sound data at different spatial levels, improving the distribution of background noise. The audio input interface supports multiple formats such as Pulse-code modulation (PCM) and Advanced Audio Coding (AAC), and has a preset buffer queue and format conversion module.

[0123] The inertial measurement unit, integrated into mobile terminals or wearable devices, can be configured with a sampling frequency between 200Hz and 1kHz, supporting data drift removal and gesture fusion to adapt to different interactive dynamic scenarios. The camera module is configured with a wide-angle CMOS image sensor and equipped with a low-light perception algorithm to support visual information acquisition in low-light conditions.

[0124] The timestamp alignment process is completed through the central control scheduling module. Based on a high-precision timer, nanosecond-level time stamps are added to each type of sensor data, and multi-channel time frame alignment is achieved through a buffer sliding window mechanism to ensure the consistency of perception timing and the feasibility of data fusion.

[0125] This embodiment uses multi-channel audio acquisition, motion state analysis and visual scene modeling to fully perceive the environmental state of the target audio and form a unified spatiotemporal perception input, so that subsequent noise recognition, enhancement strategy selection and audio personalization processing have higher accuracy, context adaptability and user feature fit.

[0126] In one embodiment, the above step S20 includes:

[0127] S201, parsing the ambient audio data through an audio processing branch, identifying noise type identifiers including transient and steady-state noise, and generating audio noise features;

[0128] S202, analyzing the multimodal sensor data through a sensor processing branch, detecting a noise source direction and predicting a noise spectrum offset, and generating a sensor positioning feature;

[0129] S203, fusing the audio noise feature and the sensor positioning feature to generate a spatiotemporal joint feature tensor;

[0130] S204 : Based on the spatiotemporal joint feature tensor, generate environmental feature information including a noise type identifier, a noise source direction, a noise spectrum offset, and a signal-to-noise ratio quantization value.

[0131] In this embodiment, the noise perception network is a perception model with a multi-branch input structure designed to extract information characterizing noise interference from heterogeneous environmental perception data. The network is divided into an audio processing branch and a sensor processing branch, corresponding to the feature extraction pathways for ambient audio data and multimodal sensor data, respectively.

[0132] The audio processing branch is dedicated to parsing raw ambient audio data, usually analyzed in frames. The audio signal is first extracted through a short-time Fourier transform (STFT) to obtain a time-frequency spectrum, and then passes through a convolution or time series modeling unit, such as a one-dimensional convolutional network, a gated recurrent unit (GRU), or a transformer structure, to extract audio noise features. In this process, the model needs to distinguish between transient noise (such as knocking sounds, whistles) and steady-state noise (such as air conditioning sounds, engine sounds), and mark the corresponding noise type as one of the features for subsequent interference intensity assessment and response strategy selection.

[0133] The sensor processing branch is responsible for parsing multimodal sensor data, which contains multidimensional data. Data from the inertial measurement unit (IMU) can be used to estimate the device's posture changes and acceleration distribution in space. Image frames captured by the camera can be combined with object detection models to determine the positional trends of suspected sound sources. This branch combines spatial position information with spectral-related signals (such as predicting the frequency range of changes by using the sound source angle and displacement) to calculate the direction of the noise source and the noise spectrum offset. The processed output is the sensor positioning feature.

[0134] Audio noise features are fused with sensor positioning features to form a unified spatiotemporal joint feature tensor. This tensor preserves information about sound variations in frequency and time, as well as information about external environmental disturbances in space and state. In terms of structural implementation, fusion paths can be constructed using feature channel splicing, attention fusion, or multi-head interaction mechanisms, enabling the tensor to maintain data consistency while improving its perceptual expressiveness.

[0135] Based on the aforementioned joint spatiotemporal feature tensor, the noise-aware network outputs a collection of information containing four types of annotated fields. First, the noise type identifier distinguishes the type of interference in the environment. Second, the noise source direction helps determine the spatial orientation of the interference. Third, the noise spectrum offset is used to model the dynamically changing frequency of noise interference. Finally, the signal-to-noise ratio quantization value represents the intensity comparison between the speech and noise components in the audio, which can be used to dynamically adjust the strength of subsequent enhancement strategies. Together, these pieces of information constitute structured environmental features, serving as an important reference for the input of the enhancement model.

[0136] This embodiment utilizes both auditory and non-auditory information to identify and model noise by performing branch parsing and spatiotemporal fusion processing on ambient audio data and multimodal sensor data within a unified network structure. The identification of transient and steady-state noise enhances the granularity of interference classification, while the fusion of spatial direction and spectral offset improves adaptability to dynamic interference scenarios. Signal-to-noise ratio quantification provides a quantitative basis for subsequent enhancement intensity control. This effectively improves the system's perception accuracy and responsiveness to complex noise scenarios, ensuring that the subsequently generated enhanced audio signal more closely reflects the actual interference state of the environment, thereby improving speech intelligibility and naturalness.

[0137] In one embodiment, the above step S30 includes:

[0138] S301, parsing the environmental characteristic information, obtaining a noise type identifier and a signal-to-noise ratio quantization value, and determining a noise suppression strategy based on the noise type identifier and the signal-to-noise ratio quantization value;

[0139] S302, selecting a speech enhancement processing mode according to the noise suppression strategy;

[0140] S303, inputting the target audio into the audio enhancement model branch corresponding to the selected speech enhancement processing mode to generate a preliminary enhanced audio signal;

[0141] S304 , processing the preliminary enhanced audio signal through a multi-scale discriminant model, optimizing waveform smoothness and frequency domain harmonic integrity of the preliminary enhanced audio signal, and generating an enhanced audio signal.

[0142] In this embodiment, the target audio needs to be jointly processed with the corresponding environmental feature information before being input into the audio enhancement model. The environmental feature information contains two key reference quantities: the noise type identifier and the signal-to-noise ratio quantization value, which are used to distinguish the degree to which different types of noise interference and quantization noise mask the speech signal. In the actual processing process, the model first parses the information structure, extracts the category attributes of the noise interference (such as steady-state and transient) and the signal-to-noise ratio level of the current speech signal (such as high SNR, medium SNR, low SNR or negative SNR), and thus determines the selection method and parameter configuration strategy of the subsequent enhancement path.

[0143] The noise suppression strategy is a set of rule configurations generated under the combined guidance of the two aforementioned parameters, designed to dynamically control the strength and processing of enhancements. For example, when detecting steady-state noise with a low SNR, the strategy might favor enhancing static spectrogram smoothing and background filtering. Conversely, when detecting sudden noises such as collisions or high-frequency whistles, it might prioritize temporal resolution and dynamic response to transient noise attenuation. This strategy serves as input to control the selection of speech enhancement processing modes.

[0144] The speech enhancement processing mode is specifically divided into two branches: steady-state noise enhancement mode and burst noise enhancement mode. The former is suitable for low-frequency steady-state interference such as air conditioning and motor background, and often uses frequency domain spectrogram restoration and low-pass filtering strategies to improve the clarity of the main speech. The latter is suitable for high-transient energy events such as alarms and keyboard tapping, focusing on high-time resolution processing and multi-channel activation response to prevent key speech features from being obscured. Based on the aforementioned strategy, the system selects the mode branch that adapts to the current scenario, inputs the target audio into the corresponding processing path, and generates a preliminary enhanced audio signal.

[0145] This preliminary enhanced audio signal has already undergone interference suppression structurally, but to further optimize its auditory quality and expressive integrity, it requires waveform refinement and spectral domain correction using a specialized multi-scale discriminant model. This model applies hierarchical frequency processing to the preliminary enhanced audio signal, capturing detailed features at different scales to optimize waveform continuity and harmonic structure integrity within the frequency domain. By combining reconstruction paths in both the time and frequency domains, the final output is an enhanced audio signal that exhibits greater clarity, naturalness, and semantic integrity than the original target audio.

[0146] This embodiment dynamically adapts to different interference scenarios by combining environmental feature information with the target audio input into the audio enhancement model and selecting a speech enhancement mode based on noise type and signal-to-noise ratio. Differentiated pattern branches enable classified responses to steady-state and burst interference, while a multi-scale discriminant model further enhances the structural continuity and detail integrity of the speech signal, achieving a balance between enhancement effect and speech quality. This processing flow not only improves speech intelligibility in complex noisy environments, but also significantly enhances the naturalness and auditory comfort of the enhanced speech.

[0147] In one embodiment, the above step S40 includes:

[0148] S401, receiving a reference audio input stream of a target speaker to obtain a reference audio sample;

[0149] S402, analyzing the frequency spectrum characteristics of the reference audio sample to generate fundamental frequency distribution characteristics and formant distribution characteristics;

[0150] S403, processing the fundamental frequency distribution features and the formant distribution features through the style adaptive normalization layer of the speaker feature extraction module to generate an initial speaker style vector;

[0151] S404, applying an incremental learning mechanism to process the initial speaker style vector to generate an updated speaker feature representation;

[0152] S405: Encode the timbre characteristic pattern in the updated speaker feature representation to generate a personalized feature vector.

[0153] In this embodiment, the system first receives a reference audio input stream of the target speaker, which can be derived from voice assistant wake-up data, recorded voice, or historical voice interaction records, and stably captures its voice samples through an audio buffer mechanism to form a reference audio sample. After receiving the audio sample, the system performs a spectrum analysis operation, mainly extracting the stable characteristic structure of the voice in the time-frequency domain. Among them, the fundamental frequency distribution feature is used to reflect the speaker's vocal cord vibration period and speech periodic behavior, which can usually be obtained through Fourier transform or short-time autocorrelation methods; the resonance peak distribution feature represents the energy concentration area of ​​the voice in different frequency bands, which is often used to express the differences in individual vocal tract structure and pronunciation habits. The acquisition method can be based on LPC (linear predictive coding) or cepstrum modeling.

[0154] These spectral features are input into the speaker feature extraction module, which includes a style-adaptive normalization layer for standardizing the speech style information of different speakers and mapping it into the style space. This layer employs a modulation mechanism that enables controllable parameter space migration of the same semantic unit across different individuals, thereby extracting a more representative initial speaker style vector. This style vector captures individual differences in timbre, rhythm, and frequency domain dynamics, serving as the foundation for subsequent personalized modeling.

[0155] To improve its adaptability to diverse speaker samples over multiple rounds of interaction, this module introduces an incremental learning mechanism to self-update the initial speaker style vector. This mechanism typically incorporates short-term memory units and a long-term style graph matching mechanism. By calculating similarities with historical samples and integrating weights, it generates a more stable and discriminative updated speaker feature representation. This representation is further processed through an encoder module to encode timbre characteristics, extracting expression patterns that are relevant to speaker identity but not semantics. Ultimately, this generates a personalized feature vector for speaker adaptation and style preservation in subsequent speech enhancement tasks.

[0156] This embodiment receives reference audio samples and sequentially performs spectral feature analysis, style normalization mapping, incremental learning updates, and timbre feature encoding to accurately extract the target individual's voice style differences and generate a stable and transferable personalized feature vector. This vector provides an individual feature reference for enhanced audio in subsequent speech processing, significantly improving timbre consistency and naturalness. At the same time, with the help of an incremental learning mechanism, dynamic adaptation of individual features is achieved under multiple rounds of interaction, effectively solving the problems of voice fidelity and difference alignment across speakers and multiple environments, and improving the flexibility and generalization of personality transfer.

[0157] In one embodiment, the above step S50 includes:

[0158] S501, determining a fundamental frequency modulation parameter and a formant modulation parameter based on the personalized feature vector;

[0159] S502, extracting frequency spectrum features of the enhanced audio signal;

[0160] S503, adjusting the fundamental frequency component of the spectrum feature according to the fundamental frequency modulation parameter to generate a first modulation feature;

[0161] S504, adjusting the formant component of the frequency spectrum feature according to the formant modulation parameter to generate a second modulation feature;

[0162] S505: Fusing the first modulation feature and the second modulation feature to generate a personalized enhanced audio signal.

[0163] In this embodiment, before generating a personalized enhanced audio signal, the system first uses a previously acquired personalized feature vector as an input reference. This vector contains the target speaker's timbre structure, articulation dynamics, and acoustic characteristics. Based on this vector information, the system parses the modulation control parameters, including fundamental frequency modulation parameters and formant modulation parameters. The fundamental frequency modulation parameters are used to simulate the periodic vibration behavior of the target individual's vocal cords, while the formant modulation parameters reflect the frequency energy distribution pattern formed by the vocal tract structure.

[0164] On this basis, the enhanced audio signal is subjected to spectral analysis to extract its spectral features, such as the amplitude spectrum and phase spectrum. This process can be accomplished through short-time Fourier transform or Mel-frequency cepstral coefficient analysis. The extracted spectral features contain the original representations of the fundamental frequency component and formant components, which are then used for modulation processing in subsequent steps.

[0165] The fundamental frequency modulation processing stage maps and transforms the fundamental frequency components in the current spectrum based on the fundamental frequency modulation parameters, so that the enhanced audio signal's pitch and intonation more closely match the target speaker's original pronunciation habits. This transformation process typically involves frequency stretching, linear pitch shifting, or nonlinear reparameterization, while maintaining phase structure continuity to generate the first modulation feature.

[0166] Subsequently, the formant components in the spectrum are reconstructed and shifted based on the formant modulation parameters, adjusting the energy concentration locations of different frequency bands to simulate the resonance structure of the target speaker, thereby improving the individual intelligibility of the speech. This processing may include spectral envelope morphology reconstruction or peak region frequency mapping to generate a second modulation feature.

[0167] After the modulation features are generated separately, the system performs feature fusion processing on the first modulation feature and the second modulation feature. The fusion method may include spectral weighted superposition, time-frequency attention combination or high-dimensional embedding space splicing, and finally forms a fused personalized enhanced audio signal, so that the output speech not only has the effect of improving clarity, but also reflects the timbre characteristics and speaking style of the target individual.

[0168] This embodiment enhances the fundamental frequency and formant components of the audio signal through personalized feature vector modulation. This system achieves personalized speech transfer and stylistic reproduction, effectively restoring the target speaker's individual characteristics while preserving semantic clarity and noise reduction. This processing mechanism avoids the issues of inconsistent timbre and cloning distortion found in traditional TTS. Furthermore, through a structured modulation path, personalized information is precisely mapped to the output speech, achieving simultaneous improvements in stylistic similarity and natural listening, enhancing the expressiveness and flexibility of speech synthesis in multi-speaker adaptation scenarios.

[0169] In one embodiment, the above step S60 includes:

[0170] S601, acquiring, through an audio acquisition device, ambient acoustic data of the environment during playback of the personalized enhanced audio signal, and generating auditory feedback data;

[0171] S602, acquiring a user's facial expression image sequence through an image acquisition device to generate visual feedback data;

[0172] S603, analyzing the visual feedback data to generate lip movement rhythm features and facial expression features;

[0173] S604, analyzing background interference components in the auditory feedback data to generate an environmental noise intensity feature;

[0174] S605, determining a playback time domain adjustment parameter based on the lip movement rhythm feature and the facial expression feature;

[0175] S606: Determine a playback frequency domain adjustment parameter based on the ambient noise intensity characteristics.

[0176] In this embodiment, while the personalized enhanced audio signal is being played, the audio collection device receives real-time external environmental echoes, background sounds, and sounds generated by the user's interaction with the audio signal, generating auditory feedback data. This data is typically collected using a microphone array to achieve sufficient temporal resolution and spatial directionality, allowing analysis of whether there are external interference, echo distortion, or sound imbalance issues during playback.

[0177] At the same time, an image acquisition device, such as an infrared camera or a structured light camera, is used to capture a sequence of facial images of the user as they listen. This sequence contains information about the user's lip movement rhythm and facial expression fluctuations. By performing expression recognition and motion tracking analysis on this sequence, key visual feedback indicators are extracted, including lip opening and closing periodicity, changes in jaw tension, and micro-expression dynamics in the eyebrow and brow area. These features reflect the user's understanding and emotional response to the current audio content.

[0178] The system performs sound source separation or frequency band energy analysis on auditory feedback data, identifying background interference components such as sudden noise, steady noise, and low-frequency rumble, and quantifying them into ambient noise intensity characteristics. This characteristic can be expressed in the spectral dimension as energy increases in specific frequency bands, and in the time domain as fluctuations in the noise duration and frequency.

[0179] Based on lip movement rhythm and facial expression characteristics, the system infers the degree of compatibility between the current playback speed and the user's actual semantic reception rate. If it detects lip movement lag or facial tension, it can generate playback time domain adjustment parameters accordingly, such as lengthening pauses between speech segments or reducing the overall playback speed, thereby improving user comprehension efficiency and comfort.

[0180] Furthermore, the system uses the characteristics of ambient noise intensity to determine playback frequency domain adjustment parameters, which are used to enhance the ability of voice playback to resist background noise. For example, when low-frequency noise interference is present in the environment, the energy output ratio of the high-frequency part of the voice can be increased or redundant low-frequency components can be suppressed to avoid voice masking.

[0181] Example description: In financial applications, the voice system is embedded in bank smart teller machines or mobile financial service terminals to provide users with personalized voice guidance and risk warning services. The system first obtains the target audio content to be processed, such as explanations of the credit card application process and risk level explanations, and collects current noise background data in the bank business hall or outdoor environment where the user is located, such as the sound of people talking and traffic noise. At the same time, the inertial measurement unit and camera module in the terminal collect multimodal sensor data to monitor whether the user is moving, whether he is facing the device, and the surrounding lighting and dynamic interference. After the above audio and non-audio information is input into the noise perception network, comprehensive feature information such as the noise type, noise source direction, and spectrum offset of the current environment is obtained.

[0182] Subsequently, the environmental feature information is fed into the audio enhancement model and combined with the original target audio. The system selects an appropriate speech enhancement mode based on the noise type and the current signal-to-noise ratio level. For example, if stable low-frequency noise (such as air conditioning sound) is detected, the steady-state enhancement model is selected for processing, and the waveform details are optimized through the subsequent multi-scale discriminant model to improve the frequency domain integrity of the speech.

[0183] To personalize the delivery of spoken content, the system extracts reference audio samples from the customer's past interactions, such as recordings from previous voice verification sessions. Using the speaker feature extraction module, the system extracts the customer's timbre feature vector, including individual parameters such as fundamental frequency and formants. Based on this personalized vector, the system modulates the enhanced speech, aligning the synthesized speech with the customer's familiar voice style in terms of intonation, speaking speed, and timbre, creating a personalized enhanced audio signal.

[0184] During the playback of personalized voice prompts, the system dynamically collects feedback information through the terminal's front-mounted microphone and camera. This includes environmental reverberation changes, the intensity of voice-masking noise, and feedback characteristics such as the customer's facial expressions (such as confusion and concentration) and lip movement rhythm. Based on the user's visual and auditory responses, the system dynamically adjusts the rhythmic characteristics and spectral energy structure of the voice output, for example, by extending pauses between key points and enhancing high-frequency clarity. This ultimately generates an optimized audio signal tailored to the current state of the customer and environment, improving the efficiency of financial business explanations and compliance reminders, and enhancing customer satisfaction.

[0185] In healthcare scenarios, intelligent voice assistant systems are deployed in rehabilitation wards or nursing homes. The system provides voice broadcast services such as medication reminders, rehabilitation guidance, and physical sign inquiries for bedridden patients, those recovering from surgery, and people with cognitive impairment. First, the system obtains the target audio to be processed through the bedside voice interaction terminal, such as "It's time to take your antihypertensive medication, please get some water and take your medicine." The system then collects the current environmental audio data of the ward in real time based on the microphone array in the terminal device, which may include the sounds of medical patrols, the operation of the infusion pump, and background conversations. At the same time, the system synchronously collects multimodal sensor data around the patient, including detecting bed vibrations and patient turning over information through an inertial measurement unit, and collecting visual elements such as indoor light intensity, changes in the patient's head posture, and surrounding activity trajectories through a camera module.

[0186] Subsequently, the system inputs the environmental audio data and sensor data into the preset noise perception network, and uses the audio processing branch to identify steady-state low-frequency noise such as the operation of an infusion pump or non-steady-state interference such as sudden human voice conversations. It also uses the vision and motion sensing branches to identify the direction of the noise source, the relative position of the patient, and the dynamic changes in the signal-to-noise ratio. Finally, it integrates them into a complete spatiotemporal joint feature tensor to generate environmental feature information of the current playback situation.

[0187] The system inputs the target audio and generated environmental characteristics into the audio enhancement module, automatically analyzing the noise type and signal-to-noise ratio level to determine the appropriate noise suppression strategy. For example, when low-frequency background noise is detected and the patient's head is facing away from the terminal, the system automatically activates the steady-state noise enhancement model and uses a multi-scale discriminant network to smooth the time-frequency waveform structure of the enhanced audio to form an enhanced audio signal.

[0188] To ensure the adaptability of voice content to individual patients, the system uses voice samples recorded by patients during the file creation period (such as daily inquiry recordings or past video call audio tracks) as reference audio samples, and extracts their fundamental frequency distribution, resonance peak distribution and intonation rhythm characteristics through the speaker feature extraction module. It combines style normalization and incremental learning mechanism to construct a patient-specific timbre model and encodes it to generate personalized feature vectors.

[0189] The system then modulates the enhanced audio signal based on this personalized feature vector, adjusting the fundamental frequency component of the speech to match the intonation contour of the patient's speaking style and the formant component to restore their timbre characteristics, thus forming a personalized enhanced audio signal. This processing makes the playback more consistent with the patient's cognitive habits. For example, by mimicking the patient's own intonation, the playback helps improve their attention and memory.

[0190] During actual playback, the system continuously monitors changes in feedback from the patient and the environment. A microphone captures ambient acoustic changes, analyzes reverberation intensity and noise masking, and generates auditory feedback data. A camera captures the patient's facial expressions and lip movements, identifying states like frowning and opening the mouth, generating visual feedback data. The system extracts lip movement rhythm and facial features, and, combined with current background noise intensity, dynamically adjusts the playback rhythm and spectral distribution.

[0191] For example, when the system detects that the patient frequently frowns or opens his mouth without giving any audible feedback during speech playback, it may determine that the current speech speed is too fast or the high-frequency energy is insufficient, and thus automatically reduce the speech speed and improve the clarity of the keywords; or when it detects that the ambient noise increases in a night environment, the system actively increases the mid-frequency energy of the speech to suppress low-frequency interference, making the final optimized audio signal more intelligible and comfortable.

[0192] This embodiment comprehensively collects and analyzes audio and visual feedback during the playback of personalized enhanced audio signals, enabling a dynamic adjustment mechanism for playback parameters based on the user's actual auditory and visual responses. This mechanism optimizes speech rate and pause rhythm using visual signals such as lip movements and facial expressions, while also adjusting the speech spectrum structure based on the distribution of ambient noise. This improves perceptual adaptability and clarity during speech playback, ensuring that the output speech better meets the cognitive and audibility requirements of real-world interactive environments.

[0193] In one embodiment, a speech enhancement device based on noise perception is provided, and the speech enhancement device based on noise perception corresponds to the speech enhancement method based on noise perception in the above embodiment. Figure 3 , Figure 3 This is a functional module diagram of a preferred embodiment of a noise-aware speech enhancement device according to the present invention. It includes a target audio acquisition module 10, a noise perception module 20, an audio enhancement module 30, a speaker feature extraction module 40, a personalized modulation module 50, a playback feedback analysis module 60, and a signal optimization module 70. Each functional module is described in detail below:

[0194] The target audio acquisition module 10 is used to acquire the target audio to be processed and collect the ambient audio data and multimodal sensor data of the environment in which the target audio is located;

[0195] The noise perception module 20 is configured to input the environmental audio data and the multimodal sensor data into a preset noise perception network to generate environmental feature information;

[0196] The audio enhancement module 30 is configured to input the environmental feature information and the target audio into an audio enhancement model to generate an enhanced audio signal;

[0197] A speaker feature extraction module 40 is configured to obtain a reference audio sample and process the reference audio sample through the speaker feature extraction module to generate a personalized feature vector;

[0198] a personalized modulation module 50, configured to process the enhanced audio signal based on the personalized feature vector to generate a personalized enhanced audio signal;

[0199] A playback feedback analysis module 60 is configured to collect feedback data when the personalized enhanced audio signal is played, and determine playback time domain adjustment parameters and playback frequency domain adjustment parameters based on the feedback data;

[0200] The signal optimization module 70 is configured to adjust the time domain characteristics and the frequency domain characteristics of the personalized enhanced audio signal according to the playback time domain adjustment parameter and the playback frequency domain adjustment parameter to generate an optimized audio signal.

[0201] In one embodiment, the target audio acquisition module 10 is specifically configured to:

[0202] Obtain the target audio signal to be processed through the audio input interface;

[0203] Collect raw ambient audio data through a multi-channel microphone array;

[0204] Obtain device motion state data including three-dimensional acceleration and angular velocity information through the inertial measurement unit;

[0205] Capture environmental visual data including scene lighting intensity and motion trajectory of dynamic objects through the camera module;

[0206] The raw ambient audio data, the device motion state data, and the ambient visual data are aligned with synchronization timestamps to generate multimodal sensor data.

[0207] In one embodiment, the noise sensing module 20 is specifically configured to:

[0208] parsing the ambient audio data through the audio processing branch, identifying noise type identifiers including transient and steady-state noise, and generating audio noise features;

[0209] Analyzing the multimodal sensor data through a sensor processing branch to detect a noise source direction and predict a noise spectrum offset to generate a sensor positioning feature;

[0210] fusing the audio noise features and the sensor positioning features to generate a spatiotemporal joint feature tensor;

[0211] Based on the spatiotemporal joint feature tensor, environmental feature information including a noise type identifier, a noise source direction, a noise spectrum offset, and a signal-to-noise ratio quantization value is generated.

[0212] In one embodiment, the audio enhancement module 30 is specifically configured to:

[0213] parsing the environmental characteristic information to obtain a noise type identifier and a signal-to-noise ratio quantization value, and determining a noise suppression strategy based on the noise type identifier and the signal-to-noise ratio quantization value;

[0214] selecting a speech enhancement processing mode according to the noise suppression strategy;

[0215] The target audio input is combined with the audio enhancement model branch corresponding to the selected speech enhancement processing mode to generate a preliminary enhanced audio signal;

[0216] The preliminary enhanced audio signal is processed by a multi-scale discriminant model to optimize the waveform smoothness and frequency domain harmonic integrity of the preliminary enhanced audio signal to generate an enhanced audio signal.

[0217] In one embodiment, the speaker feature extraction module 40 is specifically configured to:

[0218] Receive a reference audio input stream of a target speaker and obtain a reference audio sample;

[0219] Analyzing the frequency spectrum characteristics of the reference audio sample to generate fundamental frequency distribution characteristics and formant distribution characteristics;

[0220] Processing the fundamental frequency distribution features and the formant distribution features through a style adaptive normalization layer of a speaker feature extraction module to generate an initial speaker style vector;

[0221] Applying an incremental learning mechanism to process the initial speaker style vector to generate an updated speaker feature representation;

[0222] The timbre characteristic pattern in the updated speaker feature representation is encoded to generate a personalized feature vector.

[0223] In one embodiment, the personalized modulation module 50 is specifically configured to:

[0224] Determining a fundamental frequency modulation parameter and a formant modulation parameter based on the personalized feature vector;

[0225] Extracting frequency spectrum features of the enhanced audio signal;

[0226] Adjusting the fundamental frequency component of the frequency spectrum feature according to the fundamental frequency modulation parameter to generate a first modulation feature;

[0227] adjusting the formant component of the spectral feature according to the formant modulation parameter to generate a second modulation feature;

[0228] The first modulation feature and the second modulation feature are fused to generate a personalized enhanced audio signal.

[0229] In one embodiment, the playback feedback analysis module 60 is specifically configured to:

[0230] Acquiring, through an audio acquisition device, ambient acoustic data when the personalized enhanced audio signal is played, and generating auditory feedback data;

[0231] Acquire a user's facial expression image sequence through an image acquisition device to generate visual feedback data;

[0232] Analyzing the visual feedback data to generate lip movement rhythm features and facial expression features;

[0233] Analyzing background interference components in the auditory feedback data to generate an ambient noise intensity feature;

[0234] Determining playback time domain adjustment parameters based on the lip movement rhythm characteristics and facial expression characteristics;

[0235] Based on the environmental noise intensity characteristics, a playback frequency domain adjustment parameter is determined.

[0236] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a speech enhancement method based on noise perception.

[0237] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a speech enhancement method based on noise perception.

[0238] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0239] Acquire target audio to be processed, and collect ambient audio data and multimodal sensor data of an environment in which the target audio is located;

[0240] Inputting the environmental audio data and the multimodal sensor data into a preset noise perception network to generate environmental feature information;

[0241] Inputting the environmental feature information and the target audio into an audio enhancement model to generate an enhanced audio signal;

[0242] Obtaining a reference audio sample, and processing the reference audio sample through a speaker feature extraction module to generate a personalized feature vector;

[0243] Processing the enhanced audio signal based on the personalized feature vector to generate a personalized enhanced audio signal;

[0244] Collecting feedback data when the personalized enhanced audio signal is played, and determining a playback time domain adjustment parameter and a playback frequency domain adjustment parameter according to the feedback data;

[0245] The time domain characteristics and frequency domain characteristics of the personalized enhanced audio signal are adjusted according to the playback time domain adjustment parameter and the playback frequency domain adjustment parameter to generate an optimized audio signal.

[0246] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0247] Acquire target audio to be processed, and collect ambient audio data and multimodal sensor data of an environment in which the target audio is located;

[0248] Inputting the environmental audio data and the multimodal sensor data into a preset noise perception network to generate environmental feature information;

[0249] Inputting the environmental feature information and the target audio into an audio enhancement model to generate an enhanced audio signal;

[0250] Obtaining a reference audio sample, and processing the reference audio sample through a speaker feature extraction module to generate a personalized feature vector;

[0251] Processing the enhanced audio signal based on the personalized feature vector to generate a personalized enhanced audio signal;

[0252] Collecting feedback data when the personalized enhanced audio signal is played, and determining a playback time domain adjustment parameter and a playback frequency domain adjustment parameter according to the feedback data;

[0253] The time domain characteristics and frequency domain characteristics of the personalized enhanced audio signal are adjusted according to the playback time domain adjustment parameter and the playback frequency domain adjustment parameter to generate an optimized audio signal.

[0254] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0255] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0256] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0257] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A speech enhancement method based on noise perception, characterized in that: The following steps are involved: Acquire target audio to be processed, and collect ambient audio data and multimodal sensor data of an environment in which the target audio is located; Inputting the environmental audio data and the multimodal sensor data into a preset noise perception network to generate environmental feature information; Inputting the environmental feature information and the target audio into an audio enhancement model to generate an enhanced audio signal; Obtaining a reference audio sample, and processing the reference audio sample through a speaker feature extraction module to generate a personalized feature vector; Processing the enhanced audio signal based on the personalized feature vector to generate a personalized enhanced audio signal; Collecting feedback data when the personalized enhanced audio signal is played, and determining a playback time domain adjustment parameter and a playback frequency domain adjustment parameter according to the feedback data; The time domain characteristics and frequency domain characteristics of the personalized enhanced audio signal are adjusted according to the playback time domain adjustment parameter and the playback frequency domain adjustment parameter to generate an optimized audio signal.

2. The method for speech enhancement based on noise perception according to claim 1, wherein: Obtaining target audio to be processed and collecting ambient audio data and multimodal sensor data of the environment in which the target audio is located, including: Obtain the target audio signal to be processed through the audio input interface; Collect raw ambient audio data through a multi-channel microphone array; Obtain device motion state data including three-dimensional acceleration and angular velocity information through the inertial measurement unit; Capture environmental visual data including scene lighting intensity and motion trajectory of dynamic objects through the camera module; The raw ambient audio data, the device motion state data, and the ambient visual data are aligned with synchronization timestamps to generate multimodal sensor data.

3. The method for speech enhancement based on noise perception according to claim 1, wherein: Inputting the environmental audio data and the multimodal sensor data into a preset noise perception network to generate environmental feature information includes: parsing the ambient audio data through the audio processing branch, identifying noise type identifiers including transient and steady-state noise, and generating audio noise features; Analyzing the multimodal sensor data through a sensor processing branch to detect a noise source direction and predict a noise spectrum offset to generate a sensor positioning feature; fusing the audio noise features and the sensor positioning features to generate a spatiotemporal joint feature tensor; Based on the spatiotemporal joint feature tensor, environmental feature information including a noise type identifier, a noise source direction, a noise spectrum offset, and a signal-to-noise ratio quantization value is generated.

4. The method for speech enhancement based on noise perception according to claim 1, wherein: Inputting the environmental feature information and the target audio into an audio enhancement model to generate an enhanced audio signal includes: parsing the environmental characteristic information to obtain a noise type identifier and a signal-to-noise ratio quantization value, and determining a noise suppression strategy based on the noise type identifier and the signal-to-noise ratio quantization value; selecting a speech enhancement processing mode according to the noise suppression strategy; The target audio input is combined with the audio enhancement model branch corresponding to the selected speech enhancement processing mode to generate a preliminary enhanced audio signal; The preliminary enhanced audio signal is processed by a multi-scale discriminant model to optimize the waveform smoothness and frequency domain harmonic integrity of the preliminary enhanced audio signal to generate an enhanced audio signal.

5. The method for speech enhancement based on noise perception according to claim 1, wherein: Obtaining a reference audio sample and processing the reference audio sample through a speaker feature extraction module to generate a personalized feature vector, including: Receive a reference audio input stream of a target speaker and obtain a reference audio sample; Analyzing the frequency spectrum characteristics of the reference audio sample to generate fundamental frequency distribution characteristics and formant distribution characteristics; Processing the fundamental frequency distribution features and the formant distribution features through a style adaptive normalization layer of a speaker feature extraction module to generate an initial speaker style vector; Applying an incremental learning mechanism to process the initial speaker style vector to generate an updated speaker feature representation; The timbre characteristic pattern in the updated speaker feature representation is encoded to generate a personalized feature vector.

6. The method for speech enhancement based on noise perception according to claim 1, wherein: Processing the enhanced audio signal based on the personalized feature vector to generate a personalized enhanced audio signal includes: Determining a fundamental frequency modulation parameter and a formant modulation parameter based on the personalized feature vector; Extracting frequency spectrum features of the enhanced audio signal; Adjusting the fundamental frequency component of the frequency spectrum feature according to the fundamental frequency modulation parameter to generate a first modulation feature; adjusting the formant component of the spectral feature according to the formant modulation parameter to generate a second modulation feature; The first modulation feature and the second modulation feature are fused to generate a personalized enhanced audio signal.

7. The method for speech enhancement based on noise perception according to claim 1, wherein: Collecting feedback data when the personalized enhanced audio signal is played, and determining a playback time domain adjustment parameter and a playback frequency domain adjustment parameter according to the feedback data, including: Acquiring, through an audio acquisition device, ambient acoustic data when the personalized enhanced audio signal is played, and generating auditory feedback data; Acquire a user's facial expression image sequence through an image acquisition device to generate visual feedback data; Analyzing the visual feedback data to generate lip movement rhythm features and facial expression features; Analyzing background interference components in the auditory feedback data to generate an ambient noise intensity feature; Determining playback time domain adjustment parameters based on the lip movement rhythm characteristics and facial expression characteristics; Based on the environmental noise intensity characteristics, a playback frequency domain adjustment parameter is determined.

8. A speech enhancement device based on noise perception, characterized in that: The noise perception-based speech enhancement device comprises: A target audio acquisition module is used to acquire target audio to be processed and collect ambient audio data and multimodal sensor data of the environment in which the target audio is located; a noise perception module, configured to input the ambient audio data and the multimodal sensor data into a preset noise perception network to generate environmental feature information; an audio enhancement module, configured to input the environmental feature information and the target audio into an audio enhancement model to generate an enhanced audio signal; A speaker feature extraction module is used to obtain a reference audio sample and process the reference audio sample through the speaker feature extraction module to generate a personalized feature vector; a personalized modulation module, configured to process the enhanced audio signal based on the personalized feature vector to generate a personalized enhanced audio signal; a playback feedback analysis module, configured to collect feedback data when the personalized enhanced audio signal is played, and determine playback time domain adjustment parameters and playback frequency domain adjustment parameters based on the feedback data; The signal optimization module is used to adjust the time domain characteristics and frequency domain characteristics of the personalized enhanced audio signal according to the playback time domain adjustment parameter and the playback frequency domain adjustment parameter to generate an optimized audio signal.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a noise-awareness-based speech enhancement program stored in the memory and executable on the processor. When the noise-awareness-based speech enhancement program is executed by the processor, the steps of the noise-awareness-based speech enhancement method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores a noise-awareness-based speech enhancement program, which, when executed by a processor, implements the steps of the noise-awareness-based speech enhancement method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Multi-mode intelligent terminal AI voice wake-up method and device based on voice control

    CN120853549A

  • Intelligent pickup and speech recognition system based on multi-modal fusion

    CN120954408A

  • Intelligent sound pickup and speech recognition system based on multimodal fusion

    CN120954408B

  • Cloud terminal dynamic audio processing method based on AI

    CN121191528A

  • Intelligent auxiliary agent response service method of customer service center in financial industry

    CN121309727A