Audio processing system of photographic camera

Through the photographic camera audio processing system, combined with a high-sensitivity microphone array and multiple audio processing technologies, the problem that traditional cameras cannot synchronize high-quality audio processing is solved, and real-time acquisition and processing of high-quality audio is achieved, which is suitable for audio and video synchronization in complex sound fields.

CN120600041AActive Publication Date: 2025-09-05XIAMEN NANYANG UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511108431.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-09-05
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

Traditional cameras cannot directly obtain high-quality audio when recording high-quality images and audio simultaneously. Users need to carry additional recording equipment, which reduces convenience and cost advantages. Existing technologies lack solutions for synchronous high-quality audio processing.

Method used

A photographic camera audio processing system is used, including an audio acquisition module, a human voice detection module, a human voice highlighting module, a sound quality enhancement module, and a time domain storage module. Through high-sensitivity microphone arrays, pre-emphasis-framing preprocessing, human voice detection, steady-state and impact noise suppression, sound quality enhancement, and other technologies, real-time acquisition and processing of high-quality audio is achieved.

Benefits of technology

Ensures wideband audio capture with a high signal-to-noise ratio, effectively separates human voices from non-human voices, suppresses noise interference, and improves sound quality. It is suitable for complex sound fields such as outdoor live broadcasts and interviews, improving convenience and audio and video synchronization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600041A_ABST
    Figure CN120600041A_ABST
Patent Text Reader

Abstract

The invention discloses an audio processing system for a photographic camera, which relates to the technical field of audio processing and comprises an audio acquisition module, a human voice detection module, a human voice highlighting module, a tone quality enhancement module and a time domain storage module. The audio acquisition module is used for acquiring and preprocessing real-time audio; the human voice detection module is used for performing endpoint frame calculation on the spectrum data to be analyzed to obtain a suspected human voice data segment, and calculating a weighted Mel frequency cepstrum coefficient to obtain target human voice data and non-human voice data; the human voice highlighting module is used for performing first frequency band division and noise suppression processing on the target human voice data to obtain steady-state noise suppression data and impact noise suppression data; the tone quality enhancement module is used for performing second frequency band division and frequency domain enhancement processing on the non-human voice data to obtain updated non-human voice data; and the time domain storage module is used for performing time-frequency domain change on the impact noise suppression data, the steady-state noise suppression data and the updated non-human voice data to obtain time domain audio data and storing the time domain audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and in particular to an audio processing system for a photographic camera. Background Art

[0002] Cameras currently on the market primarily focus on image capture. Despite significant advances in image quality, capture speed, and intelligence, they still have limitations when it comes to multimedia recording. In particular, in situations where high-quality images and audio must be recorded and uploaded simultaneously, such as news reporting, outdoor live broadcasts, and live photography, traditional cameras and mobile phones often cannot directly meet these requirements. Users often need to carry additional recording equipment, which not only increases the burden but can also lead to audio and video mismatches due to synchronization issues between devices. In addition to this lack of timeliness, post-processing of the images and audio also requires high-quality audio material to ensure the final audiovisual work is well-presented.

[0003] Currently, a Chinese invention patent application with application number CN202110405044.2 discloses a panoramic audio processing method for panoramic cameras. The application specifically includes six parts: a pre-processing module, a panoramic sound encoding module, a mono audio sound and image placement module, a dynamic sound field positioning processing module, a virtual speaker decoding module, and a psychoacoustic headphone playback module. It uses a unified audio format to reposition the playback space sound field at any angle, solving the problem of poor end-user interactivity. The method has the function of repositioning the sound field based on angle information, which can effectively combine panoramic sound format audio technology with panoramic video technology. Summary of the Invention

[0004] The technical problem solved by the present invention is that, in situations where high-quality images and audio need to be recorded simultaneously, traditional cameras are often unable to directly obtain high-quality audio. Users usually need to carry additional recording equipment, which reduces convenience and cost advantages. The existing technology lacks a solution that can synchronously process high-quality audio and enhance human voices while the camera is shooting.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: A photographic camera audio processing system, comprising: Audio acquisition module, human voice detection module, human voice highlight module, sound quality enhancement module, time domain storage module The audio acquisition module is used to collect and preprocess real-time audio to obtain pre-emphasized time domain data and spectrum data to be analyzed; The human voice detection module is used to calculate the endpoint frames of the spectrum data to be analyzed to obtain the suspected human voice data segment, and calculate the weighted Mel frequency cepstral coefficient to obtain the target human voice data and non-human voice data; The human voice highlighting module is used to perform a first frequency band division and noise suppression process on the target human voice data to obtain steady-state noise suppression data and impact noise suppression data; The sound quality enhancement module is used to perform second frequency band division and frequency domain enhancement processing on the non-human voice data to obtain updated non-human voice data; The time domain storage module is used to perform time-frequency domain transformation on the impact noise suppression data, the steady-state noise suppression data and the updated non-human voice data to obtain time domain audio data and store the data.

[0006] As a preferred solution of the camera audio processing system of the present invention, wherein: The audio acquisition module includes a real-time acquisition unit and a pre-processing unit; The real-time acquisition unit is used to acquire real-time audio data through a high-sensitivity microphone array. The high-sensitivity microphone array is started synchronously when the camera starts recording and enters a standby state when the camera stops recording.

[0007] As a preferred solution of the camera audio processing system of the present invention, wherein: The preprocessing unit is used to preprocess the real-time audio data to obtain pre-emphasized time domain data and spectrum data to be analyzed. The processing logic includes: The real-time audio data is pre-emphasized by a first-order high-pass filter to obtain pre-emphasized time domain data. The pre-emphasized time domain data is segmented based on a preset frame window step size and a preset frame shift rate, and the spectrum data to be analyzed is obtained by Blackman-Harris windowing and FFT calculation.

[0008] As a preferred solution of the camera audio processing system of the present invention, wherein: The human voice detection module includes an audio endpoint detection unit and a spectrum feature detection unit, which is used to perform human voice analysis on the spectrum data to be analyzed to obtain target human voice data and non-human voice data; The audio endpoint detection unit is used to calculate the pre-emphasized time domain data to obtain the endpoint frame, and combine it with the spectrum data to be analyzed to obtain the suspected human voice data segment. The processing logic includes: Calculate the pre-emphasized time domain data to calculate the current frame energy and the historical average energy of the previous K frames, divide the current frame energy by the historical average energy of the previous K frames to obtain a short-time energy ratio, and calculate the zero-crossing rate of the current frame of the pre-emphasized time domain data; If and only if the short-time energy ratio of the current frame belongs to the preset first threshold interval and the zero-crossing rate of the current frame belongs to the preset second threshold interval, the current frame is marked as an endpoint frame, and the spectrum data to be analyzed is extracted based on the frame index of the endpoint frame to obtain a suspected human voice data segment.

[0009] As a preferred solution of the camera audio processing system of the present invention, wherein: The spectrum feature detection unit is used to perform spectrum feature analysis on the human voice data segment to obtain target human voice data and non-human voice data; The spectrum data to be analyzed is filtered by a Gaussian Mel filter bank, the energy integral of each filter output in the Gaussian Mel filter bank is calculated to obtain the frequency band energy distribution, and the frequency band energy distribution is statistically analyzed to obtain the Gaussian Mel energy spectrum; The Gaussian Mel energy spectrum is logarithmically compressed, and the discrete cosine transform is performed on the logarithmically compressed Gaussian Mel energy spectrum to obtain Mel frequency cepstral coefficients. The Mel frequency cepstral coefficients are decomposed into three layers using the Daubechies-4 wavelet basis to obtain low-frequency approximate components and high-frequency detail components. The low-frequency approximate components are kept with their original weights, and the high-frequency detail components are multiplied by a preset amplification weight. The weighted Mel frequency cepstral coefficients are reconstructed by inverse wavelet transform. The weighted Mel-frequency cepstral coefficients are classified and regressed through the pre-trained Gaussian mixture model GMM to obtain the target human voice data and non-human voice data.

[0010] As a preferred solution of the camera audio processing system of the present invention, wherein: The vocal prominence module includes a steady-state noise suppression unit and an impact noise suppression unit; The steady-state noise suppression unit is used to divide the target human voice data into a first frequency band, extract the 20Hz to 500Hz part of the target human voice data through low-pass filtering to obtain the low-frequency band of the human voice data, calculate the spectral energy corresponding to the low-frequency band of each frame of human voice data, divide the spectral energy corresponding to the low-frequency human voice data segment by the spectral energy corresponding to the target human voice data to obtain a low-frequency energy ratio value. When the low-frequency energy ratio value is greater than or equal to a preset background noise threshold, it is determined that steady-state noise exists in the target human voice data of the frame, and spectral subtraction and calculation are performed on the target human voice data of the frame based on a preset noise reduction factor to obtain steady-state noise suppression data.

[0011] As a preferred solution of the camera audio processing system of the present invention, wherein: The impulse noise suppression unit is used to perform dynamic threshold filtering on the target human voice data to obtain impulse noise suppression data. The processing logic includes: Perform Morlet wavelet decomposition on each frame of the target human voice data to obtain Morlet decomposition layer coefficients, calculate the noise energy estimate of each layer based on the Morlet decomposition layer coefficients, and calculate the dynamic threshold based on the noise energy estimate of each layer; When the Morlet decomposition layer coefficient is less than or equal to the dynamic threshold, the Morlet decomposition layer coefficient is assigned to 0; When the Morlet decomposition layer coefficient is greater than the dynamic threshold and less than or equal to 2 times the dynamic threshold, the Morlet decomposition layer coefficient is linearly scaled; When the Morlet decomposition layer coefficient is greater than 2 times the dynamic threshold, the original value of the Morlet decomposition layer coefficient is retained; The Morlet decomposition layer coefficients after dynamic threshold processing are reconstructed by wavelet to obtain the impact noise suppression data; The dynamic threshold calculation expression is: ; represents the dynamic threshold corresponding to the i-th layer, represents the noise energy estimate corresponding to the i-th layer, N represents the total number of Morlet decomposition layer coefficients, and ln represents the natural logarithm; The bandwidth parameter of the Morlet wavelet decomposition is set to 6, the center frequency parameter is set to 2, and the number of decomposition levels is set to 5.

[0012] As a preferred solution of the camera audio processing system of the present invention, wherein: The sound quality enhancement module includes a frequency domain division unit and a segment enhancement unit, which are used to enhance the non-human voice data to obtain updated non-human voice data; The frequency domain division unit is used to perform a second frequency band division on the frequency of the non-human voice data according to a preset non-human voice frequency band threshold, dividing the part less than the non-human voice frequency band threshold into non-human voice low frequency band data, and dividing the part greater than or equal to the non-human voice frequency band threshold into non-human voice high frequency band data.

[0013] As a preferred solution of the camera audio processing system of the present invention, wherein: The segment enhancement unit is used to enhance the non-human voice low-frequency band data and the non-human voice high-frequency band data to obtain updated non-human voice data. The processing logic includes: The non-human voice low-frequency band data is notched by comb filter to obtain wind noise removal spectrum data; Based on the AR model, the spectrum of the non-human voice high-frequency band data is extrapolated, and the high-frequency spread spectrum data is obtained through the logarithmic domain automatic gain control processing AGC of the preset gain compression ratio coefficient; The wind noise removal spectrum data and the high-frequency spread spectrum data are spliced ​​in the frequency domain to obtain updated non-human voice data.

[0014] As a preferred solution of the camera audio processing system of the present invention, wherein: The time domain storage module includes a time-frequency domain restoration unit and a data storage unit; The time-frequency domain restoration unit is used to perform inverse FFT transformation processing on the impact noise suppression data, the steady-state noise suppression data and the updated non-human voice data, and perform overlap-add processing based on the preset frame rate and Blackman-Harris window function to obtain time-domain audio data; The data storage unit is used to add a timestamp tag to the time domain audio data and complete the storage.

[0015] Beneficial effects of the present invention: This application ensures wide-band and high signal-to-noise ratio acquisition of the original audio signal through a high-sensitivity microphone array and pre-emphasis-framing preprocessing technology. The human voice detection module combines short-time energy ratio, zero-crossing rate dual-threshold endpoint detection with wavelet reconstruction-based weighted Mel-frequency cepstral coefficient feature analysis to achieve separation of human voice and non-human voice through a Gaussian mixture model. It also uses the Daubechies-4 wavelet basis to perform three-layer decomposition and reconstruction, which is beneficial to strengthening the speech resonance peak characteristics while suppressing high-frequency noise interference. Dual-modal noise suppression is adopted in the human voice prominence module. The steady-state noise suppression unit uses the low-frequency energy ratio to dynamically trigger spectral subtraction to specifically eliminate wind noise and equipment background noise in the 20-500Hz frequency band. Morlet wavelet decomposition is performed in the impact noise suppression unit, and transient noise suppression is achieved through a dynamic threshold function, which is beneficial to suppress burst noise while retaining transient features such as speech plosives. The sound quality enhancement module uses frequency division processing of comb filter notch and AR model high-frequency extrapolation to eliminate low-frequency wind noise while expanding high-frequency details of ambient sound. It is suitable for application scenarios such as outdoor live broadcasts and interview shooting that require synchronous optimization of human voices in complex sound fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic diagram of the basic flow of an audio processing system for a photographic camera provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0017] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.

[0018] Reference Figure 1 , as one embodiment of the present invention, provides a camera audio processing system, comprising: Audio acquisition module, human voice detection module, human voice highlight module, sound quality enhancement module, time domain storage module The audio acquisition module is used to collect and preprocess real-time audio to obtain pre-emphasized time domain data and spectrum data to be analyzed; The human voice detection module is used to calculate the endpoint frames of the spectrum data to be analyzed to obtain the suspected human voice data segment, and calculate the weighted Mel frequency cepstral coefficient to obtain the target human voice data and non-human voice data; The human voice highlighting module is used to perform a first frequency band division and noise suppression process on the target human voice data to obtain steady-state noise suppression data and impact noise suppression data; The sound quality enhancement module is used to perform second frequency band division and frequency domain enhancement processing on the non-human voice data to obtain updated non-human voice data; The time domain storage module is used to perform time-frequency domain transformation on the impact noise suppression data, the steady-state noise suppression data and the updated non-human voice data to obtain time domain audio data and store the data.

[0019] In this embodiment, the audio acquisition module includes a real-time acquisition unit and a pre-processing unit; The real-time acquisition unit is used to acquire real-time audio data through a high-sensitivity microphone array. The high-sensitivity microphone array is started synchronously when the camera starts recording and enters a standby state when the camera stops recording.

[0020] In this embodiment, the preprocessing unit is used to preprocess the real-time audio data to obtain pre-emphasized time domain data and spectrum data to be analyzed. The processing logic includes: The real-time audio data is pre-emphasized by a first-order high-pass filter to obtain pre-emphasized time domain data. The pre-emphasized time domain data is segmented based on a preset frame window step size and a preset frame shift rate, and the spectrum data to be analyzed is obtained by Blackman-Harris windowing and FFT calculation.

[0021] In this embodiment, the human voice detection module includes an audio endpoint detection unit and a spectrum feature detection unit, which is used to perform human voice analysis on the spectrum data to be analyzed to obtain target human voice data and non-human voice data; The audio endpoint detection unit is used to calculate the pre-emphasized time domain data to obtain the endpoint frame, and combine it with the spectrum data to be analyzed to obtain the suspected human voice data segment. The processing logic includes: Calculate the pre-emphasized time domain data to calculate the current frame energy and the historical average energy of the previous K frames, divide the current frame energy by the historical average energy of the previous K frames to obtain a short-time energy ratio, and calculate the zero-crossing rate of the current frame of the pre-emphasized time domain data; If and only if the short-time energy ratio of the current frame belongs to the preset first threshold interval and the zero-crossing rate of the current frame belongs to the preset second threshold interval, the current frame is marked as an endpoint frame, and the spectrum data to be analyzed is extracted based on the frame index of the endpoint frame to obtain a suspected human voice data segment.

[0022] In this embodiment, the spectrum feature detection unit is used to perform spectrum feature analysis on the human voice data segment to obtain target human voice data and non-human voice data; The spectrum data to be analyzed is filtered by a Gaussian Mel filter bank, the energy integral of each filter output in the Gaussian Mel filter bank is calculated to obtain the frequency band energy distribution, and the frequency band energy distribution is statistically analyzed to obtain the Gaussian Mel energy spectrum; The Gaussian Mel energy spectrum is logarithmically compressed, and the discrete cosine transform is performed on the logarithmically compressed Gaussian Mel energy spectrum to obtain Mel frequency cepstral coefficients. The Mel frequency cepstral coefficients are decomposed into three layers using the Daubechies-4 wavelet basis to obtain low-frequency approximate components and high-frequency detail components. The low-frequency approximate components are kept with their original weights, and the high-frequency detail components are multiplied by a preset amplification weight. The weighted Mel frequency cepstral coefficients are reconstructed by inverse wavelet transform. The weighted Mel-frequency cepstral coefficients are classified and regressed through the pre-trained Gaussian mixture model GMM to obtain the target human voice data and non-human voice data.

[0023] In this embodiment, the human voice highlighting module includes a steady-state noise suppression unit and an impact noise suppression unit; The steady-state noise suppression unit is used to divide the target human voice data into a first frequency band, extract the 20Hz to 500Hz part of the target human voice data through low-pass filtering to obtain the low-frequency band of the human voice data, calculate the spectral energy corresponding to the low-frequency band of each frame of human voice data, divide the spectral energy corresponding to the low-frequency human voice data segment by the spectral energy corresponding to the target human voice data to obtain a low-frequency energy ratio value. When the low-frequency energy ratio value is greater than or equal to a preset background noise threshold, it is determined that steady-state noise exists in the target human voice data of the frame, and spectral subtraction and calculation are performed on the target human voice data of the frame based on a preset noise reduction factor to obtain steady-state noise suppression data.

[0024] In this embodiment, the impulse noise suppression unit is used to perform dynamic threshold filtering on the target human voice data to obtain impulse noise suppressed data. The processing logic includes: Perform Morlet wavelet decomposition on each frame of the target human voice data to obtain Morlet decomposition layer coefficients, calculate the noise energy estimate of each layer based on the Morlet decomposition layer coefficients, and calculate the dynamic threshold based on the noise energy estimate of each layer; When the Morlet decomposition layer coefficient is less than or equal to the dynamic threshold, the Morlet decomposition layer coefficient is assigned to 0; When the Morlet decomposition layer coefficient is greater than the dynamic threshold and less than or equal to 2 times the dynamic threshold, the Morlet decomposition layer coefficient is linearly scaled; When the Morlet decomposition layer coefficient is greater than 2 times the dynamic threshold, the original value of the Morlet decomposition layer coefficient is retained; The Morlet decomposition layer coefficients after dynamic threshold processing are reconstructed by wavelet to obtain the impact noise suppression data; The dynamic threshold calculation expression is: ; represents the dynamic threshold corresponding to the i-th layer, represents the noise energy estimate corresponding to the i-th layer, N represents the total number of Morlet decomposition layer coefficients, and ln represents the natural logarithm; Morlet wavelet decomposition is a prior art known in the art. In a specific application of the present invention, the bandwidth parameter of Morlet wavelet decomposition is set to 6, the center frequency parameter is set to 2, and the number of decomposition levels is set to 5.

[0025] In this embodiment, the sound quality enhancement module includes a frequency domain division unit and a segment enhancement unit, which are used to enhance the non-human voice data to obtain updated non-human voice data; The frequency domain division unit is used to perform a second frequency band division on the frequency of the non-human voice data according to a preset non-human voice frequency band threshold, dividing the part less than the non-human voice frequency band threshold into non-human voice low frequency band data, and dividing the part greater than or equal to the non-human voice frequency band threshold into non-human voice high frequency band data.

[0026] In this embodiment, the segmented enhancement unit is used to enhance the non-human voice low-frequency data and the non-human voice high-frequency data to obtain updated non-human voice data. The processing logic includes: The non-human voice low-frequency band data is notched by comb filter to obtain wind noise removal spectrum data; Based on the AR model, the spectrum of the non-human voice high-frequency band data is extrapolated, and the high-frequency spread spectrum data is obtained through the logarithmic domain automatic gain control (AGC) with a gain compression ratio of 1:4. In this embodiment, the compression ratio of the gain adjustment is set so that when the input signal increases by 4 dB, the corresponding output signal increases by only 1 dB, thereby suppressing excessive noise and increasing the amplitude of weak signals to obtain high-frequency spread spectrum data.

[0027] The wind noise removal spectrum data and the high-frequency spread spectrum data are spliced ​​in the frequency domain to obtain updated non-human voice data.

[0028] In this embodiment, the time domain storage module includes a time-frequency domain restoration unit and a data storage unit; The time-frequency domain restoration unit is used to perform inverse FFT transformation processing on the impact noise suppression data, the steady-state noise suppression data and the updated non-human voice data, and perform overlap-add processing based on the preset frame rate and Blackman-Harris window function to obtain time-domain audio data; The data storage unit is used to add a timestamp tag to the time domain audio data and complete the storage.

[0029] A high-sensitivity microphone array and pre-emphasis-framing preprocessing technology ensure wideband, high signal-to-noise ratio (SNR) acquisition of the original audio signal. The voice detection module combines short-time energy ratio and zero-crossing rate dual-threshold endpoint detection with wavelet-based weighted Mel-frequency cepstral coefficient feature analysis to separate human voice from non-human voice using a Gaussian mixture model. A three-layer decomposition and reconstruction using the Daubechies-4 wavelet enhances speech formant characteristics while suppressing high-frequency noise interference. The voice prominence module employs dual-modal noise suppression. The steady-state noise suppression unit uses dynamic spectral subtraction triggered by low-frequency energy ratios to specifically eliminate wind noise and equipment background noise in the 20-500Hz frequency band. The impulse noise suppression unit performs Morlet wavelet decomposition and transient noise suppression using a dynamic threshold function. This helps suppress burst noise while preserving transient features such as speech plosives. The sound quality enhancement module uses frequency division processing of comb filter notch and AR model high-frequency extrapolation to eliminate low-frequency wind noise while expanding high-frequency details of ambient sound. It is suitable for application scenarios such as outdoor live broadcasts and interview shooting that require synchronous optimization of human voices in complex sound fields.

[0030] Those skilled in the art will appreciate that embodiments of the present invention may provide methods, systems, or computer program products. Therefore, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium may be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0031] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A camera audio processing system, characterized in that: include: Audio acquisition module, human voice detection module, human voice highlighting module, sound quality enhancement module and time domain storage module; The audio acquisition module is used to collect and preprocess real-time audio to obtain pre-emphasized time domain data and spectrum data to be analyzed; The human voice detection module is used to calculate the endpoint frames of the spectrum data to be analyzed to obtain the suspected human voice data segment, and calculate the weighted Mel frequency cepstral coefficient to obtain the target human voice data and non-human voice data; The human voice highlighting module is used to perform a first frequency band division and noise suppression process on the target human voice data to obtain steady-state noise suppression data and impact noise suppression data; The sound quality enhancement module is used to perform second frequency band division and frequency domain enhancement processing on the non-human voice data to obtain updated non-human voice data; The time domain storage module is used to perform time-frequency domain transformation on the impact noise suppression data, the steady-state noise suppression data and the updated non-human voice data to obtain time domain audio data and store the data.

2. The audio processing system for a photographic camera according to claim 1, wherein: The audio acquisition module includes a real-time acquisition unit and a pre-processing unit; The real-time acquisition unit is used to acquire real-time audio data through a high-sensitivity microphone array. The high-sensitivity microphone array is started synchronously when the camera starts recording and enters a standby state when the camera stops recording.

3. The audio processing system for a photographic camera according to claim 2, wherein: The preprocessing unit is used to preprocess the real-time audio data to obtain pre-emphasized time domain data and spectrum data to be analyzed. The processing logic includes: The real-time audio data is pre-emphasized by a first-order high-pass filter to obtain pre-emphasized time domain data. The pre-emphasized time domain data is segmented based on a preset frame window step size and a preset frame shift rate, and the spectrum data to be analyzed is obtained by Blackman-Harris windowing and FFT calculation.

4. The audio processing system for a photographic camera according to claim 1, wherein: The human voice detection module includes an audio endpoint detection unit and a spectrum feature detection unit, which is used to perform human voice analysis on the spectrum data to be analyzed to obtain target human voice data and non-human voice data; The audio endpoint detection unit is used to calculate the pre-emphasized time domain data to obtain the endpoint frame, and combine it with the spectrum data to be analyzed to obtain the suspected human voice data segment. The processing logic includes: Calculate the pre-emphasized time domain data to calculate the current frame energy and the historical average energy of the previous K frames, divide the current frame energy by the historical average energy of the previous K frames to obtain a short-time energy ratio, and calculate the zero-crossing rate of the current frame of the pre-emphasized time domain data; If and only if the short-time energy ratio of the current frame belongs to the preset first threshold interval and the zero-crossing rate of the current frame belongs to the preset second threshold interval, the current frame is marked as an endpoint frame, and the spectrum data to be analyzed is extracted based on the frame index of the endpoint frame to obtain a suspected human voice data segment.

5. The audio processing system for a photographic camera according to claim 4, wherein: The spectrum feature detection unit is used to perform spectrum feature analysis on the human voice data segment to obtain target human voice data and non-human voice data; The spectrum data to be analyzed is filtered by a Gaussian Mel filter bank, the energy integral of each filter output in the Gaussian Mel filter bank is calculated to obtain the frequency band energy distribution, and the frequency band energy distribution is statistically analyzed to obtain the Gaussian Mel energy spectrum; The Gaussian Mel energy spectrum is logarithmically compressed, and the discrete cosine transform is performed on the logarithmically compressed Gaussian Mel energy spectrum to obtain Mel frequency cepstral coefficients. The Mel frequency cepstral coefficients are decomposed into three layers using the Daubechies-4 wavelet basis to obtain low-frequency approximate components and high-frequency detail components. The low-frequency approximate components are kept with their original weights, and the high-frequency detail components are multiplied by a preset amplification weight. The weighted Mel frequency cepstral coefficients are reconstructed by inverse wavelet transform. The weighted Mel-frequency cepstral coefficients are classified and regressed through the pre-trained Gaussian mixture model GMM to obtain the target human voice data and non-human voice data.

6. The camera audio processing system according to claim 1, wherein: The vocal prominence module includes a steady-state noise suppression unit and an impact noise suppression unit; The steady-state noise suppression unit is used to divide the target human voice data into a first frequency band, extract the 20Hz to 500Hz part of the target human voice data through low-pass filtering to obtain the low-frequency band of the human voice data, calculate the spectral energy corresponding to the low-frequency band of each frame of human voice data, divide the spectral energy corresponding to the low-frequency human voice data segment by the spectral energy corresponding to the target human voice data to obtain a low-frequency energy ratio value. When the low-frequency energy ratio value is greater than or equal to a preset background noise threshold, it is determined that steady-state noise exists in the target human voice data of the frame, and spectral subtraction and calculation are performed on the target human voice data of the frame based on a preset noise reduction factor to obtain steady-state noise suppression data.

7. The audio processing system for a photographic camera according to claim 6, wherein: The impulse noise suppression unit is used to perform dynamic threshold filtering on the target human voice data to obtain impulse noise suppression data. The processing logic includes: Perform Morlet wavelet decomposition on each frame of the target human voice data to obtain Morlet decomposition layer coefficients, calculate the noise energy estimate of each layer based on the Morlet decomposition layer coefficients, and calculate the dynamic threshold based on the noise energy estimate of each layer; When the Morlet decomposition layer coefficient is less than or equal to the dynamic threshold, the Morlet decomposition layer coefficient is assigned to 0; When the Morlet decomposition layer coefficient is greater than the dynamic threshold and less than or equal to 2 times the dynamic threshold, the Morlet decomposition layer coefficient is linearly scaled; When the Morlet decomposition layer coefficient is greater than 2 times the dynamic threshold, the original value of the Morlet decomposition layer coefficient is retained; The Morlet decomposition layer coefficients after dynamic threshold processing are reconstructed by wavelet to obtain the impact noise suppression data; The dynamic threshold calculation expression is: ; represents the dynamic threshold corresponding to the i-th layer, represents the noise energy estimate corresponding to the i-th layer, N represents the total number of Morlet decomposition layer coefficients, and ln represents the natural logarithm; The bandwidth parameter of the Morlet wavelet decomposition is set to 6, the center frequency parameter is set to 2, and the number of decomposition levels is set to 5.

8. The camera audio processing system according to claim 1, wherein: The sound quality enhancement module includes a frequency domain division unit and a segment enhancement unit, which are used to enhance the non-human voice data to obtain updated non-human voice data; The frequency domain division unit is used to perform a second frequency band division on the frequency of the non-human voice data according to a preset non-human voice frequency band threshold, dividing the part less than the non-human voice frequency band threshold into non-human voice low frequency band data, and dividing the part greater than or equal to the non-human voice frequency band threshold into non-human voice high frequency band data.

9. The audio processing system for a photographic camera according to claim 8, wherein: The segment enhancement unit is used to enhance the non-human voice low-frequency band data and the non-human voice high-frequency band data to obtain updated non-human voice data. The processing logic includes: The non-human voice low-frequency band data is notched by comb filter to obtain wind noise removal spectrum data; Based on the AR model, the spectrum of the non-human voice high-frequency band data is extrapolated, and the high-frequency spread spectrum data is obtained through the logarithmic domain automatic gain control processing AGC of the preset gain compression ratio coefficient; The wind noise removal spectrum data and the high-frequency spread spectrum data are spliced ​​in the frequency domain to obtain updated non-human voice data.

10. The camera audio processing system according to claim 1, wherein: The time domain storage module includes a time-frequency domain restoration unit and a data storage unit; The time-frequency domain restoration unit is used to perform inverse FFT transformation processing on the impact noise suppression data, the steady-state noise suppression data and the updated non-human voice data, and perform overlap-add processing based on the preset frame rate and Blackman-Harris window function to obtain time-domain audio data; The data storage unit is used to add a timestamp tag to the time domain audio data and complete the storage.

Citation Information

Patent Citations

  • Panoramic audio processing method for panoramic camera

    CN113347530A

  • Abnormal voice detecting method based on time-domain and frequency-domain analysis

    CN102664006A

  • Audio noise reduction method and device, electronic equipment and computer readable storage medium

    CN112951259A

  • Sound processing method and system, readable storage medium and computer equipment

    CN116052726A

  • Intelligent microphone pickup and speech enhancement method, system and device

    CN120186515A