A photographic camera audio processing system

The audio acquisition, voice detection and sound quality enhancement technologies of the photographic camera audio processing system solve the problem that traditional cameras cannot synchronously process high-quality audio, and achieve high signal-to-noise ratio audio acquisition and sound quality improvement, which is suitable for application scenarios with complex sound fields.

CN120600041BActive Publication Date: 2025-10-10XIAMEN NANYANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511108431.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-10-10
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

Traditional cameras cannot directly obtain high-quality audio when recording high-quality images and audio simultaneously. Users need to carry additional recording equipment, which reduces convenience and cost advantages. Existing technologies lack solutions for synchronously processing high-quality audio.

Method used

It adopts audio acquisition module, human voice detection module, human voice highlighting module, sound quality enhancement module and time domain storage module, and realizes real-time acquisition, separation and enhancement of high-quality audio through high-sensitivity microphone array, pre-emphasis-framing preprocessing, human voice detection, dual-modal noise suppression and sound quality enhancement technology.

Benefits of technology

It ensures audio acquisition with a high signal-to-noise ratio, effectively separates human voices from non-human voices, suppresses noise, and improves sound quality. It is suitable for applications that synchronously optimize human voices in complex sound fields such as outdoor live broadcasts and interview shooting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600041B_ABST
    Figure CN120600041B_ABST
Patent Text Reader

Abstract

The application discloses a kind of photographic camera audio processing systems, it is related to the technical field of audio processing, including audio acquisition module, human voice detection module, human voice highlighting module, tone quality enhancement module and time domain storage module;Audio acquisition module is used to carry out acquisition and pre-processing to real-time audio;Human voice detection module is used to obtain suspected human voice data section by endpoint frame calculation to the spectrum data to be analyzed, obtains target human voice data and non-human voice data by calculating weighted Mel frequency cepstrum coefficient;Human voice highlighting module is used to carry out first frequency band division and noise suppression processing to target human voice data, obtains steady-state noise suppression data and impact noise suppression data;Tone quality enhancement module is used to carry out second frequency band division and frequency domain enhancement processing to non-human voice data, obtains updated non-human voice data;Time domain storage module is used to obtain time domain audio data by time-frequency domain change to impact noise suppression data, steady-state noise suppression data and updated non-human voice data and storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and in particular to an audio processing system for a photographic camera. Background Art

[0002] Cameras currently on the market primarily focus on image capture. Despite significant advances in image quality, capture speed, and intelligence, they still have limitations when it comes to multimedia recording. In particular, in situations where high-quality images and audio must be recorded and uploaded simultaneously, such as news reporting, outdoor live broadcasts, and live photography, traditional cameras and mobile phones often cannot directly meet these requirements. Users often need to carry additional recording equipment, which not only increases the burden but can also lead to audio and video mismatches due to synchronization issues between devices. In addition to this lack of timeliness, post-processing of the images and audio also requires high-quality audio material to ensure the final audiovisual work is well-presented.

[0003] Currently, a Chinese invention patent application with application number CN202110405044.2 discloses a panoramic audio processing method for panoramic cameras. The application specifically includes six parts: a pre-processing module, a panoramic sound encoding module, a mono audio sound and image placement module, a dynamic sound field positioning processing module, a virtual speaker decoding module, and a psychoacoustic headphone playback module. It uses a unified audio format to reposition the playback space sound field at any angle, solving the problem of poor end-user interactivity. The method has the function of repositioning the sound field based on angle information, which can effectively combine panoramic sound format audio technology with panoramic video technology. Summary of the Invention

[0004] The technical problem solved by the present invention is that, in situations where high-quality images and audio need to be recorded simultaneously, traditional cameras are often unable to directly obtain high-quality audio. Users usually need to carry additional recording equipment, which reduces convenience and cost advantages. The existing technology lacks a solution that can synchronously process high-quality audio and enhance human voices while the camera is shooting.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0006] A photographic camera audio processing system, comprising:

[0007] Audio acquisition module, human voice detection module, human voice highlight module, sound quality enhancement module, time domain storage module

[0008] The audio acquisition module is used to collect and preprocess real-time audio to obtain pre-emphasized time domain data and spectrum data to be analyzed;

[0009] The human voice detection module is used to calculate the endpoint frames of the spectrum data to be analyzed to obtain the suspected human voice data segment, and calculate the weighted Mel frequency cepstral coefficient to obtain the target human voice data and non-human voice data;

[0010] The human voice highlighting module is used to perform a first frequency band division and noise suppression process on the target human voice data to obtain steady-state noise suppression data and impact noise suppression data;

[0011] The sound quality enhancement module is used to perform second frequency band division and frequency domain enhancement processing on the non-human voice data to obtain updated non-human voice data;

[0012] The time domain storage module is used to perform time-frequency domain transformation on the impact noise suppression data, the steady-state noise suppression data and the updated non-human voice data to obtain time domain audio data and store the data.

[0013] As a preferred solution of the camera audio processing system of the present invention, wherein:

[0014] The audio acquisition module includes a real-time acquisition unit and a pre-processing unit;

[0015] The real-time acquisition unit is used to acquire real-time audio data through a high-sensitivity microphone array. The high-sensitivity microphone array is started synchronously when the camera starts recording and enters a standby state when the camera stops recording.

[0016] As a preferred solution of the camera audio processing system of the present invention, wherein:

[0017] The preprocessing unit is used to preprocess the real-time audio data to obtain pre-emphasized time domain data and spectrum data to be analyzed. The processing logic includes:

[0018] The real-time audio data is pre-emphasized by a first-order high-pass filter to obtain pre-emphasized time domain data. The pre-emphasized time domain data is segmented based on a preset frame window step size and a preset frame shift rate, and the spectrum data to be analyzed is obtained by Blackman-Harris windowing and FFT calculation.

[0019] As a preferred solution of the camera audio processing system of the present invention, wherein:

[0020] The human voice detection module includes an audio endpoint detection unit and a spectrum feature detection unit, which is used to perform human voice analysis on the spectrum data to be analyzed to obtain target human voice data and non-human voice data;

[0021] The audio endpoint detection unit is used to calculate the pre-emphasized time domain data to obtain the endpoint frame, and combine it with the spectrum data to be analyzed to obtain the suspected human voice data segment. The processing logic includes:

[0022] Calculate the pre-emphasized time domain data to calculate the current frame energy and the historical average energy of the previous K frames, divide the current frame energy by the historical average energy of the previous K frames to obtain a short-time energy ratio, and calculate the zero-crossing rate of the current frame of the pre-emphasized time domain data;

[0023] If and only if the short-time energy ratio of the current frame belongs to the preset first threshold interval and the zero-crossing rate of the current frame belongs to the preset second threshold interval, the current frame is marked as an endpoint frame, and the spectrum data to be analyzed is extracted based on the frame index of the endpoint frame to obtain a suspected human voice data segment.

[0024] As a preferred solution of the camera audio processing system of the present invention, wherein:

[0025] The spectrum feature detection unit is used to perform spectrum feature analysis on the human voice data segment to obtain target human voice data and non-human voice data;

[0026] The spectrum data to be analyzed is filtered by a Gaussian Mel filter bank, the energy integral of each filter output in the Gaussian Mel filter bank is calculated to obtain the frequency band energy distribution, and the frequency band energy distribution is statistically analyzed to obtain the Gaussian Mel energy spectrum;

[0027] The Gaussian Mel energy spectrum is logarithmically compressed, and the discrete cosine transform is performed on the logarithmically compressed Gaussian Mel energy spectrum to obtain Mel frequency cepstral coefficients. The Mel frequency cepstral coefficients are decomposed into three layers using the Daubechies-4 wavelet basis to obtain low-frequency approximate components and high-frequency detail components. The low-frequency approximate components are kept with their original weights, and the high-frequency detail components are multiplied by a preset amplification weight. The weighted Mel frequency cepstral coefficients are reconstructed by inverse wavelet transform.

[0028] The weighted Mel-frequency cepstral coefficients are classified and regressed through the pre-trained Gaussian mixture model GMM to obtain the target human voice data and non-human voice data.

[0029] As a preferred solution of the camera audio processing system of the present invention, wherein:

[0030] The vocal prominence module includes a steady-state noise suppression unit and an impact noise suppression unit;

[0031] The steady-state noise suppression unit is used to divide the target human voice data into a first frequency band, extract the 20Hz to 500Hz part of the target human voice data through low-pass filtering to obtain the low-frequency band of the human voice data, calculate the spectral energy corresponding to the low-frequency band of each frame of human voice data, divide the spectral energy corresponding to the low-frequency human voice data segment by the spectral energy corresponding to the target human voice data to obtain a low-frequency energy ratio value. When the low-frequency energy ratio value is greater than or equal to a preset background noise threshold, it is determined that steady-state noise exists in the target human voice data of the frame, and spectral subtraction and calculation are performed on the target human voice data of the frame based on a preset noise reduction factor to obtain steady-state noise suppression data.

[0032] As a preferred solution of the camera audio processing system of the present invention, wherein:

[0033] The impulse noise suppression unit is used to perform dynamic threshold filtering on the target human voice data to obtain impulse noise suppression data. The processing logic includes:

[0034] Perform Morlet wavelet decomposition on each frame of the target human voice data to obtain Morlet decomposition layer coefficients, calculate the noise energy estimate of each layer based on the Morlet decomposition layer coefficients, and calculate the dynamic threshold based on the noise energy estimate of each layer;

[0035] When the Morlet decomposition layer coefficient is less than or equal to the dynamic threshold, the Morlet decomposition layer coefficient is assigned to 0;

[0036] When the Morlet decomposition layer coefficient is greater than the dynamic threshold and less than or equal to 2 times the dynamic threshold, the Morlet decomposition layer coefficient is linearly scaled;

[0037] When the Morlet decomposition layer coefficient is greater than 2 times the dynamic threshold, the original value of the Morlet decomposition layer coefficient is retained;

[0038] The Morlet decomposition layer coefficients after dynamic threshold processing are reconstructed by wavelet to obtain the impact noise suppression data;

[0039] The dynamic threshold calculation expression is:

[0040] ;

[0041] represents the dynamic threshold corresponding to the i-th layer, represents the noise energy estimate corresponding to the i-th layer, N represents the total number of Morlet decomposition layer coefficients, and ln represents the natural logarithm;

[0042] The bandwidth parameter of the Morlet wavelet decomposition is set to 6, the center frequency parameter is set to 2, and the number of decomposition levels is set to 5.

[0043] As a preferred scheme of the photographic camera audio processing system, wherein:

[0044] The tonal enhancement module comprises a frequency domain division unit and a segmented enhancement unit, which are used for enhancing the non-human voice data to obtain updated non-human voice data.

[0045] The frequency domain division unit is used for performing second frequency band division on the frequency of the non-human voice data according to a preset non-human voice frequency band threshold, dividing the part less than the non-human voice frequency band threshold into non-human voice low-frequency band data, and dividing the part greater than or equal to the non-human voice frequency band threshold into non-human voice high-frequency band data.

[0046] As a preferred scheme of the photographic camera audio processing system, wherein:

[0047] The segmented enhancement unit is used for respectively enhancing the non-human voice low-frequency band data and the non-human voice high-frequency band data to obtain the updated non-human voice data, and the processing logic comprises:

[0048] The non-human voice low-frequency band data is notch processed by a comb filter to obtain wind noise removal spectrum data;

[0049] The non-human voice high-frequency band data is spectrum extrapolation processed based on an AR model, and high-frequency expansion spectrum data is obtained through logarithmic domain automatic gain control (AGC) processing of a preset gain compression ratio coefficient;

[0050] The wind noise removal spectrum data and the high-frequency expansion spectrum data are frequency domain splicing processed to obtain the updated non-human voice data.

[0051] As a preferred scheme of the photographic camera audio processing system, wherein:

[0052] The time domain storage module comprises a time-frequency domain restoration unit and a data storage unit;

[0053] The time-frequency domain restoration unit is used for performing inverse FFT change processing on the impact noise suppression data, the steady-state noise suppression data and the updated non-human voice data, performing overlap-add processing based on a preset frame shift rate and a Blackman-Harris window function to obtain time domain audio data;

[0054] The data storage unit is used for adding a time stamp label to the time domain audio data and completing storage.

[0055] The application ensures the wideband high signal-to-noise ratio collection of the original audio signal through a high-sensitivity microphone array and a pre-emphasis-frame pre-processing technology. The human voice detection module combines the short-time energy ratio, the double-threshold endpoint detection of the zero-crossing rate and the feature analysis of the weighted Mel frequency cepstrum coefficient based on wavelet reconstruction, and realizes the separation of human voice and non-human voice through a Gaussian mixture model. The Daubechies-4 wavelet basis is used for three-layer decomposition and reconstruction, which is beneficial to strengthen the speech formant feature and suppress the high-frequency noise interference. The double-mode noise suppression is used in the human voice highlighting module, the low-frequency energy ratio dynamic triggering spectral subtraction is used in the stationary noise suppression unit, the wind noise and the device bottom noise in the frequency band of 20-500 Hz are eliminated, the Morlet wavelet decomposition is carried out in the impact noise suppression unit, and the transient noise suppression is realized through a dynamic threshold function, which is beneficial to retain the transient features such as speech burst and complete the suppression of the burst noise. The frequency division processing of the comb filter notch and the AR model high-frequency extrapolation is used in the sound quality enhancement module, which is beneficial to expand the high-frequency details while eliminating the low-frequency wind noise, and is suitable for the application scenarios such as outdoor live broadcast and interview shooting which need to optimize the human voice in a complex sound field. BRIEF DESCRIPTION OF DRAWINGS

[0056] Figure 1 A basic flow diagram of a photographic camera audio processing system is provided for an embodiment of the application. DETAILED DESCRIPTION

[0057] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the specific embodiments of the application will be described in detail below with reference to the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the application, rather than all the embodiments.

[0058] REFERENCE Figure 1 For an embodiment of the application, a photographic camera audio processing system is provided, comprising:

[0059] The audio acquisition module, the human voice detection module, the human voice highlighting module, the sound quality enhancement module, and the time domain storage module

[0060] The audio acquisition module is used for collecting and pre-processing the real-time audio to obtain pre-emphasis time domain data and spectrum data to be analyzed;

[0061] The human voice detection module is used for calculating the endpoint frame of the spectrum data to be analyzed to obtain suspected human voice data segments, and calculating the weighted Mel frequency cepstrum coefficient to obtain target human voice data and non-human voice data;

[0062] The human voice highlighting module is used for dividing the target human voice data into a first frequency band and performing noise suppression processing to obtain stationary noise suppression data and impact noise suppression data;

[0063] The timbre enhancement module is configured to perform second-band division and frequency domain enhancement processing on the non-human voice data to obtain updated non-human voice data.

[0064] The time domain storage module is configured to perform time-frequency domain conversion on the impact noise suppression data, the steady-state noise suppression data and the updated non-human voice data to obtain time domain audio data and store the time domain audio data.

[0065] In this embodiment, the audio acquisition module includes a real-time acquisition unit and a preprocessing unit.

[0066] The real-time acquisition unit is configured to acquire real-time audio data through a high-sensitivity microphone array, and the high-sensitivity microphone array is started synchronously when the camera starts recording and enters a standby state when the camera stops recording.

[0067] In this embodiment, the preprocessing unit is configured to preprocess the real-time audio data to obtain pre-emphasis time domain data and to-be-analyzed frequency spectrum data, and the processing logic includes:

[0068] The real-time audio data is pre-emphasized through a first-order high-pass filter to obtain pre-emphasis time domain data, the pre-emphasis time domain data is segmented based on a preset frame window step and a preset frame shift rate, and the to-be-analyzed frequency spectrum data is obtained through Blackman-Harris windowing processing and FFT calculation.

[0069] In this embodiment, the human voice detection module includes an audio endpoint detection unit and a frequency spectrum feature detection unit, configured to analyze the to-be-analyzed frequency spectrum data to obtain target human voice data and non-human voice data.

[0070] The audio endpoint detection unit is configured to calculate the pre-emphasis time domain data to obtain an endpoint frame, and obtain a suspected human voice data segment in combination with the to-be-analyzed frequency spectrum data, and the processing logic includes:

[0071] The current frame energy and the historical average energy of the previous K frames are calculated based on the pre-emphasis time domain data, the short-time energy ratio is calculated by dividing the current frame energy by the historical average energy of the previous K frames, and the zero-crossing rate of the current frame of the pre-emphasis time domain data is calculated.

[0072] The current frame is marked as an endpoint frame only when the short-time energy ratio of the current frame belongs to a preset first threshold interval and the zero-crossing rate of the current frame belongs to a preset second threshold interval, and the suspected human voice data segment is extracted from the to-be-analyzed frequency spectrum data based on the frame index of the endpoint frame.

[0073] In this embodiment, the frequency spectrum feature detection unit is configured to analyze the human voice data segment to obtain target human voice data and non-human voice data.

[0074] The spectrum data to be analyzed is filtered by a Gaussian Mel filter bank, the energy integral of each filter output in the Gaussian Mel filter bank is calculated to obtain the frequency band energy distribution, and the frequency band energy distribution is statistically analyzed to obtain the Gaussian Mel energy spectrum;

[0075] The Gaussian Mel energy spectrum is logarithmically compressed, and the discrete cosine transform is performed on the logarithmically compressed Gaussian Mel energy spectrum to obtain Mel frequency cepstral coefficients. The Mel frequency cepstral coefficients are decomposed into three layers using the Daubechies-4 wavelet basis to obtain low-frequency approximate components and high-frequency detail components. The low-frequency approximate components are kept with their original weights, and the high-frequency detail components are multiplied by a preset amplification weight. The weighted Mel frequency cepstral coefficients are reconstructed by inverse wavelet transform.

[0076] The weighted Mel-frequency cepstral coefficients are classified and regressed through the pre-trained Gaussian mixture model GMM to obtain the target human voice data and non-human voice data.

[0077] In this embodiment, the human voice highlighting module includes a steady-state noise suppression unit and an impact noise suppression unit;

[0078] The steady-state noise suppression unit is used to divide the target human voice data into a first frequency band, extract the 20Hz to 500Hz part of the target human voice data through low-pass filtering to obtain the low-frequency band of the human voice data, calculate the spectral energy corresponding to the low-frequency band of each frame of human voice data, divide the spectral energy corresponding to the low-frequency human voice data segment by the spectral energy corresponding to the target human voice data to obtain a low-frequency energy ratio value. When the low-frequency energy ratio value is greater than or equal to a preset background noise threshold, it is determined that steady-state noise exists in the target human voice data of the frame, and spectral subtraction and calculation are performed on the target human voice data of the frame based on a preset noise reduction factor to obtain steady-state noise suppression data.

[0079] In this embodiment, the impulse noise suppression unit is used to perform dynamic threshold filtering on the target human voice data to obtain impulse noise suppressed data. The processing logic includes:

[0080] Perform Morlet wavelet decomposition on each frame of the target human voice data to obtain Morlet decomposition layer coefficients, calculate the noise energy estimate of each layer based on the Morlet decomposition layer coefficients, and calculate the dynamic threshold based on the noise energy estimate of each layer;

[0081] When the Morlet decomposition layer coefficient is less than or equal to the dynamic threshold, the Morlet decomposition layer coefficient is assigned to 0;

[0082] When the Morlet decomposition layer coefficient is greater than the dynamic threshold and less than or equal to 2 times the dynamic threshold, the Morlet decomposition layer coefficient is linearly scaled;

[0083] When the Morlet decomposition layer coefficient is greater than 2 times the dynamic threshold, the original value of the Morlet decomposition layer coefficient is reserved;

[0084] The Morlet decomposition layer coefficient after dynamic threshold processing is calculated by wavelet reconstruction to obtain impact noise suppression data;

[0085] The dynamic threshold calculation expression is:

[0086] ;

[0087] The dynamic threshold corresponding to the i-th layer is represented by The noise energy estimate corresponding to the i-th layer is represented by N, which represents the total number of Morlet decomposition layer coefficients, and ln represents the natural logarithm;

[0088] Morlet wavelet decomposition is a prior art known in the art, and in the specific application of the present application, the bandwidth parameter of Morlet wavelet decomposition is set to 6, the center frequency parameter is set to 2, and the decomposition layer number is set to 5.

[0089] In this embodiment, the sound quality enhancement module includes a frequency domain division unit and a segmented enhancement unit, which are used to enhance the non-vocal data to obtain updated non-vocal data.

[0090] The frequency domain division unit is used to perform a second frequency band division on the frequency of the non-vocal data according to a preset non-vocal frequency band threshold, divide the part less than the non-vocal frequency band threshold into non-vocal low frequency band data, and divide the part greater than or equal to the non-vocal frequency band threshold into non-vocal high frequency band data.

[0091] In this embodiment, the segmented enhancement unit is used to enhance the non-vocal low frequency band data and the non-vocal high frequency band data respectively to obtain updated non-vocal data, and the processing logic includes:

[0092] The non-vocal low frequency band data is notch processed by a comb filter to obtain wind noise removal spectrum data;

[0093] The non-vocal high frequency band data is spectrum extrapolation processed based on an AR model, and high frequency expansion spectrum data is obtained by logarithmic domain automatic gain control processing AGC with a gain compression ratio coefficient of 1:4.

[0094] In this embodiment, the compression ratio of the gain adjustment is set to 1:4, so that the input signal increases by only 1 dB when it increases by 4 dB, thereby suppressing strong noise and enhancing weak signal amplitude to obtain high frequency expansion spectrum data.

[0095] The wind noise removal spectrum data and the high frequency expansion spectrum data are frequency domain splicing processed to obtain updated non-vocal data.

[0096] In this embodiment, the time domain storage module includes a time-frequency domain restoration unit and a data storage unit;

[0097] The time-frequency domain restoration unit is used to perform inverse FFT transformation processing on the impact noise suppression data, the steady-state noise suppression data and the updated non-human voice data, and perform overlap-add processing based on the preset frame rate and Blackman-Harris window function to obtain time-domain audio data;

[0098] The data storage unit is used to add a timestamp tag to the time domain audio data and complete the storage.

[0099] A high-sensitivity microphone array and pre-emphasis-framing preprocessing technology ensure wideband, high signal-to-noise ratio (SNR) acquisition of the original audio signal. The voice detection module combines short-time energy ratio and zero-crossing rate dual-threshold endpoint detection with wavelet-based weighted Mel-frequency cepstral coefficient feature analysis to separate human voice from non-human voice using a Gaussian mixture model. A three-layer decomposition and reconstruction using the Daubechies-4 wavelet enhances speech formant characteristics while suppressing high-frequency noise interference. The voice prominence module employs dual-modal noise suppression. The steady-state noise suppression unit uses dynamic spectral subtraction triggered by low-frequency energy ratios to specifically eliminate wind noise and equipment background noise in the 20-500Hz frequency band. The impulse noise suppression unit performs Morlet wavelet decomposition and transient noise suppression using a dynamic threshold function. This helps suppress burst noise while preserving transient features such as speech plosives. The sound quality enhancement module uses frequency division processing of comb filter notch and AR model high-frequency extrapolation to eliminate low-frequency wind noise while expanding high-frequency details of ambient sound. It is suitable for application scenarios such as outdoor live broadcasts and interview shooting that require synchronous optimization of human voices in complex sound fields.

[0100] Those skilled in the art will appreciate that embodiments of the present invention may provide methods, systems, or computer program products. Therefore, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium may be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0101] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A camera audio processing system, characterized in that: include: Audio acquisition module, human voice detection module, human voice highlighting module, sound quality enhancement module and time domain storage module; The audio acquisition module is used to collect and preprocess real-time audio to obtain pre-emphasized time domain data and spectrum data to be analyzed; The human voice detection module is used to calculate the endpoint frames of the spectrum data to be analyzed to obtain the suspected human voice data segment, and calculate the weighted Mel frequency cepstral coefficient to obtain the target human voice data and non-human voice data; The human voice highlighting module is used to perform a first frequency band division and noise suppression process on the target human voice data to obtain steady-state noise suppression data and impact noise suppression data; The sound quality enhancement module is used to perform second frequency band division and frequency domain enhancement processing on the non-human voice data to obtain updated non-human voice data; The time domain storage module is used to perform time-frequency domain transformation on the impact noise suppression data, the steady-state noise suppression data and the updated non-human voice data to obtain time domain audio data and store the data.

2. The audio processing system for a photographic camera according to claim 1, wherein: The audio acquisition module includes a real-time acquisition unit and a pre-processing unit; The real-time acquisition unit is used to acquire real-time audio data through a high-sensitivity microphone array. The high-sensitivity microphone array is started synchronously when the camera starts recording and enters a standby state when the camera stops recording.

3. The audio processing system for a photographic camera according to claim 2, wherein: The preprocessing unit is used to preprocess the real-time audio data to obtain pre-emphasized time domain data and spectrum data to be analyzed. The processing logic includes: The real-time audio data is pre-emphasized by a first-order high-pass filter to obtain pre-emphasized time domain data. The pre-emphasized time domain data is segmented based on a preset frame window step size and a preset frame shift rate, and the spectrum data to be analyzed is obtained by Blackman-Harris windowing and FFT calculation.

4. The audio processing system for a photographic camera according to claim 1, wherein: The human voice detection module includes an audio endpoint detection unit and a spectrum feature detection unit, which is used to perform human voice analysis on the spectrum data to be analyzed to obtain target human voice data and non-human voice data; The audio endpoint detection unit is used to calculate the pre-emphasized time domain data to obtain the endpoint frame, and combine it with the spectrum data to be analyzed to obtain the suspected human voice data segment. The processing logic includes: Calculate the pre-emphasized time domain data to calculate the current frame energy and the historical average energy of the previous K frames, divide the current frame energy by the historical average energy of the previous K frames to obtain a short-time energy ratio, and calculate the zero-crossing rate of the current frame of the pre-emphasized time domain data; If and only if the short-time energy ratio of the current frame belongs to the preset first threshold interval and the zero-crossing rate of the current frame belongs to the preset second threshold interval, the current frame is marked as an endpoint frame, and the spectrum data to be analyzed is extracted based on the frame index of the endpoint frame to obtain a suspected human voice data segment.

5. The audio processing system for a photographic camera according to claim 4, wherein: The spectrum feature detection unit is used to perform spectrum feature analysis on the human voice data segment to obtain target human voice data and non-human voice data; The spectrum data to be analyzed is filtered by a Gaussian Mel filter bank, the energy integral of each filter output in the Gaussian Mel filter bank is calculated to obtain the frequency band energy distribution, and the frequency band energy distribution is statistically analyzed to obtain the Gaussian Mel energy spectrum; The Gaussian Mel energy spectrum is logarithmically compressed, and the discrete cosine transform is performed on the logarithmically compressed Gaussian Mel energy spectrum to obtain Mel frequency cepstral coefficients. The Mel frequency cepstral coefficients are decomposed into three layers using the Daubechies-4 wavelet basis to obtain low-frequency approximate components and high-frequency detail components. The low-frequency approximate components are kept with their original weights, and the high-frequency detail components are multiplied by a preset amplification weight. The weighted Mel frequency cepstral coefficients are reconstructed by inverse wavelet transform. The weighted Mel-frequency cepstral coefficients are classified and regressed through the pre-trained Gaussian mixture model GMM to obtain the target human voice data and non-human voice data.

6. The camera audio processing system according to claim 1, wherein: The vocal prominence module includes a steady-state noise suppression unit and an impact noise suppression unit; The steady-state noise suppression unit is used to divide the target human voice data into a first frequency band, extract the 20Hz to 500Hz part of the target human voice data through low-pass filtering to obtain the low-frequency band of the human voice data, calculate the spectral energy corresponding to the low-frequency band of each frame of human voice data, divide the spectral energy corresponding to the low-frequency human voice data segment by the spectral energy corresponding to the target human voice data to obtain a low-frequency energy ratio value. When the low-frequency energy ratio value is greater than or equal to a preset background noise threshold, it is determined that steady-state noise exists in the target human voice data of the frame, and spectral subtraction is performed on the target human voice data of the frame based on a preset noise reduction factor to calculate steady-state noise suppression data.

7. The audio processing system for a photographic camera according to claim 6, wherein: The impulse noise suppression unit is used to perform dynamic threshold filtering on the target human voice data to obtain impulse noise suppression data. The processing logic includes: Perform Morlet wavelet decomposition on each frame of the target human voice data to obtain Morlet decomposition layer coefficients, calculate the noise energy estimate of each layer based on the Morlet decomposition layer coefficients, and calculate the dynamic threshold based on the noise energy estimate of each layer; When the Morlet decomposition layer coefficient is less than or equal to the dynamic threshold, the Morlet decomposition layer coefficient is assigned to 0; When the Morlet decomposition layer coefficient is greater than the dynamic threshold and less than or equal to 2 times the dynamic threshold, the Morlet decomposition layer coefficient is linearly scaled; When the Morlet decomposition layer coefficient is greater than 2 times the dynamic threshold, the original value of the Morlet decomposition layer coefficient is retained; The Morlet decomposition layer coefficients after dynamic threshold processing are reconstructed by wavelet to obtain the impact noise suppression data; The dynamic threshold calculation expression is: ; represents the dynamic threshold corresponding to the i-th layer, represents the noise energy estimate corresponding to the i-th layer, N represents the total number of Morlet decomposition layer coefficients, and ln represents the natural logarithm; The bandwidth parameter of the Morlet wavelet decomposition is set to 6, the center frequency parameter is set to 2, and the number of decomposition levels is set to 5.

8. The camera audio processing system according to claim 1, wherein: The sound quality enhancement module includes a frequency domain division unit and a segment enhancement unit, which are used to enhance the non-human voice data to obtain updated non-human voice data; The frequency domain division unit is used to perform a second frequency band division on the frequency of the non-human voice data according to a preset non-human voice frequency band threshold, dividing the part less than the non-human voice frequency band threshold into non-human voice low frequency band data, and dividing the part greater than or equal to the non-human voice frequency band threshold into non-human voice high frequency band data.

9. The audio processing system for a photographic camera according to claim 8, wherein: The segment enhancement unit is used to enhance the non-human voice low-frequency band data and the non-human voice high-frequency band data to obtain updated non-human voice data. The processing logic includes: The non-human voice low-frequency band data is notched by comb filter to obtain wind noise removal spectrum data; Based on the AR model, the spectrum of the non-human voice high-frequency band data is extrapolated, and the high-frequency spread spectrum data is obtained through the logarithmic domain automatic gain control processing AGC of the preset gain compression ratio coefficient; The wind noise removal spectrum data and the high-frequency spread spectrum data are spliced ​​in the frequency domain to obtain updated non-human voice data.

10. The camera audio processing system according to claim 1, wherein: The time domain storage module includes a time-frequency domain restoration unit and a data storage unit; The time-frequency domain restoration unit is used to perform inverse FFT transformation processing on the impact noise suppression data, the steady-state noise suppression data and the updated non-human voice data, and perform overlap-add processing based on the preset frame rate and Blackman-Harris window function to obtain time-domain audio data; The data storage unit is used to add a timestamp tag to the time domain audio data and complete the storage.

Citation Information

Patent Citations

  • Panoramic audio processing method for panoramic camera

    CN113347530A

  • Abnormal voice detecting method based on time-domain and frequency-domain analysis

    CN102664006A

  • Audio noise reduction method and device, electronic equipment and computer readable storage medium

    CN112951259A