Signal processing device and audio input equipment
By processing the input signal through components such as demultiplexers, filters, converters, and processors in the signal processing device, the problem of speech detection under background noise interference is solved, and the accurate recognition and output of speech signals are achieved.
Patent Information
- Application Number
- CN202422584747.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Utility models(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2034-10-24
AI Technical Summary
Existing technologies struggle to accurately detect speech signals when background noise levels and input speech levels are close, leading to inaccurate output.
The input signal is received by a demultiplexer, processed by components such as filters, converters, smoothers and processors, the norm of the difference or ratio between frequency components is calculated, noise is filtered out by filters, the signal is enhanced by a combiner, and the characteristic signal is output by a decoder.
It improves the recognition ability and accuracy of voice signals, effectively removes noise, and ensures the accuracy of the output signal.
Smart Images

Figure CN223528197U_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The utility model relates to audio input device field, concretely relates to a signal processing device and audio input device. BACKGROUND
[0002] With the use of multimedia, microphone is an indispensable component of multimedia. Microphone belongs to audio acquisition device, and the microphone receives the sound of multiple objects at the same time, so if the existence of the sound of a specific object, such as voice, is determined by simply comparing the volume, the method has certain limitations; and if the noise is large, the processing system cannot detect the sound of the specific object received by the microphone.
[0003] The prior art currently provides a way to solve the above problems, which is a technology for detecting voice by determining the background noise level of the input voice frame and comparing the volume of the input voice frame with the threshold corresponding to the noise level. However, this technology has the problem that the background noise level and the input voice level are close, the volume of the input voice frame is too small to accurately compare or the comparison result is inaccurate, which makes it difficult to accurately output the input voice signal to the outside. UTILITY MODEL CONTENT
[0004] In order to overcome one of the deficiencies of the prior art, the purpose of the utility model is to provide a signal processing device and an audio input device, which can effectively process audio and accurately obtain input voice.
[0005] To solve the above problems, the technical scheme adopted by the utility model is as follows:
[0006] A signal processing device, comprising
[0007] A demultiplexer for receiving an input signal;
[0008] A filter for filtering the signal processed by the demultiplexer to obtain a filtered signal;
[0009] A transformer for transforming the filtered signal into an amplitude component signal in the frequency domain or time domain;
[0010] A smoother for smoothing the amplitude component signal along the frequency to obtain a smoothed amplitude component signal;
[0011] A processor for calculating the norm of the difference or ratio between the smoothed amplitude component signals between adjacent frequency components, and calculating the sum of the norms; the processor can obtain a feature signal according to the sum;
[0012] A decoder for generating an output signal according to the feature signal.
[0013] Further, the input signal comprises a mixed signal and a side signal.
[0014] Further, the processor comprises
[0015] an amplitude calculation unit for calculating a norm of a difference or a ratio between the smoothed amplitude component signals between adjacent frequency components;
[0016] a summing unit for calculating a sum of the norms;
[0017] an analysis processing unit for detecting a speech signal in the input signal according to the sum and converting into a feature signal according to the speech signal.
[0018] Further, the filter is a low-pass filter.
[0019] Further, between the filter and the transformer, further configured with:
[0020] a harmonic generation unit for generating at least one bass frequency or treble frequency harmonic based on available headroom in the filtered signal to generate a harmonic signal;
[0021] a combiner for merging and combining the harmonic signal and the filtered signal to obtain an enhanced signal and sending the enhanced signal to the decoder.
[0022] An audio input device comprising an antenna, a control board, a power supply and the signal processing device, the control board and the signal processing device are electrically connected, the power supply provides power for the controller, and the control board transmits signals to the outside through the antenna.
[0023] Compared with the prior art, the audio input device has the beneficial effects that:
[0024] The signal processing device of the utility model utilizes the demultiplexer to receive input signal, so that the received multiple different types of input signal can be signal separated and the original signal is recovered, which is beneficial to the processor for processing. The transformer transforms the signal processed by the demultiplexer into the amplitude component signal in the frequency domain or time domain, and then the smoother smoothes the amplitude component signal along the frequency or time to obtain the smoothed amplitude component signal. Such design can enhance the level value of the input signal, which is beneficial to the processor for accurately detecting and identifying the required feature signal in the input signal, and improves the identification ability and accuracy. Then the filter is used to filter the feature signal to obtain the filtered signal, which effectively removes the noise in the feature signal and ensures the accuracy of the output signal.
[0025] The utility model will be further explained in detail in combination with the drawings and specific embodiment. Attached Figure Description
[0026] Fig. 1 This is a schematic block diagram of the signal processing device in an embodiment of this utility model;
[0027] Fig. 2 This is a schematic block diagram of the signal processing device in an improved embodiment of the present invention;
[0028] Fig. 3 This is a configuration diagram of the audio input device in an embodiment of this utility model.
[0029] Explanation of icon numbers:
[0030] Demultiplexer 1, Filter 2, Converter 3, Smoother 4, Processor 5, Amplitude Calculation Unit 51, Summation Unit 52, Analysis and Processing Unit 53, Decoder 6, Harmonic Generation Unit 7, Combiner 8, Power Supply 9, Antenna 10, Control Board 11. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this utility model clearer, the present utility model will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this utility model and are not intended to limit this utility model.
[0032] Reference Figs. 1 to 3 The signal processing apparatus shown includes a demultiplexer 1, a converter 3, a smoother 4, a processor 5, a filter 2, and a decoder 6. The demultiplexer 1 receives an input signal; the filter 2 filters the signal processed by the demultiplexer 1 to obtain a filtered signal; the converter 3 transforms the filtered signal into amplitude component signals in the frequency or time domain; the smoother 4 smooths the amplitude component signals along the frequency to obtain smoothed amplitude component signals; the processor 5 calculates the norm of the difference or ratio between the smoothed amplitude component signals at adjacent frequency components and calculates the sum of the norms; the processor 5 can obtain a feature signal based on the sum; and the decoder 6 generates an output signal based on the feature signal.
[0033] By using the above structure, the possibility of determining the existence of speech in the input signal or the attribute of the speech can be improved to some extent. At the same time, by using the above design, the speech can be greatly changed in frequency, and the noise can be smoothed in frequency. It should be noted that by using the cumulative value of the norm in the frequency direction, it is determined that the speech exists with a higher probability as the cumulative value is larger. Hard decision can be performed by comparing the cumulative value with a threshold value, such as 0 / 1, or soft decision can be performed by rounding the cumulative value itself, which can be achieved by a round function. It should be noted that in the above embodiment, the norm is actually a function, which can be determined according to a preset parameter. Since there can be male voices, female voices and child voices in the speech, the parameter can be set according to the voice characteristics of different corresponding objects, and then the function of the norm can be determined.
[0034] In addition, in the above embodiment, there can be multiple mixed signals in the input signal, such as the presence of two object voices at the same time, which requires the processor 5 to determine according to the sum of the norms. The main purpose of the present application is to separate the speech signal from the noise and complete the output of the speech signal to the multimedia device. Therefore, whether there is mixed speech or not, the present application will not be smoothed in frequency.
[0035] It should be noted that in one embodiment of the present application, the input signal is transformed and divided into multiple frequency components after passing through the transformer 3, and the transformation used can use Fourier transform.
[0036] The signal processing device receives the input signal by using the demultiplexer 1, so that the received multiple different types of input signals can be separated, the original signal can be recovered, and the processor 5 can be processed. The transformer 3 transforms the signal processed by the demultiplexer 1 into an amplitude component signal in the frequency domain or time domain, and then the smoother 4 smoothes the amplitude component signal along the frequency or time to obtain a smoothed amplitude component signal. Such a design can enhance the level value of the input signal, which is beneficial to the processor 5 to accurately detect and identify the required feature signal in the input signal, and improve the ability and accuracy of identification. Then the filter 2 filters the feature signal to obtain a filtered signal, effectively removes the noise in the feature signal, and ensures that the output signal is not distorted and accurate.
[0037] Further, in the above embodiment, the input signal includes a mixed signal and side information, and the processor 5 further includes an extracting unit for extracting a characteristic signal included in the mixed signal, i.e., corresponding speech. Meanwhile, in some embodiments, the extracting unit can also be used to extract an extension type identifier indicating whether the extension region includes a residual signal from the side information, and when the extension type identifier indicates that the extension region includes a residual signal, the extracting unit extracts control restriction information for a residual usage mode from the side information. In order to facilitate processing of the above residual signal, the processor 5 further includes a residual processing unit for obtaining an enhanced object signal from the mixed signal using the residual signal. Such an arrangement enables the decoder 6 to accurately output a signal at a later stage.
[0038] Further, referring to Fig. 2 In one embodiment of the present application, the processor 5 includes
[0039] an amplitude calculation unit 51 for calculating a norm of a difference or a ratio between the smoothed amplitude component signals between adjacent frequency components;
[0040] a summing unit 52 for calculating a sum of the norms;
[0041] an analysis processing unit 53 for detecting speech signals in the input signal based on the sum and converting the speech signals into characteristic signals.
[0042] In the present embodiment, the analysis processing unit 53 is mainly used to identify speech signals. It should be noted that speech signals can be determined by calculating a norm of a difference or a ratio between the smoothed amplitude component signals between adjacent frequency components and calculating a sum of the norms. For details, reference can be made to Masakiyo Fujimoto, "The Fundamentals and Recent Progress of Voice Activity Detection", the Institute of Electronics, Information and Communication Engineers, IEICE Technical Report SP2010-23, June 2010, which discloses a method of obtaining an average of noisy signal amplitude spectra of frames in which no target sound is generated as an estimated noise spectrum. The specific scheme is not described in detail herein.
[0043] Referring to Fig. 2In an improved embodiment of the present application, considering that the speech signal in the input signal is generally a low frequency signal, the frequency range of the speech signal is 300-3400Hz, in order to filter out the noise in the environment, the filter 2 is a low pass filter, so that the high frequency noise can be effectively filtered out and most of the speech signal is retained. Although the speech signal in the input signal can be preliminarily determined by the transformer 3, the smoother 4 and the processor 5 in the present application, but there is a lot of noise in the input signal, in order to reduce the processing amount of the transformer 3, the smoother 4 and the processor 5, in an improved embodiment of the present application, in order to better remove high frequency noise, the filter 2 and the transformer 3 are further configured with:
[0044] The harmonic generation unit 7 is configured to generate at least one bass frequency or treble frequency harmonic based on the available headroom in the filtered signal to generate a harmonic signal.
[0045] The combiner 8 is configured to combine and mix the harmonic signal and the filtered signal to obtain an enhanced signal and send the enhanced signal to the decoder 6.
[0046] The filtered signal is enhanced by the combiner 8 and the harmonic generation unit 7, so that the loss is reduced. In fact, the technology in the patent application No. 13 / 592,182-Audio Adjustment System filed on April 22, 2012 in the United States can be referred to, which conveys bass enhancement setting changes to a device implementing a bass enhancement system. In fact, the combiner 8 is further connected to a level adjuster, which is configured to adaptively apply gain to at least a lower frequency band in the enhanced signal, and adaptively adjust the gain applied to the enhanced signal, wherein the gain depends on the available headroom in the filtered signal. In fact, these are all conventional technical solutions, and the present application only integrates and applies them.
[0047] Referring to Fig. 3 In addition, the present application also provides an audio input device, which comprises an antenna 10, a control board 11, a power supply 9 and the signal processing device, the control board 11 and the signal processing device are electrically connected, the power supply 9 provides power for the controller, and the control board 11 transmits signals to the outside through the antenna 10. The audio input device can actually be a microphone or other speech receiving device, and the antenna 10 can be directly connected to a loudspeaker, so that the signal can be transmitted to the outside. The audio input device can improve the signal receiving capability of the device, and is suitable as an audio input device of a multimedia device, which can better receive signals and improve the audio-visual effect.
[0048] The above-mentioned embodiments are only preferred embodiments of the present application, and cannot be used to limit the scope of protection of the present application. Any non-essential changes and replacements made by those skilled in the art on the basis of the present application shall fall within the scope of protection of the present application.
Claims
1. A signal processing device, characterized by, The input signal comprises a mixed signal and a side signal. The processor comprises an amplitude calculation unit for calculating the norm of the difference or ratio between the smoothed amplitude component signals between adjacent frequency components; a summing unit for calculating the sum of the norms; an analysis processing unit for detecting and analyzing a speech signal in the input signal according to the sum and converting the speech signal into a feature signal. The filter is a low-pass filter. The filter and the transformer are further configured with:
2. A signal processing device according to claim 1, characterized in that: a harmonic generation unit for generating at least one low or high frequency harmonic based on the available headroom in the filtered signal to generate a harmonic signal; 3. A signal processing device according to claim 1, characterized in that: a combiner for merging and combining the harmonic signal and the filtered signal to obtain an enhanced signal and sending the enhanced signal to the transformer. The signal processing device comprises an antenna, a control board, a power supply and the signal processing device of any one of claims 1-5, the control board and the signal processing device are electrically connected, the power supply provides power for the control board, and the control board transmits signals to the outside through the antenna. 4. The signal processing device of claim 1, wherein: 5. The signal processing device of claim 1, wherein: 6. An audio input device, characterized by: