Low power consumption recording screen system based on fft and dynamic characteristic fusion technology

This recording jamming system, which uses FFT and dynamic feature fusion technology, solves the shortcomings of existing recording jammers in terms of accuracy, battery life, and coverage, achieving low power consumption and high efficiency in recording jamming, and is suitable for various scenarios.

CN121054018BActive Publication Date: 2026-07-07SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTHEAST UNIV
Filing Date
2025-08-28
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

Existing recording jammers suffer from problems such as unstable interference effects, difficulty in accurately distinguishing human voices from background music, strong directionality, poor battery life, high price, and limited shielding range in complex environments.

Method used

A recording shielding system based on FFT and dynamic feature fusion technology is adopted. It distinguishes between human voices and non-human voices through spectrum analysis, optimizes the shielding angle and distance by using array probe layout, and combines Class D power amplifier technology to generate and transmit low-power interference signals. An analog subtractor is used to achieve signal complementarity, ensuring the adaptability and reliability of interference signals.

Benefits of technology

It achieves intelligent recognition and targeted masking of human voices and music, with wide coverage, low power consumption, and low cost. It can effectively mask recording devices in complex environments and is suitable for various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121054018B_ABST
    Figure CN121054018B_ABST
Patent Text Reader

Abstract

The application discloses an extremely low-power recording shielding system based on FFT and dynamic feature fusion technology, and aims to effectively protect information security and privacy through technical means. The system starts from a design framework, and includes sound signal collection, analog-digital conversion, digital signal processing, interference signal generation, digital-analog conversion and power amplification and the like. First, a spectrum analysis method is used to distinguish human voice signals and non-human voice signals, so as to provide a basis for generation of interference signals, and a signal processing path is further optimized, through human voice or non-human voice recognition and signal selection and processing, adaptability and reliability of the interference signals are ensured. The application provides a theoretical basis and technical support for further development of the recording shielding device, and is expected to play an important role in the field of information security and privacy protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of audio signal processing and dynamic feature recognition and fusion technology, specifically relating to an ultra-low power recording shielding system based on FFT and dynamic feature fusion technology. Background Technology

[0002] In today's digital age, information security and privacy protection have become crucial issues. With the widespread use of recording devices and continuous technological advancements, the privacy of individuals and organizations faces unprecedented challenges. Whether it's a confidential business meeting, a private personal conversation, or other situations requiring confidentiality, the risk of recording leaks is constantly increasing. To effectively address this issue, recording masking technology has emerged.

[0003] A recording jammer is a device that interferes with recording equipment using specific techniques, preventing it from clearly recording a target sound. Its core principle is to generate a specific interference signal and superimpose it onto the target sound, thereby masking the target sound within the recording equipment. Traditional recording jammers often use random signal superposition, but this method has many limitations, such as unstable interference effects and the inability to accurately distinguish between human voices and background music. With technological advancements, the market has placed higher demands on recording jammers, requiring them to process specific audio signals more accurately and achieve more efficient recording jamming effects.

[0004] Recording jammers primarily work by emitting interference signals at specific frequencies to mask the target sound, thus blocking recording. Their working principle is based on the superposition of sound waves; when the interference signal and the target sound signal are superimposed within the same frequency range, interference occurs, preventing the recording equipment from clearly recording the target sound. Common recording jammers use white noise or pink noise as the interference signal source. These noise signals have wide bandwidth characteristics, covering a broad frequency range, thus interfering with various recording devices. For example, white noise typically has a frequency range between 20Hz and 20kHz, which coincides with the frequency range of sound heard by the human ear, therefore it can effectively interfere with most recording devices.

[0005] Currently, the ultrasonic interference and frequency mixing technologies commonly used in recording jammers on the market can achieve a shielding effect in certain scenarios, but they have many limitations:

[0006] (1) The propagation characteristics of ultrasound determine that it is highly directional and attenuates quickly, making it easily blocked or reflected by obstacles (such as walls and furniture), resulting in a small actual effective shielding range (usually only 1-2 meters). In open spaces or complex environments, interference signals are difficult to cover evenly, and shielding dead zones may occur, leaving an opportunity for recording equipment.

[0007] (2) Since the signal pattern of the mixing interference is relatively fixed (such as fixed frequency combination or predictable random pattern), professionals can identify the characteristics of the interference signal through spectrum analysis, and then remove the interference through filtering and noise reduction algorithms to restore the original recording content. The reliability of the shielding effect is low, and some devices can filter out the ultrasonic signal through software filtering technology, resulting in interference failure.

[0008] (3) In order to ensure the interference effect, the ultrasonic transmitter needs to continuously output high power, resulting in poor equipment endurance. It usually requires an external power supply or frequent charging, which makes it inconvenient to use.

[0009] (4) Background noise in different scenarios (such as the sound of air conditioning in the office and traffic noise outside) may overlap with the frequency mixing interference signal, causing the interference signal to be "submerged"; while in a quiet environment, the noise generated by frequency mixing interference will be too obvious, exposing the shielding behavior and arousing vigilance.

[0010] (5) Currently, the price of recording jammers on the market varies greatly, mainly due to factors such as product performance, functions, and brand. According to market research data, low-end recording jammers are priced between 500 and 1000 yuan. These products typically have simple functions and limited interference effects, and are mainly used in small conference rooms or personal privacy protection scenarios. Mid-range recording jammers are priced between 1000 and 3000 yuan, and have stronger interference capabilities and more stable performance, suitable for medium-sized conference rooms or commercial venues. High-end recording jammers are priced above 3000 yuan. These products use advanced interference technology and high-quality hardware components, and can provide a wider frequency coverage and stronger interference effects. They are suitable for large conference rooms, important meeting venues, or scenarios with extremely high recording jamming requirements. These high-end products use multi-band interference technology, which can simultaneously cover multiple frequency bands of sound signals, and the interference effect is significantly better than ordinary products. However, the price is very high, which makes it difficult to popularize information security in society. Summary of the Invention

[0011] To address the aforementioned issues, this invention discloses an ultra-low power recording shielding system based on FFT and dynamic feature fusion technology. By using spectrum analysis, it distinguishes between human and non-human voice signals, providing a basis for generating interference signals. Then, it selects and processes human or non-human voice signals, ensuring the adaptability and reliability of interference signals. The system is low-cost, has a wide coverage, and boasts strong battery life.

[0012] To achieve the above objectives, the technical solution of the present invention is as follows:

[0013] The ultra-low power recording shielding system based on FFT and dynamic feature fusion technology includes the following hardware components: an input module, a dynamic feature recognition module, a shielding path selection module, a data fusion signal processing system module, and a shielding signal output module.

[0014] The input module acquires sound signals via a microphone, which converts them into weak electrical signals. However, these weak electrical signals are easily masked by noise and are too faint for the main control unit to acquire and analyze. Therefore, the subsequent stage of this module uses an AD620 instrumentation amplifier for signal conditioning. The AD620 features low offset voltage (50μV max) and high common-mode rejection ratio (100dB min), effectively suppressing common-mode interference in the input signal and amplifying only the differential signal component. This is particularly important for measuring weak signals in complex electromagnetic environments, significantly improving signal accuracy and reliability. Finally, the input module provides a dynamic electrical signal with a high signal-to-noise ratio for analysis by the subsequent dynamic feature recognition module.

[0015] The design scheme of the dynamic feature recognition module is as follows:

[0016] The amplitude of the electrical signal acquired by the input module in the time domain consists of human voice and non-human voice signals, and the dynamic characteristics of the two are distinguished using the frequency domain. The MCU acquires the high signal-to-noise ratio dynamic electrical signal provided by the input module through the ADC acquisition circuit, and selects...

[0017] Microcontrollers such as STM32F103RCT6 convert signals into digital signals; the acquired digital electrical signals undergo Fast Fourier Transform, spectrum analysis, and machine learning algorithms to achieve automatic recognition and feature extraction of human and non-human voice signals; based on the recognition results and preset rules, it determines whether to initiate the transmission of a shielded signal; when there is no audio signal or a non-target signal is detected, it controls the cessation of shielded signal transmission, achieving intelligent control.

[0018] The masking path selection module allows for the separate masking of human voices and non-human voices by selecting a path:

[0019] The audio signal output from the input module is connected to the shielded path selection module, which is controlled by two MOSFETs. The audio signal is connected to the drain of the two MOSFETs, and the sources of the two MOSFETs are connected to two different filters. Filter 1 is a bandpass filter with a frequency range of 85-525Hz, which preserves the dynamic characteristics of human voices. Filter 2 is a high-pass filter with a cutoff frequency of 600Hz, which preserves the dynamic characteristics of non-human voices. The MOSFETs are turned on and off by two high and low level inputs to the gate, thereby selecting which filter the audio signal enters, thus distinguishing between human voices and non-human voices, and extracting the dynamic characteristics of either human voice or non-human voice, providing the signal to be processed for the subsequent data fusion signal processing module.

[0020] The design scheme of the data fusion signal processing system module is as follows: when recording, if it is desired that a certain component of the sound signal (hereinafter referred to as human voice) will not be recorded or that the power of the human voice signal will be attenuated throughout the sound propagation, then the ratio of the dynamic characteristic E0 of the energy of that component to the total sound energy E1 should reach a constant value, let's call it... In general, non-human voices are considered ambient noise. Commercially available recording jammers amplify the energy of non-human voices (E1 is much greater than E0) (this amplification is achieved by applying large random noise). This results in the weaker human voice signal being buried in the high-energy ambient noise (generated by the recording jammer) during recording, with the energy ratio E consistently around 0. To achieve this effect, typical recording jammers require a large amount of energy to generate the "ambient noise." However, this system innovatively uses an analog subtractor to achieve signal complementarity for the extracted analog electrical signal (the output of the jamming path selection module). (Assuming the goal is to prevent human voices from being recorded, the recording jammer outputs a signal complementary to the dynamic characteristics of human voices; this signal is superimposed on the human voice signal.) Through energy superposition and the principle of weakening dynamic characteristics, the recording device captures only background music while attenuating human voices (attenuation ≥20dB). This innovative design brings the dynamic characteristic E0 of the human voice signal and the output signal of the shield to approximately zero, while the total sound energy E1 remains essentially unchanged. This achieves a constant energy ratio E of approximately 0, making it impossible to extract the dynamic characteristics of the human voice during recording. The hardware design of this module involves a high-speed analog subtractor circuit at the amplifier input. (A hypothetical example of this design (using human voice removal as an example; the same principle applies to non-human voice removal): If we want to prevent human voice from being recorded, assuming the human voice is 1 + sin(x) ≥ 0 (sin(x) is its dynamic characteristic), then the filter extracts the dynamic characteristics of the human voice signal. The analog subtractor complements this dynamic characteristic, outputting 1 - sin(x) ≥ 0. Through energy superposition, the amplitude of the received signal during recording remains constant at 2, thus losing the dynamic characteristics of the human voice and achieving the recording shielding effect.)

[0021] The invention design of the shielded signal output module is based on a Class D power amplifier and an ultrasonic probe (with a conversion efficiency of 80%-95%). The analog signal, after filtering and a high-speed subtractor, is sent to "Amplifier 2" for further amplification. A Class D power amplifier (such as TPA3116) is used to amplify the power of the shielded signal, increasing the voltage to a level sufficient to drive the ultrasonic probe (approximately 20V). The amplified shielded signal is then emitted as ultrasonic waves through the ultrasonic probe. By utilizing the high conversion efficiency of the Class D power amplifier, the hardware design of this system can meet the requirement of a power input of ≤6W. Simultaneously, the array probe layout optimizes the shielding angle (≥60°) and distance (≥10 meters), thus preventing interference with the normal recording of the recording equipment.

[0022] The ultra-low power recording shielding system based on FFT and dynamic feature fusion technology is used as follows:

[0023] (1) First, the sound electrical signal is amplified by "Amplifier 1 (AD620 circuit)". The amplified signal is then sent to "MCU's ADC (Analog-to-Digital Converter)". A microcontroller (MCU) such as STM32F103RCT6 is used to convert it into a digital signal. The digital signal is then sent to the "MCU Analysis (Human Voice or Non-Human Voice)" dynamic feature recognition module, where it is analyzed to distinguish between human voice and non-human voice.

[0024] (2) Based on the recognition results and preset rules, the shielding path selection module is activated to determine whether to start the shielding signal transmission. When there is no audio signal or a non-target signal is detected, the shielding signal transmission is stopped to achieve intelligent control. When a target signal is detected, the shielding path selection module controls the start and stop of the MOS transistor through two high and low level input gates, thereby selecting which filter the sound signal enters, and extracting the dynamic features of either human voice or non-human voice, providing the signal to be processed for the subsequent data fusion signal processing module.

[0025] (3) The electrical signal extracted by a specific filter is complemented by an analog subtractor. Through the principle of energy superposition and weakening dynamic characteristics, the recording device captures only the background music and weakens the human voice.

[0026] (4) The filtered analog signal is sent to “Amplifier 2” for amplification. The amplified shielding signal is emitted in the form of ultrasonic waves through the ultrasonic probe. The shielding angle and distance are optimized by the array probe layout to interfere with the normal recording of the recording equipment.

[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0028] (1) The present invention uses an array-type probe layout to optimize the shielding angle (≥60°) and distance (≥10 meters), and its directional coverage is wider than that of the recording jammers on the market, and its actual effective range is larger.

[0029] (2) This invention's system is based on audio signal feature analysis and integrates machine learning and acoustic interference technology to achieve intelligent recognition (92% accuracy) and directional shielding (attenuation ≥20dB) of human voice and music signals. Furthermore, due to the superposition of complementary signal energy, the recording system cannot extract human voice content through spectrum analysis, thus ensuring strong information security. Most recording shielding systems on the market achieve the effect of recording failure by superimposing random high-frequency noise onto the sound source. Although this method is simple and easy to implement, if criminals obtain the recording material, they can easily purchase the existing recording shielding device to decode the random noise and then extract the human voice content from the recording material using FFT or DFT techniques.

[0030] (3) In terms of power consumption, the present invention has extremely low power consumption (which can be set to within 6W and the power conversion ratio is as high as 90%), and has a good shielding effect under this extremely low power consumption.

[0031] (4) The present invention has dynamic feature fusion technology, which can automatically adjust feature weights according to real-time signal strength, such as enhancing the proportion of frequency domain features in low signal-to-noise ratio environments; and has lightweight deployment capability, which can compress the storage of SVM model parameters to less than 500KB through model compression technology (such as pruning and quantization), and adapt to embedded platforms such as STM32.

[0032] (5) The cost of the entire system is far lower than that of basic recording jamming systems on the market (the implementation cost of this system is less than RMB 100). Therefore, this invention demonstrates significant application value in the field of privacy protection. Its design concept provides a theoretical basis and technical support for the further development of recording jammers, and can play an important role in the fields of information security and privacy protection. Attached Figure Description

[0033] Figure 1 This is a flowchart illustrating the overall structure of the present invention.

[0034] Figure 2 This is a design diagram of the weak signal conditioning circuit of the present invention.

[0035] Figure 3 This is a design diagram of the filter channel selection module of the present invention.

[0036] Figure 4 This is a circuit design diagram of a specific filter 1 of the present invention.

[0037] Figure 5 This is a circuit design diagram of the specific filter 2 of the present invention.

[0038] Figure 6 This is a circuit design diagram of the high-speed analog subtractor of the present invention.

[0039] Figure 7 This is the low-power control output design topology of the present invention.

[0040] Figure 8 This is a flowchart of the main control algorithm design of the present invention. Detailed Implementation

[0041] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0042] As shown in the figure, this invention discloses an ultra-low power recording jamming system based on an FFT data fusion framework. This system, grounded in audio signal feature analysis, integrates machine learning and acoustic interference techniques. Starting with the system design framework, it proposes a signal processing flow, including sound signal acquisition, analog-to-digital conversion, digital signal processing, interference signal generation, digital-to-analog conversion, and power amplification. First, spectral analysis is used to distinguish between human and non-human voice signals, providing a basis for generating interference signals. Then, the signal processing path is further optimized through human or non-human voice recognition and signal selection and processing, ensuring the adaptability and reliability of the interference signals.

[0043] The whole process is as follows Figure 1 As shown, the system acquires sound signals through a microphone and uses Fast Fourier Transform (FFT), spectrum analysis, and machine learning algorithms to automatically identify human voices (85-525Hz) and music signals (accuracy 92%, recognition time ≤0.8 seconds). Signal conditioning is performed using an AD620 instrumentation amplifier, employing MOSFETs to switch shielded paths and combining a Butterworth filter (human voice bandpass 80-160Hz, non-human voice high-pass 450Hz) to extract the target signal. The electrical signal extracted by this specific filter is then used to achieve signal complementarity using an analog subtractor. Through energy superposition and the principle of weakening dynamic characteristics, an interference signal is generated, which is then driven by a Class D power amplifier (efficiency 80%-95%) to transmit the interference signal via an ultrasonic probe.

[0044] like Figure 2 As shown, the weak signal output by the microphone needs to be conditioned before it can be processed by subsequent circuits. An AD620 instrumentation amplifier is used to build the amplification circuit. The AD620 features low offset voltage (50μV max) and high common-mode rejection ratio (100dB min), effectively suppressing common-mode interference in the input signal and amplifying only the differential signal component. This is particularly important for measuring weak signals in complex electromagnetic environments, significantly improving signal accuracy and reliability. The gain is adjusted by connecting an external resistor RG. The microphone signal can be amplified to approximately 0.5V to meet the ADC sampling requirements. Simultaneously, an RC filter circuit is added to filter out high-frequency noise and improve signal purity.

[0045] like Figure 3 As shown, the shielding path selection module uses two N-channel MOSFETs to build a switching circuit. The MCU, controlled by an STM32F103RCT6, controls the MOSFETs' on / off state by outputting high and low levels through GPIO pins. When a human voice signal is detected, the MCU controls the corresponding pin to output a high level, turning on the MOSFET connected to the human voice processing path, and the signal enters the human voice shielding signal generation circuit. Similarly, when a music signal is detected, it switches to the music signal processing path. Its function allows users to manually or automatically select the mode of shielding human voices, non-human voices, or both simultaneously, according to their actual needs. Specifically, the microcontroller enables two buttons (button 1 and button 2 are NOT related). When button 1 is pressed, one pin of the MCU (Pin1=1) outputs a high level (Pin2=0). The high level is input to the gate (G) of the MOSFET, and the drain (sound signal) and source (filter 1 input) are turned on. When button 2 is pressed, another pin of the MCU (Pin2=1) outputs a high level (Pin1 becomes 0 at this time). The high level is input to the gate (G) of the MOSFET, and the drain (sound signal) and source (filter 2 input) are turned on.

[0046] like Figure 1 and 4 As shown, this bandpass filter (filter 1) is designed, which can extract human voice signals from 85Hz to 150Hz for subsequent FFT processing, thereby eliminating the human voice signal.

[0047] like Figure 1 and 5 As shown, this high-pass filter (filter 2) is designed, which can extract non-human voice signals (including white noise and music) above 500Hz for subsequent FFT processing, thereby eliminating non-human voice signals.

[0048] like Figure 1 and 6 As shown, the design scheme of the data fusion signal processing system module is that, during recording, if it is desired that a certain component of the sound signal (hereinafter referred to as human voice) will not be recorded or that the power of the human voice signal will be attenuated during the entire sound propagation, then the ratio of the dynamic characteristic E0 of the energy of that component to the total sound energy E1 should reach a constant value, let's call it... In general, non-human voices are considered ambient noise. Commercially available recording jammers amplify the energy of non-human voices (E1 is much greater than E0) (this amplification is achieved by applying large random noise). This results in the weaker human voice signal being buried in the high-energy ambient noise (generated by the recording jammer) during recording, with the energy ratio E consistently around 0. To achieve this effect, typical recording jammers require a large amount of energy to generate the "ambient noise." However, this system innovatively uses an analog subtractor to achieve signal complementarity for the extracted analog electrical signal (the output of the jamming path selection module). (Assuming the goal is to prevent human voices from being recorded, the recording jammer outputs a signal complementary to the dynamic characteristics of human voices; this signal is superimposed on the human voice signal.) Through energy superposition and the principle of weakening dynamic characteristics, the recording device captures only background music while attenuating human voices (attenuation ≥20dB). This innovative design reduces the dynamic characteristic E0 of the human voice signal and the output signal of the shield to approximately zero, while the total sound energy e1 remains essentially unchanged. This achieves a constant energy ratio e of approximately 0, making it impossible to extract the dynamic characteristics of the human voice during recording. The hardware design of this module involves a high-speed analog subtractor circuit at the amplifier input. (An embodiment of this design (taking human voice removal as an example; the same principle applies to non-human voice removal): To prevent human voice from being recorded, assume the human voice is 1 + sin(x) ≥ 0 (sin(x) is its dynamic characteristic). The filter extracts the dynamic characteristics of the human voice signal, and the analog subtractor complements this dynamic characteristic, outputting 1 - sin(x) ≥ 0. Through energy superposition, the amplitude of the received signal during recording remains constant at 2, thus losing the dynamic characteristics of the human voice and achieving the recording shielding effect.)

[0049] like Figure 1 and 7As shown, the programmable gain amplifier (PGA), as the core component for power consumption regulation, plays a crucial role in amplifying weak audio signals to specific power levels as needed. This module employs digital control to achieve dynamic gain adjustment, receiving commands from the MCU via the SPI interface to precisely control the output power within a 1-4W range in 0.1W increments. Its operating principle is based on the voltage-controlled gain amplifier (VCA) architecture, using the VCA821 chip as the core device. This chip has a 40dB dynamic gain range (corresponding to a voltage amplification factor of 1-100 times). The gain is digitally programmable by adjusting the feedback resistor network through an external digital potentiometer (such as the X9C104). In the signal processing flow, the digital audio signal converted by the ADC is first analyzed by the MCU to calculate the optimal shielding power required for the current environment. Then, a 16-bit control word is sent to the X9C104 via the SPI bus to configure the VCA821's gain value in real time. This design not only achieves fine-tuning of output power but also features extremely low static power consumption (approximately 25mW) and fast response characteristics (gain switching time <1μs), ensuring the system maintains optimal energy efficiency in various scenarios. Notably, to avoid introducing additional noise, the entire programmable gain circuit employs a differential input structure and incorporates an LC filter network at the power supply end to effectively suppress high-frequency ripple interference. Through this digital, programmable gain control mechanism, the system can dynamically adjust power consumption according to actual needs, minimizing energy consumption while ensuring shielding effectiveness, perfectly meeting the system's requirement of operating at an extremely low power consumption of ≤6W input power. In the subsequent power amplification stage, this invention selects a Class D amplifier. Class D amplifiers utilize pulse width modulation (PWM) technology to convert audio signals into high-frequency square wave signals, achieving power amplification by controlling the switching state of MOSFETs. Compared to traditional Class AB amplifiers (efficiency approximately 60%), Class D amplifiers maintain high efficiency across the entire output power range, significantly reducing power consumption. The system's overall hardware design meets the requirements of power input power ≤6W, output power adjustable from 1-4W, shielding distance ≥10 meters, and angle ≥60°.

[0050] like Figure 1 and 8As shown, this invention uses the on-chip ADC of the STM32 microcontroller to acquire the electrical signal converted by the microphone, and uses FFT to perform spectral analysis on the electrical signal. FFT converts the time-domain signal into the frequency domain, enabling the system to identify the frequency components in the signal. For complex signals, time-domain analysis may struggle to identify their frequency characteristics, while frequency-domain analysis can intuitively display the frequency content of the signal. Directly calculating the Discrete Fourier Transform (DFT) has high computational complexity, while FFT significantly improves computational efficiency through algorithm optimization, making it suitable for real-time data processing in microcontrollers. To achieve the goal of distinguishing between human voices and music, after comprehensively comparing machine learning and deep learning methods, rhythm and prosody analysis methods, dynamic range analysis methods, and spectral analysis methods, spectral analysis was selected. Spectral analysis is one of the commonly used methods to distinguish between human voices and music. By analyzing the spectral characteristics of audio signals, different features of human voices and music can be extracted. The fundamental frequency of human voices is typically between 85Hz and 185Hz (male) and 165Hz and 525Hz (female). Music frequencies cover 20Hz to 20kHz, encompassing a wider frequency range. In terms of spectral distribution, the human voice spectrum is relatively concentrated, with its main energy between 100Hz and 1kHz, and its harmonic structure is relatively regular. Non-human voices, on the other hand, have a wider spectral distribution, containing harmonics and overtones from various instruments, resulting in a more complex spectral structure. Regarding harmonic structure, the harmonic frequencies of human voices are integer multiples of the fundamental frequency, and the harmonic structure is relatively regular. Musical harmonic structures are complex, with different instruments exhibiting different harmonic characteristics. By performing FFT on audio signals and extracting their spectral information, the peak values, harmonic structure, and energy distribution of the spectrum can be effectively analyzed.

[0051] It should be noted that the above content merely illustrates the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. For those skilled in the art, various improvements and modifications can be made without departing from the principle of the present invention, and all such improvements and modifications fall within the scope of protection of the claims of the present invention.

Claims

1. An ultra-low power recording shielding system based on FFT and dynamic feature fusion technology, characterized in that: The system hardware includes: an input module, a dynamic feature recognition module, a shielding path selection module, a data fusion signal processing system module, and a shielding signal output module; The design scheme of the dynamic feature recognition module is as follows: The amplitude of the electrical signal acquired by the input module in the time domain consists of human voice and non-human voice signals, and the dynamic characteristics of the two are distinguished using the frequency domain. The MCU acquires the high signal-to-noise ratio dynamic electrical signal provided by the input module through the ADC acquisition circuit, and uses an STM32F103RCT6 microcontroller to convert it into a digital signal. The acquired digital electrical signal is subjected to fast Fourier transform, spectrum analysis and machine learning algorithms to realize automatic recognition and feature extraction of human voice and non-human voice signals. Based on the identification results and preset rules, it is determined whether to start the transmission of the shielding signal; when there is no audio signal or a non-target signal is detected, the transmission of the shielding signal is stopped, thus realizing intelligent control. The masking path selection module enables the separate masking of human voices and non-human voices by selecting a path: The audio signal output from the input module is connected to the shielded path selection module, which is controlled by two MOSFETs. The audio signal is connected to the drain of the two MOSFETs, and the sources of the two MOSFETs are connected to two different filters. Filter 1 is a bandpass filter with a frequency range of 85-525Hz, which preserves the dynamic characteristics of human voices. Filter 2 is a high-pass filter with a cutoff frequency of 600Hz, which preserves the dynamic characteristics of non-human voices. The MOSFETs are turned on and off by two high and low level inputs to the gate, thereby selecting which filter the audio signal enters, thus distinguishing between human voices and non-human voices, and extracting the dynamic characteristics of either human voice or non-human voice, providing the signal to be processed for the subsequent data fusion signal processing module. The design scheme of the data fusion signal processing system module is as follows: when recording, if it is desired that a certain component of the sound signal will not be recorded or that the power of the human voice signal will be attenuated throughout the sound propagation, then the dynamic characteristics of the energy of that component are considered. With the energy of the total sound The ratio should reach a constant value, let it be denoted as . For this system, an analog subtractor is used to achieve signal complementarity for the extracted analog electrical signal. Through energy superposition and the principle of weakening dynamic characteristics, the recording device captures only background music while attenuating vocals. In application, the dynamic characteristics of the vocal signal and the output signal of the shield are compared. Pulling it down to approximately zero, the total energy of the sound If the values ​​remain essentially unchanged, then the energy ratio is still achieved. If the value is always 0, then the dynamic characteristics of the human voice cannot be extracted during recording. The hardware design of this module involves incorporating a high-speed analog subtractor circuit at the amplifier's input. If you want your voice to not be recorded, let's assume the voice is... ,in Given its dynamic characteristics, the filter extracts the dynamic characteristics of the human voice signal, and the analog subtractor complements these dynamic characteristics to output... By superimposing energy, the amplitude of the signal received by the recording is always 2, thus losing the dynamic characteristics of the human voice and achieving the effect of recording shielding.

2. The ultra-low power recording shielding system based on FFT and dynamic feature fusion technology according to claim 1, characterized in that: The input module acquires sound signals through a microphone, which converts the sound signals into weak electrical signals. The subsequent stage of the module uses an AD620 instrumentation amplifier to perform signal conditioning, ultimately providing a high signal-to-noise dynamic electrical signal for analysis by the subsequent dynamic feature recognition module.

3. The ultra-low power recording shielding system based on FFT and dynamic feature fusion technology according to claim 1, characterized in that: The design scheme for the shielded signal output module is as follows: Based on a Class D power amplifier and an ultrasonic probe, the analog signal, after filtering and high-speed subtraction, is sent to "Amplifier 2" for further amplification. The Class D power amplifier is used to amplify the shielding signal, increasing the voltage to a level that can drive the ultrasonic probe. The amplified shielding signal is then emitted as ultrasonic waves through the ultrasonic probe. By utilizing the high-efficiency Class D power amplifier, the hardware design of this system meets the requirement of power input power ≤6W. At the same time, the array probe layout optimizes the shielding angle and distance, thus interfering with the normal recording of the recording equipment.

4. The ultra-low power recording shielding system based on FFT and dynamic feature fusion technology according to claim 1, characterized in that: The system is used as follows: (1) First, the sound electrical signal is amplified by "Amplifier 1" and the amplified signal is sent to "MCU's ADC". The STM32F103RCT6 microcontroller is selected to convert it into a digital signal. The digital signal is then sent to the MCU analysis and dynamic feature recognition module, where it is analyzed to distinguish between human voice and non-human voice. (2) Based on the recognition results and preset rules, the shielding path selection module is started to determine whether to start the shielding signal transmission; when there is no audio signal or a non-target signal is detected, the shielding signal transmission is stopped to achieve intelligent control; when a target signal is detected, the shielding path selection module controls the start and stop of the MOS transistor through two high and low level input gates, thereby selecting which filter the sound signal enters, and then extracting the dynamic features of either human voice or non-human voice, providing the signal to be processed for the subsequent data fusion signal processing module; (3) The electrical signal extracted by the specific filter is complemented by an analog subtractor. Through the principle of energy superposition and weakening dynamic characteristics, the recording device captures only the background music and weakens the human voice. (4) The filtered analog signal is sent to "Amplifier 2" for amplification. The amplified shielding signal is emitted in the form of ultrasonic waves through the ultrasonic probe. The shielding angle and distance are optimized by the array probe layout to interfere with the normal recording of the recording equipment.

Citation Information

Patent Citations

  • Concealed interference signal generating device and method based on ultrasonic parametric array

    CN111064543A

  • Recording shielding method and device based on ultrasonic interference

    CN120074737A