Audio information sensing methods, electronic devices and media based on baseband electromagnetic side channels

CN117831547BActive Publication Date: 2026-08-14HUNAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

然而,这些方法应用于现实中的感知场景时,常面临着应用局限性大、实现效果不佳、可操作难度大等问题

Benefits of technology

[0033](1)本发明利用移动设备音频电路本身辐射的基频电磁信号进行音频信息感知,与已有侧信道感知方法相比,在不需要复杂、昂贵的感知装置情况下,可以稳定地实现高精确度的语音感知,面对大规模应用需求时可大幅降低设备部署的成本。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117831547B_ABST
    Figure CN117831547B_ABST
Patent Text Reader

Abstract

This invention provides an audio sensing method, electronic device, and medium based on a fundamental frequency electromagnetic side channel. This invention utilizes the fundamental frequency electromagnetic signal radiated by the audio circuitry of a mobile device for audio information sensing. Compared to existing side channel sensing methods, this invention can stably achieve high-accuracy voice sensing without requiring complex and expensive sensing devices. This invention uses a clustering-based strategy to extract voice features from the electromagnetic signal for audio signal recovery, avoiding the data pre-training stage relied upon in existing side channel sensing technologies. This enables audio signal sensing in scenarios utilizing side channel information, and the method of this invention has high adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of side channel utilization and information perception, and in particular relates to an audio information perception method, electronic device and medium based on a baseband electromagnetic side channel. Background Technology

[0002] Currently, methods for sensing audio information from mobile devices based on side channels mainly fall into two categories: one utilizes internal motion sensors within the mobile device to detect vibration signals at the speaker, and the other actively detects vibration signals at the location of the sound source using external signal sources. However, when these methods are applied to real-world sensing scenarios, they often face limitations, poor implementation results, and operational difficulties. For example, the first type of method, based on motion sensors (accelerometers, gyroscopes), is often limited by low sampling rates, resulting in incomplete acquisition of speech frequency information, which fundamentally affects the accuracy of speech signal perception. The second type, using high-frequency signals (WiFi, LiDAR) for external detection, is easily affected by inherent environmental interference factors (obstacles, human movement) and cannot be applied to signal perception in non-ideal environments.

[0003] Therefore, the above-mentioned sensing methods are difficult to apply to complex and ever-changing real-world scenarios. Furthermore, due to limitations in their technical principles, these vibration signal-based side-channel sensing methods rely on data-driven training models to obtain semantic recognition results, resulting in high computational resource consumption and deployment costs. Therefore, there is a current need to explore and research new sensing methods to achieve low-cost, easy-to-operate, environmentally adaptable, and effective audio information sensing technology.

[0004] In recent years, research and application of electromagnetic signals from electronic devices have developed rapidly. As a byproduct of electronic device operation, electromagnetic signals contain useful information highly correlated with the device's operating state, and are therefore often utilized as an effective information sensing side channel. However, research on audio information sensing based on electromagnetic side channels is scarce. Current research reveals that audio information can be extracted from devices by receiving and decoupling high-frequency electromagnetic signals (above MHz). However, the coupling phenomenon involved in this approach only exists in certain structurally unique devices, making it unsuitable for widely used mobile devices in practical scenarios. Furthermore, the signal decoupling process involves a series of complex signal transformations, significantly increasing the computational overhead and difficulty of semantic information analysis. Therefore, a method for audio information sensing without relying on high-frequency coupling by receiving fundamental frequency electromagnetic signals (0-4kHz) directly associated with audio information has become a primary consideration. However, currently, there are no commercially available devices capable of receiving fundamental frequency electromagnetic signals, and how to semantically convert potentially audio-rich fundamental frequency electromagnetic signals is a critical problem that urgently needs to be solved. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by providing an audio sensing method, electronic device, and medium based on the fundamental frequency electromagnetic side channel. The invention designs an audio extraction method that does not rely on data training based on the characteristics of the fundamental frequency electromagnetic signal, which can realize audio information sensing in scenarios utilizing side channel information. The method of this invention has high adaptability.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] An audio information sensing method based on a fundamental frequency electromagnetic side channel includes the following steps:

[0008] S1: Acquires the fundamental frequency electromagnetic signal of the mobile device's audio system;

[0009] S2: Preprocess the fundamental frequency electromagnetic signal to remove environmental noise signals of fixed frequency from the fundamental frequency electromagnetic signal;

[0010] S3: Use wavelet transform to remove the non-steady environmental noise signal from the fundamental frequency electromagnetic signal after S2 processing;

[0011] S4: Use clustering algorithms to extract the audio signal from the fundamental frequency electromagnetic signal after processing in S3, and remove the signal distortion caused by device characteristics in the audio signal;

[0012] S5: Convert the audio signal obtained after processing in step S4 into an audio file.

[0013] This invention utilizes the fundamental frequency electromagnetic signal radiated by the audio circuitry of mobile devices for audio information perception. The fundamental frequency electromagnetic signal shares the same origin as the original audio signal and naturally contains complete and rich speech information. Compared to existing side-channel sensing methods, this invention can stably achieve high-accuracy speech perception without requiring complex and expensive sensing devices. By using a clustering-based strategy to extract speech features from the electromagnetic signal for audio signal recovery, it avoids the data pre-training step relied upon in existing side-channel sensing technologies. This enables audio signal perception in scenarios utilizing side-channel information, demonstrating high adaptability.

[0014] Furthermore, the specific implementation process of step S3 includes:

[0015] A1: Decompose the fundamental frequency electromagnetic signal obtained after processing S2 into low-frequency components and high-frequency components;

[0016] A2: Perform wavelet transform on the low-frequency components of the signal to obtain approximate coefficients, and perform wavelet transform on the high-frequency components of the signal to obtain detail coefficients;

[0017] A3: Filter the detail coefficients to remove non-steady environmental noise signals;

[0018] A4: Use inverse wavelet transform to restore the approximation coefficients and the denoised detail coefficients to time-domain values ​​to obtain the fundamental frequency electromagnetic signal after removing the non-steady environmental noise signal.

[0019] After S3 processing, a large amount of non-steady-state environmental noise contained in the fundamental frequency electromagnetic signal has been eliminated.

[0020] Furthermore, the specific implementation process of step S4 includes:

[0021] (1) Starting from point p0(t0,f0,A0) in the spectrum diagram SD of the fundamental frequency electromagnetic signal obtained after S3 processing, filter points p0(t0,f0,A0) within a predefined distance D. p All adjacent points within a given range, when the number of adjacent points is greater than or equal to a predefined number N. p At that time, a new cluster C0 is created by p0(t0,f0,A0) and the adjacent points of p0(t0,f0,A0), where A0 represents the amplitude of the fundamental frequency electromagnetic signal obtained after S3 processing at time t0 and frequency f0.

[0022] (2) Taking point p′(t′,f′,A′)∈C0 as the new starting point, filter points p′(t′,f′,A′) within a predefined distance D. p All adjacent points within a given range, when the number of adjacent points is greater than or equal to a predefined number N. pAt that time, the neighboring points of p′(t′,f′,A′) are merged into cluster C0 to expand C0, where A′ represents the amplitude of the fundamental frequency electromagnetic signal obtained after S3 processing at time t′ and frequency f′.

[0023] (3) For other points within the expanded cluster, repeat step (2) to continue expanding the cluster until the predefined distance D is reached. p All points within the cluster are merged into the expanded cluster, and the current clustering is completed, yielding the final cluster.

[0024] (4) Repeat steps (1) to (3) for the remaining points in SD to form a new cluster C. n This continues until all points in the SD have been processed, generating a signal cluster C = {C0,…C...} n ,…C N};

[0025] (5) Calculate C for each cluster n point p i (t i ,f i A i The average amplitude p i (t i ,f i A i )∈C n A i Representing time t i Frequency f i The amplitude of the fundamental frequency electromagnetic signal obtained after processing by S3;

[0026] (6) Extract the average amplitude Avg n Greater than or equal to γA max Clustering, forming new signal clusters {C0,…C d ,…C D}, this new signal cluster {C0,…C d ,…C D Let} be an audio signal, where γ is a coefficient, 1 ≥ γ ≥ 0, and A max This represents the maximum amplitude of each point in cluster C.

[0027] After S4 processing, the speech frequency components in the fundamental frequency electromagnetic signal have been completely extracted and can be directly converted into audio files for machine semantic recognition. This invention uses a clustering-based strategy to extract speech features from the fundamental frequency electromagnetic signal for audio signal recovery, avoiding the data pre-training step relied upon in existing side-channel sensing technologies, and realizing audio signal sensing in side-channel information utilization scenarios.

[0028] Based on the same concept, the present invention provides an electronic device, comprising:

[0029] One or more processors;

[0030] A memory having stored one or more programs that, when executed by one or more processors, cause the one or more processors to implement the steps of an audio information sensing method based on a baseband electromagnetic side channel.

[0031] Based on the same concept, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of an audio information sensing method based on a baseband electromagnetic side channel.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0033] (1) This invention uses the fundamental frequency electromagnetic signal radiated by the audio circuit of the mobile device itself to perceive audio information. Compared with the existing side channel perception method, it can stably achieve high-precision voice perception without the need for complex and expensive perception devices, and can significantly reduce the cost of device deployment when facing large-scale application needs.

[0034] (2) By utilizing the characteristics of strong penetration of electromagnetic signals and unrestricted spatial radiation angle, the present invention makes the sensing technology easy to apply to multiple real-world sensing scenarios, with high adaptability and low cost.

[0035] (3) The present invention uses a clustering-based strategy to extract speech features from electromagnetic signals for audio signal recovery, which avoids the data pre-training step that existing side channel sensing technologies rely on, and realizes audio signal sensing in the scenario of utilizing side channel information. Attached Figure Description

[0036] Figure 1 A schematic diagram illustrating the electromagnetic signal generation principle during the operation of an audio system.

[0037] Figure 2 This is a flowchart of audio perception processing based on fundamental frequency electromagnetic signals;

[0038] Figure 3 This is a schematic diagram of the three-level decomposition of wavelet transform;

[0039] Figure 4 The results show the correlation analysis between the fundamental frequency electromagnetic signal and the audio signal. Detailed Implementation

[0040] The present invention will be described in detail below with reference to embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other. For ease of description, the words "upper," "lower," "left," and "right" appearing below only indicate that they are consistent with the upper, lower, left, and right directions of the drawings themselves, and do not limit the structure.

[0041] Example

[0042] This embodiment provides a device for acquiring fundamental frequency electromagnetic signals. The acquisition device mainly consists of three modules: a signal receiving module, a signal amplification module, and an analog-to-digital conversion module.

[0043] Signal Receiving Module: Addressing the issue of receiving fundamental frequency electromagnetic signals, there are no readily available devices on the market. This invention presents an electromagnetic signal receiving device based on induced electromotive force. Inspired by the phenomenon that PCB circuit boards can be attacked by electromagnetic signal injection during operation, the design utilizes the ESP32 microcontroller to create a lightweight and convenient electromagnetic signal sensing prototype. The ESP32 microcontroller is commonly used in IoT development, capable of achieving analog signal acquisition and analog-to-digital conversion with a certain degree of accuracy. Its small size and low cost facilitate the construction of large-scale sensing systems.

[0044] Signal Amplification Module: This type of fundamental frequency electromagnetic signal is susceptible to environmental noise interference, and its signal strength is usually weak. A dedicated signal amplification module for this fundamental frequency electromagnetic signal is added after the signal receiving module. This signal amplification module consists of a three-stage cascaded circuit of filtering-amplification-filtering. In this embodiment, a fourth-order low-pass switched-capacitor filter (model TLC14) and an AD620 amplifier module are used. The signal passes through this cascaded circuit to achieve preliminary amplification and filtering, thereby enabling the subsequent analog-to-digital converter module to achieve higher conversion accuracy.

[0045] Analog-to-digital conversion module: Analog-to-digital conversion is a crucial step in the signal sensing process. However, the built-in analog-to-digital conversion module in the ESP32 is a SAR model, and its highest conversion accuracy is 12 bits. Under this configuration, there is still room for improvement in the signal sensing accuracy. Therefore, this embodiment also provides an upgrade option, an external high-precision ADC, which upgrades the ADC to a Σ-Δ type 32-bit ADC commonly used in medium-to-high precision sensing requirements, and mounts it on a Raspberry Pi (Raspberry Pi 4 Model B) for data sensing and acquisition.

[0046] According to such Figure 1 , Figure 2 As shown, this embodiment provides an audio information sensing method based on a baseband electromagnetic side channel, including:

[0047] When a user plays audio normally using the audio system of a mobile device, it generates corresponding electromagnetic signals that radiate outward.

[0048] Connect the antenna (ordinary wire, such as DuPont wire) to a microcontroller board (ESP32) with an analog interface. The microcontroller communicates with the host computer (computer or server) via serial port.

[0049] The serial port data reading interface (based on Python code) is called to read the changes in electromagnetic field amplitude sensed by the microcontroller port in real time and generate corresponding data files for storage.

[0050] The subsequent signal denoising preprocessing, environmental noise removal, audio extraction, and signal distortion removal were performed using Matlab software.

[0051] A. Signal preprocessing: Use a band-stop filter to filter 50Hz and its harmonics, use a high-pass filter to filter noise signals above 5000Hz and below 20Hz, and enhance the frequency components related to the audio signal by normalizing the signal.

[0052] B. Signal Denoising: In practical signal sensing scenarios, due to the presence of random noise in the environment, the audio signal and environmental noise in the received signal spectrum are often intertwined, making it impossible to directly distinguish between the two using effective means, either in the time domain or the frequency domain. However, in the wavelet domain, the wavelet coefficients corresponding to the effective signal and noise can exhibit different intensity distributions in different frequency bands. Based on this phenomenon, this embodiment of the invention introduces the wavelet transform (DWT) method to convert the signal to the wavelet domain in an attempt to separate the effective signal from the noise. The specific implementation steps of the wavelet transform-based denoising method are as follows.

[0053] Step 1. The signal will first be decomposed into three levels according to the principle of high and low frequencies (using high-pass and low-pass filters). The decomposition process is as follows: Figure 3 As shown. After decomposition, the signal is decomposed into low-frequency and high-frequency components.

[0054] Step 2. Perform wavelet transforms on the low-frequency and high-frequency components respectively. After the transform, the low-frequency components are the corresponding approximation coefficients, and the high-frequency components are the corresponding detail coefficients. The expression for the transform is shown in equation (1).

[0055]

[0056] in, (L is the wavelet decomposition series, L = 3) are the approximation coefficients. For detail coefficients E(n) is the measured value representing the high-frequency / low-frequency component of the decomposed discrete electromagnetic radiation signal, N is the length of E(n), φ and ψ are mutually orthogonal wavelet basis functions, and K... l This indicates the length of the coefficients in the l-th level decomposition.

[0057] Step 3. After the wavelet transform in the previous step, the signal needs to be filtered in the wavelet domain. Due to the approximation coefficients... This mainly corresponds to the low-frequency components of the signal, which primarily characterize the speech phonemes and therefore do not require further noise filtering. The noise filtering stage only focuses on the detail coefficients. Based on the principle that in the wavelet domain, the signal strength of noise is always less than that of the effective signal, for the detail coefficients... The dynamic threshold thr for each layer is calculated using the Birgé-Massart strategy. l And use it to update the detail factor. To remove environmental noise, the coefficients are calculated as shown in equation (2):

[0058]

[0059] iff means if and only if;

[0060] Step 4. Use inverse wavelet transform to restore the signal from the wavelet domain to the time-domain measurement form E′(n). The restoration process is based on all resulting coefficients (i.e., approximation coefficients). and the denoising detail coefficient The calculation method is shown in equation (3):

[0061]

[0062] C. Audio Extraction: After eliminating environmental noise, harmonic distortion components caused by device characteristics still exist in the signal E′(n), which manifest in the spectrum as peripheral frequency components with slightly lower signal strength surrounding the speech component. Therefore, this embodiment of the invention can solve this type of harmonic distortion problem by using a clustering-based strategy (DBSCAN).

[0063] The specific implementation of DBSCAN is described below. Assuming SD represents the STFT spectrum of the electromagnetic radiation measurement signal E′(n), each bin(t,f,A) is considered a point, denoted as p(t,f,A)∈SD, where A represents the amplitude of E′(n) at time t and frequency f. DBSCAN starts from an unlabeled point p0(t0,f0,A0) and filters through predefined distances D. pAll neighboring points within the range. If the number of neighbors is not less than the predefined parameter N. p Then, a new cluster C0 is created by grouping p0 and its neighbors. DBSCAN then uses point p′(t′,f′,A′)∈C0 as the new starting point and merges D... p Expand C0 using neighbors within a given distance. When a predefined distance D is defined... p All points within the cluster are merged into C0, thus completing the current clustering C0 and yielding the final C0. DBSCAN will then perform the previous steps on the remaining points in SD, forming new clusters, until all points have been processed. Euclidean distance is used to measure the distance between any two points (p...). i ,p j ),

[0064] DBSCAN generates a signal cluster C = {C0,…C1}. n ,…C N Then calculate C for each cluster. n point p i The average amplitude is shown in equation (4):

[0065]

[0066] And “A” max "Set to the maximum amplitude at each point in cluster C. The audio signal component is identified as the average amplitude Avg." n Not less than γA max The clusters. γ is a coefficient falling within [0,1]. Specifically, the audio signal clusters are shown in equation (5).

[0067] C audio ={0, ..., Cd, ..., C D},iff Avg n ≥γA max (5)

[0068] iff indicates that γ = 0.1 is set if and only if γ > 0.1. After multiple tests and evaluations, the optimal extraction effect can be achieved. After speech cluster extraction, the noise clusters identified in the spectrum are distorted signals and white noise. This is achieved by setting their corresponding amplitude A. i =0, It can achieve the desired cleaning effect.

[0069] D. Audio Restoration: First, perform an inverse STFT transformation on the signal from the previous step to convert the signal from the frequency domain to the time domain. Then, use the Matlab signal processing toolbox to convert the restored electromagnetic measurement values ​​into an audio file. The toolbox can directly convert electromagnetic radiation measurement values ​​as a time-varying signal into a .wav file. This ".wav" file can then be directly converted into semantic information by speech-to-text software or manual recognition.

[0070] The quality of the reproduced audio is assessed by audio evaluation indicators. The effect of electromagnetic sensing under the operation of different device audio systems is shown in the figure below. The WER (Automatic Speech-to-Text Error Rate) and STOI (Speech Awareness, a value >0.7 indicates that the test audio has good intelligibility) indicators show that the sensing scheme of the present invention can be applied to a variety of mobile devices on the market.

[0071]

[0072] Correlation analysis verified a high correlation between the electromagnetic signal and the original audio signal. Furthermore, comprehensive multi-device cross-validation experiments confirmed the basic feasibility of audio sensing based on the electromagnetic side channel. The correlation analysis results between the electromagnetic signal and the audio signal are as follows: Figure 4 As shown, Figure 4 The consistent occurrence of peak values ​​indicates that the fundamental frequency electromagnetic signal and the audio signal are highly correlated in multi-device testing experiments.

[0073] This embodiment provides an electronic device, including:

[0074] One or more processors;

[0075] A memory that stores one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the steps of an audio information sensing method based on a baseband electromagnetic side channel.

[0076] In some implementations, the memory may be high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk storage device.

[0077] In other implementations, the processor can be any type of general-purpose processor, such as a central processing unit (CPU) or a digital signal processor (DSP), and there is no limitation here.

[0078] This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of an audio information sensing method based on a baseband electromagnetic side channel.

[0079] The above embodiments should be understood as being used only to illustrate the present invention more clearly, and not to limit the scope of the present invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art fall within the scope defined by the appended claims.

Claims

1. An audio information sensing method based on a fundamental frequency electromagnetic side channel, characterized in that, Includes the following steps: S1: Acquires the fundamental frequency electromagnetic signal of the mobile device's audio system; S2: Preprocess the fundamental frequency electromagnetic signal to remove environmental noise signals of fixed frequency from the fundamental frequency electromagnetic signal; S3: Use wavelet transform to remove the non-steady environmental noise signal from the fundamental frequency electromagnetic signal after S2 processing; S4: Use a clustering algorithm to extract the audio signal from the fundamental frequency electromagnetic signal after processing in S3, and remove the signal distortion caused by device characteristics in the audio signal; S5: Convert the audio signal obtained after processing in step S4 into an audio file.

2. The audio information sensing method based on the fundamental frequency electromagnetic side channel according to claim 1, characterized in that, The specific implementation process of step S3 includes: A1: Decompose the fundamental frequency electromagnetic signal obtained after processing S2 into low-frequency components and high-frequency components; A2: Perform wavelet transform on the low-frequency components of the signal to obtain approximate coefficients, and perform wavelet transform on the high-frequency components of the signal to obtain detail coefficients; A3: Filter the detail coefficients to remove non-steady environmental noise signals; A4: Use inverse wavelet transform to restore the approximation coefficients and the denoised detail coefficients to time-domain values ​​to obtain the fundamental frequency electromagnetic signal after removing the non-steady environmental noise signal.

3. The audio information sensing method based on the fundamental frequency electromagnetic side channel according to claim 1, characterized in that, The specific implementation process of step S4 includes: (1) Starting from point p0(t0,f0,f0) in the spectrum diagram SD of the fundamental frequency electromagnetic signal obtained after S3 processing, filter points p0(t0,f0,A0) within a predefined distance D. p All adjacent points within a given range, when the number of adjacent points is greater than or equal to a predefined number N. p At that time, a new cluster C0 is created by p0(t0,f0,A0) and the adjacent points of p0(t0,f0,A0), where A0 represents the amplitude of the fundamental frequency electromagnetic signal obtained after S3 processing at time t0 and frequency f0. (2) Taking point p′(t′, f′, A′)∈C0 as the new starting point, select points p′(t′, f′, A′) within a predefined distance D. p All adjacent points within a given range, when the number of adjacent points is greater than or equal to a predefined number N. p At that time, the neighboring points of p′(t′, f′, A′) are merged into cluster C0 to expand C0, where A′ represents the amplitude of the fundamental frequency electromagnetic signal obtained after S3 processing at time t′ and frequency f′. (3) For other points within the expanded cluster, repeat step (2) to continue expanding the cluster until the predefined distance D is reached. p All points within the cluster are merged into the expanded cluster, and the current clustering is completed, yielding the final cluster. (4) Repeat steps (1) to (3) for the remaining points in SD to form a new cluster C. n This continues until all points in the SD have been processed, generating a signal cluster C = {C0,…C...} n ,…C N }; (5) Calculate C for each cluster n point p i (t i ,f i A i The average amplitude p i (t i ,f i A i )∈C n A i Representing time t i Frequency f i The amplitude of the fundamental frequency electromagnetic signal obtained after processing by S3; (6) Extract the average amplitude Avg n Greater than or equal to γA max Clustering, forming new signal clusters {C0,…C d ,…C D }, this new signal cluster {C0,…C d ,…C D Let} be an audio signal, where γ is a coefficient, 1 ≥ γ ≥ 0, and A max This represents the maximum amplitude of each point in cluster C.

4. An electronic device, characterized in that, include: One or more processors; A memory having stored one or more programs that, when executed by one or more processors, cause the one or more processors to perform the steps of the method according to any one of claims 1-3.

5. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Electromagnetic side channel attack defense method based on ultrasonic waves, electronic equipment and medium

    CN118264384A