Ultrasonic event spoofing attack defense method and apparatus

By performing signal time-domain feature processing and spectral difference analysis on the environmental audio of smart home systems, false events in ultrasonic event spoofing attacks are identified and eliminated. This solves the confusion problem of smart home systems when learning user behavior habits, improves the accuracy of analysis and user experience, and ensures life safety.

CN119814345BActive Publication Date: 2025-11-25TIANMUSHAN LABORATORY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411396037.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-08
Publication Date
2025-11-25
Estimated Expiration
2044-10-08

AI Technical Summary

Technical Problem

Smart home systems are vulnerable to ultrasonic event deception attacks when using environmental audio data to learn user behavior habits. This can lead to confusion between false and real sound events, affecting the learning and understanding of users' daily behavior patterns, resulting in incorrect intelligent decisions, reduced user experience, and security risks.

Method used

By preprocessing environmental audio based on signal time-domain characteristics, ultrasonic attacks are detected. False events are identified and eliminated by utilizing signal spectrum differences, generating clusters of real sound events to defend against ultrasonic event deception attacks.

Benefits of technology

This improves the accuracy of smart home systems in analyzing the frequency and time distribution of events, ensuring that the system's learning outcomes reflect user behavior patterns, thereby improving user experience and safeguarding life safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119814345B_ABST
    Figure CN119814345B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of information security, in particular to an ultrasonic event deception attack defense method and device, wherein the method comprises the following steps: based on signal time domain characteristics and spectral roll-off points, pre-processing ambient audio and ultrasonic attack detection are carried out to determine whether an ultrasonic attack is suffered during a recording process; if the ultrasonic attack is suffered, based on a signal spectrum false index, event deception detection is carried out on each sound event cluster to determine whether a false event injected by an attacker is mixed in the cluster; if the false event is mixed in, true and false event separation is carried out on each sound event cluster with the mixed false event to obtain a real sound event cluster after the false event is removed, and ultrasonic event deception attack is defended. The application can detect whether an ultrasonic attack is suffered during a recording process and whether a false event is mixed in a sound event cluster, and can remove the mixed false event to obtain a real sound event cluster, thereby realizing effective defense against ultrasonic event deception attack.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information security technology, and in particular to a method and device for defending against ultrasonic event deception attacks. Background Technology

[0002] In related technologies, with the rapid development of home IoT technology, smart home systems integrate various device sensors and intelligent control systems to achieve automated control and intelligent interaction of the home environment, bringing users a comfortable, convenient, safe, and efficient smart home living experience. In recent years, to fully meet users' flexible and diverse personalized lifestyle needs, smart home systems have deeply integrated with artificial intelligence technology. By analyzing and processing home environment data collected by device sensors, they learn and understand users' daily behavior patterns, thereby achieving autonomous intelligent decision-making and providing users with more accurate and detailed customized intelligent services. For example, the system can use sound sensors such as microphones to record the user's home environment audio throughout the day. By detecting all sound events that occur during this period and clustering them according to the differences in signal characteristics of different events, multiple event clusters are formed. Then, the frequency and time distribution of various events are analyzed to gain a preliminary understanding of the user's living habits and activity patterns, providing a basis for intelligent decision-making.

[0003] However, smart home systems using environmental audio data to learn user behavior can be vulnerable to ultrasonic event spoofing attacks. Attackers can exploit non-linear hardware flaws in microphones to modulate pre-recorded event sounds onto an ultrasonic carrier wave, generating inaudible attack signals, and injecting them into the microphone of the recording device. When the microphone receives these ultrasonic attack signals, it demodulates and reconstructs the original sound event signal, thus spoofing the occurrence of sound events without the user's notice. Furthermore, smart home systems typically focus on the waveform differences in the time domain between different event signals when clustering sound events. Therefore, attackers can carefully design injected fake sound events with signal waveforms similar to real events occurring in the user's home environment (such as knocking, closing doors, etc.), causing the system to mistakenly confuse these fake events with similar real sound events and group them all into the same event cluster. The presence of numerous false events within a cluster can mislead the system's analysis of event frequency and time distribution, leading to serious deviations in the system's learning and understanding of users' daily behavior patterns. This can result in flawed intelligent decisions, significantly reducing user experience and potentially posing security risks, thus impacting users' normal lives. This issue urgently needs to be addressed. Summary of the Invention

[0004] This application provides a method and apparatus for defending against ultrasonic event deception attacks, in order to solve the problem that smart home systems are easily interfered with by ultrasonic event deception attacks when using environmental audio data to learn user behavior habits. This can easily lead to confusion between fake sound events secretly injected by attackers and real events, resulting in deviations in the learning and understanding of users' daily behavior patterns, and thus making incorrect intelligent decisions. This not only significantly reduces the user experience, but may also bring security risks and affect users' normal lives.

[0005] The first aspect of this application provides a method for defending against ultrasonic event deception attacks, comprising the following steps: preprocessing environmental audio based on signal time-domain characteristics to obtain multiple sound event clusters composed of a series of event signal blocks that satisfy preset waveform similarity conditions; performing ultrasonic attack detection on the environmental audio based on the signal spectrum roll-off point to determine whether the environmental audio was subjected to an ultrasonic attack during recording; if subjected to the ultrasonic attack, performing event deception detection on each of the multiple sound event clusters based on the signal spectrum spoofing index to determine whether a spoofing event injected by an attacker is mixed into each sound event cluster; if the spoofing event injected by the attacker is mixed in, performing real and fake event separation on each sound event cluster where a spoofing event is detected based on signal frequency-domain characteristics to obtain a real sound event cluster after removing the spoofing event, thereby defending against ultrasonic event deception attacks.

[0006] Optionally, in one embodiment of this application, the preprocessing of ambient audio based on signal time-domain features to obtain multiple sound event clusters composed of a series of event signal blocks satisfying preset waveform similarity conditions includes: taking the absolute value of the ambient audio data and performing filtering and smoothing processing to obtain a filtered audio signal; performing frame segmentation processing on the filtered audio signal to obtain an audio signal frame sequence of target length, calculating the short-time energy value of each frame signal in the audio signal frame sequence, and marking each frame signal as an event signal or background noise according to the short-time energy value to determine multiple event signal frames; merging the multiple event signal frames in the audio signal frame sequence whose adjacent time interval is less than a preset threshold into continuous event signal blocks to obtain a sound event sequence; extracting time-domain features based on the absolute value of the signal amplitude of each event signal block in the sound event sequence to obtain a first feature vector satisfying a first preset condition, generating a first feature vector set of all events; performing dimensionality reduction processing and clustering on the first feature vector set of all events to obtain the multiple sound event clusters.

[0007] Optionally, in one embodiment of this application, the step of performing ultrasonic attack detection on the environmental audio based on the signal spectrum roll-off point to determine whether the environmental audio has been subjected to ultrasonic attack during the recording process includes: extracting a first signal spectrum of the environmental audio; calculating a spectrum roll-off point in a target frequency range based on the first signal spectrum; and determining that the ultrasonic attack has been subjected to the ultrasonic attack if the spectrum roll-off point is higher than a preset detection threshold.

[0008] Optionally, in one embodiment of this application, if subjected to the ultrasonic attack, the step of performing event deception detection on each of the plurality of sound event clusters based on the signal spectrum spoofing index to determine whether a spoofing event injected by the attacker is mixed into each sound event cluster includes: extracting the second signal spectrum of each event in each sound event cluster; calculating the spoofing index of each event based on the second signal spectrum; and determining that the spoofing event injected by the attacker is mixed into the sound event cluster where any one event is located if the spoofing index of any one event in each event is higher than a preset detection threshold.

[0009] Optionally, in one embodiment of this application, if a false event injected by the attacker is mixed in, then based on the signal frequency domain features, each detected sound event cluster mixed in with false events is separated into true and false events to obtain a true sound event cluster after removing false events, thus defending against ultrasonic event deception attacks. This includes: extracting frequency domain features from the signal spectrum of each event signal block in the sound event cluster mixed with false events to obtain a second feature vector that meets a second preset condition, generating a second feature vector set for all events; performing dimensionality reduction processing and clustering on the second feature vector set for all events to obtain the true sound event cluster after removing false events.

[0010] A second aspect of this application provides an ultrasonic event deception attack defense device, comprising: a processing module for preprocessing environmental audio based on signal time-domain characteristics to obtain multiple sound event clusters composed of a series of event signal blocks satisfying preset waveform similarity conditions; a detection module for performing ultrasonic attack detection on the environmental audio based on the signal spectrum roll-off point to determine whether an ultrasonic attack has occurred during the recording of the environmental audio; a judgment module for performing event deception detection on each of the multiple sound event clusters based on the signal spectrum spoofing index in the event of an ultrasonic attack to determine whether a false event injected by an attacker has been mixed into each sound event cluster; and a defense module for performing real and false event separation on each sound event cluster containing a detected false event based on signal frequency-domain characteristics in the event of the false event injected by the attacker, to obtain a real sound event cluster after removing the false event, thereby defending against ultrasonic event deception attacks.

[0011] Optionally, in one embodiment of this application, the processing module includes: a processing unit, configured to take the absolute value of the ambient audio data and perform filtering and smoothing processing to obtain a filtered audio signal; a marking unit, configured to perform frame segmentation processing on the filtered audio signal to obtain an audio signal frame sequence of a target length, calculate the short-time energy value of each frame signal in the audio signal frame sequence, and mark each frame signal as an event signal or background noise according to the short-time energy value to determine multiple event signal frames; a merging unit, configured to merge the multiple event signal frames in the audio signal frame sequence whose adjacent time interval is less than a preset threshold into a continuous event signal block to obtain a sound event sequence; a first extraction unit, configured to extract time-domain features according to the absolute value of the signal amplitude of each event signal block in the sound event sequence to obtain a first feature vector that satisfies a first preset condition, and generate a first feature vector set for all events; and a first partitioning unit, configured to perform dimensionality reduction processing and clustering on the first feature vector set for all events to obtain the multiple sound event clusters.

[0012] Optionally, in one embodiment of this application, the detection module includes: a second extraction unit for extracting a first signal spectrum of the ambient audio; a first calculation unit for calculating a spectral roll-off point within a target frequency range based on the first signal spectrum; and a first determination unit for determining that the system has been subjected to the ultrasonic attack if the spectral roll-off point is higher than a preset detection threshold.

[0013] Optionally, in one embodiment of this application, the judgment module includes: a third extraction unit, configured to extract the second signal spectrum of each event in each sound event cluster; a second calculation unit, configured to calculate the false index of each event based on the second signal spectrum; and a second determination unit, configured to determine that a false event injected by the attacker has been mixed into the sound event cluster where any one event is located if the false index of any one event in each event is higher than a preset detection threshold.

[0014] Optionally, in one embodiment of this application, the defense module includes: a fourth extraction unit, configured to extract frequency domain features from the signal spectrum of each event signal block in the sound event cluster mixed with false events, to obtain a second feature vector that satisfies a second preset condition, and generate a set of second feature vectors for all events; and a second partitioning unit, configured to perform dimensionality reduction processing and clustering on the set of second feature vectors for all events, to obtain the real sound event cluster after removing false events.

[0015] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the ultrasonic event deception attack defense method as described in the above embodiments.

[0016] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described ultrasonic event deception attack defense method.

[0017] A fifth aspect of this application provides a computer program product, including a computer program that, when executed, is used to implement the above-described ultrasonic event deception attack defense method.

[0018] This application can detect whether environmental audio recordings are subjected to ultrasonic attacks and determine whether false events are mixed into sound event clusters. It can also remove these false events to obtain genuine sound event clusters, thus defending against ultrasonic event deception attacks. Therefore, this allows smart home systems to, based on sound event detection and clustering of environmental audio using signal time-domain characteristics, preliminarily determine whether the recording process of the environmental audio has been subjected to ultrasonic attacks using the signal spectrum differences between normal and affected audio. If an ultrasonic attack is detected, it further identifies and removes false events mixed into the sound event clusters using the signal spectrum differences between genuine and false events, ultimately obtaining genuine sound event clusters free of any false events. This improves the accuracy of the smart home system's analysis of the frequency and time distribution of various events, ensuring that the system's learning results accurately reflect the user's daily behavior patterns. This provides a reliable basis for the system to make more accurate intelligent decisions, significantly improving the user's life experience and ensuring user safety. This solves the problem that smart home systems are susceptible to ultrasonic event deception attacks when using environmental audio data to learn user behavior habits. This can easily lead to confusion between fake sound events secretly injected by attackers and real events, resulting in deviations in the learning and understanding of users' daily behavior patterns and making incorrect intelligent decisions. This not only significantly reduces the user experience but may also bring security risks and affect users' normal lives.

[0019] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0020] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0021] Figure 1 This is a flowchart of a method for defending against ultrasonic event deception attacks according to an embodiment of this application;

[0022] Figure 2 This is a flowchart of an environmental audio preprocessing method according to an embodiment of this application;

[0023] Figure 3 This is a flowchart of a method for defending against ultrasonic event deception attacks based on signal spectrum analysis according to an embodiment of this application;

[0024] Figure 4 This is a schematic diagram of the ultrasonic event deception attack defense device provided according to an embodiment of this application;

[0025] Figure 5 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application.

[0026] Figure label:

[0027] 10-Ultrasonic event deception attack defense device: 100-Processing module, 200-Detection module, 300-Judgment module and 400-Defense module; 501-Memory, 502-Processor and 503-Communication interface. Detailed Implementation

[0028] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0029] The ultrasonic event deception attack defense method and apparatus of this application are described below with reference to the accompanying drawings. In the aforementioned background art, smart home systems are susceptible to ultrasonic event deception attacks when learning user behavior habits using environmental audio data. These attacks can easily confuse false sound events secretly injected by attackers with real events, leading to deviations in the learning and understanding of users' daily behavior patterns and resulting in incorrect intelligent decisions. This not only significantly reduces user experience but may also pose security risks and affect users' normal lives. This application provides an ultrasonic event deception attack defense method. In this method, it can detect whether the environmental audio recording process is subjected to ultrasonic attacks and determine whether false events are mixed into the sound event cluster. It can also remove the mixed false events to obtain the real sound event cluster, thus completing the defense against ultrasonic event deception attacks. This invention enables smart home systems to detect and cluster environmental audio events based on signal time-domain characteristics, obtaining sound event clusters. It then uses the signal spectrum difference between normal and malicious audio to initially determine whether the recording of the environmental audio was subjected to ultrasonic attacks. If an ultrasonic attack is detected, it further identifies and eliminates false events mixed into the sound event cluster by utilizing the signal spectrum difference between real and fake events, ultimately obtaining real sound event clusters free of false events. This improves the accuracy of the smart home system's analysis of the frequency and time distribution of various events, ensuring that the system's learning results accurately reflect the user's daily behavior patterns. This provides a reliable basis for the system to make more accurate intelligent decisions, significantly improving the user's life experience and ensuring user safety. This addresses the problem in related technologies where smart home systems using environmental audio data to learn user habits are susceptible to ultrasonic event deception attacks, easily confusing secretly injected false sound events with real events. This leads to deviations in the learning and understanding of the user's daily behavior patterns, resulting in incorrect intelligent decisions that not only significantly reduce user experience but may also pose security risks and disrupt the user's normal life.

[0030] Specifically, Figure 1 A flowchart illustrating a method for defending against ultrasonic event deception attacks provided in this application embodiment.

[0031] like Figure 1 As shown, the ultrasonic event deception attack defense method includes the following steps:

[0032] In step S101, based on the time-domain characteristics of the signal, the ambient audio is preprocessed to obtain multiple sound event clusters composed of a series of event signal blocks that meet preset waveform similarity conditions.

[0033] In some implementations, when smart home systems detect sound events in clustered environmental audio, they typically focus on the waveform differences in the time domain of different event signals. Therefore, attackers can carefully design and inject fake sound events with signal waveforms similar to real events occurring in the user's home environment (e.g., sounds of knocking or closing doors), causing the system to mistakenly confuse these fake events with similar real sound events and group them all into the same event cluster. A large number of fake events mixed into a cluster misleads the system's analysis of event frequency and temporal distribution, leading to serious deviations in the system's learning and understanding of users' daily behavioral patterns. This can easily result in incorrect intelligent decisions, significantly degrading the user experience and potentially creating security risks, thus affecting the user's normal life.

[0034] Therefore, in this embodiment, the collected environmental audio data can be preprocessed based on the time-domain characteristics of the event signals to obtain multiple sound event clusters composed of a series of event signal blocks that meet certain waveform similarity conditions, and then these sound event clusters can be further processed and judged.

[0035] The process will now be explained in further detail.

[0036] Optionally, in one embodiment of this application, based on the signal time-domain characteristics, the ambient audio is preprocessed to obtain multiple sound event clusters composed of a series of event signal blocks that satisfy preset waveform similarity conditions. This includes: taking the absolute value of the ambient audio data and performing filtering and smoothing processing to obtain a filtered audio signal; performing frame segmentation processing on the filtered audio signal to obtain an audio signal frame sequence of a target length; calculating the short-time energy value of each frame signal in the audio signal frame sequence; marking each frame signal as an event signal or background noise based on the short-time energy value to determine multiple event signal frames; merging multiple event signal frames in the audio signal frame sequence whose adjacent time intervals are less than a preset threshold into continuous event signal blocks to obtain a sound event sequence; extracting time-domain features based on the absolute value of the signal amplitude of each event signal block in the sound event sequence to obtain a first feature vector that satisfies a first preset condition, generating a first feature vector set for all events; and performing dimensionality reduction processing and clustering on the first feature vector set for all events to obtain multiple sound event clusters.

[0037] In actual implementation, the embodiments of this application may, but are not limited to, five steps when preprocessing environmental audio data based on the time-domain characteristics of event signals to obtain multiple sound event clusters composed of a series of event signal blocks that meet certain waveform similarity conditions. Figure 2 This is a flowchart of an environmental audio preprocessing method according to an embodiment of this application, as follows: Figure 2 As shown:

[0038] Step S201: Take the absolute value of the ambient audio data and perform filtering and smoothing to obtain the filtered audio signal. For example, the absolute value of the raw audio data S collected by the device's microphone can be taken, and the signal can be smoothed using, but is not limited to, an Exponentially Weighted Moving Average (EWMA) filter to eliminate noise interference.

[0039] Specifically, in this application embodiment, when performing EWMA filtering on the original audio data S, the method for calculating the EWMA value of each sample data point in the signal can be expressed as follows:

[0040] S EWMA (i)=α*S(i)+(1-α)*S EWMA (i-1)

[0041] Where S(i) represents the amplitude of the i-th sample data point in the signal, and α represents the smoothing factor. α can usually be, but is not limited to, 0.1. In specific applications, it can be set or adjusted by professionals in the field according to actual needs.

[0042] Step S202: Perform frame segmentation on the filtered audio signal to obtain an audio signal frame sequence of the target length. Calculate the short-time energy value of each frame in the audio signal frame sequence. Based on the short-time energy value, mark each frame as an event signal or background noise to determine multiple event signal frames. For example, the filtered audio signal S... EWMA Frame segmentation is performed to obtain a length of m. s audio signal frame sequence And calculate the signal s for each frame. i The short-time energy (STE) is used to determine whether the STE value exceeds a preset detection threshold. e The frame is identified as an event signal, and the STE value is lower than thr e The frame was determined to be background noise.

[0043] Specifically, in the embodiments of this application, the filtered audio signal S EWMA When performing frame segmentation, the frame length can typically be set to 20ms and the frame shift to 10ms, but this is not limited to the case where a skilled technician in the field sets or adjusts these settings according to actual needs. For each frame signal s... i When calculating Short Time Energy (STE), it is equal to the sum of the squares of the signal amplitudes of all sample points in the frame, and the formula can be expressed as follows:

[0044]

[0045] Among them, s i(j) represents the signal amplitude of the j-th sample point in the i-th frame of the signal. This represents the number of sample points in the i-th frame of the signal. When performing event signal detection, it is common practice, but not limited to, setting a detection threshold thr... e The maximum STE value of the first 100 frames of signal can be represented as follows:

[0046] thr e =max(STE(s1),STE(s2),…,STE(s 100 ))

[0047] Furthermore, in specific applications, the settings or adjustments can be made by professionals in the field according to actual needs.

[0048] Step S203: Merge multiple event signal frames in the audio signal frame sequence whose adjacent time intervals are less than a preset threshold into a continuous event signal block to obtain a sound event sequence. For example, for adjacent time intervals less than the preset threshold, ... m The event signal frames can be merged into consecutive event signal blocks to obtain a sound event sequence E = {e1, e2, ..., e...} consisting of n event signal blocks. n}

[0049] Specifically, in the embodiments of this application, when merging short-interval event signal frames, the interval merging threshold thr can typically be used, but is not limited to, setting the interval merging threshold thr. m The time is set to 0.5s, but can be set or adjusted by professionals in the field according to actual needs in specific applications.

[0050] Step S204: Extract time-domain features based on the absolute value of the signal amplitude of each event signal block in the sound event sequence to obtain a first feature vector that satisfies a first preset condition, and generate a set of first feature vectors for all events. For example, for each event signal block e in the sound event sequence E... i Time-domain features can be extracted using the absolute value of its signal amplitude to obtain an n T 3D eigenvectors Among them, the features include, but are not limited to, 10 features that intuitively describe the statistical characteristics of signal amplitude, such as length, area under the curve, maximum value, average value, variance, standard deviation, energy, mean square value, root mean square value, and root square amplitude; and 10 features that describe the waveform shape and energy temporal distribution characteristics of the signal, such as skewness, kurtosis, waveform factor, peak factor, impulse factor, margin factor, time centroid, time-domain stretching, time-domain skewness, and time-domain kurtosis.

[0051] Specifically, in the embodiments of this application, for each event signal block e iWhen performing time-domain feature extraction, length, area under the curve, maximum value, average value, variance, standard deviation, energy, mean square value, root mean square value, and root square amplitude can be calculated using simple statistical methods. The specific definitions and calculation methods for skewness, kurtosis, waveform factor, peak factor, impulse factor, margin factor, time-domain centroid, time-domain stretching, time-domain skewness, and time-domain kurtosis can be expressed as follows:

[0052] Skewness describes the degree of asymmetry in a signal waveform. It is equal to the ratio of the third center distance of the signal amplitude to the cube of the standard deviation. The formula can be expressed as follows:

[0053]

[0054] Where s(i) represents the amplitude of the i-th sample data point in the signal, N represents the number of sample points in the signal, Mean represents the average value, and Std represents the standard deviation.

[0055] Kurtosis describes the steepness of a signal waveform and is equal to the ratio of the fourth center distance of the signal amplitude to the fourth power of the standard deviation. The formula can be expressed as follows:

[0056]

[0057] The waveform factor can be used to describe the waveform shape of a signal. It is equal to the ratio of the root mean square value to the absolute mean value, and the formula can be expressed as follows:

[0058]

[0059] Where RMS represents the root mean square value, which is equal to the arithmetic square root of the mean square of the signal amplitude. The formula can be expressed as follows:

[0060]

[0061] The crest factor describes the extreme degree of a signal's peak value in a waveform. It is often used to detect impulse components in a signal and is equal to the ratio of the signal's peak value to its root mean square (RMS) value. The formula can be expressed as follows:

[0062]

[0063] Where Max represents the maximum value (i.e., the peak value).

[0064] The impulse factor is commonly used to detect the impulse component in a signal. It is equal to the ratio of the signal's peak value to its absolute average value, and can be expressed by the following formula:

[0065]

[0066] The margin factor is equal to the ratio of the signal peak value to the root square amplitude, and the formula can be expressed as follows:

[0067]

[0068] Wherein, SMR represents the square root amplitude, which is equal to the square of the arithmetic square root mean of the signal amplitude. The formula can be expressed as follows:

[0069]

[0070] The temporal centroid can be used to represent the time point where the signal energy is concentrated. It is equal to the weighted average of the sampling times of all sample points in the signal, weighted by the signal amplitude. The formula can be expressed as follows:

[0071]

[0072] Where t(i) represents the sampling time of the i-th sample data point in the signal, and t(1) = 0 is defined.

[0073] Temporal spread can be used to describe the degree of dispersion of signal energy distribution relative to the temporal centroid, and the formula can be expressed as follows:

[0074]

[0075] Temporal skewness can be used to describe the degree of asymmetry in the signal energy distribution relative to the temporal centroid, and the formula can be expressed as follows:

[0076]

[0077] Temporal kurtosis can be used to describe the steepness of the signal energy distribution relative to the temporal centroid, and the formula can be expressed as follows:

[0078]

[0079] Step S205: Perform dimensionality reduction and clustering on the first feature vector set of all events to obtain multiple sound event clusters. For example, but not limited to, Principal Component Analysis (PCA) can be used to analyze the feature vector set of all events. Dimensionality reduction is performed, and all events can be clustered using, but not limited to, the K-means clustering algorithm, to obtain k sound event clusters C = {c1, c2, ..., c...}. k}. Wherein, each event cluster c i Depend on It consists of event signal blocks with similar waveforms, that is

[0080] Specifically, in the embodiments of this application, when performing K-Means clustering, the optimal number of clusters k can generally be determined by using, but is not limited to, the Elbow Method or the Silhouette Coefficient evaluation index. In specific applications, those skilled in the art can select or adjust the appropriate method according to actual needs.

[0081] In step S102, ultrasonic attack detection is performed on the ambient audio based on the signal spectrum roll-off point to determine whether the ambient audio was subjected to ultrasonic attack during the recording process.

[0082] In some embodiments, after preprocessing the ambient audio, this application can detect whether the recording of the ambient audio segment has been subjected to ultrasonic attacks. Specifically, ultrasonic attack detection can be performed on the ambient audio based on the spectral roll-off point of the ambient audio signal to preliminarily determine whether the device microphone has been subjected to ultrasonic attacks during the recording of ambient audio.

[0083] The process will now be explained in more detail.

[0084] Optionally, in one embodiment of this application, ultrasonic attack detection is performed on the ambient audio based on the signal spectrum roll-off point to determine whether an ultrasonic attack has occurred during the recording of the ambient audio. This includes: extracting the first signal spectrum of the ambient audio; calculating the spectrum roll-off point within the target frequency range based on the first signal spectrum; and determining that an ultrasonic attack has occurred if the spectrum roll-off point is higher than a preset detection threshold.

[0085] In actual implementation, when determining whether an ambient audio recording process has been subjected to ultrasonic attack by performing ultrasonic attack detection on ambient audio based on the signal spectrum roll-off point, this application embodiment may, but is not limited to, first extracting the first signal spectrum of the ambient audio, then calculating the spectrum roll-off point in the target frequency range based on the first signal spectrum, and determining that the ambient audio recording process has been subjected to ultrasonic attack if the spectrum roll-off point is higher than a certain detection threshold.

[0086] For example, embodiments of this application can extract the signal spectrum of the entire audio data S and calculate its spectral roll-off point within a specific frequency range. If this value is higher than a preset detection threshold, then... uIf the microphone is subjected to an ultrasonic attack during recording, then all k sound event clusters C are sent to the next step for further processing; otherwise, all k sound event clusters C are directly used as the final output.

[0087] Specifically, in the embodiments of this application, when calculating the spectral roll-off point for the entire audio signal, it can be taken as [F]. l ,F r The signal spectrum within the frequency range is determined, and the frequency boundary point where the signal spectrum energy decreases to τ% of the total energy is found. The formula can be expressed as follows:

[0088] satisfy

[0089] Among them, F l and F r F represents the left and right boundaries of the frequency range of the signal spectrum. l Typically, but not limited to, 10kHz can be used, F r Fixed at 24kHz (i.e., half the microphone recording sampling rate), f(j) represents the j-th frequency component in the signal spectrum, r(i) represents the signal amplitude of the i-th frequency component in the signal spectrum, and N f τ% represents the number of frequency components in the signal spectrum, and can typically be, but is not limited to, 85%–95%. When performing ultrasonic attack detection, the detection threshold thr can typically be, but is not limited to... u Set to 20kHz. In specific applications, F... l τ and thr u All of these can be set or adjusted by professionals in the field according to actual needs.

[0090] It should be noted that the specific formula parameters and detection thresholds can be set or adjusted by those skilled in the art according to actual conditions and needs. The embodiments in this application are only illustrative and do not impose specific limitations.

[0091] In step S103, if an ultrasonic attack is detected, event deception detection is performed on each of the multiple sound event clusters based on the signal spectrum spoofing index to determine whether a fake event injected by the attacker is mixed into each sound event cluster.

[0092] Based on the descriptions of other embodiments, it is understood that if an ultrasonic attack is detected during the recording of ambient audio, this application can further process all sound event clusters C. Specifically, event deception detection can be performed on each sound event cluster based on the fake index of the event signal spectrum to determine whether fake events injected by the attacker have been mixed into each sound event cluster.

[0093] The process will now be explained in more detail.

[0094] Optionally, in one embodiment of this application, if subjected to an ultrasonic attack, event deception detection is performed on each of the multiple sound event clusters based on the signal spectrum spoofing index to determine whether a spoofing event injected by the attacker is mixed into each sound event cluster. This includes: extracting the second signal spectrum of each event in each sound event cluster; calculating the spoofing index of each event based on the second signal spectrum; and determining that a spoofing event injected by the attacker is mixed into the sound event cluster where any event is located if the spoofing index of any event in each event is higher than a preset detection threshold.

[0095] In actual implementation, when this application embodiment performs event deception detection on each sound event cluster based on the signal spectrum false index to determine whether a false event injected by an attacker has been mixed into the cluster, it may, but is not limited to, first extract the second signal spectrum of each event in the sound event cluster, then calculate the false index of each event based on the second signal spectrum, and if the false index of a certain event is higher than a certain detection threshold, that is, if the false index of any one of the events is higher than a certain detection threshold, it can be determined that a false event injected by an attacker has been mixed into the sound event cluster.

[0096] For example, embodiments of this application can extract each sound event cluster c i Each event The signal spectrum is analyzed, and its fake index is calculated. If the fake index of an event in the cluster is higher than a preset detection threshold, then... f If this is the case, then it can be assumed that the cluster of sound events may contain false sound events injected by an attacker using ultrasound. Therefore, all k sound event clusters C = {c1, c2, ..., c...} can be considered as... k The categories are divided into two types: one is the cluster of sound events C mixed with false events. f Those will be sent to subsequent steps for further processing; the other category is the cluster C of real sound events that have not been mixed with false events. t This will be directly used as part of the final output.

[0097] Specifically, in the embodiments of this application, for each event signal block When calculating the fake index, it can be set as the logarithm of the ratio of the area occupied by high-frequency components to that occupied by low-frequency components in the signal spectrum. The formula can be expressed as follows:

[0098]

[0099] Here, f1 can typically be, but is not limited to, 12kHz (i.e., 1 / 4 of the microphone recording sampling rate), and f2 is fixed at 24kHz (i.e., 1 / 2 of the microphone recording sampling rate). When performing event spoofing detection, the detection threshold thr can typically be, but is not limited to, […]. f Set to -0.5. In practical applications, f1 and thr f All of these can be set or adjusted by professionals in the field according to actual needs.

[0100] It should be noted that the specific formula parameters and detection thresholds can be set or adjusted by those skilled in the art according to actual conditions and needs. The embodiments in this application are only illustrative and do not impose specific limitations.

[0101] In step S104, if a false event injected by an attacker is mixed in, then based on the frequency domain characteristics of the signal, the true and false events are separated for each cluster of sound events that is detected to be mixed in with false events, so as to obtain the true sound event cluster after removing the false events and defend against ultrasonic event deception attacks.

[0102] As one possible approach, after detecting the presence of fake events injected by an attacker in a cluster of sound events, this application can separate real and fake events in each cluster of sound events that has been detected by the attacker based on the frequency domain characteristics of the event signal, so as to eliminate all fake events and obtain the real sound event clusters, thereby achieving defense against ultrasonic event deception attacks.

[0103] The process will now be explained in more detail.

[0104] Optionally, in one embodiment of this application, if a false event injected by an attacker is mixed in, then based on the signal frequency domain characteristics, each sound event cluster in which a false event is detected is separated into true and false events to obtain a true sound event cluster after removing the false event, thus defending against ultrasonic event deception attacks. This includes: extracting frequency domain features from the signal spectrum of each event signal block in the sound event cluster in which the false event is mixed in, to obtain a second feature vector that meets a second preset condition, generating a set of second feature vectors for all events; performing dimensionality reduction processing and clustering on the set of second feature vectors for all events to obtain a true sound event cluster after removing the false event.

[0105] In some embodiments, in the process of eliminating false events to obtain real sound event clusters, this application needs to utilize the signal spectrum difference between real sound events and false events. Therefore, this application may, but is not limited to, first extracting the frequency domain features of each event signal block in the sound event cluster mixed with false events, obtaining a second feature vector that meets certain conditions, and then performing dimensionality reduction and clustering on the set of second feature vectors of all events, thereby obtaining the real sound event clusters.

[0106] For example, embodiments of this application address the sound event cluster c that is mixed with false events. f,i Each event signal block Frequency domain features can be extracted using its signal spectrum to obtain an n F 3D eigenvectors The features include, but are not limited to, five characteristics describing the shape and energy distribution of the signal's spectral profile: spectral centroid, spectral stretch, spectral skewness, spectral kurtosis, and spectral roll-off point. Then, but not limited to, the PCA (Principal Component Analysis) method can be used to analyze the feature vector set of all events in the cluster. Dimensionality reduction is performed, and the K-Means clustering algorithm can be used, but is not limited to, to separate all real and fake events in the clusters, resulting in real sound event clusters after removing fake events. Finally, all clusters of real sound events, excluding fake events, will be removed. As another part of the final output.

[0107] Specifically, in the embodiments of this application, for each event signal block When performing frequency domain feature extraction, the specific definitions and calculation methods for spectral centroid, spectral stretch, spectral skewness, spectral kurtosis, and spectral roll-off point can be expressed as follows:

[0108] The spectral centroid can be used to represent the frequency points where energy is concentrated in the signal spectrum. It is equal to the weighted average of all frequency components in the signal spectrum, weighted by the signal amplitude. The formula can be expressed as follows:

[0109]

[0110] Where f(i) represents the i-th frequency component in the signal spectrum, r(i) represents the signal amplitude of that frequency component, and N f It indicates the number of frequency components in the signal spectrum.

[0111] Spectral spread can be used to describe the degree of dispersion of the signal's spectral energy distribution relative to the spectral centroid. The formula can be expressed as follows:

[0112]

[0113] Spectral skewness describes the degree of asymmetry between the shape of a signal's spectrum and its centroid. The formula can be expressed as follows:

[0114]

[0115] Spectral kurtosis describes the steepness of a signal's spectral shape relative to its centroid. The formula can be expressed as follows:

[0116]

[0117] The spectral roll-off point can be used to represent the frequency boundary where the signal's spectral energy drops to a specific percentage of the total energy. The formula can be expressed as follows:

[0118] satisfy

[0119] Among them, F l and F r F represents the left and right boundaries of the frequency range of the signal spectrum. l It is generally possible, but not limited to, to use 0Hz, F. r Fixed at 24kHz (i.e., half the microphone recording sampling rate), τ% can typically, but is not limited to, 85%–95%. In specific applications, F... l Both τ and τ can be set or adjusted by professionals in the field according to actual needs.

[0120] It should be noted that the specific formula parameter values ​​can be set or adjusted by those skilled in the art according to actual conditions and needs. The embodiments in this application are only illustrative and do not impose specific limitations.

[0121] The present application will be described in detail below with reference to a specific embodiment.

[0122] Figure 3 This is a flowchart of a method for defending against ultrasonic event deception attacks based on signal spectrum analysis, according to an embodiment of this application. Figure 3 As shown:

[0123] Step S301: Acquire ambient audio collected by the device's microphone.

[0124] Step S302: Perform preprocessing on the collected environmental audio based on the signal time domain characteristics.

[0125] Step S303: Preprocessing is completed, resulting in multiple sound event clusters composed of a series of event signal blocks with similar waveforms.

[0126] Step S304: Perform ultrasonic attack detection on the entire environmental audio segment based on the signal spectrum roll-off point.

[0127] Step S305: Determine whether the recording of the ambient audio was subjected to an ultrasonic attack. If no ultrasonic attack was detected, output all sound event clusters directly as the final result.

[0128] Step S306: If an ultrasonic attack is detected, perform event deception detection based on the signal spectrum spoofing index for each sound event cluster.

[0129] Step S307: Determine whether a fake event injected by the attacker has been mixed into each sound event cluster. If no fake event is detected in the cluster, the sound event cluster is directly output as the final result.

[0130] Step S308: If a false event is detected in the cluster, the sound event cluster is processed further.

[0131] Step S309: Separate real and fake events for each cluster of sound events that has been detected to have false events mixed in, based on the frequency domain features of the signal.

[0132] Step S310: Obtain the cluster of real sound events after removing false events, and output it as the final result to complete the defense against ultrasonic event deception attacks.

[0133] The ultrasonic event deception attack defense method proposed in this application can detect whether the recording of environmental audio is subjected to ultrasonic attacks and determine whether false events are mixed into the sound event cluster. It can also remove the mixed false events to obtain the real sound event cluster, thus completing the defense against ultrasonic event deception attacks. Therefore, the smart home system, based on sound event detection and clustering of environmental audio according to signal time-domain characteristics, can initially determine whether the recording of environmental audio is subjected to ultrasonic attacks by using the signal spectrum difference between normal audio and the victim audio. If an ultrasonic attack is detected, it further uses the signal spectrum difference between real and false events to identify and remove false events mixed into the sound event cluster, ultimately obtaining a real sound event cluster without any false events. This improves the accuracy of the smart home system's analysis of the frequency and time distribution of various events, ensuring that the system's learning results accurately reflect the user's daily behavior patterns. This provides a reliable basis for the system to make more accurate intelligent decisions, significantly improving the user's life experience and ensuring the user's safety. This solves the problem that smart home systems are susceptible to ultrasonic event deception attacks when using environmental audio data to learn user behavior habits. This can easily lead to confusion between fake sound events secretly injected by attackers and real events, resulting in deviations in the learning and understanding of users' daily behavior patterns and making incorrect intelligent decisions. This not only significantly reduces the user experience but may also bring security risks and affect users' normal lives.

[0134] Secondly, the ultrasonic event deception attack defense device proposed according to the embodiments of this application is described with reference to the accompanying drawings.

[0135] Figure 4 This is a schematic diagram of the ultrasonic event deception attack defense device according to an embodiment of this application.

[0136] like Figure 4 As shown, the ultrasonic event deception attack defense device 10 includes: a processing module 100, a detection module 200, a judgment module 300, and a defense module 400.

[0137] The processing module 100 is used to preprocess the ambient audio based on the signal time-domain characteristics to obtain multiple sound event clusters composed of a series of event signal blocks that meet preset waveform similarity conditions.

[0138] The detection module 200 is used to perform ultrasonic attack detection on environmental audio based on the signal spectrum roll-off point, so as to determine whether the environmental audio was subjected to ultrasonic attack during the recording process.

[0139] The judgment module 300 is used to perform event deception detection on each of multiple sound event clusters based on the signal spectrum spoofing index in the event of an ultrasonic attack, so as to determine whether a fake event injected by the attacker is mixed into each sound event cluster.

[0140] The defense module 400 is used to separate real and fake events for each cluster of sound events that is detected to be mixed in with fake events, based on the frequency domain characteristics of the signal, in the case of fake events injected by attackers, so as to obtain the real sound event cluster after removing the fake events, thus defending against ultrasonic event deception attacks.

[0141] Optionally, in one embodiment of this application, the processing module 100 includes: a processing unit, a marking unit, a merging unit, a first extraction unit, and a first division unit.

[0142] The processing unit is used to take the absolute value of the ambient audio data and perform filtering and smoothing to obtain the filtered audio signal.

[0143] The marking unit is used to perform frame segmentation processing on the filtered audio signal to obtain an audio signal frame sequence of target length, calculate the short-time energy value of each frame signal in the audio signal frame sequence, and mark each frame signal as an event signal or background noise based on the short-time energy value, so as to determine multiple event signal frames.

[0144] The merging unit is used to merge multiple event signal frames in an audio signal frame sequence whose adjacent time intervals are less than a preset threshold into a continuous event signal block to obtain a sound event sequence.

[0145] The first extraction unit is used to extract time-domain features based on the absolute value of the signal amplitude of each event signal block in the sound event sequence, so as to obtain a first feature vector that satisfies the first preset condition and generate a set of first feature vectors for all events.

[0146] The first partitioning unit is used to reduce the dimensionality of the first feature vector set of all events and perform clustering to obtain multiple sound event clusters.

[0147] Optionally, in one embodiment of this application, the detection module 200 includes: a second extraction unit, a first calculation unit, and a first determination unit.

[0148] The second extraction unit is used to extract the first signal spectrum of the ambient audio.

[0149] The first calculation unit is used to calculate the spectral roll-off point within the target frequency range based on the spectrum of the first signal.

[0150] The first determination unit is used to determine whether an ultrasonic attack has occurred when the spectral roll-off point is higher than a preset detection threshold.

[0151] Optionally, in one embodiment of this application, the judgment module 300 includes: a third extraction unit, a second calculation unit, and a second determination unit.

[0152] The third extraction unit is used to extract the second signal spectrum of each event in each sound event cluster.

[0153] The second calculation unit is used to calculate the spurious index for each event based on the spectrum of the second signal.

[0154] The second determination unit is used to determine whether a fake event injected by an attacker has been mixed into the audio event cluster where any event is located, if the fake index of any event in each event is higher than the preset detection threshold.

[0155] Optionally, in one embodiment of this application, the defense module 400 includes: a fourth extraction unit and a second division unit.

[0156] The fourth extraction unit is used to extract frequency domain features from the signal spectrum of each event signal block in the sound event cluster mixed with false events, so as to obtain a second feature vector that meets the second preset condition and generate a set of second feature vectors for all events.

[0157] The second partitioning unit is used to reduce the dimensionality of the second feature vector set of all events and perform clustering to obtain the real sound event clusters after removing false events.

[0158] It should be noted that the foregoing explanation of the embodiment of the ultrasonic event deception attack defense method also applies to the ultrasonic event deception attack defense device of this embodiment, and will not be repeated here.

[0159] The ultrasonic event deception attack defense device proposed in this application can detect whether the recording of environmental audio is subjected to ultrasonic attacks and determine whether false events are mixed into the sound event cluster. It can also remove the mixed false events to obtain the real sound event cluster, thus completing the defense against ultrasonic event deception attacks. Therefore, the smart home system, based on sound event detection and clustering of environmental audio according to signal time-domain characteristics, can initially determine whether the recording of environmental audio is subjected to ultrasonic attacks by using the signal spectrum difference between normal audio and the victim audio. If an ultrasonic attack is detected, it further uses the signal spectrum difference between real and false events to identify and remove false events mixed into the sound event cluster, ultimately obtaining a real sound event cluster without any false events. This improves the accuracy of the smart home system's analysis of the frequency and time distribution of various events, ensuring that the system's learning results accurately reflect the user's daily behavior patterns. This provides a reliable basis for the system to make more accurate intelligent decisions, significantly improving the user's life experience and ensuring the user's safety. This solves the problem that smart home systems are susceptible to ultrasonic event deception attacks when using environmental audio data to learn user behavior habits. This can easily lead to confusion between fake sound events secretly injected by attackers and real events, resulting in deviations in the learning and understanding of users' daily behavior patterns and making incorrect intelligent decisions. This not only significantly reduces the user experience but may also bring security risks and affect users' normal lives.

[0160] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include:

[0161] The memory 501, the processor 502, and the computer program stored on the memory 501 and capable of running on the processor 502.

[0162] When the processor 502 executes the program, it implements the ultrasonic event deception attack defense method provided in the above embodiments.

[0163] Furthermore, electronic devices also include:

[0164] Communication interface 503 is used for communication between memory 501 and processor 502.

[0165] The memory 501 is used to store computer programs that can run on the processor 502.

[0166] The memory 501 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0167] If the memory 501, processor 502, and communication interface 503 are implemented independently, then the communication interface 503, memory 501, and processor 502 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0168] Optionally, in a specific implementation, if the memory 501, processor 502, and communication interface 503 are integrated on a single chip, then the memory 501, processor 502, and communication interface 503 can communicate with each other through an internal interface.

[0169] Processor 502 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0170] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described ultrasonic event deception attack defense method.

[0171] This application also provides a computer program product, including a computer program that can run computer instructions. When the computer instructions are executed by a processor, they implement the ultrasonic event deception attack defense method provided in this application.

[0172] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0173] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0174] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0175] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0176] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or more of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0177] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0178] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0179] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for defending against ultrasonic event deception attacks, characterized in that, Includes the following steps: Based on the time-domain characteristics of the signal, the ambient audio is preprocessed to obtain multiple sound event clusters composed of a series of event signal blocks that meet preset waveform similarity conditions; Based on the signal spectrum roll-off point, ultrasonic attack detection is performed on the environmental audio to determine whether the environmental audio was subjected to ultrasonic attack during the recording process. If subjected to the ultrasonic attack, event deception detection is performed on each of the multiple sound event clusters based on the signal spectrum spoofing index to determine whether a fake event injected by the attacker is mixed into each sound event cluster. If false events injected by the attacker are mixed in, then based on the frequency domain characteristics of the signal, the true and false events are separated for each cluster of sound events that are detected to be mixed in with false events, so as to obtain the true sound event clusters after removing false events, thus defending against ultrasonic event deception attacks.

2. The method according to claim 1, characterized in that, The preprocessing of ambient audio based on signal time-domain characteristics yields multiple sound event clusters composed of a series of event signal blocks that satisfy preset waveform similarity conditions, including: The absolute value of the ambient audio data is taken and then filtered and smoothed to obtain the filtered audio signal. The filtered audio signal is segmented into frames to obtain an audio signal frame sequence of target length. The short-time energy value of each frame in the audio signal frame sequence is calculated. Each frame is marked as an event signal or background noise based on the short-time energy value to determine multiple event signal frames. The multiple event signal frames in the audio signal frame sequence whose adjacent time intervals are less than a preset threshold are merged into a continuous event signal block to obtain a sound event sequence; Temporal features are extracted based on the absolute value of the signal amplitude of each event signal block in the sound event sequence to obtain a first feature vector that satisfies the first preset condition, and a set of first feature vectors for all events is generated. The first feature vector set of all events is subjected to dimensionality reduction and clustering to obtain the multiple sound event clusters.

3. The method according to claim 1, characterized in that, The step of performing ultrasonic attack detection on the environmental audio based on the signal spectrum roll-off point to determine whether the environmental audio was subjected to ultrasonic attack during recording includes: Extract the first signal spectrum of the ambient audio; Calculate the spectral roll-off point within the target frequency range based on the spectrum of the first signal; If the spectral roll-off point is higher than a preset detection threshold, it is determined that the system has been subjected to the ultrasonic attack.

4. The method according to claim 1, characterized in that, If the system is subjected to the ultrasonic attack, then based on the signal spectrum spoofing index, event deception detection is performed on each of the plurality of sound event clusters to determine whether spoofed events injected by the attacker are mixed into each sound event cluster, including: Extract the second signal spectrum of each event in each sound event cluster; The spurious index for each event is calculated based on the second signal spectrum; If the false index of any one of the events exceeds a preset detection threshold, it is determined that the false event injected by the attacker has been mixed into the audio event cluster to which the event belongs.

5. The method according to claim 1, characterized in that, If a false event injected by the attacker is mixed in, then based on the signal frequency domain characteristics, each detected false event mixed in is separated into true and false events to obtain the true sound event cluster after removing the false events, thus defending against ultrasonic event deception attacks, including: Frequency domain features are extracted from the signal spectrum of each event signal block in the sound event cluster mixed with the false event to obtain a second feature vector that meets the second preset condition, and a set of second feature vectors for all events is generated. The second feature vector set of all events is subjected to dimensionality reduction and clustering to obtain the real sound event cluster after removing false events.

6. A device for defending against ultrasonic event deception attacks, characterized in that, include: The processing module is used to preprocess the ambient audio based on the signal time-domain characteristics to obtain multiple sound event clusters composed of a series of event signal blocks that meet preset waveform similarity conditions; The detection module is used to perform ultrasonic attack detection on the environmental audio based on the signal spectrum roll-off point, so as to determine whether the environmental audio was subjected to ultrasonic attack during the recording process. The judgment module is used to perform event deception detection on each of the multiple sound event clusters based on the signal spectrum spoofing index in the event of the ultrasonic attack, so as to determine whether the spoofing events injected by the attacker are mixed into each sound event cluster. The defense module is used to separate real and fake events for each cluster of sound events detected by the attacker when false events are mixed in, based on the frequency domain characteristics of the signal, so as to obtain the real sound event cluster after removing the false events, thus defending against ultrasonic event deception attacks.

7. The apparatus according to claim 6, characterized in that, The processing module includes: The processing unit is used to take the absolute value of the ambient audio data and perform filtering and smoothing processing to obtain the filtered audio signal; The marking unit is used to perform frame segmentation processing on the filtered audio signal to obtain an audio signal frame sequence of target length, calculate the short-time energy value of each frame signal in the audio signal frame sequence, and mark each frame signal as an event signal or background noise according to the short-time energy value to determine multiple event signal frames; The merging unit is used to merge multiple event signal frames in the audio signal frame sequence whose adjacent time intervals are less than a preset threshold into a continuous event signal block to obtain a sound event sequence. The extraction unit is used to extract time-domain features based on the absolute value of the signal amplitude of each event signal block in the sound event sequence, so as to obtain a first feature vector that satisfies a first preset condition and generate a set of first feature vectors for all events. The partitioning unit is used to perform dimensionality reduction processing and clustering on the first feature vector set of all events to obtain the multiple sound event clusters.

8. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the ultrasonic event deception attack defense method as described in any one of claims 1-5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the ultrasonic event deception attack defense method as described in any one of claims 1-5.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed, it is used to implement the ultrasonic event deception attack defense method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Replay attack living body detection method based on ultrasonic time domain processing

    CN111739560A

  • Ultrasonic attack prevention for speech enabled devices

    US20190043471A1