An audio interference method and system based on dynamic analysis

By using dynamic analysis and multi-objective optimization methods, an interference signal matching the audio signal is generated, which solves the problem of unstable interference effect in existing technologies and achieves precise interference and semantic destruction in complex audio environments.

CN120544608BActive Publication Date: 2025-11-11GUANGZHOU BOCHENG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510950713.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-11-11
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

Existing audio jamming techniques lack flexibility and adaptability, making it difficult to cope with complex and ever-changing audio signal environments, resulting in insufficiently accurate and unstable jamming effects.

Method used

A dynamic analysis-based approach is adopted to extract audio temporal features through convolutional neural networks and long short-term memory networks, combine them with conditional generative adversarial networks to generate interference signals, and optimize the interference effect and semantic destruction degree through multi-objective optimization algorithms and real-time feedback mechanisms.

Benefits of technology

It achieves precise interference in complex and dynamic audio environments, ensuring that the interference signal matches the audio content, maximizing the interference effect and enhancing the degree of semantic disruption, and adapting to different types of audio signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544608B_ABST
    Figure CN120544608B_ABST
Patent Text Reader

Abstract

This invention discloses an audio interference method and system based on dynamic analysis, belonging to the field of audio processing technology. The method includes converting the original audio signal into a spectrogram, extracting audio temporal features from the spectrogram, acquiring random noise, combining the audio temporal features and random noise to generate a first interference signal, judging the first interference signal, and proceeding to the next step if it meets the criteria; optimizing the first interference signal to obtain a second interference signal, synthesizing the second interference signal with the original audio signal to obtain a third interference signal, evaluating the interference effect and semantic destruction degree of the third interference signal, obtaining an evaluation result, and adjusting the third interference signal based on the evaluation result until the final interference signal is obtained. This invention considers the temporal and variability of the audio signal and generates a matching interference signal according to the content characteristics of the audio, ensuring the accuracy of the interference signal and maximizing the interference effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of audio processing technology, and particularly relates to an audio interference method and system based on dynamic analysis. Background Technology

[0002] With the rapid development of audio technology, audio signal processing, analysis, and interference have been widely applied in various fields. Audio signal interference technology is commonly used in speech recognition, audio privacy protection, intelligent voice assistants, and anti-interference communication. In these fields, the goal of audio interference methods is often to interfere with the normal transmission, analysis, or recognition of audio by modifying the audio signal or embedding specific noise.

[0003] However, existing audio jamming techniques mostly focus on generating noise with fixed frequency, amplitude, and duration, lacking flexibility and adaptability. This results in inaccurate jamming effects and difficulty in handling complex and ever-changing audio signal environments. Traditional jamming methods include white noise-based overlay and random noise superposition. While these methods can jam audio signals to some extent, they often fail to achieve efficient and accurate jamming due to the lack of dynamic adaptation between the jamming signal and the target audio content, and the resulting semantic disruption is insufficient. Furthermore, existing audio jamming methods typically lack the ability to analyze the temporal changes of audio signals, failing to adjust the jamming signal in real time according to changes in the audio signal. This leads to unstable jamming effects in long-duration or dynamic audio, making them ineffective in dealing with complex audio signals.

[0004] To address this issue, we propose a dynamic analysis-based audio interference method and system. Summary of the Invention

[0005] The purpose of this invention is to solve the problem that the interference effect is unstable in the prior art and cannot effectively deal with complex audio signals, and to propose an audio interference method and system based on dynamic analysis.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] An audio interference method based on dynamic analysis includes:

[0008] The original audio signal is input, converted into a frequency domain signal, and a spectrogram is output. The spectrogram is a two-dimensional matrix, with different dimensions representing time frames and frequency components, respectively. The spectrogram is then input into a convolutional neural network and a long short-term memory network to extract audio temporal features. The convolutional neural network is used to extract local features, and the long short-term memory network is used to extract temporal dependencies.

[0009] Random noise is acquired, and audio temporal features are combined with random noise through a convolutional neural network to generate a first interference signal. The first interference signal is generated by combining adversarial loss and temporal similarity regularization. The first interference signal is compared with the original audio signal to output a judgment result. The judgment result extracts features of the original audio signal and the first interference signal through multi-layer convolutional operations, and outputs a judgment on whether it is an interference signal through a fully connected layer.

[0010] The optimization objectives of the first interference signal are achieved through a multi-objective optimization algorithm. The optimization objectives include maximizing the difference between the first interference signal and the original audio signal, maximizing the deceptiveness and irreversibility of the interference effect, and minimizing the generation time of the first interference signal. After achieving the optimization objectives, the second interference signal is obtained.

[0011] The second interference signal and the original audio signal are combined by weighted summation to obtain the third interference signal. The third interference signal is optimized by guiding audio rhythm disorder and signal energy reconstruction to enhance semantic masking and content destruction capabilities. The third interference signal introduces audio rhythm structure similarity loss to constrain rhythm consistency with the second interference signal.

[0012] The interference effect and semantic destruction degree of the third interference signal are evaluated at each preset interval to obtain the evaluation result. The interference effect is measured by the L2 norm to measure the degree of interference of the third interference signal on the original audio signal. When the L2 norm is the largest, the interference effect is the strongest. The semantic destruction degree is measured by the signal-to-noise ratio regularization term. When the signal-to-noise ratio regularization term is the lowest, it means that the original speech signal is covered by the unstructured interference signal to a greater extent, the semantic recognition difficulty is increased, and thus the interference effect is enhanced. Based on the evaluation result, the third interference signal is adjusted according to the preset strategy until the optimal interference effect and semantic destruction degree are obtained, and the final interference signal is output.

[0013] Preferably, the first interference signal is generated by converting the audio temporal features into a format suitable for grid processing through an embedding layer and then performing multi-layer convolution operations.

[0014] Preferably, the adversarial loss is used to measure the similarity between the generated interference signal and the real interference signal, so that the generator can generate interference signals that can deceive the discriminator.

[0015] Preferably, the temporal similarity regularization term measures the temporal difference between the first interference signal and the original audio signal using the L2 norm, thereby avoiding excessive distortion of the temporal structure of the original audio signal by the generated interference signal.

[0016] Preferably, the multi-objective optimization algorithm is performed by particle swarm optimization, where each particle represents a second interference signal configuration, the position of which contains the characteristics of the second interference signal in the time domain, the dimension of the particle corresponds to the time length of the second interference signal, and the optimal solution, i.e. the second interference signal, is obtained by updating the position and velocity of the particles.

[0017] Preferably, a weighting coefficient is introduced in the weighted synthesis process of the second interference signal and the original audio signal. When the value of the weighting coefficient is the largest, the proportion of the original audio signal retained is the largest. When the value of the weighting coefficient is the smallest, the influence of the second interference signal is the largest.

[0018] Preferably, the audio rhythm structure similarity loss is obtained by calculating the similarity of the rhythm patterns of the third interference signal and the original audio signal in the time domain.

[0019] Preferably, the signal-to-noise ratio regularization term is used to minimize the noise generated during the synthesis process, which is accomplished by minimizing the signal-to-noise ratio.

[0020] The preferred preset strategy is:

[0021] If the interference effect and / or semantic disruption level are below the threshold, increase the strength of the third interference signal; if the semantic disruption level meets the requirements but the interference effect is below the threshold, increase the proportion of the third interference signal.

[0022] An audio interference system based on dynamic analysis includes:

[0023] The feature extraction module is configured to take the original audio signal as input, convert the original audio signal into a frequency domain signal and output a spectrogram, wherein the spectrogram is a two-dimensional matrix, and its different dimensions represent time frames and frequency components, respectively; the spectrogram is input into a convolutional neural network and a long short-term memory network to extract audio temporal features, wherein the convolutional neural network is used to extract local features and the long short-term memory network is used to extract temporal dependencies;

[0024] An interference generation module is configured to acquire random noise and combine audio temporal features with random noise through a convolutional neural network to generate a first interference signal. The first interference signal is generated by combining adversarial loss and temporal similarity regularization. The first interference signal is compared with the original audio signal to output a judgment result. The judgment result extracts features of the original audio signal and the first interference signal through multi-layer convolutional operations, and outputs a judgment on whether it is an interference signal through a fully connected layer.

[0025] The target optimization module is configured to optimize the first interference signal using a multi-objective optimization algorithm. The optimization objectives include maximizing the difference between the first interference signal and the original audio signal, maximizing the deceptiveness and irreversibility of the interference effect, and minimizing the generation time of the first interference signal. After achieving the optimization objectives, a second interference signal is obtained.

[0026] The signal synthesis module is configured to synthesize a third interference signal by weighted summation of the second interference signal and the original audio signal. The third interference signal is optimized by guiding audio rhythm disorder and signal energy reconstruction to enhance semantic masking and content destruction capabilities. The third interference signal incorporates audio rhythm structure similarity loss to constrain rhythm consistency with the second interference signal.

[0027] The feedback optimization module is configured to evaluate the interference effect and semantic destruction degree of the third interference signal at each preset interval, and obtain the evaluation result. The interference effect is measured by the L2 norm to measure the degree of interference of the third interference signal on the original audio signal. When the L2 norm is the largest, the interference effect is the strongest. The semantic destruction degree is measured by the signal-to-noise ratio regularization term. When the signal-to-noise ratio regularization term is the lowest, it means that the original speech signal is covered by the unstructured interference signal to the greatest extent, the semantic recognition difficulty is the highest, and thus the interference effect is the strongest. Based on the evaluation result, the third interference signal is adjusted according to a preset strategy until the optimal interference effect and semantic destruction degree are obtained, and the final interference signal is output.

[0028] In summary, the technical effects and advantages of this invention are as follows: This audio interference method and system based on dynamic analysis considers the temporal and variability of audio signals and can generate matching interference signals according to the content characteristics of the audio, ensuring the accuracy of the interference signal and maximizing the interference effect. Simultaneously, this invention also considers the degree of semantic destruction, ensuring that the interference signal achieves maximum semantic destruction while achieving effective interference. Through this method, this invention can adapt to complex dynamic audio environments and generate effective and accurate interference in different types of audio signals. Attached Figure Description

[0029] Figure 1 This is a flowchart of the steps in this invention;

[0030] Figure 2 This is a schematic diagram of the system structure in this invention. Detailed Implementation

[0031] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0032] like Figure 1 As shown, an audio interference method based on dynamic analysis includes:

[0033] The original audio signal is input, converted into a frequency domain signal, and a spectrogram is output. The spectrogram is a two-dimensional matrix, with different dimensions representing time frames and frequency components, respectively. The spectrogram is then input into a convolutional neural network and a long short-term memory network to extract audio temporal features. The convolutional neural network is used to extract local features, and the long short-term memory network is used to extract temporal dependencies.

[0034] Random noise is acquired, and audio temporal features are combined with random noise through a convolutional neural network to generate a first interference signal. The first interference signal is generated by combining adversarial loss and temporal similarity regularization term during the generation process. The first interference signal is compared with the original audio signal to output a judgment result. The judgment result extracts features of the original audio signal and the first interference signal through multi-layer convolutional operations, and outputs a judgment on whether it is an interference signal through a fully connected layer.

[0035] The optimization objectives of the first interference signal are achieved through a multi-objective optimization algorithm. The optimization objectives include maximizing the difference between the first interference signal and the original audio signal, maximizing the deceptiveness and irreversibility of the interference effect, and minimizing the generation time of the first interference signal. After achieving the optimization objectives, the second interference signal is obtained.

[0036] The third interference signal is synthesized by weighted summation of the second interference signal and the original audio signal. The third interference signal is optimized by guiding audio rhythm disorder and signal energy reconstruction to enhance semantic masking and content destruction capabilities.

[0037] The interference effect and semantic destruction degree of the third interference signal are evaluated at each preset interval to obtain the evaluation results. The interference effect is measured by the L2 norm to measure the degree of interference of the third interference signal on the original audio signal. When the L2 norm is the largest, the interference effect is the strongest. The audio quality of the semantic destruction degree is measured by the signal-to-noise ratio. When the signal-to-noise ratio is the lowest, it means that the original speech signal is covered by the unstructured interference signal to a greater extent, the semantic recognition difficulty is increased, and thus the interference effect is enhanced. Based on the evaluation results, the third interference signal is adjusted according to a preset strategy until the final interference signal with the best interference effect and semantic destruction degree is obtained and output.

[0038] While closely related, the "degree of semantic disruption" and "interference effect" proposed in this invention are fundamentally different. The degree of semantic disruption primarily measures the extent to which the interfering signal disrupts the semantic content, language structure, and information coherence of the original speech, emphasizing the listener's (or recognition system's) "failure to understand" the speech content. We indirectly measure this feature through metrics such as signal-to-noise ratio (SNR) reduction. The "interference effect," on the other hand, more broadly describes the ability of the interfering signal to block the overall recognition and reconstruction process of the original audio signal. It includes both semantic disruption and acoustic structural perturbations, such as temporal rhythm interruption, spectral masking, and continuity disruption. Therefore, in this invention, we designed multiple regularization terms and evaluation mechanisms to jointly improve interference performance from both the dimensions of semantic understanding failure and acoustic interference intensity. This distinction ensures that the interference system possesses both structural attack capabilities and semantic masking capabilities, thereby achieving the technical goals of audio interference more comprehensively.

[0039] The specific steps of this plan are as follows:

[0040] Step 1: Extraction of temporal features of audio signals:

[0041] Audio signal preprocessing:

[0042] Input: Raw audio signal .

[0043] raw audio signal It is recorded by a device (such as a microphone) or obtained from an existing audio database, and the unit is usually seconds. The sampling rate (such as 44.1kHz) determines the number of samples per second.

[0044] Operation: The short-time Fourier transform (STFT) is used to convert the time-domain signal into a frequency-domain signal. STFT extracts frequency-domain features by dividing the audio signal into short time windows, with each window generating a spectrogram to capture local frequency variations.

[0045] Example: By dividing the audio signal into 20-millisecond windows with a 10-millisecond overlap, the signal in each time window is converted into a frequency domain representation.

[0046]

[0047] Where t is the time point and f is the frequency;

[0048] Output: Spectrum It is a two-dimensional matrix, where Represents a time frame. Represents frequency components.

[0049] Feature extraction model:

[0050] Input: Spectrum .

[0051] Each spectrogram represents the frequency characteristics of the audio signal within a time window. This data will be further processed using a neural network model.

[0052] Operation: The temporal features of the audio signal are extracted by combining a 1D convolutional neural network (CNN) and a long short-term memory network (LSTM).

[0053] CNN: Uses 1D convolutional layers to extract local temporal features of audio signals and capture important features in the spectrogram, such as pitch changes and frequency abrupt changes.

[0054] LSTM: Based on the features extracted by CNN, LSTM is used to model the long-term temporal dependencies of audio, capturing long-term dependency features such as rhythm changes and intonation fluctuations in audio signals.

[0055] Neural network structure:

[0056] CNN layer: Input spectrogram Then, it goes through multiple convolutional layers. The size of each convolutional kernel is set to 3, and the stride is 1. These convolutional layers automatically extract local temporal features from the spectrogram.

[0057] LSTM layer: The features extracted by CNN are fed into the LSTM layer for temporal modeling. The LSTM layer can capture long-term temporal dependencies in audio signals, such as rhythm fluctuations and emotional changes in audio.

[0058] Output: After processing by the LSTM layer, the temporal characteristics of the audio signal are obtained. It is a vector that contains the changes of the audio signal at different points in time.

[0059] The formula is expressed as:

[0060]

[0061] Variable description:

[0062] The spectrum obtained through STFT conversion represents the distribution of audio signals at different time frames and frequency components.

[0063] The temporal features extracted after processing by CNN and LSTM contain the dynamic features of the audio signal at each time point and are used as input for the generation of subsequent interference signals.

[0064] Data execution instructions:

[0065] Input audio data: Time-domain signal acquired from the audio device. Alternatively, an existing audio file can be loaded from a database. The audio signal will first be converted into a frequency domain signal via STFT. .

[0066] Model Execution: Spectrum The audio signal is fed into a 1D-CNN model, where the CNN layers extract local features. These features are then passed to the LSTM layers to capture the temporal dependencies of the audio, ultimately outputting temporal features. .

[0067] In this way, we are able to extract rich temporal features from the original audio signal. These features accurately reflect the dynamic changes of the audio signal, providing a basis for the generation of subsequent interference signals. The innovation of this process lies in the use of a combination of CNN and LSTM to simultaneously handle short-term and long-term temporal dependencies, ensuring that the temporal features in the audio signal can be fully captured.

[0068] Step 2: Generation of interference signals based on time-series characteristics:

[0069] In step 2, our goal is to leverage the audio temporal features extracted in step 1. Generate audio signals Adaptive matching interference signal To achieve effective interference with audio signals while maximizing semantic disruption, we employ a Conditional Generative Adversarial Network (cGAN), which uses a generator and a discriminator to generate and optimize the interference signal.

[0070] First, the generator's task is to generate interference signals. And ensure that the generated signal has the same timing characteristics as the audio signal. Adaptation. The generator's input includes not only random noise. It also includes the temporal characteristics of audio signals. We use a convolutional neural network (CNN) to extract temporal features. and noise Combined, an interference signal matching the audio timing characteristics is generated. Specifically, after receiving the input, the generator first embeds the temporal feature vector through an embedding layer. Converted to a format suitable for network processing, and then subjected to several layers of convolutional operations (e.g., using...). (using convolution kernels and convolution with a stride of 1) to gradually generate interference signals that are compatible with the audio signal.

[0071] The discriminator's task is to determine whether the generated interference signal is a real interference signal. Its input includes the original audio signal. and generated interference signals The discriminator extracts features from the audio and interference signals through multiple convolutional operations and outputs the judgment result through a fully connected layer.

[0072] In the generator optimization objective, we combine adversarial loss and temporal similarity regularization. The generator's optimization goal is to maximize the realism of the generated interference signal while minimizing the temporal difference between the generated interference signal and the original audio signal. We use adversarial loss to incentivize the generator to produce interference signals that can pass the discriminator's judgment, achieving an effect as close to the real signal as possible. The temporal similarity regularization term is added to the generator's loss function to maintain the temporal consistency between the generated interference signal and the audio signal, preventing the generation of interference signals from disrupting the basic structure of the audio.

[0073] Loss function of generator It consists of two parts:

[0074] Adversarial loss: used to measure the similarity between the generated interference signal and the real interference signal, prompting the generator to generate interference signals that can deceive the discriminator.

[0075] Temporal similarity regularization term: The L2 norm is used to measure the temporal difference between the generated interference signal and the original audio signal, ensuring that the generated interference signal does not excessively distort the temporal structure of the audio. The generator's final loss function is:

[0076]

[0077] in, It is the generator's adversarial loss. It is a time-series similarity regularization term. Controlling the strength of regularization, It is the generated interference signal. This is the original audio signal. The introduction of a temporal similarity regularization term helps maintain the structural consistency of the audio signal during the generation of interference signals. Specifically, this regularization term measures the temporal difference between the generated interference signal and the original audio signal using the L2 norm, thereby ensuring that the generated interference signal does not disrupt the rhythm and structure of the audio signal.

[0078] The goal of the discriminator is to maximize its ability to distinguish between generated and real signals. The discriminator's loss function... The calculation measures the error of the discriminator in determining the authenticity of interference signals.

[0079]

[0080] Through this adversarial training, the generator and discriminator are continuously optimized. The generator gradually produces more realistic and effective interference signals, while the discriminator optimizes its ability to distinguish between real and generated signals. The interference signal is generated by combining temporal features with a Conditional Generative Adversarial Network (cGAN), and by introducing a temporal similarity regularization term, the interference signal is ensured to be temporally consistent with the audio signal, maximizing both the interference effect and the semantic disruption effect. This innovative approach provides a precise and dynamic interference signal generation solution for this patent, ensuring adaptive matching and effective interference between the audio and interference signals.

[0081] Step 3: Multi-target optimization of interference signals:

[0082] In step 3, our goal is to modify the already generated interference signal. Further optimization is performed to ensure that the signal maximizes both the interference and semantic disruption effects, while also considering the real-time requirements of the generation process. The core task of optimization is to adjust the generated interference signal to achieve an optimal balance using a multi-objective optimization method. This step focuses on modeling and optimizing the interference signal so that it conforms to the interference target, maximizes the degree of semantic disruption, and possesses real-time generation capabilities. In step 2, the interference signal was generated using a conditional generative adversarial network (cGAN). The signal has already correlated with the timing characteristics of the audio signal. Adaptation. At this point, we have a preliminary interference signal, but the interference effect may need further optimization. In this step, we will use a multi-objective optimization algorithm to optimize these interference signals, ensuring that the generated signal can not only effectively interfere with the audio signal, but also maintain the clarity and intelligibility of the audio, and meet real-time requirements.

[0083] We divide the objectives of interference signal optimization into the following aspects:

[0084] Maximizing the interference effect: This is one of the core objectives of this invention. We hope that the generated interference signal can significantly alter the characteristics of the audio signal, ensuring that the audio signal is effectively interfered with. This objective can be measured by calculating the difference between the interference signal and the original audio signal. We use the L2 norm to measure the interference signal. and the original audio signal The difference between them, and to maximize this difference, to enhance the interference effect.

[0085] Maximizing Semantic Disruption: This invention aims to construct interference signals with unstructured, non-periodic, and semantically fragmenting characteristics to mask the original audio content audibly, disrupt its structure, and make it semantically difficult to recognize. We use the signal-to-noise ratio (SNR) as one of the indirect metrics for interference effectiveness. A decrease in SNR typically indicates an increased proportion of non-semantic components in the overall signal, significantly increasing the difficulty of speech recognition and understanding. Therefore, this invention strategically reduces the SNR, guiding the signal towards a direction of high interference and high confusion to effectively mask and disrupt the target speech.

[0086] Optimizing generation time delay: Real-time performance is crucial in practical applications, especially during interference signal generation. We aim to optimize the speed of interference signal generation to ensure it occurs within a specified timeframe, meeting the demands of real-time applications. By minimizing the time required to generate the interference signal, we can improve the system's response speed.

[0087] To optimize multiple objectives simultaneously, we combine the objective functions using a weighted summation approach, forming a comprehensive fitness function. This function can simultaneously consider the optimization objectives of interference effects, semantic corruption degree loss, and generation latency:

[0088]

[0089] It is the objective function of the interference effect, which measures the difference between the interference signal and the original audio signal (usually using the L2 norm).

[0090] It is the objective function for the degree of semantic corruption loss, and the quality of the audio signal is measured by SNR.

[0091] It is the objective function for generating delay, which measures the time it takes to generate interference signals.

[0092] These are the weights of each objective function, used to balance the importance of different objectives.

[0093] To address the aforementioned multi-objective optimization problem, we chose the Particle Swarm Optimization (PSO) algorithm. PSO finds the optimal solution by simulating the foraging behavior of a flock of birds and updating the positions of particles. Here, each particle represents a perturbation signal configuration, and its position contains the characteristics of the perturbation signal in the time domain.

[0094] Particle representation: The position of each particle corresponds to an interference signal. It exhibits different characteristics in the time domain (such as amplitude and frequency), and the particle dimension is related to the duration of the interference signal. Correspondingly.

[0095] Fitness evaluation: Each particle is evaluated based on the comprehensive objective function. Through evaluation, particles with higher fitness are more likely to meet our optimization goals. By calculating the fitness value of each particle, we can select the optimal interference signal configuration.

[0096] Particle Update: The PSO algorithm searches for the optimal solution by updating the position and velocity of particles. The particle's position represents the generated configuration of the disturbance signal, while the velocity determines how the particle updates within the search space. The particle's velocity and position are updated using the following formula:

[0097] Speed ​​updates:

[0098]

[0099] Location update:

[0100]

[0101] in, It is inertial weight. It is the acceleration constant. It is a random number. It is the local optimal position of the particle. It is the globally optimal position.

[0102] Fitness evaluation and particle update: The fitness of a particle is calculated by weighting the sum of each objective function. This is obtained by iterating through the updating of particle positions and velocities to find the optimal interference signal.

[0103] After multiple iterations and optimizations, the particle swarm will eventually find the optimal interference signal. This interference signal maximizes both the interference effect and the degree of semantic disruption, while ensuring the efficiency of the generation process. The output signal will serve as the final interference signal for use in subsequent steps.

[0104] Step 4: Synthesis of interference signal and audio:

[0105] In step 3, we performed multi-objective optimization on the generated interference signal using the Particle Swarm Optimization (PSO) algorithm to obtain the final optimized interference signal. At this point, the interference signal already has a good interference effect, and semantic loss and real-time requirements have been considered. However, in practical applications, the synthesis of the interference signal and the original audio signal requires precise control. It is necessary to ensure the influence of the interference signal on the audio while avoiding damage to important features such as the emotion and rhythm of the audio. Therefore, the core task of this step is to synthesize the interference signal and the original audio signal and optimize the synthesis effect.

[0106] In this process, we propose an innovative weighted summation synthesis method. By adjusting the weights of each part in the synthesis, we can maximize the interference effect while maintaining the audio quality and coherence as much as possible. We divide the synthesis process into several objectives to optimize the interference effect, the degree of semantic disruption, and rhythmic coherence, respectively.

[0107] Weighted synthesis of interference signal and audio signal:

[0108] First, interference signals and the original audio signal We employ a weighted summation method for synthesis. To balance the interference effect and semantic disruption during the synthesis process, we introduce weighting coefficients. By adjusting By controlling the value of , we can manage the relative weights of the original audio signal and the interference signal, thereby optimizing the balance between interference effect and semantic destruction.

[0109] The synthesis formula is as follows:

[0110]

[0111] It is the final synthesized signal, which contains the result of combining the interference signal and the original audio signal.

[0112] It is the original audio signal.

[0113] It is an optimized interference signal.

[0114] It is a weighting coefficient that controls the ratio between the original audio signal and the interference signal. The larger the value, the greater the proportion of the original audio signal retained; while The smaller the value, the greater the impact of the interference signal.

[0115] By adjusting We can flexibly control the strength of the interference signal, balancing the loss of interference effect and semantic destruction.

[0116] Optimize the degree of semantic disruption and emotional structure during the synthesis process:

[0117] To enhance the destructive effect and camouflage capability of the interference signal, this invention designs a rhythmic scrambling regularization term for semantic masking and structural interference. This term guides the generated interference signal to exhibit rhythmic variation patterns similar to natural speech, thereby enhancing its deceptiveness as "pseudo-speech" in audio systems. We introduce an audio rhythmic structure scrambling loss (RDL) to control the degree to which the interference signal disrupts the original rhythmic structure in the time domain. The audio rhythmic structure similarity loss (RSL) can be achieved by calculating the similarity of the rhythmic patterns between the synthesized audio signal and the original audio signal in the time domain. Our designed regularization term is as follows:

[0118]

[0119] Audio rhythm structure similarity loss, representing the synthesized signal. With the original audio signal Differences in rhythmic structure.

[0120] and : These represent the rhythmic characteristics of the original audio signal and the synthesized audio signal, respectively. Rhythmic information can be extracted by calculating the periodic fluctuations of the audio signal in the time domain.

[0121] L2 norm: representing the rhythmic differences in an audio signal.

[0122] By minimizing We ensure the synthesized signal Disrupting the original audio with mimicry of speech rhythm The temporal structure is improved to enhance the degree of semantic destruction and prevent the interference signal from appearing as completely unstructured white noise.

[0123] To further enhance the semantic distortion capability of the synthesized signal, we also designed a signal-to-noise ratio (SNR) regularization term to constrain the degree of masking of the original audio's semantic message during synthesis. We achieve a higher degree of semantic ambiguity and incomprehension disruption by minimizing the SNR.

[0124]

[0125] Signal-to-noise ratio loss represents the quality difference between the synthesized signal and the original audio signal.

[0126] Signal-to-noise ratio (SNR) is a measure of the quality of a synthesized signal. and the original audio signal The difference in quality. The higher the SNR, the better the audio signal quality.

[0127] By minimizing the signal-to-noise ratio, we effectively enhance the semantic disruption of the audio signal, ensuring that the interference signal disrupts and remains unrecognizable to the original content.

[0128] Ultimately, we obtained the synthesized interference audio signal. This signal can maximize the interference with the target speech. While maintaining the semantic destruction level, the interference also maintains its misleading nature at the auditory camouflage level. Synthetic signal This will be used as input for subsequent steps, further processed and applied. In step 4, we use a weighted summation method to process the interference signal. With the original audio signal The synthesis process is optimized by combining audio rhythmic structure perturbation loss and signal-to-noise ratio regularization term.

[0129] Step 5: Real-time feedback and adaptive adjustment:

[0130] In step 4, we use weighted summation to separate the interference signal. With the original audio signal The final interfering audio signal is obtained by synthesis. However, due to the dynamic changes in audio signals and the environment, the synthesized interference signal may no longer be suitable for new scenarios or real-time requirements. Therefore, the goal of this step is to optimize the generation process of the interference audio signal through real-time feedback and adaptive adjustment, ensuring that its interference effect maximizes the degree of semantic disruption.

[0131] Real-time feedback mechanism: Real-time monitoring of synthesized interference audio signals We will assess its interference effect and the degree of semantic disruption. First, we will judge its interference effect by measuring its interference effect. For the original audio signal The degree of interference is measured using the L2 norm to determine the difference:

[0132]

[0133] here, It measures the difference between the interference signal and the original audio signal in the time domain; the larger the difference, the stronger the interference effect.

[0134] Next, we evaluate the degree of semantic corruption using the signal-to-noise ratio (SNR). A higher SNR indicates a worse degree of semantic corruption. We ensure audio quality by minimizing the SNR loss.

[0135]

[0136] here, Indicates the synthesized signal With the original audio signal The quality difference between them is that the higher the SNR, the worse the semantic corruption.

[0137] Adaptive Adjustment Strategy: Based on real-time feedback, the system will automatically adjust the parameters during the interference signal generation process. Specifically, the system will adjust according to the following strategy:

[0138] If the interference effect is insufficient or the semantic disruption is poor, the system will increase the strength of the interference signal.

[0139] If the semantic disruption is high but the interference effect is weak, the system will appropriately increase the proportion of the interference signal by adjusting the weighting coefficients. This is used to control the ratio between the interference signal and the original audio signal.

[0140] Weighting coefficients The adjustment formula is:

[0141]

[0142] By controlling the relative weights between the original audio signal and the interference signal, the system dynamically adjusts the weights through real-time feedback. The value of ensures the optimal balance between the interference effect of the synthesized signal and the degree of semantic destruction.

[0143] Ultimately, the interference audio signal is optimized through real-time feedback and adaptive adjustment. In step 5, we optimized the interfering audio signal through real-time feedback and adaptive adjustment mechanisms. The system generates interference signals in real time. Through real-time monitoring and adjustment, it ensures that the interference signals can always adapt to environmental changes and maintain the best interference effect and semantic disruption level, providing an efficient and flexible audio interference generation scheme to meet various needs in practical applications.

[0144] The technical solutions described in the above embodiments of this application have at least the following technical effects or advantages: This invention proposes a technical solution capable of dynamically generating interference signals in real time. This solution not only considers the temporal and variability of audio signals but also generates matching interference signals based on the content characteristics of the audio, ensuring the accuracy of the interference signal and maximizing its interference effect. Simultaneously, this invention also considers the degree of semantic disruption, ensuring that the interference signal maximizes the degree of semantic disruption while achieving effective interference. Through this method, this invention can adapt to complex dynamic audio environments and generate effective and accurate interference in different types of audio signals.

[0145] This application also provides an audio interference system based on dynamic analysis, such as... Figure 2 As shown, it includes:

[0146] The feature extraction module is configured to take the original audio signal as input, convert the original audio signal into a frequency domain signal and output a spectrogram, wherein the spectrogram is a two-dimensional matrix, and its different dimensions represent time frames and frequency components, respectively; the spectrogram is input into a convolutional neural network and a long short-term memory network to extract audio temporal features, wherein the convolutional neural network is used to extract local features and the long short-term memory network is used to extract temporal dependencies;

[0147] An interference generation module is configured to acquire random noise and combine audio temporal features with random noise through a convolutional neural network to generate a first interference signal. During the generation process, the first interference signal is generated by combining adversarial loss and temporal similarity regularization. The first interference signal is compared with the original audio signal, and a judgment result is output. The judgment result extracts features from the original audio signal and the first interference signal through multiple convolutional operations, and outputs a judgment on whether it is an interference signal through a fully connected layer.

[0148] The target optimization module is configured to optimize the first interference signal using a multi-objective optimization algorithm. The optimization objectives include maximizing the difference between the first interference signal and the original audio signal, maximizing the deceptiveness and irreversibility of the interference effect, and minimizing the generation time of the first interference signal. After achieving the optimization objectives, a second interference signal is obtained.

[0149] The signal synthesis module is configured to synthesize a third interference signal by weighted summation of the second interference signal and the original audio signal. The third interference signal is optimized by guiding audio rhythm disorder and signal energy reconstruction to enhance semantic masking and content destruction capabilities.

[0150] The feedback optimization module is configured to evaluate the interference effect and semantic destruction degree of the third interference signal at each preset interval, and obtain the evaluation result. The interference effect is measured by the L2 norm to measure the degree of interference of the third interference signal on the original audio signal. When the L2 norm is the largest, the interference effect is the strongest. The semantic destruction degree is measured by the signal-to-noise ratio. When the signal-to-noise ratio is the lowest, it means that the original speech signal is covered by the unstructured interference signal to the greatest extent, the semantic recognition difficulty is the highest, and thus the interference effect is the strongest. Based on the evaluation result, the third interference signal is adjusted according to a preset strategy until the optimal interference effect and semantic destruction degree are obtained, and the final interference signal is output.

[0151] The working principle is as follows: After extracting audio features, an interference signal is generated by combining temporal features with a conditional generative adversarial network (cGAN). By introducing a temporal similarity regularization term, the interference signal is ensured to be consistent with the audio signal in time, maximizing the interference effect while maximizing the degree of semantic destruction. After optimization, the second interference signal is synthesized with the original audio signal by weighted summation. The synthesis process is optimized by combining audio rhythm structure similarity loss and signal-to-noise ratio regularization term. Through real-time monitoring and adjustment, the system can ensure that the interference signal can always adapt to environmental changes and maintain the best interference effect and degree of semantic destruction, providing an efficient and flexible audio interference generation scheme.

[0152] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. An audio interference method based on dynamic analysis, characterized in that, include: S1: Input the original audio signal, convert the original audio signal into a frequency domain signal and output a spectrogram, wherein the spectrogram is a two-dimensional matrix, and its different dimensions represent time frames and frequency components respectively; input the spectrogram into a convolutional neural network and a long short-term memory network to extract audio temporal features, wherein the convolutional neural network is used to extract local features and the long short-term memory network is used to extract temporal dependencies; S2: Obtain random noise, and combine the audio temporal features with the random noise through a convolutional neural network to generate a first interference signal; wherein, the first interference signal is generated by combining adversarial loss and temporal similarity regularization term; compare the first interference signal with the original audio signal and output a judgment result, wherein the judgment result extracts features of the original audio signal and the first interference signal through multi-layer convolutional operations, and outputs a judgment on whether it is an interference signal through a fully connected layer; the adversarial loss is used to measure the similarity between the generated interference signal and the real interference signal, so that the generator generates an interference signal that can deceive the discriminator; S3: The optimization objectives of the first interference signal are achieved through a multi-objective optimization algorithm, wherein the optimization objectives include maximizing the difference between the first interference signal and the original audio signal, maximizing the deceptiveness and irreversibility of the interference effect, and minimizing the generation time of the first interference signal; after achieving the optimization objectives, the second interference signal is obtained. S4: The second interference signal and the original audio signal are combined by weighted summation to obtain the third interference signal. The third interference signal is optimized by guiding audio rhythm disorder and signal energy reconstruction to enhance semantic masking and content destruction capabilities. The third interference signal incorporates an audio rhythm structure similarity loss to constrain rhythm consistency with the second interference signal. The audio rhythm structure similarity loss is obtained by calculating the similarity of the rhythm patterns of the third interference signal and the original audio signal in the time domain. S5: Evaluate the interference effect and semantic destruction degree of the third interference signal at each preset interval to obtain the evaluation result. The interference effect is measured by the L2 norm to measure the degree of interference of the third interference signal on the original audio signal. When the L2 norm is the largest, the interference effect is the strongest. The semantic destruction degree is measured by the signal-to-noise ratio regularization term. When the signal-to-noise ratio regularization term is the lowest, it means that the original speech signal is covered by the unstructured interference signal to a greater extent, the semantic recognition difficulty is increased, and thus the interference effect is enhanced. Based on the evaluation result, adjust the third interference signal according to the preset strategy until the optimal interference effect and semantic destruction degree are obtained, and output the final interference signal.

2. The audio interference method based on dynamic analysis according to claim 1, characterized in that, The first interference signal is generated by converting the audio temporal features into a format suitable for grid processing through an embedding layer and then performing multi-layer convolution operations.

3. The audio interference method based on dynamic analysis according to claim 1, characterized in that, The temporal similarity regularization term measures the temporal difference between the first interference signal and the original audio signal using the L2 norm, thus preventing the generated interference signal from excessively distorting the temporal structure of the original audio signal.

4. The audio interference method based on dynamic analysis according to claim 1, characterized in that, The multi-objective optimization algorithm is performed using a particle swarm optimization algorithm. Each particle represents a second interference signal configuration, and its position contains the characteristics of the second interference signal in the time domain. The dimension of the particle corresponds to the time length of the second interference signal. The optimal solution, i.e., the second interference signal, is obtained by updating the position and velocity of the particles.

5. The audio interference method based on dynamic analysis according to claim 1, characterized in that, In the process of weighted synthesis of the second interference signal and the original audio signal, a weighting coefficient is introduced. When the value of the weighting coefficient is the largest, the proportion of the original audio signal is retained is the largest. When the value of the weighting coefficient is the smallest, the influence of the second interference signal is the largest.

6. The audio interference method based on dynamic analysis according to claim 1, characterized in that, The signal-to-noise ratio regularization term is used to minimize the noise generated during the synthesis process, which is accomplished by minimizing the signal-to-noise ratio.

7. The audio interference method based on dynamic analysis according to claim 1, characterized in that, The preset strategy is as follows: If the interference effect and / or semantic disruption level are below the threshold, increase the strength of the third interference signal; if the semantic disruption level meets the requirements but the interference effect is below the threshold, increase the proportion of the third interference signal.

8. An audio interference system based on dynamic analysis, characterized in that, include: The feature extraction module is configured to take the original audio signal as input, convert the original audio signal into a frequency domain signal and output a spectrogram, wherein the spectrogram is a two-dimensional matrix, and its different dimensions represent time frames and frequency components, respectively; the spectrogram is input into a convolutional neural network and a long short-term memory network to extract audio temporal features, wherein the convolutional neural network is used to extract local features and the long short-term memory network is used to extract temporal dependencies; An interference generation module is configured to acquire random noise and combine audio temporal features with the random noise through a convolutional neural network to generate a first interference signal. This first interference signal is generated by combining adversarial loss and temporal similarity regularization. The first interference signal is compared with the original audio signal, and a judgment result is output. This judgment result extracts features from both the original audio signal and the first interference signal through multiple convolutional operations, and outputs a judgment on whether it is an interference signal through a fully connected layer. The adversarial loss is used to measure the similarity between the generated interference signal and the real interference signal, enabling the generator to generate interference signals that can deceive the discriminator. The target optimization module is configured to optimize the first interference signal using a multi-objective optimization algorithm. The optimization objectives include maximizing the difference between the first interference signal and the original audio signal, maximizing the deceptiveness and irreversibility of the interference effect, and minimizing the generation time of the first interference signal. After achieving the optimization objectives, a second interference signal is obtained. The signal synthesis module is configured to synthesize a third interference signal by weighted summation of a second interference signal and the original audio signal. The third interference signal is optimized through guided audio rhythm distortion and signal energy reconstruction to enhance semantic masking and content destruction capabilities. An audio rhythm structure similarity loss is introduced into the third interference signal to constrain its rhythmic consistency with the second interference signal. This audio rhythm structure similarity loss is obtained by calculating the similarity of the rhythmic patterns of the third interference signal and the original audio signal in the time domain. The feedback optimization module is configured to evaluate the interference effect and semantic destruction degree of the third interference signal at each preset interval, and obtain the evaluation result. The interference effect is measured by the L2 norm to measure the degree of interference of the third interference signal on the original audio signal. When the L2 norm is the largest, the interference effect is the strongest. The semantic destruction degree is measured by the signal-to-noise ratio regularization term. When the signal-to-noise ratio regularization term is the lowest, it means that the original speech signal is covered by the unstructured interference signal to the greatest extent, the semantic recognition difficulty is the highest, and thus the interference effect is the strongest. Based on the evaluation result, the third interference signal is adjusted according to a preset strategy until the optimal interference effect and semantic destruction degree are obtained, and the final interference signal is output.

Citation Information

Patent Citations

  • Communication interference method based on generative adversarial network

    CN112183352A

  • Method and device for generating interference signal of target voice signal

    CN114337908A