Multi-modal data security analysis technology based on depth fuzzy wavelet learning model
Through the deep fuzzy wavelet learning model, the problem of feature representation heterogeneity when data fusion is combined with different modalities is solved, and more accurate and robust security threat detection is achieved.
Patent Information
- Application Number
- CN202510323937.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The existing multimodal data security analysis model has feature representation heterogeneity when dealing with the fusion of different modal data, resulting in limited information sharing and joint analysis effects, affecting the ability to identify complex attack behaviors.
The deep fuzzy wavelet learning model is adopted to enhance data resolution through wavelet transformation, model uncertainty of fuzzy membership function, and optimize network parameters through adaptive optimization algorithms, select key features in combination with information gain, reduce the impact of redundant features, and improve the learning efficiency and detection accuracy of the model.
Effectively integrate data on different modalities, improve the accuracy of security threat detection, enhance the ability to identify complex attack behaviors, and improve the generalization ability and robustness of the model.
Smart Images

Figure CN120223379A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data security analysis, and particularly to a method for deep fuzzy wavelet learning. Background Art
[0002] Multi-modal data security analysis is one of the important applications of artificial intelligence in the field of network security. It identifies potential security threats by analyzing and processing different modal data, thereby enhancing the system's defense capabilities and reducing security risks. In recent years, with the rapid development of technologies such as deep learning, computer vision, and natural language processing, multi-modal data security analysis has made significant progress. However, the network security threats in the real world are highly complex and concealed. When faced with new attack methods or changes in the data environment, existing detection models may be difficult to accurately identify abnormal patterns therein because the new data may contain attack features or malicious behavior patterns that the model has never encountered. This requires the model to be continuously updated and adjusted, resulting in increased consumption of computing resources. The deep fuzzy wavelet learning model has broad application prospects in the field of multi-modal data analysis due to its powerful feature extraction and adaptive learning capabilities. However, traditional methods still face challenges in processing the fusion of different modal data. One of the main problems is the heterogeneity of feature representations between different modal data, which limits the effect of information sharing and joint analysis, thereby affecting the model's ability to identify complex attack behaviors. Summary of the Invention
[0003] The purpose of the present invention is to overcome the above technical drawbacks and deficiencies, and provide a multi-modal data security analysis technology based on a deep fuzzy wavelet learning model.
[0004] The technical steps of the present invention are as follows:
[0005] 1. A multi-modal data security analysis technology based on a deep fuzzy wavelet learning model, comprising the following steps:
[0006] Step 1: Obtain the multi-modal security analysis dataset. The experimental data uses the MIDAS dataset and the AVID dataset. The MIDAS dataset contains text logs, network traffic, and industrial control system (ICS) data, which can be used to detect malicious attacks; the AVID dataset contains adversarial samples for images and videos and is suitable for adversarial attack detection. In the experiment, preprocessing and feature extraction are performed on different modal data respectively. The text data is tokenized, and semantic features are extracted using TF-IDF and BERT embedding methods; for image data, multi-scale wavelet transform (DWT) is used to decompose the images, and low-frequency and high-frequency features are extracted, and deep features are extracted in combination with the ResNet-50 convolutional neural network (CNN); for audio data, Mel-frequency cepstral coefficients (MFCC) are calculated to obtain time-frequency features; for video data, key frames are extracted, and temporal features are extracted using a 3D convolutional network (C3D).
[0007] Step 2: Construct a deep fuzzy wavelet learning model based on the fusion of wavelet neural network (WNN) and fuzzy inference system (FIS). The information gain method is used to select key features to avoid the influence of invalid features on network training. In the model, wavelet transform is used to enhance the data discrimination ability, the fuzzy membership function is used to model uncertainty, and the network parameters are optimized through an adaptive optimization algorithm to improve the accuracy of multi-modal data security analysis.
[0008] Step 3: Optimize the model using a spatial search optimization algorithm to reduce fuzzy overlap between different classes and improve the classification accuracy. In the specific optimization process, the initial network is used as a benchmark, and global search, local search, and opposition search strategies are used to optimize and iterate the population, thereby enhancing the generalization ability and robustness of the model and strengthening the detection ability for malicious samples.
[0009] Step 4: Output the multi-modal data security analysis results. Input the multi-modal data to be detected, and use the optimized deep fuzzy wavelet learning model for classification and anomaly detection to identify potential security threats, including data poisoning, adversarial attacks, etc., and output the analysis results.
[0010] The specific steps of Step 1 are as follows:
[0011] Step 1.1: The datasets used in this experiment are the MIDAS dataset and the AVID dataset. The MIDAS dataset contains 12,000 ICS logs, 8,500 network traffic data, and 6,300 text logs, covering normal operation data and various attack samples. The AVID dataset contains 5,000 adversarial sample images and 3,200 adversarial sample videos, involving security threats such as adversarial noise, steganography attacks, and data poisoning.
[0012] Step 1.2, Data Preprocessing: For different modal data, perform denoising, format conversion, and normalization respectively: For text data, remove stop words, perform stemming, and use the BERT tokenizer for sub-word decomposition; For image data: Unify the resolution to 224×224, perform normalization, and convert to the Lab color space to enhance edge features; For audio data, perform time-frequency analysis using the Short-Time Fourier Transform (STFT) and remove silent segments; For video data, extract key frames and calculate motion information using the optical flow method.
[0013] Step 1.3, Text Data Feature Extraction: Use the TF-IDF and BERT embedding methods to obtain text semantic features.
[0014] TF-IDF calculation features:
[0015]
[0016] where f t is the number of occurrences of word t in document d, N is the total number of documents, and n t is the number of documents containing t.
[0017] Step 1.4, Image Data Feature Extraction: Use the Discrete Wavelet Transform (DWT) and ResNet-50 for feature extraction.
[0018] Wavelet transform features:
[0019] Perform multi-scale wavelet decomposition on the image to extract the low-frequency component (LL) and high-frequency components (LH, HL, HH):
[0020]
[0021] where s[m,n] is the input image, is the wavelet basis function.
[0022] Step 1.5, Audio Data Feature Extraction: Calculate the Mel Frequency Cepstral Coefficients (MFCC) and spectral centroid.
[0023] MFCC calculation:
[0024]
[0025] where X m is the spectrum after Mel filtering, and k is the coefficient index.
[0026] Spectral centroid calculation:
[0027]
[0028] where S(f) is the power spectrum at frequency f.
[0029] Step 1.6: Extract video data features. Calculate the inter-frame motion information using the optical flow method and extract spatio-temporal attention features.
[0030] Optical flow calculation features:
[0031]
[0032] where u and v represent the components of the optical flow in the x and y directions respectively, and I is the image pixel intensity.
[0033] Step 1.7: Feature normalization and dimensionality reduction
[0034] Normalization: Use the Z-score standardization method to make the feature distribution have a mean of 0 and a variance of 1:
[0035]
[0036] where μ is the mean and σ is the standard deviation.
[0037] Then use principal component analysis to reduce the feature dimension and improve the calculation efficiency.
[0038] The specific steps of Step 2 are as follows:
[0039] Step 2.1: Feature selection based on information gain
[0040] Screen the features extracted from the MIDAS and AVID datasets, calculate the information gain of each feature, and eliminate redundant or invalid features.
[0041] The information gain calculation formula is as follows:
[0042] IG(T,X) = H(T) - H(T|X)
[0043] where IG(T,X) represents the information gain of feature X for the target variable T, H(T) is the entropy of the target variable, and H(T|X) is the conditional entropy of T given feature X.
[0044] Step 2.2: Construct a fuzzy wavelet neural network
[0045] Use fuzzy C-means (FCM) clustering to partition the feature space and calculate the fuzzy membership degree:
[0046]
[0047] where u ij is the membership degree of sample x j to cluster i, v i is the cluster center, and m is the fuzzy exponent.
[0048] Then, wavelet transform is used to enhance data features, and the obtained fuzzy features are decomposed into multi-scale features:
[0049]
[0050] Among them, W ψ (a, b) is the wavelet coefficient, is the wavelet basis function, and a and b represent the scale and translation parameters respectively.
[0051] Step 2.3: Optimize the fuzzy rules and model parameters using an adaptive optimization algorithm
[0052] Optimizing the fuzzy rules and network parameters using an adaptive optimization algorithm to improve the learning efficiency and adaptability of the model specifically includes the following steps:
[0053] Adaptive rule adjustment:
[0054] Using a data-driven method, dynamically adjust the fuzzy rules in the fuzzy wavelet neural network to adapt to the distribution changes of different categories, and the importance is represented by I r :
[0055]
[0056] Among them, u rj is the membership degree of the sample X j to the rule r, and IG(T, X j ) is the information gain of this feature.
[0057] Then, use the Adam optimization algorithm to optimize the adaptive learning rate:
[0058] Adam optimization algorithm:
[0059]
[0060] Among them, θ t is the current model parameter, L is the loss function, α is the initial learning rate, m t and v t are the first-order and second-order moment estimates respectively, and β1 and β2 are the exponential decay rates.
[0061] Then, dynamically adjust the learning rate according to the loss change, enhance the adaptability to multi-modal data, and perform effective feature screening. The adjustment formula is as follows:
[0062]
[0063] Among them, η t is the current learning rate, and λ is the decay coefficient.
[0064] The specific steps of Step 3 are as follows:
[0065] Step 3.1: Using the initial model obtained in Step 2 as a benchmark, construct an optimization population and initialize multiple possible parameter combinations to ensure that the optimization search range covers a sufficient feature space.
[0066] Assume the model parameter vector is:
[0067] Θ = {θ1, θ2,..., θ n}
[0068] where θ i represents the i-th parameter of the model. Initialize the population {Θ 1 , Θ 2 ,..., Θ M} as the starting point for optimization, where M is the population size.
[0069] Step 3.2: Use a global search strategy to randomly sample in the high-dimensional feature space. The goal is to minimize the classification error loss function:
[0070]
[0071] where N is the training samples, x i is the training data, y i is the true label, f(x i ; Θ) is the model prediction output, is the cross-entropy loss function, and an optimization is performed using a search strategy based on gradient estimation:
[0072]
[0073] where η is the learning rate, is the gradient of the loss function with respect to the parameters.
[0074] Step 3.3: Based on the excellent individuals obtained from the global search, introduce local search and use gradient information or heuristic methods to fine-tune the model parameters to reduce the randomness in the optimization process and improve the classification stability. Use dynamic gradient descent to accelerate convergence:
[0075]
[0076] Θ t+1 = Θ t + v t+1
[0077] β is the momentum factor, which makes the parameter update affected by the gradient at the previous moment, thereby smoothing the optimization path.
[0078] Step 3.4: Introduce the adversarial search strategy, construct adversarial features or data augmentation strategies, perform robustness testing on the model, and optimize the adversarial loss to improve the model's detection ability for abnormal samples and attack samples.
[0079] To improve the model's detection ability against adversarial attacks, an adversarial perturbation δ is added to generate adversarial samples:
[0080]
[0081] where ∈ is the perturbation amplitude and sign(·) represents the gradient direction. The generated x adv can be used to evaluate the adversarial robustness of the model. The optimization objective becomes:
[0082]
[0083] where the max operation finds the most difficult-to-detect adversarial samples, and the min operation optimizes the model's robustness against them.
[0084] Step 3.5: Perform performance evaluation on the optimized model, compare the classification accuracy, robustness metrics, computational overhead, etc. before and after optimization, and finally select the optimal parameter combination to complete the optimization process. The evaluation metrics are as follows:
[0085] Classification accuracy:
[0086]
[0087] where 1(·) is the indicator function, is the predicted value.
[0088] Adversarial sample detection success rate:
[0089]
[0090] where is the prediction result of the adversarial sample x adv and ASR reflects the robustness of the model.
[0091] Finally, select the optimized parameter Θ that maximizes the classification accuracy and minimizes the ASR as the final model parameter.
[0092] Step 4: Output the security analysis result, detect the input multi-modal data, and output the data content security analysis result.
[0093] Advantages and beneficial effects of the present invention
[0094] The present invention proposes a multi-modal data security analysis technology based on a deep fuzzy wavelet learning model, which can effectively fuse data of different modalities and improve the accuracy of security threat detection. First, the present invention uses a deep fuzzy wavelet learning model to extract and fuse features from multi-modal data such as text, images, audio, and video, thereby enhancing the ability to identify complex security threats. Second, during the model training process, an information gain method is used to select key features, avoiding the influence of invalid features on network training and improving the learning efficiency and detection accuracy of the model. In addition, the present invention uses a spatial search optimization algorithm to optimize the model, reducing the fuzzy overlap phenomenon between different categories and improving the classification accuracy and generalization ability. Experimental results on the MIDAS and AVID datasets show that the present invention exhibits superior detection performance in multi-modal data security analysis tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] Figure 1 is the overall implementation flowchart of the deep fuzzy wavelet learning model for multi-modal data security;
[0096] Figure 2 is the overall flowchart of the deep fuzzy wavelet learning model; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0097] The present invention will be described in detail below with reference to the accompanying drawings and examples.
[0098] Step 1: Obtain a multi-modal security analysis dataset. The experimental data uses the MIDAS dataset and the AVID dataset. In the experiment, preprocessing and feature extraction are performed on data of different modalities respectively. The text data is tokenized, and semantic features are extracted using the TF-IDF and BERT embedding methods; the image data is decomposed using multi-scale wavelet transform (DWT) to extract low-frequency and high-frequency features, and deep features are extracted in combination with the ResNet-50 convolutional neural network (CNN); the audio data calculates the Mel frequency cepstral coefficients (MFCC) to obtain time-frequency features; the video data extracts key frames, and temporal features are extracted using a 3D convolutional network (C3D). The specific steps are as follows:
[0099] Step 1.1: Randomly extract samples from the MIDAS and AVID datasets respectively. The MIDAS dataset contains 12,000 ICS logs, 8,500 network traffic data, and 6,300 text logs; the AVID dataset contains 5,000 adversarial sample images and 3,200 adversarial sample videos. The training set and the test set are divided, with 70% used for training and 30% used for testing.
[0100] Step 1.2: Preprocess the text data, including removing stop words, stemming, and performing sub-word decomposition using the BERT tokenizer. Finally, convert it into TF-IDF feature vectors (dimension 5000) and BERT embedding vectors (dimension 768). The training set is 9800×(5000 + 768), and the test set is 4200×(5000 + 768).
[0101] Step 1.3: Preprocess the image data. First, resize the images to 224×224 and perform color space conversion to the Lab color space to enhance edge features. Subsequently, apply multi-scale wavelet transform (DWT) for image decomposition, extract low-frequency and high-frequency components, and combine with ResNet-50 for feature extraction. The final feature vector dimension is 2048, the training set dimension is 3500×2048, and the test set dimension is 1500×2048.
[0102] Step 1.4: Preprocess the audio data. Use short-time Fourier transform (STFT) for time-frequency analysis and remove silent segments. Subsequently, calculate Mel-frequency cepstral coefficients (MFCC) to obtain 40-dimensional features, and combine with 5 additional features such as spectral centroid. The final feature dimension is 45, the training set is 2240×45, and the test set is 960×45.
[0103] Step 1.5: Process the video data. First, use the optical flow method to calculate the inter-frame motion information and extract 16 key frames from each video segment, with each frame size of 112×112×3. Then, use a 3D convolutional network (C3D) to extract temporal features. The final feature vector dimension is 4096, the training set is 2240×4096, and the test set is 960×4096.
[0104] Step 1.6: Perform Z-score normalization on all features to make the mean 0 and the variance 1, and use PCA for dimensionality reduction. Finally, unify the feature dimension to 512 to improve the computational efficiency. The training set dimension is 9800×512, and the test set dimension is 4200×512.
[0105] Step 2: Construct a deep fuzzy wavelet learning model based on the fusion of wavelet neural network (WNN) and fuzzy inference system (FIS). Use the information gain method to select key features to avoid the influence of invalid features on network training. In the model, wavelet transform is used to enhance the data discrimination ability, fuzzy membership functions are used to model uncertainties, and network parameters are optimized through an adaptive optimization algorithm to improve the accuracy of multi-modal data security analysis. The specific steps are as follows:
[0106] Step 2.1: Calculate the information gain of all features and select the top 128 features with the highest contribution to reduce the influence of redundant features on the training process.
[0107] Step 2.2: Using the Fuzzy C-Means (FCM) clustering method, divide the selected 128-dimensional features into 5 clusters and calculate the fuzzy membership degrees. Subsequently, apply wavelet transform to perform multi-scale decomposition on the features, extract high-frequency and low-frequency features, and finally combine them into a new feature representation. The dimension of the training set is 9800×256, and the dimension of the test set is 4200×256.
[0108] Step 2.3: Train the WNN model using the Adam optimization algorithm, set the initial learning rate to 0.001, reduce it by 0.5 times every 10 rounds, set the maximum number of training rounds to 100 rounds, set the batch size to 128, and use the mean squared error (MSE) as the loss function.
[0109] Step 3: Optimize the model using the spatial search optimization algorithm, reduce the fuzzy overlap of different categories, and improve the classification accuracy. In the specific optimization process, taking the initial network as the benchmark, use the global search, local search, and opposition search strategies to optimize and iterate the population, so as to improve the generalization ability and robustness of the model and enhance the detection ability for malicious samples. The specific steps are as follows:
[0110] Step 3.1: Initialize 50 different combinations of hyperparameters, including the learning rate (0.001~0.01), the number of fuzzy rules (5~10), and the wavelet basis functions (Haar, Daubechies, Symlets), and construct a population for search optimization.
[0111] Step 3.2: Adopt the global search strategy for preliminary parameter screening, calculate the cross-entropy loss, and the optimal parameter combination is: learning rate 0.005, wavelet basis function Daubechies-4, number of fuzzy rules 7, and the classification accuracy reaches 92.4%.
[0112] Step 3.3: On the basis of the global search, apply the local search strategy, use the momentum factor 0.9 for gradient update, further optimize the classification model, and finally the classification accuracy is increased to 93.8%.
[0113] Step 3.4: Introduce the opposition search strategy, generate adversarial samples, add perturbations to the input data using the FGSM method (perturbation amplitude ε = 0.01), calculate the detection success rate of the model on the adversarial samples, optimize the adversarial loss, and finally the ASR (attack success rate of adversarial samples) drops to 6.2%.
[0114] Step 3.5: Evaluate the final model, with a classification accuracy of 94.2% and a detection success rate of adversarial samples of 93.1%, finally determine the optimal combination of hyperparameters and complete the model training.
[0115] Step 4: Output the security analysis results, detect the input multi-modal data, and output the data content security analysis results.
[0116] The deep fuzzy wavelet learning model proposed by the present invention was subjected to 8 repeated experiments using the MIDAS dataset and the AVID dataset, and the average results were taken for analysis. The experiments used classification accuracy (%), robustness index (RI), and adversarial attack error rate (AER, %) as the main evaluation indicators, with FMNN (fuzzy neural network) as the comparison baseline. The experimental results are shown in Table 1.
[0117] Comparison of data results in Table 1
[0118]
[0119] The above is only a preferred example of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A multimodal data security analysis technology based on a deep fuzzy wavelet learning model, comprising the following steps: Step 1. Obtain a multimodal security analysis dataset. The experimental data uses the MIDAS dataset and the AVID dataset. The MIDAS dataset contains text logs, network traffic, and industrial control system (ICS) data, which can be used to detect malicious attacks; the AVID dataset contains adversarial samples for images and videos, which is suitable for adversarial attack detection. In the experiment, different modal data are preprocessed and feature extracted respectively, text data is segmented, and semantic features are extracted using TF-IDF and BERT embedding methods; image data uses multi-scale wavelet transform (DWT) to decompose the image, extract low-frequency and high-frequency features, and combine ResNet-50 convolutional neural network (CNN) to extract deep features; Calculate the Mel Spectral Coefficient (MFCC) of audio data to obtain time-frequency features; Key frames are extracted from video data, and temporal features are extracted using a 3D convolutional network (C3D). Step 2: Build a deep fuzzy wavelet learning model based on the fusion of wavelet neural network (WNN) and fuzzy inference system (FIS). Use the information gain method to select key features to avoid invalid features affecting network training. In the model, wavelet transform is used to enhance data resolution, fuzzy membership function is used to model uncertainty, and the network parameters are optimized through adaptive optimization algorithm to improve the accuracy of multimodal data security analysis. Step 3: Use the spatial search optimization algorithm to optimize the model, reduce the fuzzy overlap of different categories, and improve the classification accuracy. In the specific optimization process, the initial network is used as the benchmark, and the global search, local search, and adversarial search strategies are used to optimize the population iteratively, thereby improving the generalization and robustness of the model and enhancing the detection ability of malicious samples. Step 4: Output the multimodal data security analysis results. Input the multimodal data to be tested, use the optimized deep fuzzy wavelet learning model for classification and anomaly detection, identify potential security threats, including data poisoning, adversarial attacks, etc., and output the analysis results.
2. The multimodal data security analysis method based on deep fuzzy wavelet learning model according to claim 1 is characterized by: The specific steps for obtaining multimodal data and extracting features as described in step 1 are as follows: Step 1.
1. The datasets used in this experiment are MIDAS dataset and AVID. The MIDAS dataset contains 12,000 ICS logs, 8,500 network traffic data, and 6,300 text logs, covering normal operation data and various attack samples. The AVID dataset contains 5,000 adversarial sample images and 3,200 adversarial sample videos, involving security threats such as adversarial noise, steganographic attacks, and data poisoning. Step 1.2: Data preprocessing: De-noising, format conversion and standardization are performed for different modal data: for text data, stop words are removed, stemming is performed, and subword decomposition is performed using the BERT word segmenter; for image data, the resolution is unified to 224×224, normalized, and converted to Lab color space to enhance edge features; For audio data, short-time Fourier transform (STFT) is used for time-frequency analysis and silent segments are removed; Video data, key frames are extracted, and motion information is calculated using the optical flow method. Step 1.3: Extract text data features and use TF-IDF and BERT embedding methods to obtain text semantic features. TF-IDF calculation features: Among them, f t is the number of occurrences of word t in document d, N is the total number of documents, n t is the number of documents containing t. Step 1.4: Image data feature extraction, using wavelet transform (DWT) and ResNet-50 for feature extraction. Wavelet transform features: Perform multi-scale wavelet decomposition on the image to extract low-frequency components (LL) and high-frequency components (LH, HL, HH): Among them, s[m,n] is the input image, is the wavelet basis function. Step 1.5: Extract audio data features and calculate Mel frequency spectrum coefficients (MFCC) and spectral centroid. MFCC calculation: Among them, X m is the spectrum after Mel filtering, and k is the coefficient index. Spectral centroid calculation: Where S(f) is the power spectrum at frequency f. Step 1.6: Extract video data features, use the optical flow method to calculate inter-frame motion information, and extract spatiotemporal attention features. Optical flow calculation features: Among them, u and v represent the components of optical flow in the x and y directions respectively, and I is the image pixel intensity. Step 1.7: Feature Normalization and Dimensionality Reduction Normalization: Use the Z-score normalization method to make the feature distribution mean 0 and variance 1: Among them, μ is the mean and σ is the standard deviation. Then use principal component analysis to reduce feature dimensions and improve computational efficiency.
3. The multimodal data security analysis method of the deep fuzzy wavelet learning model according to claim 1 is characterized by: The deep fuzzy wavelet learning model described in step 2 is constructed, and the information gain theory is used to screen important features to avoid the interference of invalid features on model learning. At the same time, wavelet transform is combined in the fuzzy reasoning process to improve the feature representation ability, and the adaptive optimization algorithm is used to optimize the model parameters. The specific steps are as follows: Step 2.1: Feature selection based on information gain The features extracted from the MIDAS and AVID datasets were screened, the information gain of each feature was calculated, and redundant or invalid features were eliminated. The information gain calculation formula is as follows: IG(T,)=H(T)-H(T|X) Among them, IG(T,X) represents the information gain of feature X on the target variable T, H(T) is the entropy of the target variable, and H(T|X) is the conditional entropy of T given feature X. Step 2.2: Constructing fuzzy wavelet neural network Fuzzy C-means (FCM) clustering is used to divide the feature space and calculate the fuzzy membership: Among them, u ij For sample x j The membership degree to cluster i, v i is the cluster center, and m is the fuzzy index. Then use wavelet transform to enhance data features and decompose the obtained fuzzy features into multi-scale features: Among them, W ψ (a,b) are wavelet coefficients, is the wavelet basis function, a and b represent the scale and translation parameters respectively. Step 2.3: Optimize fuzzy rules and model parameters using adaptive optimization algorithm Adopting adaptive optimization algorithm to optimize fuzzy rules and network parameters and improve the learning efficiency and adaptability of the model specifically includes the following steps: Adaptive rule adjustment: Using data-driven methods, the fuzzy rules in the fuzzy wavelet neural network are dynamically adjusted to adapt to the distribution changes of different categories. r express: Among them, u rj For sample X j The membership degree to rule r, IG(T,X j ) is the information gain of this feature. Then use the Adam optimization algorithm to optimize the adaptive learning rate: Adam optimization algorithm: Among them, θ t is the current model parameter, L is the loss function, α is the initial learning rate, m t and v t are the first-order and second-order moment estimates, β1 and β2 are exponential decay rates. Then, the learning rate is adjusted dynamically according to the loss changes to enhance the adaptability to multimodal data and perform effective feature screening. The adjustment formula is as follows: Among them, η t is the current learning rate, and λ is the decay coefficient.
4. The multimodal data security analysis method of the deep fuzzy wavelet learning model according to claim 1 is characterized by: The optimization process described in step 3 uses a spatial search optimization algorithm to optimize model parameters and feature representations through global search, local search, and adversarial search strategies to improve generalization and adversarial robustness. The specific steps are as follows: Step 3.1: Using the initial model trained in step 2 as a benchmark, construct an optimization population and initialize multiple possible parameter combinations to ensure that the optimization search range covers enough feature space. Assume that the model parameter vector is: Θ={θ1,θ2,...,θ n } Among them, θ i represents the i-th parameter of the model, and initializes the population {Θ 1 ,Θ 2 ,...,Θ M } as the optimization starting point, where M is the population size. Step 3.2: Use a global search strategy to randomly sample in the high-dimensional feature space, with the goal of minimizing the classification error loss function: Among them, N is the training sample, x i is the training data, y i is the true label f(x i ; Θ) is the model prediction output, l(·) is the cross entropy loss function, and the search strategy based on gradient estimation is used for optimization: Where η is the learning rate, is the gradient of the loss function with respect to the parameters. Step 3.3: Based on the excellent individuals obtained by global search, local search is introduced, and gradient information or heuristic methods are used to fine-tune model parameters to reduce randomness in the optimization process and improve classification stability. Use dynamic gradient descent to accelerate convergence: I t+1 =Θ t +v t+1 β is the momentum factor, which makes the parameter update affected by the gradient of the previous moment, thus smoothing the optimization path. Step 3.4: Introduce an adversarial search strategy, construct adversarial features or data enhancement strategies, test the robustness of the model, and optimize the adversarial loss to improve the model's detection capabilities for abnormal samples and attack samples. In order to improve the detection ability of the model against adversarial attacks, adversarial perturbation δ is added to generate adversarial samples: Among them, ∈ is the perturbation amplitude, sign(·) represents the gradient direction, and the generated x adv It can be used to evaluate the adversarial robustness of the model. The optimization objective becomes: Among them, the max operation finds the most difficult to detect adversarial examples, and the min operation optimizes the model's robustness to them. Step 3.5: Evaluate the performance of the optimized model, compare the classification accuracy, robustness index, computational overhead, etc. before and after optimization, and finally select the optimal parameter combination to complete the optimization process. The evaluation indicators are as follows: Classification accuracy: Among them, 1(·) is the indicator function, is the predicted value. Adversarial sample detection success rate: in, is the adversarial sample x adv The prediction results of ASR reflect the robustness of the model. Finally, the optimized parameter Θ that achieves the highest classification accuracy and the lowest ASR is selected as the final model parameter.
5. The multimodal data security analysis method of the deep fuzzy wavelet learning model according to claim 1 is characterized by: Step 4 outputs the security analysis results, detects the input multimodal data, and outputs the data content security analysis results.
Citation Information
Patent Citations
Public opinion risk discovery method based on multi-modal fusion algorithm
CN116756688A
Traffic big data acquisition and intelligent monitoring method and Internet of Things system thereof
CN117440017A
Regional sluice system scheduling optimization method and system
CN119005064A
Dangerous behavior identification and early warning method based on multi-modal analysis
CN119360278A
Wasserstein distance and difference metric-combined chest radiograph anomaly identification domain adaptation method and system
US20240153243A1
Cited By
Multi-modal large model confrontation safety detection method and system
CN120639526A
A multimodal large model confrontation security detection method and system
CN120639526B