An AI data processing and encryption method and device based on deep learning

Through the AI ​​data processing and encryption method based on deep learning, multimodal data is processed and complex dynamic key seeds are generated, and the problem of difficulty in capturing the complex associations between multimodal data and the limitations of the key generation mechanism in the prior art is solved, and efficient data encryption and security improvement are achieved.

CN119519949BActive Publication Date: 2025-06-13SUZHOU MENREN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411576790.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2025-06-13
Estimated Expiration
2044-11-06

AI Technical Summary

Technical Problem

The prior art is difficult to process multimodal data, cannot effectively capture complex associations between data, and the fixed key generation mechanism has limitations when facing complex and changing network environments.

Method used

Using AI data processing and encryption method based on deep learning, multimodal data is collected for preprocessing, features are extracted and features are fused through cross attention mechanisms to generate multimodal comprehensive feature vectors. Then, a graph convolutional network GCN relationship graph is constructed, node embedding vectors are generated, and multimodal comprehensive feature vectors are combined to generate enhancement nodes. Finally, complex dynamic key seeds are generated through the autoregressive AR model, the final key is derived through HKDF, and the data is encrypted using the AES-256 encryption algorithm.

Benefits of technology

It effectively captures the complex relationship between data, improves the expressiveness of node embedding vectors, significantly improves the randomness and unpredictability of keys, and enhances the security of data encryption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119519949B_ABST
    Figure CN119519949B_ABST
Patent Text Reader

Abstract

The present invention discloses an AI data processing and encryption method and device based on deep learning, which relates to the technical field of multi-modal graph embedding. The method includes collecting multi-modal data and preprocessing the multi-modal data; extracting the features of the multi-modal data, fusing the multi-modal data features through a cross-attention mechanism to generate a multi-modal comprehensive feature vector; constructing a relationship graph of an application graph convolutional network GCN to generate node embedding vectors, and combining the multi-modal comprehensive feature vector to obtain enhanced nodes; after data encryption, protecting the data through secure transmission and distributed storage, and using a reinforcement learning model to select an optimal encryption strategy. The present invention realizes diversified data collection and efficient preprocessing through multi-modal graph embedding technology, constructs a data relationship graph and applies the graph convolutional network GCN to update the node embedding vectors, combines the multi-modal comprehensive feature vector to generate enhanced node embedding vectors, effectively captures the complex relationships between data, and improves the expressiveness of the node embedding vectors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal graph embedding, and particularly to an AI data processing and encryption method and device based on deep learning. Background Art

[0002] With the development of big data and artificial intelligence technologies, the security and privacy protection of data have become increasingly important issues; the original data encryption methods mainly rely on fixed key generation mechanisms, such as symmetric encryption and asymmetric encryption based on cryptographic algorithms; however, these methods have certain limitations in the face of complex and changing network environments; for example, fixed keys are easily cracked and cannot dynamically adapt to the constantly changing threats. In recent years, deep learning technologies have demonstrated powerful capabilities in feature extraction, data processing, and pattern recognition, providing new ideas for data encryption; by combining deep learning and graph convolutional network GCN, more efficient data feature fusion and node embedding vector update can be achieved, thereby generating more complex dynamic key seeds and improving the security of data encryption.

[0003] The existing technologies usually can only process single-type data and cannot fully utilize the information of multiple data sources; moreover, the existing technologies mainly rely on fixed key generation mechanisms, such as symmetric encryption and asymmetric encryption based on cryptographic algorithms, and the existing data encryption methods often lack effective processing and feature fusion of multimodal data and are difficult to capture the complex associations between data. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an AI data processing and encryption method and device based on deep learning to solve the problem of difficult-to-capture complex associations between data.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides an AI data processing and encryption method based on deep learning, which includes collecting multimodal data and preprocessing the multimodal data; extracting the features of the multimodal data, fusing the multimodal data features through a cross-attention mechanism to generate a multimodal comprehensive feature vector; constructing a graph convolutional network GCN relationship graph to generate node embedding vectors, and combining the multimodal comprehensive feature vector to obtain enhanced nodes; generating complex key seeds from the enhanced nodes, deriving the final key through HKDF, and encrypting the data through an encryption algorithm; after completing data encryption, protecting the data through secure transmission and distributed storage, and using a reinforcement learning model to select the optimal encryption strategy.

[0008] As a preferred solution of the AI data processing and encryption method based on deep learning according to the present invention, wherein: collecting multi-modal data and preprocessing the multi-modal data, the specific steps are as follows,

[0009] Use web crawlers to collect text data, image data, audio data, and communication records from server logs, user behavior records, and network traffic;

[0010] Perform word segmentation on the text data;

[0011] Convert the text after word segmentation of the text data to lowercase, remove stop words, and extract stems to obtain clean text data;

[0012] Convert the clean text data into vectors;

[0013] For all images in the image data, scale them to a unified size, use the piecewise linear transformation method to retain the image edge information, and randomly rotate, flip, and crop the images;

[0014] Use Mel-frequency cepstral coefficients to extract audio features from the audio data, use a frequency-domain filter to denoise the audio data, and segment the audio signal;

[0015] Clean invalid entries and duplicates in the communication record data.

[0016] As a preferred solution of the AI data processing and encryption method based on deep learning according to the present invention, wherein: extracting the features of the multi-modal data, fusing the multi-modal data features through a cross-attention mechanism to generate a multi-modal comprehensive feature vector, the specific steps are as follows,

[0017] Use the BERT model to process the preprocessed text data to generate text feature vectors;

[0018] Use the ResNet-50 model to extract image feature vectors from the preprocessed image data;

[0019] Use the GRU model to process the preprocessed audio data to obtain hidden state vectors;

[0020] Fuse the features of the three-modal data through the cross-attention mechanism in multi-modal feature fusion to generate a multi-modal comprehensive feature vector.

[0021] As a preferred solution of the AI data processing and encryption method based on deep learning according to the present invention, wherein: constructing a graph convolutional network GCN relationship graph, generating node embedding vectors, and combining with the multi-modal comprehensive feature vector to obtain enhanced nodes, the specific steps are as follows,

[0022] Construct a graph convolutional network GCN relationship graph according to the preprocessed communication records;

[0023] Create a graph using the NetworkX library, add nodes and edges, and define node v and edge attributes;

[0024] Perform a non-linear transformation through the activation function σ to obtain the node embedding vector The expression is:

[0025]

[0026] Among them, is the embedding vector of node v at the l-th layer, is the embedding vector of neighbor node u at the l-th layer, N(v) is the neighbor set of node v, is the cosine similarity;

[0027] Combine the node embedding vectors in the generated graph convolutional network GCN relationship graph with the multi-modal comprehensive feature vectors to generate enhanced nodes.

[0028] As a preferred solution of the AI data processing and encryption method based on deep learning described in the present invention, wherein: the enhanced nodes are used to generate complex key seeds, the final key is derived through HKDF, and the data is encrypted through an encryption algorithm. The specific steps are as follows,

[0029] Introduce the autoregressive AR model in the time series to generate dynamic key seeds. The expression is:

[0030]

[0031] Among them, S t is the key seed at the current time point, a i is the autoregressive coefficient, S t-i is the past key seed, p is the autoregressive order, i is the lag order index in the time series, ∈ t is the error term;

[0032] Determine the autoregressive order and autoregressive coefficient through training with historical data;

[0033] Initialize the past key seed;

[0034] Convert the enhanced nodes into vectors of fixed length and input them into the autoregressive AR model to update the key seed at the current time point to obtain the complex dynamic key seed S' t , the expression is:

[0035]

[0036] Among them, is the result after being processed by the hash function, It is an enhanced node;

[0037] Derive the final key from the complex dynamic key seed through the HKDF key derivation function;

[0038] Select AES-256 as the encryption algorithm, set the encryption mode to CBC, and generate and use the initialization vector IV;

[0039] Dynamically adjust the encryption algorithm parameters according to the enhanced node;

[0040] According to the confidentiality level of the data, select the corresponding key and initialization vector for data encryption, and use the AES-256-CBC mode during the encryption process.

[0041] As a preferred solution of the AI data processing and encryption method and device based on deep learning according to the present invention, wherein: after the data encryption is completed, protect the data through secure transmission and distributed storage, and use a reinforcement learning model to select the optimal encryption strategy. The specific steps are as follows:

[0042] Securely transmit the encrypted data through the TLS / SSL protocol;

[0043] Use HDFS distributed file storage to disperse and store the encrypted data in cloud servers, and apply redundant coding to improve data reliability;

[0044] Train a reinforcement learning agent through a reinforcement learning model based on the PPO algorithm, so that the reinforcement learning agent selects the optimal encryption strategy according to the current environmental state.

[0045] As a preferred solution of the AI data processing and encryption method based on deep learning according to the present invention, wherein: the reinforcement learning agent selects the optimal encryption strategy according to the current environmental state. The specific steps are as follows:

[0046] Define the state space, action space, and reward function of the encryption working environment;

[0047] Construct a reinforcement learning model based on the PPO algorithm;

[0048] The reinforcement learning model agent selects an action according to the current encryption working environment state through a neural network;

[0049] Apply the selected action to the encryption working environment for data encryption;

[0050] Give rewards according to the security score and performance overhead in the defined reward function;

[0051] If the complex encryption successfully resists threat attacks, give a positive reward;

[0052] If the simple encryption causes performance degradation, give a negative reward;

[0053] Set a threshold. When the reinforcement learning model agent reaches and exceeds the threshold, stop the training.

[0054] In a second aspect, the present invention provides an AI data processing and encryption device based on deep learning, including a data collection module, a data extraction module, a construction module, a key generation module, and a transmission optimization module; the data collection module is used to collect multimodal data and preprocess the multimodal data; the data extraction module is used to extract the features of the multimodal data, fuse the multimodal data features through a cross-attention mechanism, and generate a multimodal comprehensive feature vector; the construction module is used to construct a graph convolutional network GCN relationship graph, generate node embedding vectors, and combine the multimodal comprehensive feature vector to obtain enhanced nodes; the key generation module is used to generate a complex key seed from the enhanced nodes, derive the final key through HKDF, and encrypt the data through an encryption algorithm; the transmission optimization module is used to protect the data through secure transmission and distributed storage after data encryption, and use a reinforcement learning model to select an optimal encryption strategy.

[0055] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the AI data processing and encryption method and device based on deep learning as described in the first aspect of the present invention is implemented.

[0056] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the AI data processing and encryption method based on deep learning as described in the first aspect of the present invention is implemented.

[0057] The beneficial effects of the present invention are as follows: Through the multimodal graph embedding technology, the diversified collection and efficient preprocessing of data are realized, and the quality and consistency of the data are improved; the graph convolutional network GCN relationship graph is constructed to update the node embedding vectors, and the enhanced nodes are generated by combining the multimodal comprehensive feature vectors, effectively capturing the complex relationships between the data and improving the expressiveness of the node embedding vectors; then, a complex dynamic key seed is generated through an autoregressive AR model, and the final key is derived through HKDF, and the AES-256 encryption algorithm is selected to encrypt the data, significantly improving the randomness and unpredictability of the key and enhancing the security of data encryption. Description of the Drawings

[0058] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0059] Figure 1 It is a flowchart of the AI data processing and encryption method based on deep learning in Embodiment 1.

[0060] Figure 2 It is a flowchart of the distributed storage and optimized encryption data strategy in Embodiment 1. Detailed implementation manners

[0061] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific implementation manners of the present invention in conjunction with the accompanying drawings of the specification.

[0062] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0063] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or selectively exclusive embodiment with other embodiments.

[0064] Embodiment 1, referring to Figure 1 and Figure 2 , is the first embodiment of the present invention. This embodiment provides an AI data processing and encryption method based on deep learning, including the following steps:

[0065] S1. Collect multi-modal data and preprocess the multi-modal data.

[0066] Furthermore, use web crawlers to collect text data, image data, audio data, and communication records from server logs, user behavior records, and network traffic;

[0067] For Chinese text in the text data, use the jieba library for word segmentation, and for English text, use the nltk library for word segmentation;

[0068] It should be noted that word segmentation is the basis of natural language processing. It divides continuous text into meaningful lexical units, facilitating subsequent text processing and analysis;

[0069] The text data after word segmentation is converted to lowercase, stop words are removed, and the stems are extracted to obtain the processed text data;

[0070] It should be noted that the text after word segmentation is uniformly converted to lowercase, common stop words are removed, and the stemming technique is used to extract the roots;

[0071] The clean text data is converted into vectors using the Word2Vec model;

[0072] It should be noted that the text data after word segmentation is converted into a high-dimensional vector representation using the Word2Vec model. Word2Vec is a commonly used word embedding technique that can map words into a high-dimensional vector space, where words with similar semantics are close in the vector space;

[0073] For all images in the image data, they are scaled to a unified size, the piecewise linear transformation method is used to retain the image edge information, and the images are randomly rotated, flipped, and cropped;

[0074] For the audio data, Mel Frequency Cepstral Coefficients are used to extract audio features, a frequency domain filter is used to denoise the audio data, and the audio signal is segmented into multiple segments;

[0075] It should be noted that the long audio signal is decomposed into smaller segments that are easier to process;

[0076] Invalid entries and duplicates in the communication record data are cleared.

[0077] It should be noted that through the multi-modal data processing method of the present invention, efficient, comprehensive, and high-quality data collection and preprocessing are achieved; a web crawler is used to automatically collect text, images, audio, and communication records from multiple data sources, improving the efficiency and coverage of data collection, and ensuring the diversity and comprehensiveness of the data; the text data is segmented, converted to lowercase, stop words are removed, and the stems are extracted, standardizing the text data, reducing noise, and improving the accuracy and efficiency of feature extraction; the clean text data is converted into vectors using the Word2Vec model, retaining the semantic information and enhancing the performance of text analysis; the image data is unified in size, edge information is retained, and random transformations are performed, increasing data diversity and improving the generalization ability and robustness of the model; MFCC features are extracted from the audio data and denoised and segmented, effectively capturing the acoustic characteristics and improving the quality of audio processing; invalid entries and duplicates in the communication records are cleared, improving the consistency and reliability of the data.

[0078] S2. Extract the features of the multi-modal data, fuse the multi-modal data features through the cross-attention mechanism, and generate a multi-modal comprehensive feature vector.

[0079] Furthermore, the preprocessed text data is processed using the BERT model to generate 768-dimensional text feature vectors;

[0080] It should be noted that BERT is a pre-trained language model based on the Transformer architecture, capable of capturing context information and semantic relationships in text;

[0081] Each 768-dimensional vector represents the features of a text segment (such as a sentence or a paragraph); these vectors are obtained by encoding the text, and each dimension represents the features of the text in a specific aspect;

[0082] The preprocessed image data is used with the ResNet-50 model to extract 2048-dimensional feature vectors of the image;

[0083] It should be noted that ResNet-50 is a deep convolutional neural network with multiple convolutional layers and residual connections, capable of capturing high-level abstract features in images;

[0084] Each 2048-dimensional vector represents the features of an image; these vectors are output from the last fully connected layer of the image, and each dimension represents the features of the image in a specific aspect.

[0085] The preprocessed audio data is processed using the GRU model to obtain 128-dimensional hidden state vectors;

[0086] It should be noted that GRU is a recurrent neural network, especially suitable for processing sequential data such as audio signals;

[0087] Each 128-dimensional vector represents the hidden state of the audio data. These vectors are output from the last time step of the GRU model, and each dimension represents the features of the audio in a specific aspect.

[0088] The features of the three modal data are fused through the cross-attention mechanism in multi-modal feature fusion to generate a multi-modal comprehensive feature vector M;

[0089] It should be noted that through the multi-modal data processing method, high-quality feature extraction and effective fusion are achieved; the BERT model is used to generate 768-dimensional text feature vectors, retaining rich semantic information and improving the quality and expressiveness of text features; the ResNet-50 model is used to extract 2048-dimensional image feature vectors, comprehensively capturing the key features of images and enhancing the richness and accuracy of image features; the GRU model is used to process audio data, generating 128-dimensional hidden state vectors, effectively capturing the temporal dynamic features of audio and enhancing the representativeness of audio features; through the cross-attention mechanism, the data features of the three modalities are fused to generate a multi-modal comprehensive feature vector M, which not only combines the advantages of each modality but also captures the correlation and complementarity between them; finally, it jointly ensures the effective processing of multi-modal data and the generation of high-quality features, significantly improving the performance and accuracy of the system and providing a solid foundation for subsequent data encryption and secure transmission.

[0090] S2.1. Fusion of the data features of the three modalities through the cross-attention mechanism in multi-modal feature fusion to generate a multi-modal comprehensive feature vector M, the expression of which is:

[0091] M = Attn(T, I, A);

[0092] where T represents the text feature vector, I represents the image feature vector, and A represents the audio feature vector.

[0093] S3. Construct a graph convolutional network GCN relationship graph, generate node embedding vectors, and combine them with the multi-modal comprehensive feature vector to obtain enhanced nodes.

[0094] Furthermore, construct a graph convolutional network GCN relationship graph based on the preprocessed communication records;

[0095] Use the NetworkX library to create a graph, add nodes and edges, and define node v and edge attributes;

[0096] Initialize an embedding vector for each node in the graph convolutional network GCN relationship graph

[0097] Calculate the cosine similarity between each node in the graph convolutional network GCN relationship graph and all its neighbor nodes u, the expression of which is:

[0098]

[0099] where is the cosine similarity, is the embedding vector of node v at the l-th layer, is the embedding vector of neighbor node u at the l-th layer;

[0100] Using cosine similarity as the weight, the neighbor node embedding vectors are weighted to obtain a weighted sum vector, expressed as:

[0101]

[0102] where N(v) is the set of neighbors of node v;

[0103] The weighted sum vector is non-linearly transformed through the activation function σ to obtain a new node embedding vector The expression is:

[0104]

[0105] The new node embedding vector in the generated graph convolutional network GCN relationship graph and the multi-modal comprehensive feature vector are combined to generate an enhanced node The expression is:

[0106]

[0107] It should be noted that by constructing the graph convolutional network GCN relationship graph, the entities and their complex associations in the communication records are structurally represented, enhancing the interpretability and processing ability of the data; providing a basis for graph convolutional operations for initializing the embedding vectors of each node, enabling the system to gradually optimize the node features; using cosine similarity to quantify the relationship strength between nodes and neighbors, effectively improving the quality and accuracy of node embeddings; finally, by fusing the multi-modal comprehensive feature vector to generate enhanced node embeddings, not only the graph structure information is retained, but also rich features from different data sources are integrated, significantly improving the expressiveness and discrimination of node representations.

[0108] S3.1. The node attributes include IP address, geographical location, and device type;

[0109] The edge attributes include the number of communications, communication time, and communication protocol type.

[0110] IP address of node attribute: One attribute of a node is its corresponding IP address;

[0111] Geographical location of node attribute: Another attribute of a node is its geographical location information;

[0112] Device type of node attribute: The third attribute of a node is the device type (such as server, mobile device, etc.);

[0113] Number of communications of edge attribute: One attribute of an edge is the number of communications between two nodes;

[0114] Communication time of edge attribute: Another attribute of an edge is the timestamp of the most recent communication;

[0115] Communication protocol type of edge attribute: The third attribute of the edge is the communication protocol type used (such as TCP, UDP, etc.).

[0116] S4. The enhanced node generates a complex key seed, derives the final key through HKDF, and encrypts the data through an encryption algorithm.

[0117] Furthermore, introduce the autoregressive AR model in the time series to generate a dynamic key seed, and the expression is:

[0118]

[0119] where S t is the key seed at the current time point, a i is the autoregressive coefficient, S t-i is the past key seed, p is the autoregressive order, i is the lag order index in the time series, and ∈ t is the error term;

[0120] Determine the autoregressive order and autoregressive coefficient through training with historical data;

[0121] Initialize the past key seed;

[0122] Convert the enhanced node into a fixed-length vector and input it into the autoregressive AR model to update the key seed at the current time point to obtain a complex dynamic key seed S' t , and the expression is:

[0123]

[0124] where is the result after being processed by the hash function;

[0125] Derive the final key from the complex dynamic key seed through the HKDF key derivation function, and the expression is:

[0126] K = HKDF(S' t , L);

[0127] where K is the final key and L is the output key length;

[0128] Select AES-256 as the encryption algorithm, set the encryption mode to CBC, and generate and use the initialization vector IV;

[0129] Dynamically adjust the encryption algorithm parameters according to the enhanced node;

[0130] Select the corresponding key and initialization vector for data encryption according to the confidentiality level of the data, and use the AES-256-CBC mode during the encryption process to ensure data security;

[0131] It should be noted that by introducing an autoregressive AR model to generate a dynamic key seed, the time dependence and dynamics of the key are increased, improving the security and anti-prediction ability of the key; adding enhanced nodes to the key seed generation process further increases the complexity and diversity of the key, enhancing the security of the key; using HKDF to derive the final key from the complex dynamic key seed ensures the high quality and randomness of the key, providing additional security; selecting AES-256 as the encryption algorithm and setting the encryption mode to CBC, generating and using the initialization vector IV, ensures the high security and integrity of the data during the encryption process, preventing data tampering or leakage.

[0132] S4.1. Dynamically adjust the key length according to certain characteristics (such as norms, values of specific dimensions, etc.) of ;

[0133] If has a large norm, a longer key can be used to improve security;

[0134] Dynamically select or generate different initialization vectors according to certain characteristics of ; for example, can generate the IV through a hash function;

[0135] Select different encryption modes (such as CBC, CTR, etc.) according to certain characteristics of ; for example, if certain dimension values of are relatively high, select a more secure encryption mode.

[0136] S4.2. It should be noted that an autoregressive AR model is used to generate a complex dynamic key seed with enhanced nodes, and the final key K is derived in combination with the HKDF algorithm to ensure the high security and randomness of the key; subsequently, the AES-256 encryption algorithm is used to encrypt the data in CBC mode with the initialization vector IV to achieve data security protection; this process not only enhances the unpredictability of the key generation mechanism but also adapts to different confidentiality requirements by dynamically adjusting the encryption parameters, thus providing a powerful and flexible data encryption solution.

[0137] S5. After completing data encryption, protect the data through secure transmission and distributed storage, and use a reinforcement learning model to select the optimal encryption strategy.

[0138] Furthermore, the encrypted data is securely transmitted through the TLS / SSL protocol, and digital signatures and certificates are used to ensure data integrity and trusted origin;

[0139] The encrypted data is dispersed and stored in cloud servers using HDFS distributed file storage, and redundant coding is applied to improve data reliability;

[0140] The reinforcement learning agent is trained through a reinforcement learning model based on the PPO algorithm, enabling the reinforcement learning agent to select the optimal encryption strategy according to the current environmental state;

[0141] It should be noted that among them, the reinforcement learning agent dynamically adjusts the encryption algorithm and parameters to adapt to environmental changes and improve overall security and performance.

[0142] It should be noted that the combination of the TLS / SSL protocol, digital signatures, and certificates ensures the confidentiality, integrity, and trusted origin of data during transmission, significantly improving the security of data transmission; the use of HDFS distributed file storage and redundant coding technology disperses and stores the encrypted data in cloud servers, enhancing data storage reliability and fault tolerance, and ensuring high data availability and recovery capabilities; through the training of an agent using a reinforcement learning model based on the PPO algorithm, the agent can dynamically select the optimal encryption strategy according to the current environmental state, improving adaptability and robustness; the reinforcement learning agent can adjust the encryption algorithm and parameters in real time to adapt to changing security threats and performance requirements, not only improving the security of the system but also optimizing overall performance.

[0143] S5.1. The reinforcement learning agent selects the optimal encryption strategy according to the current environmental state. The specific steps are as follows:

[0144] Define the state space, action space, and reward function of the encryption working environment;

[0145] Build a reinforcement learning model based on the PPO algorithm;

[0146] The reinforcement learning model agent selects actions according to the current encryption working environment state through a neural network;

[0147] Apply the selected action to the encryption working environment for data encryption;

[0148] Give rewards according to the security score and performance overhead in the defined reward function;

[0149] If complex encryption successfully resists threat attacks, give a positive reward;

[0150] If simple encryption leads to performance degradation, give a negative reward;

[0151] Set a threshold. When the reinforcement learning model agent reaches and exceeds the threshold, the training stops.

[0152] S5.2. It should be noted that HDFS is mainly used for big data processing, especially in the fields of batch data processing and data analysis; it is usually used in combination with distributed computing frameworks such as MapReduce and Apache Spark to handle tasks such as log analysis, data warehousing, machine learning, and large-scale data processing.

[0153] S5.3. Train the reinforcement learning agent using the PPO algorithm. The specific steps are as follows:

[0154] Define the state space, action space, and reward function;

[0155] State space: The state of the environment includes network traffic characteristics, user behavior characteristics, system resource usage, etc. This state information can help the agent understand the current security threat level and performance requirements;

[0156] Action space: The action space includes the selection of different encryption algorithms (such as AES-256), encryption modes (such as CBC), key lengths, and initialization vectors. The agent needs to select an action from this action space to execute;

[0157] Reward function: Design a reward function to evaluate the agent's actions; the reward function usually considers the security score (for example, the ability to resist attacks) and performance overhead (for example, the speed of encryption / decryption). The goal is to find a balance that ensures data security without affecting the system's performance;

[0158] Build a reinforcement learning model based on the PPO algorithm;

[0159] Select the agent as the model architecture. The agent is a neural network that takes the environment state as input and outputs an action probability distribution; for a discrete action space, we can use a network with a classification output;

[0160] Implement policy adjustment through the PPO algorithm. PPO is a policy gradient method that maximizes the cumulative reward by optimizing the policy; PPO improves the training stability and efficiency by introducing a clipped objective function to limit the step size of each update; the PPO algorithm calculates the Advantage Function at each time step and adjusts the policy based on this;

[0161] The agent interacts with the environment to generate a series of state-action-reward-new state sequences;

[0162] Calculate the advantage function for each action based on the reward and value function estimation;

[0163] Use the objective function of PPO to update the parameters of the policy network; PPO gradually improves the policy through multiple mini-batch updates;

[0164] Repeat the above process until the agent learns to select appropriate encryption policies in different environments;

[0165] After training is completed, the agent can select the optimal encryption policy in real time according to the current environmental state during actual operation;

[0166] The agent selects the most suitable encryption algorithm, key length, and initialization vector according to the current environmental state (such as network traffic characteristics, user behavior characteristics, etc.). For example, when detecting a high-risk attack, the agent may select a more complex encryption algorithm and a longer key length; while in the case of low risk and high performance requirements, the agent may select a simpler encryption scheme;

[0167] The agent can also continue to learn from new interactions, continuously adjust and optimize its policy to adapt to changing environmental conditions;

[0168] Deploy the trained agent to the production environment to ensure that it can access the required state information and can perform corresponding encryption operations;

[0169] Continuously monitor the performance of the agent, including security and performance metrics. If the agent's performance is found to be poor, it may be necessary to retrain or adjust the reward function.

[0170] This embodiment also provides a deep learning-based AI data processing and encryption device, including: a data collection module, a data extraction module, a construction module, a key generation module, and a transmission optimization module; the data collection module is used to collect multimodal data and preprocess the multimodal data; the data extraction module is used to extract the features of the multimodal data, fuse the multimodal data features through a cross-attention mechanism, and generate a multimodal comprehensive feature vector; the construction module is used to construct a graph convolutional network GCN relationship graph, generate node embedding vectors, and combine the multimodal comprehensive feature vectors to obtain enhanced nodes; the key generation module is used to generate complex key seeds from the enhanced nodes, derive the final key through HKDF, and encrypt the data through an encryption algorithm; the transmission optimization module is used to complete data encryption, protect the data through secure transmission and distributed storage, and use a reinforcement learning model to select the optimal encryption policy.

[0171] This embodiment also provides a computer device applicable to the case of the deep learning-based AI data processing and encryption method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the deep learning-based AI data processing and encryption method as proposed in the above embodiment.

[0172] The computer device may be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, trackball, or touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0173] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the AI data processing and encryption method based on deep learning proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0174] In summary, through the multi-modal graph embedding technology, the present invention realizes the diversified collection and efficient preprocessing of data, improves the quality and consistency of data; constructs a graph convolutional network GCN relationship graph to update the node embedding vector, combines the multi-modal comprehensive feature vector to generate enhanced nodes, effectively captures the complex relationships between data, and improves the expressiveness of the node embedding vector; then, generates a complex dynamic key seed through an autoregressive AR model, derives the final key through HKDF, and selects the AES-256 encryption algorithm to encrypt the data, significantly improving the randomness and unpredictability of the key, and enhancing the security of data encryption.

[0175] Example 2. Referring to Table 1, this is the second example of the present invention. To further verify the technical solution of the present invention, experimental simulation data of the AI data processing and encryption method based on deep learning is given.

[0176] First, use web crawlers to collect multimodal data such as text, images, audio, and communication records from server logs, user behavior records, and network traffic, and perform preprocessing; including: using the jieba library to segment Chinese texts, and the nltk library to segment English texts; converting texts to lowercase, removing stop words and extracting stems; using the Word2Vec model to convert text data into vectors; uniformly scaling image data to 224x224 pixels, using the fractional linear transformation method to retain edge information, and performing random rotation, flipping, and cropping; using Mel Frequency Cepstral Coefficients (MFCC) to extract features from audio data, and denoising through frequency domain filters.

[0177] Secondly, use the BERT model to process text data to generate 768-dimensional text feature vectors, the ResNet-50 model to extract 2048-dimensional image feature vectors, and the GRU model to process audio data to obtain 128-dimensional hidden state vectors; fuse the data features of the three modalities through the cross-attention mechanism to generate multimodal comprehensive feature vectors.

[0178] Next, create a graph structure using the NetworkX library according to the preprocessed communication records, define node and edge attributes, and initialize the embedding vectors of each node; calculate the cosine similarity between a node and its neighbor nodes, and perform weighted summation based on this, and obtain new node embedding vectors through non-linear transformation of the activation function, and then combine them with the multimodal comprehensive feature vectors to generate enhanced nodes.

[0179] Finally, generate a complex dynamic key seed through the autoregressive (AR) model, and derive the final key in combination with the HKDF algorithm; select AES-256 as the encryption algorithm, and encrypt the data in CBC mode in cooperation with the initialization vector (IV); securely transmit the encrypted data through the TLS / SSL protocol, and use HDFS to disperse and store the encrypted data in the cloud, and apply redundant coding to improve data reliability; use a reinforcement learning model based on the Proximal Policy Optimization (PPO) algorithm to train a reinforcement learning agent so that it can dynamically adjust the encryption algorithm and parameters according to the current environmental state, improving the overall security and performance.

[0180] Specifically, as shown in Table 1 below:

[0181] Table 1 Data Processing and Encryption Analysis Table

[0182]

[0183]

[0184] Demonstrated significant improvements in multi-modal data processing and encryption technologies; in the generation of comprehensive feature vectors, the present invention uses a cross-attention mechanism to fuse multi-modal features. Compared with the simple splicing method in the prior art, the efficiency has increased from 7.5 to 8.9, an increase of 18.7%; in terms of updating node embedding vectors, the present invention combines multi-modal comprehensive feature vectors through GCN, replacing the traditional static node embedding vectors. The efficiency has increased from 6.8 to 8.4, an increase of 23.5%; in terms of key generation, the present invention uses an autoregressive AR model and the HKDF algorithm, replacing the fixed key length and static encryption parameters. The efficiency has increased from 7.2 to 9.1, an increase of 26.4%; in terms of image feature extraction, the present invention uses a ResNet-50 model. Compared with the traditional convolutional neural network CNN, the accuracy has increased from 85.2% to 90.7%, an increase of 6.2%.

[0185] These technologies not only improve the processing efficiency and accuracy in various aspects, but also enhance the overall performance and security of the system, providing a more efficient and secure solution for data processing and encryption.

[0186] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An AI data processing encryption method based on deep learning, characterized in that: include, Collect multimodal data and preprocess the multimodal data; Extract the features of multimodal data, fuse the features of multimodal data through the cross-attention mechanism, and generate a multimodal comprehensive feature vector; Construct a graph convolutional network (GCN) relationship graph, generate node embedding vectors, and combine multimodal comprehensive feature vectors to obtain enhanced nodes; The enhanced node generates a complex key seed, the final key is derived through HKDF, and the data is encrypted through the encryption algorithm; After completing data encryption, the data is protected through secure transmission and distributed storage, and the optimal encryption strategy is selected using a reinforcement learning model; Construct a graph convolutional network GCN relationship graph, generate node embedding vectors, and combine multimodal comprehensive feature vectors to obtain enhanced nodes. The specific steps are as follows: Construct a graph convolutional network (GCN) relationship graph based on the preprocessed communication records; Use the NetworkX library to create a graph, add nodes and edges, and define node v and edge attributes; Through the nonlinear transformation of the activation function σ, the node embedding vector is obtained The expression is: in, is the embedding vector of node v at layer l, is the embedding vector of neighbor node u at layer l, N(v) is the neighbor set of node v, is the cosine similarity; Combine the node embedding vector and the multimodal comprehensive feature vector in the generated graph convolutional network GCN relationship graph to generate enhanced nodes; The enhanced node generates a complex key seed, the final key is derived through HKDF, and the data is encrypted through the encryption algorithm. The specific steps are as follows: The autoregressive AR model introduced into the time series generates a dynamic key seed, which is expressed as: Among them, S t is the key seed at the current time point, a i is the autoregressive coefficient, S t-i is the past key seed, p is the autoregressive order, i is the lag order index in the time series, ∈ t is the error term; Determine the autoregressive order and autoregressive coefficient through historical data training; Initialize the past key seed; Convert the enhanced node into a vector of fixed length and input it into the autoregressive AR model to update the key seed at the current time point to obtain the complex dynamic key seed S' t , the expression is: in, yes The result after hash function processing is It is an enhancement node; The final key is derived from the complex dynamic key seed through the HKDF key derivation function; Select AES-256 as the encryption algorithm, set the encryption mode to CBC, generate and use the initialization vector IV; Dynamically adjust encryption algorithm parameters based on enhanced nodes; According to the confidentiality level of the data, select the corresponding key and initialization vector for data encryption, and use the AES-256-CBC mode during the encryption process.

2. The AI ​​data processing and encryption method based on deep learning as claimed in claim 1, characterized in that: The multimodal data is collected and preprocessed, and the specific steps are: Use web crawlers to collect text data, image data, audio data, and communication records from server logs, user behavior records, and network traffic; Perform word segmentation on text data; Convert the text after word segmentation to lowercase, remove stop words, and extract stems to obtain clean text data; Convert clean text data into vectors; All images in the image data are scaled to a uniform size, the image edge information is retained using the linear transformation method, and the images are randomly rotated, flipped, and cropped; Use Mel-frequency cepstral coefficients to extract audio features from audio data, use frequency domain filters to reduce noise on audio data, and segment audio signals; Cleans up invalid entries and duplicates in communication log data.

3. The AI ​​data processing and encryption method based on deep learning as claimed in claim 2, characterized in that: The method of extracting the features of multimodal data and fusing the features of multimodal data through a cross-attention mechanism to generate a multimodal comprehensive feature vector comprises the following specific steps: Use the BERT model to process the preprocessed text data and generate text feature vectors; Use the ResNet-50 model to extract image feature vectors from the preprocessed image data; The preprocessed audio data is processed using the GRU model to obtain a hidden state vector; The three modal data features are fused through the cross-attention mechanism in multimodal feature fusion to generate a multimodal comprehensive feature vector.

4. The AI ​​data processing and encryption method based on deep learning as claimed in claim 3, characterized in that: After completing the data encryption, the data is protected through secure transmission and distributed storage, and the optimal encryption strategy is selected using the reinforcement learning model. The specific steps are as follows: Securely transmit encrypted data via TLS / SSL protocol; Use HDFS distributed file storage to store encrypted data in cloud servers and apply redundant coding to improve data reliability; The reinforcement learning agent is trained through a reinforcement learning model based on the PPO algorithm, so that the reinforcement learning agent can select the optimal encryption strategy according to the current environment state.

5. The AI ​​data processing and encryption method based on deep learning as claimed in claim 4, characterized in that: The reinforcement learning agent selects the optimal encryption strategy according to the current environment state. The specific steps are: Define the state space, action space, and reward function of the encrypted working environment; Build a reinforcement learning model based on the PPO algorithm; The reinforcement learning model agent selects actions based on the current state of the encrypted working environment through a neural network; Apply the selected action to the encryption working environment to encrypt data; Give rewards based on the safety score and performance cost in the defined reward function; If complex encryption successfully defends against threat attacks, positive rewards will be given; If simple encryption leads to performance degradation, negative rewards will be given; Set a threshold, and stop training when the reinforcement learning model agent reaches and exceeds the threshold.

6. An AI data processing and encryption device based on deep learning, based on the AI ​​data processing and encryption method based on deep learning according to any one of claims 1 to 5, characterized in that: Including, data collection module, data extraction module, construction module, key generation module and transmission optimization module; The data acquisition module is used to collect multimodal data and preprocess the multimodal data; The data extraction module is used to extract the features of multimodal data, fuse the features of multimodal data through a cross-attention mechanism, and generate a multimodal comprehensive feature vector; The construction module is used to construct a graph convolutional network (GCN) relationship graph, generate node embedding vectors, and obtain enhanced nodes in combination with multimodal comprehensive feature vectors; The key generation module is used to generate a complex key seed from the enhanced node, derive the final key through HKDF, and encrypt the data through the encryption algorithm; The transmission optimization module is used to protect data through secure transmission and distributed storage after completing data encryption, and to select the optimal encryption strategy using a reinforcement learning model.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the deep learning-based AI data processing encryption method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the deep learning-based AI data processing encryption method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Financial behavior pattern analysis and prediction method based on time sequence diagram representation learning

    CN116541755A

  • Software development management method based on artificial intelligence

    CN118295640A