Data classification method and system based on deep learning

Through deep learning-based data classification methods and distributed storage of blockchain networks, the problems of insufficient accuracy, insufficient processing capabilities and lack of customization in the existing technology are solved, and efficient, accurate and customized data classification is achieved, which improves user experience and data security.

CN120011878AInactive Publication Date: 2025-05-16BEIJING JUNDE INTELLIGENT COMPUTING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510064891.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing data classification technology has problems such as insufficient classification accuracy, insufficient data processing capabilities and lack of customization, which cannot effectively meet the needs of users.

Method used

Using a data classification method based on deep learning, we use user portrait generation model and data classification model to build distributed storage with blockchain network to achieve efficient and customized data classification.

Benefits of technology

It improves the accuracy and efficiency of data classification, meets users' needs for high accuracy, and improves user experience through customized classification, ensuring data security and traceability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011878A_ABST
    Figure CN120011878A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data classification, and discloses a data classification method and system based on deep learning. The method comprises the following steps: constructing a user portrait generation model and a data classification model by using a deep learning algorithm according to a plurality of pieces of historical user basic information and a plurality of pieces of historical to-be-classified data; generating a user portrait by using a user portrait generation model according to the real-time user basic information to obtain a real-time user portrait; performing data classification by using a data classification model according to the real-time user portrait and the real-time to-be-classified data to obtain a real-time data classification result; and performing distributed storage on the real-time user portrait, the real-time to-be-classified data and the real-time data classification result by using a block chain network. According to the method, the problems of insufficient classification accuracy, insufficient data processing capability and lack of customization in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data classification, and specifically relates to a data classification method and system based on deep learning. Background Art

[0002] With the development of information technology, work and life have entered the digital information age. The generation of a large amount of data has put forward an urgent need for data management. When managing data, it is necessary to classify different data, which requires a lot of cost. Therefore, how to classify data efficiently and quickly has become the main limitation for the popularization and development of information technology.

[0003] Existing data classification technology has the following defects: 1) Insufficient classification accuracy: Most existing data classifications rely on simple programs or tools based on file suffixes or additional information, which cannot mine the deep features of the data and are not suitable for multimodal data, resulting in low classification accuracy and failure to meet user needs; 2) Insufficient data processing capabilities: The existing data classification technology has a low level of intelligence, resulting in low data processing efficiency and is not suitable for large-scale data processing; 3) Lack of customization: Most existing data classification technologies use preset classification rules, ignoring the fact that different users have large differences in classification standards, resulting in low user experience. Summary of the invention

[0004] In order to solve the problems of insufficient classification accuracy, insufficient data processing capabilities and lack of customization in the prior art, the present invention aims to provide a data classification method and system based on deep learning.

[0005] The technical solution adopted by the present invention is: A data classification method based on deep learning includes the following steps: Based on some historical user basic information and some historical data to be classified, a deep learning algorithm is used to build a user portrait generation model and a data classification model; According to the real-time user basic information, the user portrait generation model is used to generate the user portrait to obtain the real-time user portrait; According to the real-time user portrait and the real-time data to be classified, the data classification model is used to classify the data and obtain the real-time data classification results; Use the blockchain network to distribute the real-time user portraits, real-time data to be classified, and real-time data classification results.

[0006] Further, the user portrait generation model is constructed based on the RF-MLP algorithm, and the user portrait generation model includes a key feature screening module constructed based on the RF algorithm and a user portrait generation module constructed based on the MLP algorithm, which are connected in sequence; The data classification model includes a data classification model constructed based on the GCN-logfBank-CNN-LSTM-Attention-DBN algorithm, and the data classification model includes a graph structure feature extraction module constructed based on the GCN algorithm, an audio feature extraction module constructed based on the logfBank algorithm, an image feature extraction module constructed based on the CNN algorithm, a sequence feature extraction module constructed based on the LSTM algorithm, an attention weight module constructed based on the Attention mechanism, and a data classification module constructed based on the DBN algorithm. The graph structure feature extraction module, the audio feature extraction module, the image feature extraction module, and the sequence feature extraction module are all connected to the attention weight module, and the attention weight module is connected to the data classification module.

[0007] Furthermore, based on some historical user basic information and some historical data to be classified, a deep learning algorithm is used to construct a user portrait generation model and a data classification model, including the following steps: Collecting some historical user basic information and some historical data to be classified, and preprocessing them to obtain some preprocessed historical user basic information and some preprocessed historical data to be classified; Performing data analysis on the preprocessed historical data to be classified to obtain corresponding preprocessed historical audio data, preprocessed historical image data, and preprocessed historical sequence data; Based on some pre-processed historical user basic information, a deep learning algorithm is used to build a user portrait generation model and generate several historical user portraits; A data classification model is constructed using a deep learning algorithm based on several historical user portraits, several preprocessed historical audio data to be classified, preprocessed historical image data, and preprocessed historical sequence data.

[0008] Furthermore, based on some pre-processed historical user basic information, a deep learning algorithm is used to construct a user portrait generation model, and generate some historical user portraits, including the following steps: Use the RF-MLP algorithm to build an initial user portrait generation model; the initial user portrait generation model includes an initial key feature screening module and an initial user portrait generation module; According to a number of pre-processed historical user basic information, an initial key feature screening module is trained to obtain a number of key feature indicators, a number of historical key features of each pre-processed historical user basic information, and a final key feature screening module; According to several historical key features of all pre-processed historical user basic information, the initial user portrait generation module is trained to obtain several historical user portraits and the final user portrait generation module; The final key feature screening module and the final user portrait generation module are integrated to obtain the final user portrait generation model.

[0009] Furthermore, based on the plurality of historical user portraits, the preprocessed historical audio data of the plurality of preprocessed historical data to be classified, the preprocessed historical image data, and the preprocessed historical sequence data, a data classification model is constructed using a deep learning algorithm, including the following steps: Use the GCN-logfBank-CNN-LSTM-Attention-DBN algorithm to build an initial data classification model; Taking minimizing mean square error as the optimization goal, a swarm intelligence optimization algorithm is used to optimize the initial model parameters of the initial data classification model to obtain an optimized data classification model; the optimized data classification model includes an optimized graph structure feature extraction module, an optimized audio feature extraction module, an optimized image feature extraction module, an optimized sequence feature extraction module, an optimized attention weight module and an optimized data classification module; Pre-training the optimized data classification module according to a number of pre-processed historical data to be classified, thereby obtaining a pre-trained data classification module; According to several historical user portraits, the optimized graph structure feature extraction module is trained to obtain several historical graph structure features and the final graph structure feature extraction module; According to a number of pre-processed historical audio data, the optimized audio feature extraction module is trained to obtain a number of historical audio features and a final audio feature extraction module; According to a number of pre-processed historical image data, the optimized image feature extraction module is trained to obtain a number of historical image features and a final image feature extraction module; According to a number of pre-processed historical sequence data, the optimized sequence feature extraction module is trained to obtain a number of historical sequence features and a final sequence feature extraction module; According to several historical graph structure features, several historical audio features, several historical image features and several historical sequence features, the optimized attention weight module is trained to obtain several historical fusion features and the final attention weight module; According to several historical fusion features, the optimized data classification module is trained to obtain the final data classification module; The final graph structure feature extraction module, the final image feature extraction module, the final sequence feature extraction module, the final attention weight module and the final data classification module are integrated to obtain the final data classification model.

[0010] Further, based on the real-time user basic information, a user portrait generation model is used to generate a user portrait to obtain a real-time user portrait, including the following steps: Collecting real-time user basic information, preprocessing the real-time user basic information, and obtaining preprocessed real-time user basic information; According to several key feature indicators, a key feature screening module is used to screen key features to obtain several real-time key features of real-time user basic information after preprocessing; According to several real-time key features, a user portrait generation module is used to generate a user portrait to obtain a real-time user portrait.

[0011] Furthermore, according to the real-time user portrait and the real-time data to be classified, the data classification model is used to classify the data to obtain the real-time data classification result, including the following steps: Collecting real-time data to be classified, preprocessing the real-time data to be classified, and obtaining real-time data to be classified after preprocessing; Performing data analysis on the preprocessed real-time data to be classified to obtain corresponding preprocessed real-time audio data, preprocessed real-time image data, and preprocessed real-time sequence data; According to the real-time user portrait, the pre-processed real-time audio data, the pre-processed real-time image data and the pre-processed real-time sequence data, a data classification model is used to perform data classification to obtain real-time data classification results.

[0012] Further, according to the real-time user portrait, the pre-processed real-time audio data, the pre-processed real-time image data and the pre-processed real-time sequence data, a data classification model is used to perform data classification to obtain a real-time data classification result, including the following steps: Inputting the real-time user portrait, the pre-processed real-time audio data, the pre-processed real-time image data, and the pre-processed real-time sequence data into the data classification model; According to the real-time user portrait, use the graph structure feature extraction module to extract graph structure features and obtain real-time graph structure features; According to the pre-processed real-time audio data, an audio feature extraction module is used to extract audio features to obtain real-time audio features; According to the pre-processed real-time image data, an image feature extraction module is used to extract image features to obtain real-time image features; According to the pre-processed real-time sequence data, a sequence feature extraction module is used to extract sequence features to obtain real-time sequence features; According to the preset attention weight, the attention weight module is used to fuse the real-time graph structure features, real-time audio features, real-time image features and real-time sequence features to obtain real-time fusion features; According to the real-time fusion features, the data classification module is used to classify the data and obtain the real-time data classification results.

[0013] Furthermore, the blockchain network is used to perform distributed storage of real-time user portraits, real-time data to be classified, and real-time data classification results, including the following steps: The real-time user portrait, real-time data to be classified and real-time data classification results of the same user are associated to obtain real-time associated data; The real-time associated data is stored in the IPFS system of the blockchain network to obtain the real-time data hash value, and the smart contract is called to generate the corresponding real-time transaction data according to the real-time data hash value; Call smart contracts to convert real-time transaction data into real-time blocks, and use several distributed nodes of the blockchain network to store real-time blocks on the chain and generate corresponding real-time transaction records; Update the distributed ledger of the blockchain network based on the real-time storage address, real-time retrieval tag and real-time transaction records of the real-time associated data in the IPFS system.

[0014] A data classification system based on deep learning, used to implement a data classification method, characterized in that the system includes a model building unit, a user portrait generation unit, a data classification unit and a distributed storage unit connected in sequence.

[0015] The beneficial effects of the present invention are: The present invention provides a data classification method and system based on deep learning, which realizes efficient data classification by constructing a data classification model, and the deep learning algorithm can mine the deep features of the data as the basis for data classification, thereby improving the classification accuracy and meeting the user's requirements for data classification accuracy; an automated data classification function is provided, which can extract, process and classify features of multimodal data, thereby improving the degree of intelligence and the efficiency of data processing, especially in the face of massive data processing environments, thereby improving practicality; in the data classification process, considering that different users have large differences in classification standards, a user portrait model is used to generate user portraits based on user basic information, and customized classification is performed in combination with user preferences and habits, thereby improving user experience; a blockchain network is used to perform distributed storage of data, thereby ensuring the security of the data and the traceability of data classification.

[0016] Other beneficial effects of the present invention will be further described in the specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart of the data classification method based on deep learning in the present invention.

[0018] Figure 2 It is a structural block diagram of the data classification system based on deep learning in the present invention. DETAILED DESCRIPTION

[0019] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments.

[0020] Embodiment 1: like Figure 1 As shown, this embodiment provides a data classification method based on deep learning, including the following steps: S1: Based on some historical user basic information and some historical data to be classified, a deep learning algorithm is used to build a user portrait generation model and a data classification model, including the following steps: S1-1: Collect some historical user basic information and some historical data to be classified, and pre-process them to obtain some pre-processed historical user basic information and some pre-processed historical data to be classified; User basic information includes user personal information, behavior data, preference settings, etc. The data to be classified is multimodal data, including sequence data such as images, audio, and text; The preprocessing of training data includes cleaning the data to remove noise and irrelevant information and format standardization to maintain consistency, which improves the quality of the data and provides data support for subsequent model construction; The preprocessing of audio data also includes denoising and short-time Fourier transform (STFT) processing, and the preprocessed audio data obtained is the corresponding spectrogram data; the preprocessing of image data also includes image enhancement and size cropping to enhance the texture, boundary and other feature representations in the image; the preprocessing of sequence data also includes magnitude normalization of data and processing of missing values ​​and outliers, which improves the credibility of data and the accuracy of model prediction; S1-2: Perform data analysis on the preprocessed historical data to be classified to obtain corresponding preprocessed historical audio data, preprocessed historical image data, and preprocessed historical sequence data; S1-3: Based on some pre-processed historical user basic information, use deep learning algorithm to build a user portrait generation model and generate some historical user portraits; The user portrait generation model is built based on the Random Forest (RF)-Attention-Multilayer Perceptron (MLP) algorithm, and the user portrait generation model includes a key feature screening module built based on the RF algorithm and a user portrait generation module built based on the MLP algorithm, which are connected in sequence; The key feature screening module screens the feature components of the input user basic information through the internal Classification And Regression Tree (CART), which can process a large number of feature components, generate the key feature importance score of each feature component, and select the most stable and discriminative key feature components according to the key feature importance score. The trained key feature screening module can directly screen the newly input user basic information according to the selected key features to obtain the corresponding key features; the MLP network, as a fully connected network, can perform accurate and efficient label prediction based on the key features; Based on some pre-processed historical user basic information, a user portrait generation model is constructed using a deep learning algorithm, and several historical user portraits are generated, including the following steps: S1-3-1: Use the RF-MLP algorithm to build an initial user portrait generation model; the initial user portrait generation model includes an initial key feature screening module and an initial user portrait generation module; S1-3-2: Based on a number of pre-processed historical user basic information, the initial key feature screening module is trained to obtain a number of key feature indicators, a number of historical key features of each pre-processed historical user basic information, and a final key feature screening module; S1-3-3: Based on several historical key features of all pre-processed historical user basic information, the initial user portrait generation module is trained to obtain several historical user portraits and the final user portrait generation module; Key features include the user’s age, gender, occupation, behavior, preferences, and other characteristics; S1-3-4: Integrate the final key feature screening module and the final user portrait generation module to obtain the final user portrait generation model; S1-4: Based on several historical user portraits, several preprocessed historical audio data, preprocessed historical image data and preprocessed historical sequence data, a data classification model is constructed using a deep learning algorithm; The data classification model includes a data classification model constructed based on a graph convolutional network (GCN)-logfBank-convolutional neural networks (CNN)-long short-term memory network (LSTM)-attention-deep belief network (DBN) algorithm, and the data classification model includes a graph structure feature extraction module constructed based on the GCN algorithm, an audio feature extraction module constructed based on the logfBank algorithm, an image feature extraction module constructed based on the CNN algorithm, a sequence feature extraction module constructed based on the LSTM algorithm, an attention weight module constructed based on the Attention mechanism, and a data classification module constructed based on the DBN algorithm. The graph structure feature extraction module, the audio feature extraction module, the image feature extraction module, and the sequence feature extraction module are all connected to the attention weight module, and the attention weight module is connected to the data classification module; The graph structure feature extraction module propagates features on user portraits through operations similar to convolution, extracts node features of several nodes in the graph structure data, edge features of the positional relationship between nodes, and forms complex, high-dimensional global graph structure features; the image feature extraction module is used to extract image features from image data; the audio feature extraction module is used to extract spectral features from spectrograms; the sequence feature extraction module is used to extract sequence features from sequence data; the attention weight module performs weighted fusion of multimodal features to enhance the model's attention to important features, thereby improving the accuracy and efficiency of model prediction; the data classification module learns fused features through multiple hidden layers to achieve data classification label prediction; Based on several historical user portraits, several preprocessed historical audio data, preprocessed historical image data and preprocessed historical sequence data of preprocessed historical data to be classified, a data classification model is constructed using a deep learning algorithm, including the following steps: S1-4-1: Use the GCN-logfBank-CNN-LSTM-Attention-DBN algorithm to build an initial data classification model; S1-4-2: Taking minimizing the mean square error as the optimization goal, the swarm intelligence optimization algorithm is used to optimize the initial model parameters of the initial data classification model to obtain an optimized data classification model; the optimized data classification model includes an optimized graph structure feature extraction module, an optimized audio feature extraction module, an optimized image feature extraction module, an optimized sequence feature extraction module, an optimized attention weight module and an optimized data classification module; Taking minimizing the mean square error as the optimization goal, the swarm intelligence optimization algorithm is used to optimize the initial model parameters of the initial data classification model to obtain an optimized data classification model, including the following steps: S1-4-2-1: Taking minimizing the mean square error as the optimization goal, set the fitness function and IFWA population parameters of the swarm intelligence optimization algorithm, and set the individual encoding format of the swarm intelligence optimization algorithm according to the initial model parameters of the initial data classification model; The swarm intelligence optimization algorithm is the Improved Fireworks Optimization Algorithm (IFWA) algorithm, and the fitness function formula is:

[0021] In the formula, is the fitness function; is the mean square error function; S1-4-2-2: Initialize using the Circle chaos mapping sequence according to the IFWA population parameters and individual vectors to generate several initial IFWA individuals of the initial IFWA population; The formula is:

[0022] In the formula, is the initial IFWA individual of Circle chaos mapping; is the randomly generated initial IFWA individual; It is the IFWA individual indicator; S1-4-2-3: According to the fitness function, the explosion radius, number of sparks and fitness value of the initial IFWA individual are obtained; The formula is:

[0023] In the formula, For the initial IFWA individuals The number of sparks; is a numerical constant; is the maximum fitness value in the initialized IFWA population; For the initial IFWA individuals The fitness value of is an infinitesimal constant; is the convergence factor; is a positive real number that is not 0;

[0024] In the formula, For the initial IFWA individuals Explosion radius; Adjusted constant for explosion radius; is the minimum fitness value in the initialized IFWA population; is the total number of IFWA individuals;

[0025] In the formula, is the convergence factor; tanh(.) is the hyperbolic tangent function; is the iteration indicator; is the maximum number of iterations; a max , a min are the maximum and minimum values ​​of the convergence factor respectively; λ is the deceleration rate parameter, is the decreasing cycle parameter, λ =-2 π , = π ; The number of sparks determines the number of sub-fireworks produced after each firework explodes, and the explosion radius determines the distribution range of the sparks produced after the firework explodes in the solution space. In the early stage of iteration, a When the value of is large, the number of sparks of IFWA individuals is small and the explosion radius is large, which helps to reduce the computational burden and is more widely distributed, which helps to explore more solution spaces. In the later stages of iteration, a smaller explosion radius helps to perform a fine search in a local area, and a larger number of sparks helps to increase the diversity of the search. S1-4-2-4: Perform fireworks explosion according to the explosion radius, the number of sparks and the fitness value to obtain a number of updated IFWA individuals of the updated IFWA population; The formula is:

[0026] In the formula, For updated IFWA individuals; A random number between -1 and 1; S1-4-2-5: Use the Gaussian mutation algorithm to perform Gaussian mutation on the initialized IFWA population to generate several Gaussian mutated IFWA individuals of the Gaussian mutated IFWA population; The formula is:

[0027] In the formula, is the IFWA individual of Gaussian variation; is a Gaussian distributed random number with mean and variance both 1; S1-4-2-6: Use the dynamic reverse learning algorithm to perform dynamic reverse learning on the initialized IFWA population to obtain several reverse IFWA individuals of the reverse IFWA population; The formula is:

[0028] In the formula, For reverse IFWA individuals; γ is the decreasing inertia coefficient; L max , L min are the maximum and minimum values ​​of the vector space respectively; S1-4-2-7: Obtain the fitness value of each updated IFWA individual, Gaussian mutated IFWA individual and reverse IFWA individual, and take the IFWA individual with the minimum fitness value as the optimal individual; S1-4-2-8: If the number of iterations reaches the maximum number of iterations or the fitness value of the optimal individual meets the requirements, the individual encoding vector of the optimal individual is decoded to obtain the optimal initial model parameters of the initial data classification model; S1-4-2-9: Optimize the initial data classification model according to the optimal initial model parameters to obtain an optimized data classification model; S1-4-3: Pre-training the optimized data classification module according to a number of pre-processed historical data to be classified, to obtain a pre-trained data classification module; S1-4-4: Based on several historical user portraits, the optimized graph structure feature extraction module is trained to obtain several historical graph structure features and the final graph structure feature extraction module; S1-4-5: training the optimized audio feature extraction module according to a number of pre-processed historical audio data to obtain a number of historical audio features and a final audio feature extraction module; S1-4-6: training the optimized image feature extraction module according to a number of pre-processed historical image data to obtain a number of historical image features and a final image feature extraction module; S1-4-7: According to a number of preprocessed historical sequence data, the optimized sequence feature extraction module is trained to obtain a number of historical sequence features and a final sequence feature extraction module; S1-4-8: According to a number of historical graph structure features, a number of historical audio features, a number of historical image features, and a number of historical sequence features, the optimized attention weight module is trained to obtain a number of historical fusion features and a final attention weight module; S1-4-9: According to a number of historical fusion features, the optimized data classification module is trained to obtain the final data classification module; S1-4-10: Integrate the final graph structure feature extraction module, the final image feature extraction module, the final sequence feature extraction module, the final attention weight module and the final data classification module to obtain the final data classification model; S2: Based on the real-time user basic information, a user portrait generation model is used to generate a user portrait to obtain a real-time user portrait, including the following steps: S2-1: Collecting real-time user basic information, preprocessing the real-time user basic information, and obtaining preprocessed real-time user basic information; S2-2: Based on several key feature indicators, a key feature screening module is used to screen key features to obtain several real-time key features of real-time user basic information after preprocessing; S2-3: Based on several real-time key features, a user portrait generation module is used to generate a user portrait to obtain a real-time user portrait; S3: Based on the real-time user portrait and the real-time data to be classified, the data classification model is used to classify the data to obtain the real-time data classification results, including the following steps: S3-1: collect real-time data to be classified, pre-process the real-time data to be classified, and obtain the real-time data to be classified after pre-processing; S3-2: performing data analysis on the preprocessed real-time data to be classified to obtain corresponding preprocessed real-time audio data, preprocessed real-time image data, and preprocessed real-time sequence data; S3-4: Based on the real-time user portrait, the pre-processed real-time audio data, the pre-processed real-time image data, and the pre-processed real-time sequence data, a data classification model is used to perform data classification to obtain a real-time data classification result, including the following steps: S3-4-1: input the real-time user portrait, the pre-processed real-time audio data, the pre-processed real-time image data, and the pre-processed real-time sequence data into the data classification model; S3-4-2: Based on the real-time user portrait, use the graph structure feature extraction module to extract graph structure features and obtain real-time graph structure features; S3-4-3: Using an audio feature extraction module to extract audio features based on the pre-processed real-time audio data, to obtain real-time audio features; S3-4-4: Using an image feature extraction module to extract image features based on the preprocessed real-time image data, to obtain real-time image features; S3-4-5: Using a sequence feature extraction module to extract sequence features based on the preprocessed real-time sequence data, to obtain real-time sequence features; S3-4-6: According to the preset attention weight, use the attention weight module to fuse the real-time graph structure features, real-time audio features, real-time image features and real-time sequence features to obtain real-time fusion features; S3-4-7: Based on the real-time fusion features, use the data classification module to classify the data and obtain the real-time data classification results; S4: Use the blockchain network to perform distributed storage of real-time user portraits, real-time data to be classified, and real-time data classification results, including the following steps: S4-1: associate the real-time user portrait, real-time data to be classified, and real-time data classification results of the same user to obtain real-time associated data; S4-2: Store the real-time associated data in the InterPlanetary File System (IPFS) system of the blockchain network, obtain the real-time data hash value, and call the smart contract to generate the corresponding real-time transaction data according to the real-time data hash value; The IPFS system generates a unique data hash value for real-time associated data. The data hash value represents the content and structure of the real-time associated data. Transaction data includes the data hash value, timestamp, and retrieval tag (such as data or description for later retrieval). S4-3: Call the smart contract to convert the real-time transaction data into real-time blocks, and use several distributed nodes of the blockchain network to store the real-time blocks on the chain and generate corresponding real-time transaction records; S4-4: Update the distributed ledger of the blockchain network based on the real-time storage address, real-time retrieval tag and real-time transaction record of the real-time associated data in the IPFS system.

[0029] Embodiment 2: like Figure 2 As shown, this embodiment provides a data classification system based on deep learning, which is used to implement a data classification method, characterized in that: the system includes a model building unit, a user portrait generation unit, a data classification unit and a distributed storage unit connected in sequence; A model building unit, used to build a user portrait generation model and a data classification model using a deep learning algorithm based on a number of historical user basic information and a number of historical data to be classified; A user portrait generation unit, used to generate a user portrait based on real-time user basic information using a user portrait generation model to obtain a real-time user portrait; A data classification unit is used to classify data based on real-time user portraits and real-time data to be classified using a data classification model to obtain real-time data classification results; The distributed storage unit is used to use the blockchain network to perform distributed storage of real-time user portraits, real-time data to be classified, and real-time data classification results.

[0030] The present invention provides a data classification method and system based on deep learning, which realizes efficient data classification by constructing a data classification model, and the deep learning algorithm can mine the deep features of the data as the basis for data classification, thereby improving the classification accuracy and meeting the user's requirements for data classification accuracy; an automated data classification function is provided, which can extract, process and classify features of multimodal data, thereby improving the degree of intelligence and the efficiency of data processing, especially in the face of massive data processing environments, thereby improving practicality; in the data classification process, considering that different users have large differences in classification standards, a user portrait model is used to generate user portraits based on user basic information, and customized classification is performed in combination with user preferences and habits, thereby improving user experience; a blockchain network is used to perform distributed storage of data, thereby ensuring the security of the data and the traceability of data classification.

[0031] The present invention is not limited to the above optional implementations, and anyone can derive other various forms of products under the enlightenment of the present invention. The above specific implementations should not be understood as limiting the scope of protection of the present invention. The scope of protection of the present invention should be based on the definition in the claims, and the description can be used to interpret the claims.

Claims

1. A data classification method based on deep learning, characterized in that: The steps include: Based on some historical user basic information and some historical data to be classified, a deep learning algorithm is used to build a user portrait generation model and a data classification model; According to the real-time user basic information, the user portrait generation model is used to generate the user portrait to obtain the real-time user portrait; According to the real-time user portrait and the real-time data to be classified, the data classification model is used to classify the data and obtain the real-time data classification results; Use the blockchain network to distribute the real-time user portraits, real-time data to be classified, and real-time data classification results.

2. The data classification method based on deep learning according to claim 1, characterized in that: The user portrait generation model is constructed based on the RF-MLP algorithm, and the user portrait generation model includes a key feature screening module constructed based on the RF algorithm and a user portrait generation module constructed based on the MLP algorithm, which are connected in sequence; The data classification model includes a data classification model constructed based on the GCN-logfBank-CNN-LSTM-Attention-DBN algorithm, and the data classification model includes a graph structure feature extraction module constructed based on the GCN algorithm, an audio feature extraction module constructed based on the logfBank algorithm, an image feature extraction module constructed based on the CNN algorithm, a sequence feature extraction module constructed based on the LSTM algorithm, an attention weight module constructed based on the Attention mechanism, and a data classification module constructed based on the DBN algorithm. The graph structure feature extraction module, the audio feature extraction module, the image feature extraction module, and the sequence feature extraction module are all connected to the attention weight module, and the attention weight module is connected to the data classification module.

3. The data classification method based on deep learning according to claim 2, characterized in that: Based on some historical user basic information and some historical data to be classified, a deep learning algorithm is used to build a user portrait generation model and a data classification model, including the following steps: Collecting some historical user basic information and some historical data to be classified, and preprocessing them to obtain some preprocessed historical user basic information and some preprocessed historical data to be classified; Performing data analysis on the preprocessed historical data to be classified to obtain corresponding preprocessed historical audio data, preprocessed historical image data, and preprocessed historical sequence data; Based on some pre-processed historical user basic information, a deep learning algorithm is used to build a user portrait generation model and generate several historical user portraits; A data classification model is constructed using a deep learning algorithm based on several historical user portraits, several preprocessed historical audio data to be classified, preprocessed historical image data, and preprocessed historical sequence data.

4. The data classification method based on deep learning according to claim 2, characterized in that: Based on some pre-processed historical user basic information, a user portrait generation model is constructed using a deep learning algorithm, and several historical user portraits are generated, including the following steps: Using the RF-MLP algorithm, an initial user portrait generation model is constructed; the initial user portrait generation model includes an initial key feature screening module and an initial user portrait generation module; According to a number of pre-processed historical user basic information, an initial key feature screening module is trained to obtain a number of key feature indicators, a number of historical key features of each pre-processed historical user basic information, and a final key feature screening module; According to several historical key features of all pre-processed historical user basic information, the initial user portrait generation module is trained to obtain several historical user portraits and the final user portrait generation module; The final key feature screening module and the final user portrait generation module are integrated to obtain the final user portrait generation model.

5. The data classification method based on deep learning according to claim 2, characterized in that: Based on several historical user portraits, several preprocessed historical audio data, preprocessed historical image data and preprocessed historical sequence data of preprocessed historical data to be classified, a data classification model is constructed using a deep learning algorithm, including the following steps: Use the GCN-logfBank-CNN-LSTM-Attention-DBN algorithm to build an initial data classification model; Taking minimizing mean square error as the optimization goal, a swarm intelligence optimization algorithm is used to optimize the initial model parameters of the initial data classification model to obtain an optimized data classification model; the optimized data classification model includes an optimized graph structure feature extraction module, an optimized audio feature extraction module, an optimized image feature extraction module, an optimized sequence feature extraction module, an optimized attention weight module and an optimized data classification module; Pre-training the optimized data classification module according to a number of pre-processed historical data to be classified, thereby obtaining a pre-trained data classification module; According to several historical user portraits, the optimized graph structure feature extraction module is trained to obtain several historical graph structure features and the final graph structure feature extraction module; According to a number of pre-processed historical audio data, the optimized audio feature extraction module is trained to obtain a number of historical audio features and a final audio feature extraction module; According to a number of pre-processed historical image data, the optimized image feature extraction module is trained to obtain a number of historical image features and a final image feature extraction module; According to a number of pre-processed historical sequence data, the optimized sequence feature extraction module is trained to obtain a number of historical sequence features and a final sequence feature extraction module; According to several historical graph structure features, several historical audio features, several historical image features and several historical sequence features, the optimized attention weight module is trained to obtain several historical fusion features and the final attention weight module; According to several historical fusion features, the optimized data classification module is trained to obtain the final data classification module; The final graph structure feature extraction module, the final image feature extraction module, the final sequence feature extraction module, the final attention weight module and the final data classification module are integrated to obtain the final data classification model.

6. The data classification method based on deep learning according to claim 4, characterized in that: Based on the real-time user basic information, the user portrait generation model is used to generate the user portrait to obtain the real-time user portrait, including the following steps: Collecting real-time user basic information, preprocessing the real-time user basic information, and obtaining preprocessed real-time user basic information; According to several key feature indicators, a key feature screening module is used to screen key features to obtain several real-time key features of real-time user basic information after preprocessing; According to several real-time key features, a user portrait generation module is used to generate a user portrait to obtain a real-time user portrait.

7. The data classification method based on deep learning according to claim 5, characterized in that: According to the real-time user portrait and the real-time data to be classified, the data classification model is used to classify the data and obtain the real-time data classification results, including the following steps: Collecting real-time data to be classified, preprocessing the real-time data to be classified, and obtaining real-time data to be classified after preprocessing; Performing data analysis on the preprocessed real-time data to be classified to obtain corresponding preprocessed real-time audio data, preprocessed real-time image data, and preprocessed real-time sequence data; According to the real-time user portrait, the pre-processed real-time audio data, the pre-processed real-time image data and the pre-processed real-time sequence data, a data classification model is used to perform data classification to obtain real-time data classification results.

8. The data classification method based on deep learning according to claim 7, characterized in that: According to the real-time user portrait, the pre-processed real-time audio data, the pre-processed real-time image data, and the pre-processed real-time sequence data, a data classification model is used to perform data classification to obtain a real-time data classification result, including the following steps: Inputting the real-time user portrait, the pre-processed real-time audio data, the pre-processed real-time image data, and the pre-processed real-time sequence data into the data classification model; According to the real-time user portrait, use the graph structure feature extraction module to extract graph structure features and obtain real-time graph structure features; According to the pre-processed real-time audio data, an audio feature extraction module is used to extract audio features to obtain real-time audio features; According to the pre-processed real-time image data, an image feature extraction module is used to extract image features to obtain real-time image features; According to the pre-processed real-time sequence data, a sequence feature extraction module is used to extract sequence features to obtain real-time sequence features; According to the preset attention weight, the attention weight module is used to fuse the real-time graph structure features, real-time audio features, real-time image features and real-time sequence features to obtain real-time fusion features; According to the real-time fusion features, the data classification module is used to classify the data and obtain the real-time data classification results.

9. The data classification method based on deep learning according to claim 8, characterized in that: Using the blockchain network, real-time user portraits, real-time data to be classified, and real-time data classification results are distributedly stored, including the following steps: The real-time user portrait, real-time data to be classified and real-time data classification results of the same user are associated to obtain real-time associated data; The real-time associated data is stored in the IPFS system of the blockchain network to obtain the real-time data hash value, and the smart contract is called to generate the corresponding real-time transaction data according to the real-time data hash value; Call smart contracts to convert real-time transaction data into real-time blocks, and use several distributed nodes of the blockchain network to store real-time blocks on the chain and generate corresponding real-time transaction records; Update the distributed ledger of the blockchain network based on the real-time storage address, real-time retrieval tag and real-time transaction records of the real-time associated data in the IPFS system.

10. A data classification system based on deep learning, used to implement the data classification method according to any one of claims 1 to 9, characterized in that: The system comprises a model building unit, a user portrait generating unit, a data classification unit and a distributed storage unit which are connected in sequence.