Data asset intelligent label and automatic classification algorithm

Through intelligent labeling of data assets and automatic classification algorithms, machine learning algorithms are used to automate data classification, and the problem of traditional data classification methods relying on manual operations is solved, achieving efficient and accurate data management and storage.

CN119989113AInactive Publication Date: 2025-05-13SHANGHAI ZHEJIN INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510047898.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional data classification methods rely on manual operations, consume a lot of time and energy, and are susceptible to personal subjective factors, resulting in inconsistent classification results and inefficiency.

Method used

Data asset intelligent labeling and automatic classification algorithm are adopted, including data preprocessing, feature extraction, model training, intelligent label generation and classification result optimization, and data classification is automatically generated through machine learning algorithms to reduce manual intervention.

Benefits of technology

It significantly improves the efficiency and accuracy of data management, reduces dependence on manual intervention, optimizes the data storage structure, and improves data retrieval efficiency and convenience of use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989113A_ABST
    Figure CN119989113A_ABST
Patent Text Reader

Abstract

The invention discloses a data asset intelligent label and automatic classification algorithm, and the algorithm comprises a data receiving module which is used for receiving to-be-classified data assets; the feature extraction module is used for extracting feature information from the data assets; the tag generation module is used for generating an intelligent tag according to the extracted feature information; the classification algorithm module is used for automatically classifying the data assets according to the intelligent labels; the classification result output module is used for outputting a classification result; by integrating a support vector machine (SVM) algorithm and a natural language processing technology (NLP), the efficiency and accuracy of data asset processing are remarkably improved; the data receiving module in the system can process data assets of various formats and sources, and wide compatibility of data is ensured; through combination of the feature extraction module and the label generation module, key information can be extracted from data, intelligent labels which accurately describe data contents can be generated, and the retrievability and the understandability of the data are greatly enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to data asset intelligent labels and automatic classification algorithms. Background Art

[0002] In today's era of rapid development of information technology, data assets have become one of the most valuable resources for enterprises. With the explosive growth of data volume, how to efficiently manage and utilize these data assets has become a major challenge facing enterprises. Enterprises must adopt advanced technologies and strategies to ensure the security, integrity and accessibility of data. At the same time, enterprises also need to cultivate professional data management teams to cope with increasingly complex business needs and market changes. In addition, with the continuous advancement of big data analysis and artificial intelligence technologies, enterprises can use these technologies to extract valuable information from massive data, thereby making more informed business decisions and enhancing competitiveness.

[0003] Therefore, it is particularly important to develop an intelligent labeling and automatic classification algorithm for data assets. The algorithm must not only be able to efficiently process massive text data, but also have powerful feature extraction and model training capabilities to ensure the accuracy and stability of classification; at the same time, in order to adapt to the business needs of different enterprises, the algorithm should also be highly flexible and scalable, and be able to dynamically adjust and optimize according to different application scenarios.

[0004] Traditional data classification methods usually rely on manual operations, which not only consumes a lot of time and energy, but is also easily affected by personal subjective factors, resulting in inconsistent and inefficient classification results; therefore, it is particularly important to develop an algorithm that can automatically generate and classify data assets for intelligent labels. Such an algorithm can significantly improve the efficiency and accuracy of data processing and bring huge competitive advantages to enterprises. Summary of the invention

[0005] The purpose of the present invention is to propose a data asset intelligent labeling and automatic classification algorithm to solve the problem that traditional data classification methods usually rely on manual operations, which not only consumes a lot of time and energy, but is also easily affected by personal subjective factors, thus leading to inconsistency and low efficiency of classification results.

[0006] In order to achieve the above objectives, the present invention adopts the following technology data asset intelligent label and automatic classification algorithm: the data asset intelligent label and automatic classification algorithm include:

[0007] S1. Data preprocessing: Perform preprocessing operations such as cleaning, deduplication, and formatting on the original data to ensure data quality;

[0008] S2. Feature extraction: Extract data features based on data content and context information, including but not limited to keywords, topics, and sentiment tendencies;

[0009] S3, model training: use machine learning algorithms, such as support vector machine (SVM), random forest, deep learning, etc., to train the extracted features and build a classification model;

[0010] S4, Intelligent label generation: Based on the trained classification model, new data assets are automatically labeled with intelligent labels and classified and stored according to the label content;

[0011] S5. Optimization of classification results: Through manual review and feedback mechanism, the classification results are optimized to improve the classification accuracy.

[0012] As a further description of the above technical data asset intelligent labeling and automatic classification algorithm:

[0013] The data preprocessing also includes standardizing the data to eliminate dimensional differences between different data sources; the feature extraction also includes using natural language processing technology to perform word segmentation, part-of-speech tagging and named entity recognition on text data; the model training also includes cross-validation of the classification model to evaluate the generalization ability of the model and prevent overfitting; the intelligent label generation also includes dynamically adjusting label weights and classification rules according to the business logic and application scenarios of the data assets; the classification result optimization also includes using machine learning algorithms to learn from manual review results and automatically adjust classification model parameters to achieve self-optimization.

[0014] The present invention adopts the following technical data asset intelligent labeling and automatic classification system, including:

[0015] Data preprocessing module: responsible for preprocessing the original data;

[0016] Feature extraction module: responsible for extracting data features;

[0017] Model training module: responsible for training classification models using machine learning algorithms;

[0018] Intelligent label generation module: responsible for intelligent labeling and classification of data assets based on the classification model formed by the model training module;

[0019] Classification result optimization module: responsible for manual review and optimization of the classification results generated by the intelligent label generation module.

[0020] In summary, due to the adoption of the above-mentioned technical data asset intelligent label and automatic classification algorithm, the beneficial effects of the present invention are:

[0021] The present invention aims to significantly improve the efficiency of data management through a series of innovative technical means. First, by introducing automated intelligent tags and classification algorithms, we have successfully reduced dependence on human intervention, thereby greatly improving the speed of data processing; secondly, in order to further improve the accuracy of data management, we have adopted advanced machine learning algorithms, which ensure the high accuracy of data classification and the rationality of intelligent tags; finally, through the application of intelligent tags and classification, we have optimized the data storage structure, making the storage structure of data assets more rational, and greatly improving the data retrieval efficiency and ease of use. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 Shown is an overall module diagram provided according to the present invention. DETAILED DESCRIPTION

[0023] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical data asset intelligent label and automatic classification algorithm in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] The data asset intelligent label and automatic classification algorithm provided by the present invention is a data asset intelligent label and automatic classification algorithm, and the specific steps are as follows:

[0025] S1. Data preprocessing: Perform preprocessing operations such as cleaning, deduplication, and formatting on the original data to ensure data quality;

[0026] S2. Feature extraction: Extract data features based on data content and context information, including but not limited to keywords, topics, and sentiment tendencies;

[0027] S3, model training: use machine learning algorithms, such as support vector machine (SVM), random forest, deep learning, etc., to train the extracted features and build a classification model;

[0028] The steps of model training can be divided into the following stages, each of which contains specific operations and formulas:

[0029] Step 1: Import the required machine learning models and tools; import the required Python libraries and tools, such as classifiers such as support vector machine (SVM), random forest, multi-layer perceptron (MLP), as well as tools for data set partitioning, classification reporting, and model saving; code: from sklearn.svm import SVC

[0030] from sklearn.ensemble import RandomForestClassifier

[0031] from sklearn.neural_network import MLPClassifier

[0032] from sklearn.model_selection import train_test_split

[0033] from sklearn.metrics import classification_report

[0034] import joblib

[0035] Step 2: Load or generate feature data and label data; Assume X is feature data and y is label data. There should be code to load or generate feature data and label data;

[0036] Code: X, y = ... # Load or generate feature data and label data here

[0037] Step 3: Divide the data set into training set and test set; Divide the data set into training set and test set, with the test set accounting for 20%;

[0038] Code: X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=0.2,random_state=42)

[0039] Step 4: Create and train a support vector machine model; Use the support vector machine (SVM) algorithm and the radial basis function as the kernel function to create and train the model. Then use the model to make predictions and print a classification report; Code: svm_model = SVC (kernel = 'rbf', C = 1.0, gamma = 'auto') # Use the radial basis function as the kernel function

[0040] svm_model.fit(X_train,y_train)#training model

[0041] svm_predictions = svm_model.predict(X_test) #Use the model for prediction print("SVMClassification Report:")

[0042] print(classification_reporty_test,svm_predictions)#Print classification report

[0043] Step 5: Create and train a random forest model; Use the random forest algorithm and a random forest of 100 trees to create and train the model. Then use the model to make predictions and print a classification report.

[0044] Code: rf_model = RandomForestClassifier(n_estimators = 100, random_state = 42) # Use a random forest with 100 trees

[0045] rf_model.fit(X_train,y_train)#training model

[0046] rf_predictions = rf_model.predict(X_test)#Use the model to make predictions print("Random Forest Classification Report:")

[0047] print(classification_reporty_test,rf_predictions)#Print classification report

[0048] Step 6: Create and train a deep learning model (multilayer perceptron); Use the deep learning algorithm, multilayer perceptron (MLP), create a hidden layer with 100 neurons, create and train the model. Then use the model to make predictions and print a classification report;

[0049] Code: mlp_model = MLPClassifier(hidden_layer_sizes = 100), max_iter = 500, alpha = 1e-4, solver = 'sgd', verbose = 10, random_state = 1, learning_rate_init = .1) # Use one hidden layer, 100 neurons

[0050] mlp_model.fit(X_train,y_train)#training model

[0051] mlp_predictions = mlp_model.predict(X_test)#Use the model to make predictions print("MLPClassification Report:")

[0052] print(classification_reporty_test,mlp_predictions)#Print classification report

[0053] Step 7: Save the trained model to a file; Save the trained support vector machine, random forest and multilayer perceptron models to a file;

[0054] Code: joblib.dump(svm_model,'svm_model.pkl')#Save support vector machine model joblib.dump(rf_model,'rf_model.pkl')#Save random forest model joblib.dump(mlp_model,'mlp_model.pkl')#Save multi-layer perceptron model;

[0055] S4. Smart label generation: Using the trained and optimized classification model, we can automatically add smart labels to new data assets. This process involves analyzing data assets through the model to identify key features of the data and assigning corresponding smart labels to data assets based on these features. These labels not only help to quickly identify data content, but also can classify and store data assets according to the categories indicated by the labels, thereby achieving more efficient data management.

[0056] S5. Optimization of classification results: In order to further improve the accuracy and reliability of classification, a manual review and feedback mechanism is introduced. Through the careful review of the classification results by professionals, errors that may occur in the classification process can be discovered and corrected. At the same time, by collecting user feedback information, the classification model is continuously adjusted and optimized to ensure that the accuracy of the classification results continues to improve, ultimately achieving the goal of improving the overall classification accuracy.

[0057] The data preprocessing step is not limited to cleaning and organizing the data, but also includes data standardization, a process that aims to eliminate the dimensional differences between different data sources and ensure the consistency and comparability of the data; standardization can make subsequent analysis and processing more accurate and efficient. In addition, the feature extraction step uses advanced natural language processing technology to conduct in-depth analysis of text data, including key links such as word segmentation, part-of-speech tagging, and named entity recognition;

[0058] The application of these technologies enables the system to understand and process text data more effectively, thereby extracting more valuable information features; in the model training stage, the cross-validation method is used to evaluate the classification model, which can effectively evaluate the generalization ability of the model and help prevent model overfitting, ensuring that the model performs equally well on unknown data; the intelligent label generation step takes into account the business logic and application scenarios of data assets, and dynamically adjusts label weights and classification rules to adapt to different business needs and scenario changes; finally, in terms of optimizing classification results, the present invention uses machine learning algorithms to learn from manual review results, and automatically adjusts classification model parameters to achieve self-optimization of the model, thereby continuously improving classification accuracy and efficiency.

[0059] like Figure 1 As shown in the figure, the data asset intelligent labeling and automatic classification system specifically includes:

[0060] Data preprocessing module: The main responsibility of this part is to conduct in-depth cleaning and formatting of the large amount of collected text data to ensure the purity and consistency of the data; through meticulous preprocessing work, text and inconsistencies in the data can be removed, thereby providing accurate and high-quality basic information for subsequent data analysis and processing steps.

[0061] Feature extraction module: In this module, our main task is to extract key words and phrases that are crucial for text analysis and classification from the preprocessed text data; these extracted features will serve as the basis for subsequent model training and classification work, helping to improve the accuracy and efficiency of text analysis.

[0062] Model training module: The core work of this module is to use the random forest algorithm to train the text classification model; through carefully designed algorithms and parameter optimization, we can significantly improve the performance and accuracy of the model in text classification tasks, ensuring that the model can efficiently and accurately classify text data.

[0063] Intelligent label generation module: The main function of this module is to perform intelligent analysis on newly collected text data and automatically assign corresponding classification labels to them; through this automated process, rapid classification and processing of text data can be achieved, greatly improving the efficiency and response speed of data processing.

[0064] Classification result optimization module: In this module, our work focuses on manual review of classification results to ensure the accuracy of classification. Through manual review, we can discover and correct errors that may occur in the model during the classification process, and fine-tune and optimize the model based on the review results to continuously improve the classification effect and quality.

[0065] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, can make equivalent replacements or changes to the technical data asset smart tags and automatic classification algorithms and their inventive concepts according to the present invention, which should be covered by the protection scope of the present invention.

Claims

1. Intelligent labeling and automatic classification algorithm for data assets, characterized by: The data asset intelligent label and automatic classification algorithm include: S1. Data preprocessing: Perform preprocessing operations such as cleaning, deduplication, and formatting on the original data to ensure data quality; S2. Feature extraction: Extract data features based on data content and context information, including but not limited to keywords, topics, and sentiment tendencies; S3, model training: use machine learning algorithms, such as support vector machine (SVM), random forest, deep learning, etc., to train the extracted features and build a classification model; S4, Intelligent label generation: Based on the trained classification model, new data assets are automatically labeled with intelligent labels and classified and stored according to the label content; S5. Optimization of classification results: Through manual review and feedback mechanism, the classification results are optimized to improve the classification accuracy.

2. The data asset intelligent labeling and automatic classification algorithm according to claim 1 is characterized in that: The data preprocessing also includes standardizing the data to eliminate dimensional differences between different data sources.

3. The data asset intelligent labeling and automatic classification algorithm according to claim 1 is characterized in that: The feature extraction also includes using natural language processing technology to perform word segmentation, part-of-speech tagging and named entity recognition on text data.

4. The data asset intelligent labeling and automatic classification algorithm according to claim 1 is characterized in that: The model training also includes cross-validation of the classification model to evaluate the generalization ability of the model and prevent overfitting.

5. The data asset intelligent labeling and automatic classification algorithm according to claim 1 is characterized in that: The smart tag generation also includes dynamically adjusting tag weights and classification rules according to the business logic and application scenarios of the data assets.

6. The data asset intelligent labeling and automatic classification algorithm according to claim 1 is characterized in that: The classification result optimization also includes using a machine learning algorithm to learn from manual review results and automatically adjust classification model parameters to achieve self-optimization.

7. A data asset intelligent label and automatic classification system based on the data asset intelligent label and automatic classification algorithm according to any one of claims 1 to 6, characterized in that: include: Data preprocessing module: responsible for preprocessing the original data; Feature extraction module: responsible for extracting data features from the data processed by the data preprocessing module; Model training module: responsible for training classification models using machine learning algorithms; Intelligent label generation module: responsible for intelligent labeling and classification of data assets based on the classification model formed by the model training module; Classification result optimization module: responsible for manual review and optimization of the classification results generated by the intelligent label generation module.

Citation Information

Cited By

  • Enterprise data asset intelligent management system for server research, development and manufacturing

    CN120765199A

  • An enterprise data asset intelligent management system for server research and development

    CN120765199B