Method and system for realizing data classification and grading based on large language model technology
Through large language model technology, combined with deep learning feature extraction and multi-dimensional evaluation system, various shortcomings of traditional data classification and grading methods are solved, and efficient and intelligent multimodal data classification and grading are achieved, adapting to data changes and reducing labor costs.
Patent Information
- Application Number
- CN202510392252.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
Traditional data classification and grading methods have shortcomings in accuracy, intelligence, adaptability, multi-dimensional analysis capabilities, labor costs and multi-field expansion, and are unable to effectively process multi-modal data and quickly adapt to data changes, resulting in inefficient classification and grading and limited accuracy.
Using technology based on large language models, data classification and grading is achieved through deep learning feature extraction, attention mechanism, transfer learning, multi-modal data processing, automated anomaly detection and real-time feedback iteration, combined with a multi-dimensional evaluation system.
It improves the accuracy and intelligence of data classification and grading, enhances the adaptability of the system, reduces labor costs, improves large-scale data processing performance and multi-dimensional analysis capabilities, supports multi-modal data processing, realizes predictive classification and grading, and significantly improves the scientificity and efficiency of classification and grading.
Smart Images

Figure CN120277467A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data management, and particularly relates to a method and system for realizing data classification and grading based on large language model technology. Background Art
[0002] With the advent of the digital age, the amount of data has increased explosively, and the types of data have become increasingly diverse. How to effectively classify and grade data has become a key issue. Traditional data classification and grading methods mainly rely on manual judgment and manual operations, with low efficiency and easy to make mistakes. In order to improve efficiency and accuracy, automated data classification and grading technologies have gradually developed.
[0003] Among them, the rule-based classification and grading method classifies and grades data by presetting rules and conditions. It requires manual detailed definition of various rules to clarify the characteristics and classification criteria of data. However, this method has poor flexibility and is difficult to adapt to the dynamic changes and complex semantic relationships of data.
[0004] The machine learning-based method trains a model by learning the labeled data, so as to classify and grade new data. For example, using the recurrent neural network (RNN) model in neural networks, it can process sequence data. However, the RNN model needs to process input data in order, has insufficient parallelism, high computational cost, and can only capture the dependencies of shorter sequences.
[0005] Based on Transformer-based large language models, especially the BERT model, it breaks through the original limitations, can run at a faster speed, and remember the input data for a longer time. BERT is a Transformer model with an encoder-decoder (or encoder-only) architecture, which enables language models to achieve good results in many tasks.
[0006] The above-mentioned existing technologies have the following disadvantages:
[0007] In terms of accuracy and intelligence, traditional classification and grading systems rely on predefined rules and features, lack the ability of autonomous learning and optimization, resulting in limited accuracy and intelligence.
[0008] For the support of multi-modal data, traditional systems mainly process structured data, have insufficient support for unstructured data, and have a single processing method, unable to process various types of data.
[0009] In terms of adaptability, traditional systems completely rely on manual adjustment of rule configurations, with poor flexibility and difficult to quickly adapt to new changes and new models of data.
[0010] In terms of large-scale data processing performance, traditional systems are inefficient and time-consuming when dealing with large-scale data, and cannot complete classification and grading tasks quickly and efficiently.
[0011] In terms of multi-dimensional analysis capabilities, traditional systems only analyze single-field dimensions, making it difficult to comprehensively understand data, resulting in incomplete and inaccurate classification and grading results.
[0012] Due to human errors, when data cannot match the preset rules, manual intervention is required for judgment, and the error rate of human judgment is relatively high, affecting the accuracy of classification and grading results.
[0013] In terms of labor costs, traditional systems highly rely on manual configuration and intervention, requiring a large amount of labor costs, which increases the operating costs of enterprises.
[0014] In terms of predictive classification, traditional systems only support classification and grading within the configured rules and cannot perform forward-looking predictive classification and grading on future new data.
[0015] In terms of multi-domain expansion, although traditional systems can cover multiple domains, the integration degree is limited, and they cannot deeply integrate multi-domain knowledge and experience, affecting the scientificity and accuracy of classification and grading.
[0016] The patent application document with the publication number CN116732895A discloses "A Method for Automatic Classification and Grading of Government Affairs Data Based on a Large Model". This method constructs a large language model in the government affairs field to learn classification and grading rules, realizing the automatic classification and grading of government affairs data. The comparative document is innovative in the field of government affairs data classification and grading, but has problems such as domain limitations, rule dependence, and insufficient multi-modal support.
[0017] The patent application document with the publication number CN117851860A discloses "A Method for Automatically Generating Data Classification and Grading Templates". This method constructs classification templates using hierarchical clustering algorithms and traditional machine learning models, but has three technical limitations: ① The clustering process depends on preset parameters (N / θ1 / M), resulting in insufficient dynamic adaptability; ② The classification naming based on traditional NLP technology has weak ability to capture complex semantic relationships; ③ The classification accuracy of the TextCNN model for long-tail data is only 89.2%. Summary of the Invention
[0018] To overcome the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to provide a method and system for data classification and grading based on large language model technology. Through the technical collaboration of deep learning feature extraction (word frequency, TF-IDF value, semantic vector), attention mechanism (multi-head attention using the Transformer architecture), transfer learning (model fine-tuning and adaptation to specific tasks), multi-modal data processing (supporting various data types such as text and databases), automated anomaly detection (removing noise and outliers), real-time feedback iteration (evaluation metrics and confusion matrix analysis to optimize the model), and multi-dimensional evaluation system (accuracy, precision, recall, F1 value), the accuracy and intelligence of data classification and grading are improved, the processing of multi-modal data is supported, the adaptive ability of the system is enhanced, the large-scale data processing performance is improved, the multi-dimensional analysis ability is enhanced, human errors are reduced, labor costs are lowered, predictive classification and grading are achieved, multi-field expansion and integration are strengthened, thereby enhancing the scientificity, accuracy, and efficiency of classification and grading, and meeting the needs of enterprises and organizations for data classification and grading management.
[0019] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0020] A method for data classification and grading based on large language model technology, comprising the following steps:
[0021] Step 1, collect relevant data from multiple data sources;
[0022] Step 2, preprocess the data collected in Step 1, including: cleaning, denoising, and format conversion operations to form a standardized preprocessed data set;
[0023] Step 3, obtain text data from the standardized preprocessed data set obtained in Step 2, extract the features of word frequency, TF-IDF value, and semantic vector to obtain a comprehensive feature vector;
[0024] Step 4, randomly select a part of the standardized preprocessed data set formed in Step 2 as training set data, and use the remaining part as test set data; perform manual category annotation on the divided training set data and test set data to obtain the annotated training set data and test set data;
[0025] Step 5, use the comprehensive feature vector obtained in Step 3 and the annotated training set in Step 4 to train a classification and grading model;
[0026] Step 6, input the annotated test set in Step 4 into the classification and grading model trained in Step 5 for testing to obtain a prediction result, compare the prediction result of the classification and grading model with the annotated test set in Step 4, evaluate the accuracy and performance of the classification and grading model trained in Step 5, and select the optimal classification and grading model;
[0027] Step 7: Input the data to be classified and graded into the optimal classification and grading model selected in Step 6 to obtain the classification and grading results.
[0028] The specific method of Step 1 is as follows:
[0029] Step 1.1: Determine the data sources from which data needs to be collected, including: databases and file systems;
[0030] Step 1.2: For different data sources: use SQL query statements to extract data from relational databases for databases; use file reading functions to collect the required data for file systems.
[0031] The specific method of Step 2 is as follows:
[0032] Step 2.1: Use data cleaning algorithms to remove missing values, duplicate values, and outliers from the data collected in Step 1 to obtain a preliminary cleaned dataset;
[0033] Step 2.2: For the preliminary cleaned dataset obtained in Step 2.1, implement filtering algorithms and statistical methods to remove noise and detect and process outliers, specifically including:
[0034] Smooth time series data through the moving average filtering algorithm, process volatile data using the exponentially weighted moving average, and apply Gaussian filtering to eliminate random noise;
[0035] Use the Z-score method to identify numerical outliers and the interquartile range (IQR) method to detect outliers in non-normally distributed data;
[0036] Step 2.3: Implement unified conversion for various formats of data: generate semantic vectors for text data through the BERT model, uniformly convert the date format to the ISO 8601 standard format, uniformly desensitize sensitive data, and finally generate a standardized preprocessed dataset.
[0037] The specific method of Step 3 is as follows:
[0038] Step 3.1: Use the natural language processing toolkit NLTK library to obtain text data from the standardized preprocessed dataset obtained in Step 2 and perform segmentation to obtain word and phrase units; perform word frequency (TF) statistics on the word and phrase units;
[0039] Step 3.2: Based on the word frequency statistics results in Step 3.1, first, calculate the inverse document frequency (IDF) value for each word and phrase unit, that is, the general importance of the word and phrase unit in the entire document set; secondly, multiply the word frequency (TF) value by the inverse document frequency (IDF) value to obtain the frequency of each word in a single document and its distribution in the entire document set, that is, the TF-IDF value;
[0040] In step 3.3, the semantic vector of the text is generated by using the BERT model, and the TF-IDF value obtained in step 3.2 is used as the weight to perform weighted average on the semantic vector to obtain a comprehensive feature vector.
[0041] The specific method of step 5 is as follows:
[0042] In step 5.1, for the comprehensive feature vector obtained in step 3, data cleaning is performed to remove outliers and missing values, and data standardization or normalization processing is carried out to ensure that its format and structure meet the input requirements of the selected classification and grading model, and the processed comprehensive feature vector data is obtained.
[0043] In step 5.2, the labeled training set in step 4 and the processed comprehensive feature vector data in step 5.1 are integrated according to the sample-feature-label correspondence relationship.
[0044] In step 5.3, for the classification and grading model algorithm, parameters are set: learning rate, batch size, and number of training epochs.
[0045] In step 5.4, the integrated feature vector data and the labeled training set in step 5.2 are used to train the classification and grading model with the parameters set in step 5.3. During the training process, the classification and grading model learns the mapping relationship between features and class labels.
[0046] In step 5.5, monitor the error or loss function value during the training process in step 5.4 to evaluate the learning effect of the classification and grading model. If the classification and grading model shows overfitting or underfitting, improve the model by adjusting the regularization parameter, increasing data augmentation, and optimizing the feature combination.
[0047] The specific method of step 6 is as follows:
[0048] In step 6.1, the labeled test set in step 4 is input into the classification and grading model trained in step 5 for testing to obtain the prediction results, and the prediction results are compared with the labeled test set in step 4.
[0049] In step 6.2, based on the comparison results in step 6.1, evaluation metrics are calculated, including: accuracy, precision, recall, and F1-score. The calculation methods are as follows:
[0050] Accuracy = (Number of correctly predicted samples / Total number of samples) × 100%
[0051] Precision = (True positives / (True positives + False positives)) × 100%
[0052] Recall rate = (True positives / (True positives + False negatives)) × 100%
[0053] F1 score = 2 × (Precision × Recall) / (Precision + Recall);
[0054] Step 6.3, draw a Confusion Matrix to show the prediction of the classification and grading model for different categories, which complements the evaluation metrics in Step 6.2 to obtain the evaluation results;
[0055] Step 6.4, according to the evaluation results obtained in Step 6.3, analyze whether the classification and grading model has overfitting or underfitting for certain categories; if so, adjust the parameters of the classification and grading model, increase the amount of training data, improve the feature extraction method, and retrain and evaluate; if not, select the optimal classification and grading model according to the evaluation results and actual requirements.
[0056] The present invention provides a data classification and grading system based on large language model technology, including:
[0057] A data collection module for collecting relevant data from various data sources;
[0058] A data processing module for preprocessing the collected data, including: cleaning, denoising, and format conversion operations to form a standardized preprocessed data set;
[0059] A feature extraction module for obtaining text data from the standardized preprocessed data set, extracting features such as word frequency, TF-IDF value, and semantic vector to obtain a comprehensive feature vector;
[0060] A data annotation module for randomly extracting a part of the standardized preprocessed data set as training set data and using the remaining part as test set data; performing manual category annotation on the divided training set data and test set data to obtain the annotated training set data and test set data;
[0061] A model training module for training a classification and grading model using the comprehensive feature vector and the annotated training set;
[0062] A model evaluation module for inputting the annotated test set into the trained classification and grading model for testing to obtain prediction results, comparing the prediction results of the classification and grading model with the annotated test set, evaluating the accuracy and performance of the trained classification and grading model, and selecting the optimal classification and grading model;
[0063] A data classification and grading module for inputting the data to be classified and graded into the optimal classification and grading model to obtain the classification and grading results.
[0064] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0065] 1) The accuracy and intelligence are improved: The traditional classification and grading system completely relies on predefined rules and features, lacks the ability of autonomous learning and optimization, and has relatively limited accuracy and intelligence. The present invention utilizes a large language model trained with a large amount of data, and through extracting features such as word frequency, TF-IDF value, and semantic vector in step 3, realizes deep learning feature extraction, thereby improving the accuracy and intelligence of classification and grading.
[0066] 2) It has strong multi-modal data support ability: The traditional system mainly processes structured data and has obvious deficiencies in supporting unstructured data, with a single processing method. The present invention can synchronously process the data including in the database and the file system through step 1 and perform integrated classification and grading. This enables the system to better adapt to the diversity of modern data, be able to process various types of data, and provide users with a more comprehensive data classification and grading service.
[0067] 3) It has excellent adaptability: The traditional system completely relies on manual adjustment of rule configuration, with poor flexibility and difficulty in quickly adapting to new changes. The present invention realizes automatic anomaly detection and real-time feedback iteration through step 6 (steps 6.1 to 6.4 evaluate the accuracy and performance of the classification and grading model and select the optimal model). This method can automatically adapt to new data types and features, and continuously self-optimize the classification and grading results as the data dynamically changes and new patterns emerge, thereby significantly improving the adaptability and accuracy of the model.
[0068] 4) It has outstanding multi-dimensional analysis ability: The traditional system only analyzes from the single-field dimension, with a narrow perspective and difficulty in comprehensively understanding the data. The present invention comprehensively considers multiple dimensions and attributes of the data through data preprocessing (removing noise and outliers) in step 2, feature extraction (word frequency, TF-IDF value, and semantic vector) in step 3, and model training (using multiple features for classification and grading) in step 5. Specifically, step 2 ensures the data quality, step 3 extracts rich features, and step 5 uses these features to train a classification and grading model capable of processing multi-dimensional data. Step 6 verifies the model performance through evaluation metrics and confusion matrix, and finally applies the optimal model in step 7 to ensure the comprehensiveness and accuracy of the classification and grading results. This method significantly improves the analysis ability of complex data and provides deeper data insights.
[0069] 5) Reduction of human error: When data cannot match the preset rules, traditional systems require manual intervention for judgment, but the error rate of human judgment is relatively high. In the present invention, the training set and test set labeled in step 4 are used for manual category labeling, and in step 6, by calculating evaluation metrics such as accuracy, precision, recall, and F1 value (step 6.2) and drawing a confusion matrix (step 6.3), the classification and grading errors caused by human judgment errors are reduced, and the reliability of the results is improved.
[0070] 6) Reduction of labor costs: Traditional systems highly rely on manual configuration and intervention, requiring a large amount of labor costs. In the present invention, through an automated classification and grading process (such as training a classification and grading model in step 5 and evaluating the model performance in step 6), the dependence on manual labor is reduced, and the operation costs of enterprises are lowered.
[0071] 7) Strong predictive classification ability: Traditional systems only support classification and grading within the rule configuration. Based on rich historical data and learning results, the present invention trains a classification and grading model in step 5 and inputs the data to be classified and graded into the optimal model in step 7, enabling forward-looking predictive classification and grading of future new data and providing strong support for decision-making.
[0072] 8) Good multi-domain expansion ability: Although traditional systems can cover multiple domains, the integration degree is limited. The present invention deeply integrates the knowledge and experience of multiple domains to further improve the scientificity and accuracy of classification and grading (such as collecting data from multiple data sources in step 1 and uniformly converting data in multiple formats in step 2). This enables the system to provide high-quality classification and grading services in the applications of different domains and meet the needs of different users.
[0073] 9) Aiming at the technical problems existing in the patent application document with the publication number CN116732895A in the prior art, through technological innovations such as multi-modal feature fusion, dynamic model optimization, and cross-domain adaptability, the present invention effectively solves the above problems, achieving a breakthrough improvement in classification accuracy (the experimental results show that the accuracy rate is increased by 35%), processing efficiency (the processing speed of abnormal data is increased by 4 times), and scenario adaptability, and is more suitable for cross-industry data governance scenarios.
[0074] 10) Aiming at the technical problems existing in the patent application document with the publication number CN117851860A in the prior art, by introducing the deep semantic understanding ability of large language models and a dynamic feature weighting mechanism (TF-IDF + BERT vector fusion), the present invention realizes dynamic clustering without preset parameters, multi-dimensional semantic feature fusion, and real-time model optimization, improving the classification accuracy of unstructured data to 94.2%, and effectively breaking through the technical bottlenecks of traditional methods in terms of the depth of semantic understanding, dynamic adaptability, and model generalization.
[0075] In summary, compared with the prior art, through the technical collaboration of deep learning feature extraction, attention mechanism optimization, transfer learning enhancement, multi-modal data processing, automated anomaly detection, real-time feedback iteration, and multi-dimensional evaluation system, the present invention has achieved a systematic innovation in the methodology of data classification and grading. Its innovation is reflected in three aspects: at the technical integration level, for the first time, the large language model is combined with the multi-modal feature weighting fusion technology, breaking through the limitations of traditional single-modal classification; at the process optimization level, a full-link closed-loop system from data cleaning to dynamic model optimization is constructed, significantly improving the processing efficiency and system robustness; at the application expansion level, through the interpretable feature engineering and domain adaptation fine-tuning mechanism, the system has the ability to quickly migrate across industries and scenarios. This multi-dimensional technological innovation has achieved a breakthrough improvement in key indicators such as classification accuracy, processing efficiency, and scenario adaptability. Compared with traditional methods, the accuracy has increased by more than 35%, the processing speed of abnormal data has increased by 4 times, and the model iteration cycle has been shortened by 60%, providing a new generation of intelligent solutions for data governance in various industries. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 It is a technical architecture diagram of the classification and grading system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0078] A method for realizing data classification and grading based on large language model technology includes the following steps:
[0079] Step 1, collect relevant data from multiple data sources;
[0080] The specific method of step 1 is as follows:
[0081] Step 1.1, determine the data sources from which data needs to be collected, including: databases and file systems;
[0082] Step 1.2, for different data sources: use SQL query statements to extract data from relational databases for databases; use file reading functions to collect the required data for file systems.
[0083] Step 2, preprocess the data collected in step 1, including: cleaning, denoising, and format conversion operations to form a standardized preprocessed data set;
[0084] The specific method of step 2 is as follows:
[0085] Step 2.1, use a data cleaning algorithm to remove missing values, duplicate values, and outliers from the data collected in step 1 to obtain a preliminary cleaned data set;
[0086] Step 2.2: Apply filtering algorithms and statistical methods to the preliminarily cleaned dataset obtained in Step 2.1 to remove noise and detect and process outliers, specifically including:
[0087] Smooth the time series data through the moving average filtering algorithm (time window set to 3 data points), process the volatile data using the exponentially weighted moving average (α coefficient set to 0.5), and apply Gaussian filtering (σ parameter set to 1.0) to eliminate random noise;
[0088] Use the Z-score method (threshold set to ±2.5 standard deviations) to identify numerical outliers, and adopt the interquartile range (IQR) method (expansion coefficient set to 1.5 times) to detect outliers in non-normal distribution data;
[0089] Step 2.3: Implement unified conversion for various formats of data: Generate semantic vectors for text data through the BERT model, uniformly convert the date format to the ISO 8601 standard format, and uniformly desensitize sensitive data, finally generating a standardized preprocessed dataset.
[0090] Step 3: Obtain text data from the standardized preprocessed dataset obtained in Step 2, extract the features of word frequency, TF-IDF value, and semantic vector to obtain a comprehensive feature vector;
[0091] The specific method of Step 3 is as follows:
[0092] Step 3.1: Use the natural language processing toolkit NLTK library to obtain text data from the standardized preprocessed dataset obtained in Step 2 and perform segmentation to obtain word and phrase units; perform word frequency (TF) statistics on the word and phrase units;
[0093] Step 3.2: Based on the word frequency statistics results in Step 3.1, first, calculate the inverse document frequency (IDF) value of each word and phrase unit, that is, the general importance of the word and phrase unit in the entire document collection; second, multiply the word frequency (TF) value by the inverse document frequency (IDF) value to obtain the frequency of each word in a single document and its distribution in the entire document collection, that is, the TF-IDF value;
[0094] Step 3.3: Use the BERT model to generate semantic vectors of the text, and use the TF-IDF value in Step 3.2 as the weight to perform weighted averaging on the semantic vectors to obtain a comprehensive feature vector.
[0095] Step 4: Randomly select a part of the standardized preprocessed dataset formed in Step 2 as the training set data, and use the remaining part as the test set data; perform manual category annotation on the divided training set data and test set data to obtain the labeled training set data and test set data;
[0096] Step 5: Use the comprehensive feature vectors obtained in Step 3 and the labeled training set in Step 4 to train a classification and grading model; the classification and grading model is a large Transformer model, including but not limited to: Qwen1.5-72B, Qwen2-7B, Baichuan-13B, ChatGLM3-6B.
[0097] The specific method of Step 5 is as follows:
[0098] Step 5.1: Clean the comprehensive feature vectors obtained in Step 3, remove outliers and missing values, and perform data standardization or normalization to ensure that their format and structure meet the input requirements of the selected classification and grading model, obtaining the processed comprehensive feature vector data.
[0099] Step 5.2: Integrate the labeled training set in Step 4 and the processed comprehensive feature vector data in Step 5.1 according to the sample-feature-label correspondence.
[0100] Step 5.3: Set parameters for the classification and grading model algorithm: learning rate (such as 2e-5), batch size (such as 32), number of training rounds (3-5 rounds).
[0101] Step 5.4: Use the integrated feature vector data and the labeled training set in Step 5.2 to train the classification and grading model with the parameters set in Step 5.3. During the training process, the classification and grading model learns the mapping relationship between features and class labels.
[0102] Step 5.5: Monitor the error or loss function value during the training process in Step 5.4 to evaluate the learning effect of the classification and grading model. If the classification and grading model shows overfitting or underfitting, improve the model by adjusting the regularization parameter, increasing data augmentation, and optimizing feature combinations.
[0103] Step 6: Input the labeled test set in Step 4 into the classification and grading model trained in Step 5 for testing to obtain prediction results. Compare the prediction results of the classification and grading model with the labeled test set in Step 4 to evaluate the accuracy and performance of the classification and grading model trained in Step 5, and select the optimal classification and grading model.
[0104] The specific method of Step 6 is as follows:
[0105] Step 6.1: Input the labeled test set in Step 4 into the classification and grading model trained in Step 5 for testing to obtain prediction results, and compare the prediction results with the labeled test set in Step 4.
[0106] Step 6.2, based on the comparison results in Step 6.1, calculate evaluation metrics, including: Accuracy, Precision, Recall, and F1-score. The calculation methods are as follows:
[0107] Accuracy = (Number of correctly predicted samples / Total number of samples) × 100%
[0108] Precision = (True positives / (True positives + False positives)) × 100%
[0109] Recall = (True positives / (True positives + False negatives)) × 100%
[0110] F1-score = 2 × (Precision × Recall) / (Precision + Recall);
[0111] Step 6.3, draw a Confusion Matrix to show the prediction situation of the classification and grading model for different categories, which complements the evaluation metrics in Step 6.2 to obtain the evaluation results;
[0112] Step 6.4, according to the evaluation results obtained in Step 6.3, analyze whether the classification and grading model has overfitting or underfitting for certain categories; if so, adjust the parameters of the classification and grading model, increase the amount of training data, improve the feature extraction method, and retrain and evaluate again; if not, select the optimal classification and grading model according to the evaluation results and actual requirements.
[0113] Step 7, input the data to be classified and graded into the optimal classification and grading model selected in Step 6 to obtain the classification and grading results.
[0114] The present invention provides a data classification and grading system based on large language model technology, including:
[0115] A data collection module for collecting relevant data from various data sources in Step 1;
[0116] A data processing module for preprocessing the data collected in Step 1 in Step 2, including cleaning, denoising, and format conversion operations to form a standardized preprocessed data set;
[0117] A feature extraction module for obtaining text data from the standardized preprocessed data set in Step 2 in Step 3, extracting features such as word frequency, TF-IDF value, and semantic vector to obtain a comprehensive feature vector;
[0118] A data annotation module, which is used to implement randomly selecting a part of the standardized data set preprocessed in Step 2 in Step 4 as the training set data, and using the remaining part as the test set data; performing manual category annotation on the divided training set data and test set data to obtain the annotated training set data and test set data;
[0119] A model training module, which is used to implement training a classification and grading model in Step 5 using the comprehensive feature vectors obtained in Step 3 and the training set annotated in Step 4;
[0120] A model evaluation module, which is used to implement inputting the test set annotated in Step 4 into the classification and grading model trained in Step 5 for testing to obtain a prediction result, comparing the prediction result of the classification and grading model with the test set annotated in Step 4, evaluating the accuracy and performance of the classification and grading model trained in Step 5, and selecting the optimal classification and grading model;
[0121] A data classification and grading module, which is used to implement inputting the data to be classified and graded in Step 7 into the optimal classification and grading model selected in Step 6 to obtain a classification and grading result.
[0122] As Figure 1 shown, a data classification and grading system implemented based on large language model technology has the following technical architecture:
[0123] Data layer: The data layer mainly includes two steps: data collection and data preprocessing:
[0124] 1. Data collection: It is used to implement collecting relevant data from multiple data sources in Step 1, including structured data, text data, audio data, video data, and image data.
[0125] 2. Data preprocessing: It is used to implement preprocessing the data collected in Step 1 in Step 2, including: cleaning, denoising, and format conversion operations to form a standardized preprocessed data set.
[0126] (1) Data cleaning: It is used to implement removing missing values, duplicate values, and outliers in Step 2.1.
[0127] (2) Text data processing: It is used to implement text data segmentation, term frequency (TF) statistics, etc. in Step 3.1.
[0128] (3) Data desensitization: It is used to implement unified desensitization of sensitive data in Step 2.3 to protect sensitive information.
[0129] (4) Data marking: It is used to implement manual category annotation on the divided training set and test set in Step 4 to obtain the annotated training set and test set.
[0130] (5) Data splitting: It is used to randomly select a part of the standardized preprocessed dataset formed in step 2 in step 4 as the training set data and divide it into a training set and a test set according to a ratio.
[0131] Model layer: It mainly involves feature engineering and the Transformer large language model:
[0132] 1. Feature engineering:
[0133] (1) Feature selection and construction: It is used to obtain text data from the standardized preprocessed dataset obtained in step 3 in step 3, extract features such as word frequency, TF-IDF values, and semantic vectors, and obtain a comprehensive feature vector.
[0134] (2) Feature scaling and standardization: It is used to perform data standardization or normalization processing in step 5.1.
[0135] (3) Dimensionality reduction and optimization: It is used to optimize the dimensions of the comprehensive feature vector obtained in step 3 to improve the model training efficiency.
[0136] (4) Vectorization: It is used to generate semantic vectors of text using the BERT model in step 3.3, take the TF-IDF values in step 3.2 as weights, and perform weighted averaging on the semantic vectors to obtain a comprehensive feature vector.
[0137] 2. Transformer large language model: It is used to select a suitable large Transformer model, such as Qwen2-7 / 72B, in step 5 for training the classification and grading model.
[0138] Optimization layer: It is mainly responsible for post-processing and classification and grading:
[0139] 1. Post-processing:
[0140] (1) Result verification and denoising: It is used to input the labeled test set in step 4 into the classification and grading model trained in step 5 through step 6 for testing, obtain the prediction result, and perform result verification and denoising.
[0141] (2) Category merging and subdivision: It is used to adjust the parameters of the classification and grading model and improve the feature extraction method in step 6.4.
[0142] (3) Adjustment of grading criteria: It is used to adjust the parameters of the classification and grading model and improve the feature extraction method in step 6.4.
[0143] (4) Enhancement of interpretability: It is used to monitor the error during the training process in step 5.4 in step 5.5, adjust parameters, increase the amount of data, and perform feature engineering measures to improve the model.
[0144] (5) Feedback iteration: Used to implement Step 6.4. According to the evaluation results obtained in Step 6.3, analyze whether the classification and grading model has overfitting or underfitting for certain categories; if so, adjust the parameters of the classification and grading model, increase the amount of training data, improve the feature extraction method, and retrain and evaluate; if not, select the optimal classification and grading model according to the evaluation results and actual requirements.
[0145] 2. Classification and grading:
[0146] (1) Result analysis: Used to implement Step 6.1. Input the test set labeled in Step 4 into the classification and grading model trained in Step 5 for testing to obtain the prediction results, and compare the parsed prediction results with the test set labeled in Step 4.
[0147] (2) Feature mapping: Used to implement Step 5.4. Use the feature vector data integrated in Step 5.2 and the labeled training set to train the classification and grading model set in Step 5.3. During the training process, the classification and grading model learns the relationship between features and class labels.
[0148] (3) Rule application: Apply specific rules for multi-classification fusion, and identify and handle abnormal situations.
[0149] (4) Multi-classification fusion: Used during the classification and grading process to integrate multi-dimensional features and classification rules and output comprehensive classification results.
[0150] (5) Abnormal handling: Used to implement Step 6.4. According to the evaluation results obtained in Step 6.3, analyze whether the classification and grading model has overfitting or underfitting for certain categories; if so, adjust the parameters of the classification and grading model, increase the amount of training data, improve the feature extraction method, and retrain and evaluate.
[0151] Application layer:
[0152] 1. Application interface
[0153] Mainly achieve application integration and decision support through functions such as API design and definition, permission verification and authorization, data encapsulation and conversion, API monitoring, and API version management:
[0154] (1) API design and definition: Define the API interface to ensure the secure transmission of data.
[0155] (2) Permission verification and authorization: Ensure that only authorized users can access specific data and functions.
[0156] (3) Data encapsulation and conversion: Encapsulate the data into a standard format for easy interaction between different systems.
[0157] (4) API Monitoring: Real-time monitoring of the running status of the API to ensure the stable operation of the system.
[0158] (5) API Version Management: Manage different versions of the API to ensure compatibility between old and new versions.
[0159] 2. Storage and Management
[0160] (1) Data Storage Planning: Store structured data in a relational database and unstructured data in a distributed file system.
[0161] (2) Data Writing and Updating: Support batch and real-time writing, and trace data updates through version management to ensure consistency.
[0162] (3) Data Index Establishment: Establish database indexes for structured data and construct vector indexes for unstructured data to optimize the retrieval rate.
[0163] (4) Data Resource Management: Establish a data resource catalog, label metadata, and visually manage data assets for easy retrieval and invocation.
[0164] (5) Data Security and Permission Control: Encrypt and store sensitive data, control access based on a role-based permission model, and record logs for security auditing.
[0165] Examples
[0166] Experimental Environment
[0167] Hardware Configuration: 1 NVIDIA A100 GPU, Intel Xeon Platinum 8375C CPU, 256GB of memory.
[0168] Software Framework: Python 3.10, PyTorch 2.1.0, Hugging Face Transformers 4.31.0.
[0169] Large Language Models: Qwen2-7B (open source version), BERT-base-uncased.
[0170] Dataset: Mixed dataset (including classified and graded specification texts, database table structure data), with a total of 100,000 samples, divided into 10 classification levels.
[0171] Experimental Steps:
[0172] 1. Data Collection and Preprocessing
[0173] Extract transaction records (structured data) from a MySQL database and read user comments (text data) from a CSV file.
[0174] Process the time-series data noise using the filtering algorithm in Step 2.2 (such as Gaussian filtering with σ = 1.5), and detect outliers using the IQR method (expansion coefficient 2.0).
[0175] Convert the text data into 768-dimensional semantic vectors through the BERT model, unify the date format to ISO 8601, and desensitize sensitive information (such as replacing the ID number with "***").
[0176] 2. Feature Extraction
[0177] Tokenize the preprocessed text data using NLTK and calculate the TF-IDF values (stop word filtering, maximum number of features 5000).
[0178] Combine the BERT semantic vectors with the TF-IDF weights (weight coefficient 0.6) to generate comprehensive feature vectors (768-dimensional).
[0179] 3. Model Training
[0180] Divide the dataset: 80% training set (80,000 samples), 20% test set (20,000 samples), and manually annotate the classification levels (1 - 10 levels).
[0181] Select the Qwen2-7B model and set the parameters: learning rate 1e-5, batch size 64, and number of training epochs 5.
[0182] Apply the regularization strategy in Step 5.5 (L2 regularization coefficient 0.01), and adopt the early stopping mechanism (patience value 3) to prevent overfitting.
[0183] 4. Model Evaluation
[0184] Calculate the evaluation metrics using the test set: accuracy 94.2%, F1-score 93.5%, and recall 94.8%.
[0185] Plot the confusion matrix and find that the recall rate for "high-risk scientific research projects" (level 10) reaches 97.3%, but the precision rate for "teaching and research reports" (level 6) is only 89.2%, which is improved to 93.1% after adjusting the class weights through Step 6.4.
[0186] Comparative Experiments
[0187] Comparative Model 1: Rule-based classifier (regular expression matching keywords), accuracy 78.6%, ineffective for complex semantic classification.
[0188] Comparative Model 2: BERT-base-uncased model (without TF-IDF weighting), F1-score 90.1%, with a 12% decrease in long text classification performance.
[0189] Method of the present invention: Comprehensive features + Qwen2-7B, with the accuracy rate increased by 15.6% and the multi-modal data processing speed increased by 40%.
[0190] Result analysis
[0191] Advantages of multi-modal: Experiments on the mixed dataset show that the fusion of structured data (such as teaching experiment funds) and text semantic vectors improves the classification accuracy rate by 8.7%.
[0192] Effect of dynamic optimization: Through the confusion matrix analysis in step 6.4, after adding 20% labeled samples for "medical imaging reports", the F1 value of this category increases from 89.2% to 93.1%.
[0193] Performance efficiency: The classification time for a single piece of data is 0.04 seconds (A100 GPU), supporting the processing of millions of data per day.
[0194] The key points and points to be protected in the present invention are:
[0195] (1) The adaptive ability to automatically adapt to new data types and features, capable of self-optimizing according to the dynamic changes and new patterns of data.
[0196] (2) The multi-dimensional analysis ability to classify and grade by comprehensively considering multiple dimensions and attributes of data, making the results more comprehensive and accurate.
[0197] (3) The ability to reduce human errors, reducing classification and grading errors caused by human judgment mistakes.
[0198] (4) The ability to significantly reduce labor costs, improving work efficiency by virtue of the advantages of automation and intelligence.
[0199] (5) The ability to perform forward-looking predictive classification and grading based on rich historical data and learning results.
[0200] (6) The ability to deeply integrate knowledge and experience in multiple fields, improving the scientificity and accuracy of classification and grading.
[0201] (7) The specific implementation methods and functions of the model layer, data layer, and optimization layer in the technical architecture, as well as the interaction methods between the application layer and other layers.
[0202] (8) The specific steps and methods of data preprocessing, large model training and optimization, classification and grading service engine, and application integration and decision support in the implementation process.
[0203] Glossary of relevant technical terms:
[0204] Transformer: A Transformer is a deep learning model architecture based on the attention mechanism. It has achieved remarkable results in natural language processing tasks, with advantages such as strong parallel computing capabilities and the ability to capture long-range dependencies. A Transformer usually consists of an encoder and a decoder, and encodes and decodes the input sequence through the multi-head attention mechanism to achieve language understanding and generation.
[0205] Large language model: A large language model is a language model based on deep learning. It is trained using large-scale text data to learn the statistical laws and semantic representations of language. Large language models have powerful language understanding and generation capabilities, can generate natural and fluent text, answer various questions, conduct conversations, etc. Common large language models include GPT, BERT, etc.
[0206] Classification and grading: Classification and grading is a data management method aimed at dividing data into different categories and levels according to the characteristics, attributes, or uses of the data. Classification is to group data according to certain criteria or rules so that data in the same category has similar characteristics or attributes. Grading is to further divide the data on the basis of classification, usually dividing the data into different levels according to factors such as the importance, sensitivity, or security of the data. The purpose of classification and grading is to better manage and protect data and improve the availability and security of data.
Claims
1. A method for implementing data classification and grading based on large language model technology, characterized in that, It includes the following steps: Step 1: Collect relevant data from multiple data sources; Step 2: Preprocess the data collected in Step 1, including: cleaning, denoising, and format conversion operations to form a standardized preprocessed dataset; Step 3: Obtain text data from the standardized preprocessed dataset obtained in Step 2, extract features such as word frequency, TF-IDF value, and semantic vector to obtain a comprehensive feature vector; Step 4: Randomly select a part of the standardized preprocessed dataset formed in Step 2 as the training set data, and use the remaining part as the test set data; perform manual category annotation on the divided training set data and test set data to obtain the labeled training set data and test set data; Step 5: Use the comprehensive feature vector obtained in Step 3 and the labeled training set in Step 4 to train a classification and grading model; Step 6: Input the labeled test set in Step 4 into the classification and grading model trained in Step 5 for testing to obtain a prediction result. Compare the prediction result of the classification and grading model with the labeled test set in Step 4 to evaluate the accuracy and performance of the classification and grading model trained in Step 5, and select the optimal classification and grading model; Step 7: Input the data to be classified and graded into the optimal classification and grading model selected in Step 6 to obtain the classification and grading result.
2. The method for implementing data classification and grading based on large language model technology according to claim 1, wherein, The specific method of Step 1 is as follows: Step 1.1: Determine the data sources from which data needs to be collected, including: databases and file systems; Step 1.2: For different data sources: use SQL query statements to extract data from relational databases for databases; use file reading functions to collect the required data for file systems.
3. A method for implementing data classification and grading based on large language model technology according to claim 1, characterized in that, The specific method of Step 2 is as follows: Step 2.1: Use a data cleaning algorithm to remove missing values, duplicate values, and outliers in the data collected in Step 1 to obtain a preliminary cleaned dataset; Step 2.2: Apply a filtering algorithm and statistical methods to the preliminary cleaned dataset obtained in Step 2.1 to remove noise and detect and process outliers, specifically including: Smooth time series data through a moving average filtering algorithm, process volatile data using an exponentially weighted moving average, and apply Gaussian filtering to eliminate random noise; Use the Z-score method to identify numerical outliers and the interquartile range (IQR) method to detect outliers in non-normally distributed data; Step 2.3: Implement unified conversion for various formats of data: generate semantic vectors for text data through the BERT model, uniformly convert the date format to the ISO 8601 standard format, and uniformly desensitize sensitive data to finally generate a standardized preprocessed dataset.
4. A method for implementing data classification and grading based on large language model technology according to claim 1, characterized in that, The specific method of Step 3 is as follows: Step 3.1: Use the natural language processing toolkit NLTK library to obtain text data from the standardized preprocessed dataset obtained in Step 2 and perform segmentation to obtain word and phrase units; perform word frequency (TF) statistics on the word and phrase units; Step 3.2: Based on the word frequency statistics result in Step 3.1, first, calculate the inverse document frequency (IDF) value of each word and phrase unit, that is, the general importance of the word and phrase unit in the entire document collection; Secondly, by multiplying the term frequency (TF) value with the inverse document frequency (IDF) value, the frequency of each word appearing in a single document and its distribution across the entire document collection are obtained, i.e., the TF-IDF value; Step 3.3: Use the BERT model to generate semantic vectors of the text. Take the TF-IDF values from Step 3.2 as weights and perform weighted averaging on the semantic vectors to obtain comprehensive feature vectors.
5. A method for implementing data classification and grading based on large language model technology according to claim 1, characterized in that, The specific method for Step 5 is as follows: Step 5.1: For the comprehensive feature vectors obtained in Step 3, perform data cleaning to remove outliers and missing values, and perform data standardization or normalization processing to ensure that their format and structure meet the input requirements of the selected classification and grading model, resulting in processed comprehensive feature vector data; Step 5.2: Integrate the labeled training set from Step 4 and the processed comprehensive feature vector data from Step 5.1 according to the sample-feature-label correspondence relationship; Step 5.3: For the classification and grading model algorithm, set parameters: learning rate, batch size, number of training rounds; Step 5.4: Use the integrated feature vector data from Step 5.2 and the labeled training set to train the classification and grading model with the parameters set in Step 5.
3. During the training process, the classification and grading model learns the mapping relationship between features and class labels; Step 5.5: Monitor the error or loss function value during the training process in Step 5.4 to evaluate the learning effect of the classification and grading model. If the classification and grading model shows overfitting or underfitting, improve the model by adjusting the regularization parameter, increasing data augmentation, and optimizing feature combinations.
6. A method for implementing data classification and grading based on large language model technology according to claim 1, characterized in that, The specific method for Step 6 is as follows: Step 6.1: Input the labeled test set from Step 4 into the classification and grading model trained in Step 5 for testing to obtain prediction results, and compare the prediction results with the labeled test set from Step 4; Step 6.2: Based on the comparison results in Step 6.1, calculate evaluation metrics, including: Accuracy, Precision, Recall, and F1-score. The calculation methods are as follows: Accuracy = (Number of correctly predicted samples / Total number of samples) × 100% Precision = (True positives / (True positives + False positives)) × 100% Recall = (True positives / (True positives + False negatives)) × 100% F1-score = 2 × (Precision × Recall) / (Precision + Recall); Step 6.3: Plot a confusion matrix (Confusion Matrix) to show the prediction situation of the classification and grading model for different classes, which complements the evaluation metrics in Step 6.2 to obtain the evaluation results; Step 6.4: According to the evaluation results obtained in Step 6.3, analyze whether the classification and grading model has overfitting or underfitting for certain classes; if so, adjust the parameters of the classification and grading model, increase the amount of training data, improve the feature extraction method, and retrain and evaluate; if not, select the optimal classification and grading model based on the evaluation results and actual requirements.
7. A data classification and grading system implemented based on large language model technology, characterized in that, Including: A data collection module for collecting relevant data from various data sources; A data processing module for preprocessing the collected data, including cleaning, denoising, and format conversion operations to form a standardized preprocessed dataset; A feature extraction module for obtaining text data from the standardized preprocessed dataset, extracting features such as word frequency, TF-IDF values, and semantic vectors to obtain a comprehensive feature vector; A data annotation module for randomly selecting a part of the preprocessed standardized dataset as training set data and using the remaining part as test set data; performing manual category annotation on the divided training set data and test set data to obtain the annotated training set data and test set data; A model training module for training a classification and grading model using the comprehensive feature vector and the annotated training set; A model evaluation module for inputting the annotated test set into the trained classification and grading model for testing to obtain a prediction result, comparing the prediction result of the classification and grading model with the annotated test set, evaluating the accuracy and performance of the trained classification and grading model, and selecting the optimal classification and grading model; A data classification and grading module for inputting the data to be classified and graded into the optimal classification and grading model to obtain a classification and grading result.
Citation Information
Patent Citations
Horizontal tunnel anchoring main cable erecting method
CN116732895A
Method for automatically generating data classification and grading template
CN117851860A
Cited By
Internet meteorological data resource dynamic discovery method and system fused with large language model
CN121478971A
Material main data intelligent classification method
CN121561644A