Method and system for automatically classifying and grading data

By obtaining and classifying data in the data processing system of the new enterprise, generating hierarchical coefficients, and using the recurrent neural network model for training, the problem of data processing system inability to be effectively applied due to the lack of historical data in the new enterprise is solved, and the effective classification and grading of data is achieved, which improves processing efficiency and analysis accuracy.

CN120067786AInactive Publication Date: 2025-05-30ZHILIN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510012398.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing data processing systems cannot be effectively applied in newly established enterprises because there is a lack of large amounts of historical data to train pre-models.

Method used

By obtaining existing data in the device, classifying processing based on metadata and file suffix names, hierarchical coefficients are generated, and training and testing is used for recurrent neural network models to automatically classify and classify data.

Benefits of technology

It realizes effective data classification and grading in new enterprises, has better applicability, and improves data processing efficiency, providing more comprehensive and accurate data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067786A_ABST
    Figure CN120067786A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic data classification and grading method and system, and relates to the technical field of data processing, file contents, metadata and business rules in each classification are scanned, after the file contents, the metadata and the business rules are comprehensively analyzed through an automatic algorithm, grading coefficients are generated for data, and in each data classification, the file contents, the metadata and the business rules are classified. And performing multi-stage division on the data according to a comparison result of a classification coefficient and a classification gradient threshold, periodically obtaining a classification result of the data and a classification result in each classification, after a period of time, taking the classification result and the classification result in each classification as a data set, training and testing a recurrent neural network model, and obtaining a classification result of the data. And deploying the trained recurrent neural network model into equipment. The grading system can comprehensively analyze file contents, metadata and business rules through an automatic algorithm and then generate grading coefficients for the data, so that the data in a new enterprise is effectively classified and graded, and the applicability is better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a method and system for automatically classifying and grading data. Background Art

[0002] A data classification and grading system is a system used to organize and manage data, and its goal is to classify and grade data according to certain criteria so as to store, retrieve, and protect such data more effectively. Such systems are usually widely used in organizations, enterprises, government agencies, and other organizations to ensure the security, availability, and compliance of data;

[0003] The prior art has the following deficiencies: Existing data processing systems usually rely on pre-trained models to classify and grade data. However, when a data processing system is applied to a newly established enterprise, since the enterprise does not have a large amount of historical data for model training, it will cause the data processing system to be unable to be effectively applied in the early stage of enterprise development, and there are great limitations in use. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for automatically classifying and grading data to solve the deficiencies in the background art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A method for automatically classifying and grading data, the grading method includes the following steps:

[0006] The system obtains the existing data in the device and classifies the data based on the metadata of the data and the data file suffix name;

[0007] After obtaining the classification result of the data, by scanning the file content, metadata, and business rules in each classification, and comprehensively analyzing the file content, metadata, and business rules through an automated algorithm, a grading coefficient is generated for the data;

[0008] In each data classification, the data is divided into multiple levels according to the comparison result between the grading coefficient and the grading gradient threshold;

[0009] Regularly obtain the classification result of the data and the grading result in each classification. After a period of time, use the classification result and the grading result in each classification as a data set to train and test a recurrent neural network model;

[0010] Deploy the trained recurrent neural network model to the device. When new data is generated in the subsequent device, the trained recurrent neural network model automatically classifies and grades the data.

[0011] In a preferred embodiment, the system obtains the existing data in the device and classifies the data based on the metadata of the data and the data file suffix name, including the following steps:

[0012] Collect the existing data from the device, including files, documents, images, and audio. Obtain the metadata of the data, including the creation time, modification time, and file size. Extract the suffix name of each data file, map different suffix names to specific data categories, analyze the metadata of the data, including the file size, creation time, and modification time, classify the data based on the file suffix name and metadata, and based on the data classification result, formulate corresponding access control policies by moving the data files to the corresponding folders, modifying the file names, and adding tags.

[0013] In a preferred embodiment, after obtaining the data classification result, scan the file content, metadata, and business rules in each classification, including the following steps:

[0014] For text files, scan the file content through text parsing technology, analyze the metadata of the file, including the creation time, modification time, and file size, further process the file content and metadata based on the business rules, extract key information from the file, and add tags to the file according to the extracted information. Archive or clean the file according to the business rules and the data life cycle management policy.

[0015] In a preferred embodiment, after comprehensively analyzing the file content, metadata, and business rules through an automated algorithm, generate a grading coefficient for the data, including the following steps:

[0016] Obtain the access frequency, data correlation degree, data volume, and security alarm frequency of the data;

[0017] Obtain the access frequency, data correlation degree, data volume, and security alarm frequency of the data;

[0018] Obtain the grading coefficient by comprehensively calculating the access frequency, data correlation degree, data volume, and security alarm frequency through an automated algorithm. The expression is:

[0019]

[0020] In the formula, fjx is the grading coefficient, fwc is the number of accesses, ΔT1 is the first monitoring period, aqj is the number of alarms, ΔT2 is the second monitoring period, sjl is the data volume, gld is the data correlation degree, and a1, a2, and a3 are the proportionality coefficients of the access frequency, security alarm frequency, and data correlation degree respectively, and a1 > a2 > a3, 0 < a1 ≤ 1, 0 < a2 ≤ 1, 0 < a2 ≤ 1.

[0021] In a preferred embodiment, in each data classification, the data is divided into multiple levels according to the comparison result of the classification coefficient and the classification gradient threshold, including the following steps:

[0022] The grading gradient threshold includes a first grading threshold and a second grading threshold. The larger the grading coefficient, the greater the importance of the data. After obtaining the grading coefficient, the grading coefficient is compared with the first grading threshold and the second grading threshold.

[0023] If the classification coefficient is ≥ the second classification threshold, the importance of the classified data is high, and the classified data is placed in the first-level area;

[0024] If the classification coefficient is less than the second classification threshold, and the classification coefficient is greater than or equal to the first classification threshold, the importance of the classified data is medium, and the classified data is placed in the second level area;

[0025] If the classification coefficient is less than the first classification threshold, the importance of the classified data is low, and the classified data is classified into the third level area.

[0026] In a preferred embodiment, the classification results and the grading results in each classification are used as data sets to train and test a recurrent neural network model, including the following steps:

[0027] The classification results and the grading results in each classification are used as training and testing data sets. The data are preprocessed, including standardization, normalization, and processing of missing values. The data set is divided into training and test sets. The RNN model is constructed using a deep learning framework. The structure of the model is defined, including the input layer, hidden layer, and output layer. A loss function is selected to measure the performance of the model, and an optimizer is selected to update the model parameters. The model is trained using the training set, iterated for multiple cycles, and the model parameters are updated through the back propagation algorithm to minimize the loss function. The performance of the model is verified using the test set. The generalization ability of the model on unseen data is evaluated. According to the verification results, the parameters of the model, including the learning rate, number of layers, and number of neurons, are adjusted.

[0028] In a preferred embodiment, the data volume acquisition logic is: using the attribute information of the file system or database to obtain the size of each data file, and using the monitoring tool of the storage system to track the storage capacity of the entire data set in real time;

[0029] The logic for obtaining the frequency of security alarms is as follows: enable security event log recording, record system security events, including potential threats and attack attempts, regularly analyze security event logs, and count the frequency of security alarms related to each data file.

[0030] In a preferred embodiment, the logic for obtaining the access frequency of data is as follows: Enable access log recording in the system to record the time, user, and file information of each data access, regularly analyze the access log, and count the access frequency of each data;

[0031] The logic for obtaining the data correlation degree is as follows: Identify the correlation information in the data, and obtain the quantity or identifier of other data associated with specific data by querying the system or database.

[0032] The present invention also provides a data automatic classification and grading system, including a data acquisition module, a data classification module, a coefficient generation module, a data grading module, a data collection module, and a model training module;

[0033] Data acquisition module: Obtain the existing data in the device;

[0034] Data classification module: Classify the data based on the metadata of the data and the data file suffix name;

[0035] Coefficient generation module: By scanning the file content, metadata, and business rules in each classification, comprehensively analyze the file content, metadata, and business rules through an automated algorithm, and generate a grading coefficient for the data;

[0036] Data grading module: In each data classification, perform multi-level division of the data according to the comparison result between the grading coefficient and the grading gradient threshold;

[0037] Data collection module: Regularly obtain the classification results of the data and the grading results in each classification;

[0038] Model training module: After a period of time, use the classification results and the grading results in each classification as a data set to train and test a recurrent neural network model;

[0039] Automatic processing module: Deploy the trained recurrent neural network model to the device. When new data is generated in the subsequent device, automatically classify and grade the data through the trained recurrent neural network model.

[0040] In the above technical solution, the technical effects and advantages provided by the present invention:

[0041] 1. The present invention scans the file content, metadata, and business rules in each classification. After comprehensively analyzing the file content, metadata, and business rules through an automated algorithm, a grading coefficient is generated for the data. In each data classification, the data is divided into multiple levels based on the comparison result between the grading coefficient and the grading gradient threshold. The classification results of the data and the grading results in each classification are regularly obtained. After a period of time, the classification results and the grading results in each classification are used as a data set to train and test a recurrent neural network model, and then the trained recurrent neural network model is deployed to the device. This grading system can generate a grading coefficient for the data by comprehensively analyzing the file content, metadata, and business rules through an automated algorithm, thereby effectively classifying and grading the data in a new enterprise, with better applicability;

[0042] 2. The present invention obtains the access frequency, data correlation degree, data volume, and security alarm frequency of the data, and comprehensively calculates the access frequency, data correlation degree, data volume, and security alarm frequency through an automated algorithm to obtain a grading coefficient, which not only effectively improves the processing efficiency of the data, but also analyzes the data more comprehensively and accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0044] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0046] Embodiment 1: Please refer to Figure 1 As shown, a method for automatically classifying and grading data in this embodiment includes the following steps:

[0047] The system obtains the existing data in the device, classifies the data based on the metadata of the data and the data file suffix, and after obtaining the classification result of the data, by scanning the file content, metadata, and business rules in each classification, and comprehensively analyzing the file content, metadata, and business rules through an automated algorithm, a grading coefficient is generated for the data. And in each data classification, the data is divided into multiple levels according to the comparison result between the grading coefficient and the grading gradient threshold. Regularly obtain the classification result of the data and the grading result in each classification. After a period of time, use the classification result and the grading result in each classification as a data set, train and test a recurrent neural network model, and then deploy the trained recurrent neural network model to the device. When new data is generated in the subsequent device, the trained recurrent neural network model automatically classifies and grades the data.

[0048] In this application, by scanning the file content, metadata, and business rules in each classification, and comprehensively analyzing the file content, metadata, and business rules through an automated algorithm, a grading coefficient is generated for the data. And in each data classification, the data is divided into multiple levels according to the comparison result between the grading coefficient and the grading gradient threshold. Regularly obtain the classification result of the data and the grading result in each classification. After a period of time, use the classification result and the grading result in each classification as a data set, train and test a recurrent neural network model, and then deploy the trained recurrent neural network model to the device. This grading system can generate a grading coefficient for the data by comprehensively analyzing the file content, metadata, and business rules through an automated algorithm, so as to effectively classify and grade the data in a new enterprise, and has better applicability.

[0049] Embodiment 2: The system obtains the existing data in the device and classifies the data based on the metadata of the data and the data file suffix, including the following steps:

[0050] Data collection: Collect the existing data from the device, including files, documents, images, audio, etc. Ensure that the metadata of the obtained data, such as creation time, modification time, file size, etc., is obtained.

[0051] File suffix extraction: Extract the suffix of each data file. The suffix usually reflects the type of the file. For example,.docx represents a document and.jpg represents an image, etc.

[0052] Formulate suffix rules: Formulate a set of suffix rules to map different suffixes to specific data categories. For example,.xlsx and.csv are mapped to "spreadsheet data", and.jpg and.png are mapped to "image data".

[0053] Metadata analysis: Analyze the metadata of the data, including file size, creation time, modification time, etc. This metadata is used for further data classification and grading.

[0054] Classification standard formulation: Formulate the standards for data classification based on file suffix names and metadata. For example, files with a size exceeding a certain threshold are marked as "large files", and those created within a certain time range are marked as "new data".

[0055] Data classification: Classify the data according to the formulated rules and standards. This is achieved by moving data files to corresponding folders, modifying file names, adding tags, etc.

[0056] Access control policy: Based on the data classification results, formulate corresponding access control policies. For example, for sensitive data categories, set more stringent access permissions.

[0057] Lifecycle management: According to the data classification, formulate data lifecycle management policies. For example, for obsolete or no longer needed data categories, automatically archive or delete them.

[0058] Monitoring system implementation: Establish a monitoring system to track the data classification and processing process in real time. The monitoring system issues alerts to remind administrators to handle abnormal situations.

[0059] Regular review and update: Regularly review the data classification rules and standards, and update them according to changes in business requirements or new data types.

[0060] After obtaining the data classification results, scan the file content, metadata, and business rules in each classification, including the following steps:

[0061] File content scanning: For text files (such as documents, spreadsheets, text files, etc.), scan the file content through text parsing techniques. This includes using natural language processing (NLP) techniques to identify keywords, entities, topics, and other information.

[0062] Metadata analysis: Further analyze the metadata of the file, including creation time, modification time, file size, etc. Metadata provides more background information about the data.

[0063] Business rule application: Based on business rules, further process the file content and metadata. For example, for contract files, apply business rules to extract key contract terms and check compliance, etc.

[0064] Data extraction and tag application: Extract key information from the file and add tags to the file based on the extracted information. This helps with subsequent retrieval and organization.

[0065] Sensitive information detection: For files containing sensitive information, apply sensitive information detection algorithms to ensure the secure processing of sensitive information. This involves the detection of, for example, identity information, financial information, etc.

[0066] Filing and cleaning: Files are filed or cleaned according to business rules and data lifecycle management strategies. Obsolete or no-longer-needed data is automatically filed or deleted.

[0067] Access control policy update: The access control policy is updated based on the results of further file content analysis. For example, for files containing sensitive information, access permissions need to be adjusted.

[0068] Automation process: An automated processing process is established based on classification results, content analysis, and business rules. For example, for certain files, the workflow needs to be automatically triggered to notify relevant personnel or systems for further processing.

[0069] Monitoring and auditing: A monitoring system is established to track the file content scanning and processing process in real time. In addition, an auditing mechanism is implemented to record the key steps and results during the processing.

[0070] Regular review and optimization: The processing results are regularly reviewed, and the system is optimized according to business requirements and feedback. This includes updating business rules, improving text parsing algorithms, etc.

[0071] After comprehensively analyzing the file content, metadata, and business rules through an automated algorithm, a grading coefficient is generated for the data, including the following steps:

[0072] Obtain the access frequency, data correlation, data volume, and security alarm frequency of the data;

[0073] Obtain the access frequency, data correlation, data volume, and security alarm frequency of the data;

[0074] The grading coefficient is obtained by comprehensively calculating the access frequency, data correlation, data volume, and security alarm frequency through an automated algorithm. The expression is:

[0075]

[0076] In the formula, fjx is the grading coefficient, fwc is the number of accesses, ΔT1 is the first monitoring period, aqj is the number of alarms, ΔT2 is the second monitoring period, sjl is the data volume, gld is the data correlation, and a1, a2, and a3 are the proportionality coefficients of the access frequency, security alarm frequency, and data correlation respectively, and a1 > a2 > a3, 0 < a1 ≤ 1, 0 < a2 ≤ 1, 0 < a2 ≤ 1;

[0077] This application not only effectively improves the data processing efficiency by obtaining the access frequency, data correlation, data volume, and security alarm frequency of the data and comprehensively calculating the access frequency, data correlation, data volume, and security alarm frequency through an automated algorithm to obtain the grading coefficient, but also makes the data analysis more comprehensive and accurate.

[0078] The logic for obtaining the access frequency of data is as follows: enable access log recording in the system to record the time, user, file, and other information of each data access. Regularly analyze the access log, count the access frequency of each data, and use indicators such as the average number of accesses or access frequency within a time period.

[0079] The logic of obtaining data association is to identify association information in the data, such as unique identifiers, foreign keys, etc., in order to track the association between data. By querying the system or database, the number or identifier of other data associated with specific data is obtained.

[0080] The logic of obtaining data volume is: use the attribute information of the file system or database to obtain the size of each data file. Use the monitoring tool of the storage system to track the storage capacity of the entire data set in real time and understand the overall situation of the data volume.

[0081] The logic for obtaining the frequency of security alarms is as follows: Enable security event log recording to record system security events, including potential threats, attack attempts, etc. Regularly analyze the security event log, count the frequency of security alarms related to each data file, and use indicators such as the number of alarms or average alarm frequency within a time period.

[0082] In each data classification, the data is divided into multiple levels according to the comparison result between the classification coefficient and the classification gradient threshold, including the following steps:

[0083] The grading gradient threshold includes a first grading threshold and a second grading threshold. It can be seen from the calculation expression of the grading coefficient that the larger the grading coefficient is, the greater the importance of the data is. Therefore, after obtaining the grading coefficient, the grading coefficient is compared with the first grading threshold and the second grading threshold.

[0084] If the classification coefficient is ≥ the second classification threshold, the importance of the classified data is high, and the classified data is placed in the first-level area;

[0085] If the classification coefficient is less than the second classification threshold, and the classification coefficient is greater than or equal to the first classification threshold, the importance of the classified data is medium, and the classified data is placed in the second level area;

[0086] If the classification coefficient is less than the first classification threshold, the importance of the classified data is low, and the classified data is classified into the third level area.

[0087] The classification results of the data and the grading results in each category are obtained regularly, including the following steps. After a period of time, the classification results and the grading results in each category are used as data sets to train and test the recurrent neural network model, including the following steps:

[0088] Prepare the dataset: Use the classification results and the grading results within each classification as the training and testing datasets. Ensure that the label and feature formats of the dataset conform to the input requirements of the RNN model.

[0089] Data preprocessing: Preprocess the data, including standardization, normalization, handling missing values, etc. Ensure the consistency and usability of the input data.

[0090] Divide the training set and the test set: Divide the dataset into a training set and a test set. Usually, most of the data is used for training, and a small part is used to test the performance of the model.

[0091] Design the RNN model: Design the RNN model according to the characteristics of the problem and the structure of the dataset. This includes selecting appropriate number of layers, number of neurons, activation functions, etc.

[0092] Build the model: Build the RNN model using a deep learning framework (such as TensorFlow, PyTorch, etc.). Define the structure of the model, including the input layer, hidden layer, output layer, etc.

[0093] Select the loss function and optimizer: Select an appropriate loss function to measure the performance of the model, and select a suitable optimizer to update the model parameters. For classification problems, the cross-entropy loss function is usually used.

[0094] Train the model: Use the training set to train the model. Iterate through multiple epochs and update the model parameters through the backpropagation algorithm to minimize the loss function.

[0095] Validate the model: Use the test set to validate the performance of the model. Evaluate the generalization ability of the model on unseen data and check for overfitting or underfitting issues.

[0096] Adjust the model parameters: Adjust the parameters of the model according to the validation results. This includes the learning rate, number of layers, number of neurons, etc.

[0097] Repeat training and testing: If necessary, repeat the training and testing steps until a satisfactory performance level is achieved.

[0098] Deploy the trained recurrent neural network model to the device. When new data is generated in the subsequent device, automatically classify and grade the data through the trained recurrent neural network model, including the following steps:

[0099] Model export: After training is completed, export the trained model in a format suitable for device deployment. This includes saving the model parameters as a file or in a specific format, such as ONNX (Open Neural Network Exchange).

[0100] Model Integration: Integrate the exported model into the device's application. Adaptation to the device's operating system and hardware architecture is required.

[0101] Data Preprocessing: Implement the same data preprocessing steps on the device as during training to ensure that new data meets the input requirements of the model. This includes operations such as standardization and normalization.

[0102] Inference: Implement the inference process of the model on the device, i.e., use the model for prediction. Achieved by calling the model's inference interface or using the corresponding inference engine.

[0103] New Data Acquisition: The device acquires new data through various means, such as sensors, user input, etc. Ensure that the data is accessible and processed on the device.

[0104] Apply Model: Input the new data into the deployed model for classification and grading prediction. Obtain the output results of the model.

[0105] Result Processing: Perform corresponding result processing based on the output results of the model. This includes displaying the results to the user, triggering specific actions, or sending the results to a remote server, etc.

[0106] Real-time Monitoring: Implement a real-time monitoring mechanism after deployment to track the performance of the model in the actual environment. Monitor metrics such as the accuracy, performance, and memory usage of the model.

[0107] Remote Update (if needed): If there are model updates or performance improvements, update the deployed model through the remote update mechanism without physically accessing the device.

[0108] Exception Handling and Logging: Implement an exception handling mechanism to handle abnormal situations that occur during model inference. Record logs for subsequent analysis and troubleshooting.

[0109] Regular Maintenance and Update: Regularly check the performance of the model to ensure its generalization ability under different data distributions. Update the model, adjust parameters, or perform other maintenance work as needed.

[0110] Embodiment 3: A data automatic classification and grading system described in this embodiment includes a data acquisition module, a data classification module, a coefficient generation module, a data grading module, a data collection module, and a model training module;

[0111] Data Acquisition Module: The system acquires the existing data in the device;

[0112] Data Classification Module: Classify the data based on the metadata of the data and the data file suffix name. After obtaining the classification results of the data;

[0113] Coefficient generation module: By scanning the file content, metadata, and business rules in each classification, and comprehensively analyzing the file content, metadata, and business rules through an automated algorithm, a hierarchical coefficient is generated for the data;

[0114] Data classification module: And in each data classification, the data is divided into multiple levels according to the comparison result between the classification coefficient and the classification gradient threshold;

[0115] Data acquisition module: Regularly obtain the classification results of the data and the classification results in each classification;

[0116] Model training module: After a period of time, using the classification results and the classification results in each classification as a data set, after training and testing a recurrent neural network model;

[0117] Automatic processing module: Deploy the trained recurrent neural network model to the device. When new data is generated in the subsequent device, the data is automatically classified and graded through the trained recurrent neural network model.

[0118] The above formulas are all dimensionless and take their numerical values for calculation. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.

[0119] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described are combined in a suitable manner in any one or more embodiments or examples.

[0120] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not elaborate on all the details, nor do they limit the present invention to only the specific implementation manners. Obviously, many modifications and variations can be made according to the content of this specification. This specification selects and specifically describes some embodiments to better explain the principle and practical application of the present invention, so that those skilled in the art in the relevant technical field can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A method for automatic classification and grading of data, characterized by: The classification method comprises the following steps: The system obtains the existing data in the device and classifies the data based on the data metadata and data file suffixes; After obtaining the classification results of the data, the file content, metadata and business rules in each category are scanned, and the file content, metadata and business rules are comprehensively analyzed by an automated algorithm to generate a classification coefficient for the data; In each data classification, the data is divided into multiple levels according to the comparison results between the classification coefficient and the classification gradient threshold; Regularly obtain the classification results of the data and the grading results in each category. After a period of time, use the classification results and the grading results in each category as data sets to train and test the recurrent neural network model; The trained recurrent neural network model is deployed to the device. When new data is generated in the subsequent device, the trained recurrent neural network model is used to automatically classify and grade the data.

2. The method for automatic data classification and grading according to claim 1, characterized in that: The system obtains the existing data in the device and classifies the data based on the metadata of the data and the data file suffix, including the following steps: Collect existing data from the device, including files, documents, images, and audio, obtain the data's metadata, including creation time, modification time, and file size, extract the suffix of each data file, map different suffixes to specific data categories, analyze the data's metadata, including file size, creation time, and modification time, classify the data based on the file suffix and metadata, and formulate corresponding access control policies based on the data classification results by moving data files to corresponding folders, modifying file names, and adding tags.

3. The method for automatic data classification and grading according to claim 2, characterized in that: After obtaining the classification results of the data, the following steps are included by scanning the file content, metadata, and business rules in each classification: For text files, text parsing technology is used to scan the file content and analyze the file metadata, including creation time, modification time, and file size. The file content and metadata are further processed based on business rules, key information is extracted from the file, and tags are added to the file based on the extracted information. Files are archived or cleaned up according to business rules and data lifecycle management strategies.

4. The method for automatic data classification and grading according to claim 3, characterized in that: After comprehensive analysis of file content, metadata, and business rules through automated algorithms, a classification coefficient is generated for the data, including the following steps: Obtain data access frequency, data relevance, data volume, and security alarm frequency; The classification coefficient is obtained by comprehensively calculating the access frequency, data relevance, data volume and security alarm frequency through an automated algorithm. The expression is: Where fjx is the classification coefficient, fwc is the number of accesses, ΔT1 is the first monitoring period, aqj is the number of alarms, ΔT2 is the second monitoring period, sjl is the amount of data, gld is the data association, a1, a2, a3 are the proportional coefficients of access frequency, security alarm frequency and data association, respectively, and a1>a2>a3, 0 <a1≤1,0<a2≤1,0<a2≤1。 5. The method for automatic data classification and grading according to claim 4, characterized in that: In each data classification, the data is divided into multiple levels according to the comparison result between the classification coefficient and the classification gradient threshold, including the following steps: The grading gradient threshold includes a first grading threshold and a second grading threshold. The larger the grading coefficient, the greater the importance of the data. After obtaining the grading coefficient, the grading coefficient is compared with the first grading threshold and the second grading threshold. If the classification coefficient is ≥ the second classification threshold, the importance of the classified data is high, and the classified data is placed in the first-level area; If the classification coefficient is less than the second classification threshold, and the classification coefficient is greater than or equal to the first classification threshold, the importance of the classified data is medium, and the classified data is placed in the second level area; If the classification coefficient is less than the first classification threshold, the importance of the classified data is low, and the classified data is classified into the third level area.

6. A method for automatic data classification and grading according to claim 5, characterized in that: The classification results and the grading results in each category are used as data sets to train and test the recurrent neural network model, including the following steps: The classification results and the grading results in each classification are used as training and testing data sets. The data are preprocessed, including standardization, normalization, and processing of missing values. The data set is divided into training and test sets. The RNN model is built using a deep learning framework. The structure of the model is defined, including the input layer, hidden layer, and output layer. A loss function is selected to measure the performance of the model, and an optimizer is selected to update the model parameters. The model is trained using the training set, iterated for multiple cycles, and the model parameters are updated through the back propagation algorithm to minimize the loss function. The performance of the model is verified using the test set. The generalization ability of the model on unseen data is evaluated. According to the verification results, the parameters of the model, including the learning rate, number of layers, and number of neurons, are adjusted.

7. A method for automatic data classification and grading according to claim 6, characterized in that: The logic of data volume acquisition is as follows: use the attribute information of the file system or database to obtain the size of each data file, and use the monitoring tool of the storage system to track the storage capacity of the entire data set in real time; The logic for obtaining the frequency of security alarms is as follows: enable security event log recording, record system security events, including potential threats and attack attempts, regularly analyze security event logs, and count the frequency of security alarms related to each data file.

8. The method for automatic data classification and grading according to claim 7, characterized in that: The logic for obtaining the access frequency of data is as follows: enable access log recording in the system, record the time, user, and file information of each data access, regularly analyze the access log, and count the access frequency of each data; The logic of obtaining data association is: identifying association information in the data, and obtaining the number or identifiers of other data associated with specific data by querying the system or database.

9. A data automatic classification and grading system, used to implement the grading method according to any one of claims 1 to 8, characterized in that: It includes data acquisition module, data classification module, coefficient generation module, data classification module, data collection module and model training module; Data acquisition module: obtains existing data in the device; Data classification module: classifies data based on data metadata and data file suffixes; Coefficient generation module: generates a classification coefficient for the data by scanning the file content, metadata and business rules in each category and comprehensively analyzing the file content, metadata and business rules through an automated algorithm; Data classification module: In each data classification, the data is divided into multiple levels according to the comparison results between the classification coefficient and the classification gradient threshold; Data collection module: regularly obtains data classification results and grading results in each category; Model training module: After a period of time, the classification results and the grading results in each category are used as data sets to train and test the recurrent neural network model; Automatic processing module: Deploy the trained recurrent neural network model to the device. When new data is generated in the subsequent device, the trained recurrent neural network model is used to automatically classify and grade the data.