Intelligent classification and grading method and system for structured data

By combining the intelligent classification and grading method of the rule system and the model system, the problem of the difficulty in balancing accuracy, efficiency and scalability in the classification and grading of structured data in the existing technology is solved, and high-precision and high-efficiency data classification and grading are achieved, which is suitable for a variety of structured data scenarios.

CN120632649AInactive Publication Date: 2025-09-12JIANGSU CIMER INFORMATION SECURITY TECH

Patent Information

Application Number
CN202511131173.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing structured data classification and grading methods are difficult to balance accuracy, efficiency and scalability. Manual methods are inefficient and highly subjective, rule-based methods have high maintenance costs, and model system-based methods have high requirements for training data quality and are difficult to start in scenarios with unlabeled data.

Method used

Combining the rule system and model system, intelligent classification and grading of structured data is achieved through data preprocessing, rule identification, model prediction, credibility fusion and dynamic weight adjustment.

Benefits of technology

It improves the overall accuracy of classification and grading, and is suitable for scenarios with fuzzy boundaries or scarce labels. The rule set can be continuously expanded, the model system can be continuously trained, and it supports different types of structured data sources and changing business needs. The system supports self-feedback and optimization, has low maintenance costs, and moderate human involvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632649A_ABST
    Figure CN120632649A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent classification and grading method and system for structured data, and relates to the technical field of artificial intelligence and data security, and the method comprises the steps: obtaining to-be-classified structured data, and carrying out the data preprocessing; performing preliminary identification based on the rule system, and outputting rule system credibility; inputting the preprocessed structured data into the model system to obtain the credibility of the model system; the credibility of the rule system and the credibility of the model system are subjected to weighted fusion, credibility fusion scores are obtained through calculation, and the type with the maximum fusion score is selected as a fusion classification result of the structured data; dynamically adjusting the weight of the fusion score according to the credibility difference; outputting a grading label of the to-be-classified structured data; and after user feedback is received, the rule system rule and the model system parameters are updated. The method supports different types of structured data sources and constantly changing business requirements, and has wide application prospects and industrial popularization values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and data security technology, and in particular to a method and system for intelligent classification and grading of structured data. Background Art

[0002] With the continuous advancement of informatization, enterprises and institutions are accumulating vast amounts of structured data in their daily operations. This data contains rich business information and sensitive elements, and possesses significant value and security attributes. To effectively manage and protect data and improve the efficiency of data resource utilization, it is imperative to classify and manage structured data in a hierarchical manner.

[0003] Currently, common data classification and grading methods include purely manual classification and grading, automated classification and grading based on rule engines, and intelligent classification and grading methods based on machine learning or deep learning model systems. While manual methods offer high accuracy, they are inefficient, subjective, and have poor scalability. Rule-based approaches are more efficient for processing structured fields, but they are costly to maintain and struggle to cope with changes in data formats and semantics. Model-based approaches, while capable of generalization, require high-quality training data and are difficult to implement in scenarios with unlabeled data.

[0004] In practical applications, a single approach often fails to achieve a balanced balance of accuracy, efficiency, and scalability. Therefore, a fusion solution that combines the controllability of a rule engine with the generalization capabilities of a model system is urgently needed to achieve high-precision, high-efficiency, and highly scalable data classification and grading capabilities. Summary of the Invention

[0005] The present invention is proposed in view of the problems existing in the existing intelligent classification and grading method for structured data. Therefore, the problem to be solved by the present invention is how to provide an intelligent classification and grading method and system for structured data.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a method for intelligent classification and grading of structured data, comprising: obtaining structured data to be classified and performing data preprocessing;

[0008] Perform preliminary identification based on the rule system, use manually formulated rules to parse the pre-processed structured data, and output one or more categories of matching results and the credibility of the rule system;

[0009] Input pre-processed structured data into the model system to obtain prediction classification results and model system credibility;

[0010] Conduct credibility fusion and preliminary classification decisions. For each candidate type, perform weighted fusion of the credibility of the rule system and the model system, calculate the credibility fusion score, and select the type with the largest fusion score as the fusion classification result of the structured data.

[0011] Adaptive weight update based on credibility difference. If the prediction results of the rule system and the model system are inconsistent, the credibility difference between the rule system and the model system is calculated, and the weight of the fusion score is dynamically adjusted according to the credibility difference.

[0012] According to the fusion classification results and the confidence interval, the predefined sensitivity level mapping rules are matched and the hierarchical labels of the structured data to be classified are output;

[0013] After receiving user feedback or relying on expert manual annotation, the rule system rules and model system parameters are updated to optimize the intelligent classification and grading of structured data.

[0014] As a preferred solution of the intelligent classification and grading method of structured data described in the present invention, the structured data is data with clear field definitions and fixed formats, which is stored in a relational database in a tabular form, including a user information table, a transaction record table and a business log table.

[0015] As a preferred solution of the intelligent classification and grading method for structured data described in the present invention, the data preprocessing includes format standardization, missing value completion, field unification, and noise data removal of the structured data to be classified;

[0016] The manually formulated rules include field name mapping rules, regular expression matching rules, domain dictionary matching rules, threshold rules, and semantic labeling rules coupled with structured fields.

[0017] As a preferred solution of the intelligent classification and grading method for structured data described in the present invention, the model system includes the following contents:

[0018] Load a structured dataset containing data samples and their corresponding category labels from the local file system. Each data sample is represented in JSON format.

[0019] For the loaded structured dataset, the system extracts the label values ​​of all datasets and builds a label-index mapping dictionary;

[0020] Use the word segmenter corresponding to the pre-trained Chinese language model system to perform word segmentation, encoding, truncation, and padding operations on the original text field, converting the text data into a tensor form acceptable to the model system. At the same time, the original text label is converted into an integer label index through a mapping dictionary;

[0021] Divide the processed structured dataset into training set and validation set;

[0022] Based on the labeled data, the deep learning model system is trained to understand the semantics and extract features of the structured data. The model system error is minimized through the loss function, and the cross entropy is used. As the classification loss function, the formula is:

[0023] ;

[0024] in, is the true label, predict probabilities for model systems;

[0025] The system loads a general Chinese pre-trained language model system and dynamically adjusts the output dimension of the classification layer according to the number of labels to adapt to the current task. During the training process, the strategy of saving the model system in each round and loading the optimal model system at the end of training is adopted;

[0026] Use the advanced training control interface to complete forward propagation, loss function calculation, gradient backpropagation, and optimizer updates during the training process. At the same time, evaluate the model system performance on the validation set after each round of training, and save the optimal parameter configuration based on the evaluation results.

[0027] After the model system training is completed, save the final model system parameters and word segmentation configuration;

[0028] The output of the model system is the predicted probability distribution of each classification category , expressed as:

[0029] ;

[0030] in, is the activation function, is the encoded context semantic representation, and is the classification layer weight of the model system;

[0031] The category with the largest prediction probability is selected as the final prediction result, and the probability value of the category is used as the credibility of the model system.

[0032] As a preferred solution of the intelligent classification and grading method for structured data described in the present invention, the calculation formula of the credibility fusion score is:

[0033] ;

[0034] ;

[0035] in, is the credibility fusion score, For the rule system type The current weight of For the model system type The current weight of For rules Pair Type credibility; For the model system type credibility; It represents a rule-based system. Represents a model-based system, Indicates the current pair type The number of valid rules, that is, the number of rules that meet the conditions.

[0036] As a preferred solution of the intelligent classification and grading method for structured data described in the present invention, the calculation formula for the credibility difference between the rule system and the model system is:

[0037] ;

[0038] in, is the credibility difference between the rule system and the model system, Indicates that the rule system is correct; Indicates that the model system is correct; For the rule system type credibility.

[0039] As a preferred solution of the intelligent classification and grading method for structured data described in the present invention, the weight updating method for dynamically adjusting the weight of the fusion score according to the credibility difference is:

[0040] ;

[0041] ;

[0042] in, represents the limit function used to define the upper and lower bounds, is the step size adjustment coefficient.

[0043] In a second aspect, the present invention provides an intelligent classification and grading system for structured data, comprising:

[0044] A data input module, used to input structured data to be classified;

[0045] A data preprocessing module is used to preprocess the structured data to be classified;

[0046] A rule recognition module is used to execute the rule set in the rule system and generate the rule system classification output and credibility;

[0047] Model prediction module, used to obtain the prediction classification results of the structured data to be classified and the credibility of the model system;

[0048] The fusion judgment module is used to calculate and output the final fusion classification result based on the credibility of the rule system and the model system;

[0049] Dynamic weight update module, used to dynamically adjust the weight of the fusion score according to the credibility difference between the rule system and the model system;

[0050] A classification strategy module is used to determine the sensitivity level of the structured data to be classified based on the fusion classification results and confidence interval;

[0051] Feedback optimization module, used to update rule system rules and model system parameters based on user feedback.

[0052] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein the processor implements the steps of a method for intelligent classification and grading of structured data when executing the computer program.

[0053] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of a method for intelligent classification and grading of structured data.

[0054] The beneficial effects of the present invention are as follows: the present invention improves the overall accuracy of classification and grading through the integration of rules and model systems, and is particularly suitable for scenarios with fuzzy boundaries or scarce labels; the rule set can be continuously expanded, and the model system can be continuously trained to support different types of structured data sources and ever-changing business needs; the system supports self-feedback and optimization, with low maintenance costs and moderate human involvement; it can achieve efficient automation while meeting compliance, auditing, and regulatory requirements; it is applicable to various structured data scenarios such as government affairs, finance, energy, manufacturing, and education, and does not rely on the knowledge structure of a specific industry. It breaks through the bottleneck of traditional structured data classification and grading methods, and provides intelligent, efficient, and sustainably evolving data classification and grading technology solutions with broad application prospects and industrial promotion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0056] Figure 1 A structural diagram of an intelligent classification and grading system for structured data; Figure 2 A dynamic weight update flow chart for an intelligent classification and grading method for structured data; Figure 3 A weighted adaptive graph for an intelligent classification and grading method for structured data; Figure 4 The figure shows the comparison between the fused credibility and the true credibility. DETAILED DESCRIPTION

[0057] To make the above-mentioned objects, features, and advantages of the present invention more easily understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.

[0058] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0059] Secondly, an embodiment or embodiments herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The appearance of an embodiment in different places in this specification does not necessarily refer to the same embodiment, nor is it an embodiment that is exclusive or selectively mutually exclusive of other embodiments.

[0060] Example 1

[0061] Reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, provides a method for intelligent classification and grading of structured data, comprising:

[0062] S1: Obtain structured data to be classified and perform data preprocessing;

[0063] Specifically, data preprocessing: First, the structured data to be processed is cleaned, deduplicated, and standardized to ensure data quality and consistency. Missing values ​​and outliers that may exist in the data are also processed.

[0064] Structured data refers to data with clearly defined fields and a fixed format, typically stored in tabular form in relational databases such as MySQL, Oracle, and PostgreSQL. This data includes, but is not limited to, user information tables (e.g., containing fields such as ID number, name, mobile phone number, and address), transaction record tables (e.g., containing fields such as bank card number, amount, payee and payee information), and business log tables. Structured data originates from various business systems generated by organizations in their daily operations, such as customer relationship management, enterprise resource planning, financial systems, and hospital information systems.

[0065] This method is suitable for automatically classifying and grading sensitive fields in data from the above sources, and supports refined governance and protection of data assets in different application scenarios. It can be widely used in data asset management in government agencies, finance, medical care, operators, e-commerce and other fields to solve problems such as difficult data classification, high manual labeling costs, and insufficient rule coverage. It is particularly suitable for automatically identifying and grading sensitive data in databases, data lakes, and data warehouses. For example, it can identify ID numbers and income fields in employee information tables, automatically label them as identity information and financial information, and assign labels such as high sensitivity and medium sensitivity to assist in scenarios such as data outbound review and permission control.

[0066] Structured data usually contains missing values, outliers, and format inconsistencies. It must be normalized first to provide high-quality input for rule matching and model system predictions. First, in the data preprocessing stage, the system performs cleaning, deduplication, and standardization on the original structured data to ensure the effectiveness and consistency of subsequent processing. During the cleaning process, dirty data and non-standard records are removed; deduplication is used to delete duplicate data items; and standardization unifies the data format into a form that the system can recognize, such as unifying the date into the YYYY-MM-DD format. In addition, missing values ​​are processed using methods such as mean filling, median filling, and interpolation, while outliers are identified and eliminated using the z-score method. The z-score formula is as follows:

[0067] ;

[0068] in, is the original value, is the field mean, is the field standard deviation, Data with values ​​greater than the set threshold are considered abnormal.

[0069] S2: Perform preliminary identification based on the rule system, use manually formulated rules to parse the pre-processed structured data, and output one or more categories of matching results and the credibility of the rule system ;

[0070] Specifically, the rule engine defines: by defining rules applicable to structured data fields, manually formulated rules include field name mapping rules, regular expression matching rules, domain dictionary matching rules, threshold rules, and semantic labeling rules coupled with structured fields;

[0071] For example, formatting rules for ID card numbers, bank card numbers, and mobile phone numbers are used to initially categorize data. These rules can be dynamically updated and adjusted based on actual business needs. During the rule engine identification phase, the system pre-sets and dynamically maintains a set of rules to quickly identify sensitive field types within structured data. Simple rules are written based on regular expressions, while complex rules are implemented using programming code.

[0072] The system matches each record's field values ​​against these rules and performs a preliminary classification based on the categories found. This method offers high accuracy and processing speed, making it suitable for identifying common sensitive information such as bank card numbers, mobile phone numbers, email addresses, and addresses. The rule system is configurable, allowing users to flexibly adjust it to suit different business scenarios.

[0073] During the matching process, the system performs rule matching operations on each field value in each record in turn. For a certain field value, the system will traverse the rule set, record all successfully matched rule types, and return the credibility of the rule based on the matching rule. The credibility can be preset by the rule designer, or obtained through testing of a large number of historical samples, or adjusted in combination with the rule type (regular / code). For example, in the examples of the present invention, the credibility returned by the regular expression rule is generally slightly lower than the credibility returned by the code. The system will sort the results from high to low in terms of credibility, and use the three results with the highest credibility as the results of the rule system match. If there are results of the same data type in the final result, their credibility will be merged;

[0074] Output the matching results and rule system credibility of one or more categories .

[0075] The matching results will serve as preliminary classification results for subsequent fusion classification to achieve joint judgment based on rules and model systems.

[0076] S3: Input preprocessed structured data into the model system to obtain prediction classification results and model system credibility ;

[0077] Specifically, when performing intelligent classification and prediction on structured data, a pre-trained language model system (BERT) based on the Transformer architecture is used to perform semantic modeling and feature extraction on the structured field content.

[0078] Select the bert-base-chinese model system, which contains multiple layers of bidirectional Transformer encoders. It models the input text through the self-attention mechanism and automatically captures the semantic relationships between words. The input of the model system is the field text information in the structured data (such as Zhang San, the bank card number starting with 6222). The input fields are tokenized and encoded through the tokenizer corresponding to the pre-trained model system, and converted into the tensor form acceptable to the model system. In actual operation, AutoTokenizer.from_pretrained(bert-base-chinese) is used to automatically load the Chinese BERT tokenizer, and truncation and padding processing are performed on each text sample. The system constructs a label dictionary (such as address, ID number, name, etc.) by analyzing all the categories appearing in the training samples and maps them to corresponding integers. This label mapping is input into the model system as the supervision signal during training. The classification task is defined as a multi-class classification problem. The output dimension of the model system is equal to the total number of categories, and the predicted probability of each category is output through the softmax layer.

[0079] Loading of labeled data. First, load the structured data set containing data samples and their corresponding category labels from the local file system. Each data sample is represented in JSON format, and the system completes the parsing and loading of the data through traversal reading operations.

[0080] Construction of label mapping. For the loaded structured data set, the system extracts the label values of all data sets and constructs a label-index mapping dictionary for numerical label processing during the subsequent training process of the model system. This mapping relationship is permanently stored as a JSON file to support the label restoration operation in the deployment and inference phases of the model system.

[0081] Text preprocessing and tokenization encoding. The system uses the tokenizer corresponding to the pre-trained Chinese language model system (such as AutoTokenizer provided with the BERT model system) to perform tokenization, encoding, truncation, and padding operations on the original text fields, thereby converting the text data into the tensor form acceptable to the model system. At the same time, the original text labels are converted into label indices in integer form through the mapping dictionary.

[0082] Division of training set and validation set. The system automatically divides the processed structured data set into a training set and a validation set, with a division ratio of 80% and 20%. This division strategy ensures the generalization ability of the model system training and provides a basis for performance evaluation during the training process.

[0083] Model system training: Based on a certain amount of labeled data, a deep learning model system is trained, and methods such as neural networks are used to perform in-depth feature extraction and classification of the data. The model system can learn the complex relationships between data and improve classification accuracy. During the model system training phase, the system introduces a deep learning model system (Transformer neural network) to perform semantic understanding and feature extraction on structured data. The model system performs supervised learning training based on the labeled structured data set, and minimizes the model system error through the loss function. In this invention, cross entropy is used As the classification loss function, its formula is as follows:

[0084] ;

[0085] in, is the true label, Predict probabilities for the model system.

[0086] Through iterative optimization, the model system can capture complex relationships between fields, identify boundary samples or fuzzy types that are difficult to accurately judge through rules, and improve the overall classification accuracy.

[0087] Model system construction and parameter configuration. The system loads a general-purpose Chinese pre-trained language model system (such as BERT) and dynamically adjusts the output dimensions of the classification layer based on the number of labels to adapt to the current task. Training parameters include but are not limited to: learning rate, number of training rounds, batch size, weight decay coefficient, logging frequency, and verification strategy. During training, the model system is saved per round and the optimal model system is loaded at the end of training to improve the stability and accuracy of the final model system.

[0088] Training process control and automatic evaluation: Utilizing advanced training control interfaces (such as the Trainer module in the Transformers library), the system automatically completes forward propagation, loss function calculation, gradient backpropagation, and optimizer updates during training. Furthermore, after each round of training, the system automatically evaluates model performance on a validation set and saves the optimal parameter configuration based on the evaluation results.

[0089] Model system persistent storage. After model system training is complete, the system saves the final model system parameters and tokenizer configuration to a specified path for subsequent model system deployment and inference. The saved format is compatible with mainstream frameworks, facilitating cross-platform migration and integration.

[0090] The output of the model system is the predicted probability distribution of each classification category , expressed as:

[0091] ;

[0092] in, is the activation function, is the encoded context semantic representation, and is the classification layer weight of the model system. For example, the output result is [0.1, 0.3, 0.6].

[0093] Based on the softmax output, the category with the highest probability is selected as the final prediction result, and the probability value of this category is used as the credibility of the model system's prediction result. In this paper, the softmax operation is not explicitly defined in the model system output layer, but is implicitly completed by the deep learning framework when calling the cross-entropy loss function. This loss function automatically performs the softmax operation and the negative log-likelihood calculation, thereby avoiding the problem of numerical instability and improving training efficiency.

[0094] S4: Perform credibility fusion and preliminary classification decision. For each candidate type, perform weighted fusion of the credibility of the rule system and the model system, calculate the credibility fusion score, and select the type with the largest fusion score as the fusion classification result of the structured data.

[0095] Specifically, the fusion decision mechanism: a fusion algorithm is used to merge the output of the rule engine with the prediction results of the model system, and a weighted score is calculated to generate the final classification label;

[0096] Taking a single piece of structured data as input, the system comprehensively considers the credibility of each type returned by the rule engine and the deep learning model system, and dynamically balances the influence of the two types of information sources, the rule and the model system, by building an adaptive weight adjustment mechanism, thereby achieving accurate, stable and automatic type recognition. The specific steps include:

[0097] (1) Obtaining credibility of rule system: Each rule Matching a certain type Returns a credibility , indicating the rule Consider the data to be of type credibility.

[0098] (2) Model system credibility acquisition: The deep model system returns various types of The predicted probability , indicating that the data is of type The credibility of the prediction.

[0099] (3) Credibility fusion calculation: For each candidate type , the credibility of the rule system and the model system is weighted and integrated to obtain the fusion score , used for the final classification decision. The fusion score formula is as follows:

[0100] ;

[0101] in, For the rule system type The current weight of For the model system type The current weight of For rules Pair Type credibility; For the model system type The symbol r represents a rule-based system, and the symbol m represents a model-based system. For all applicable types The sum of the credibility of the rule system; Indicates the current pair type The number of valid rules (i.e. the number of rules that meet the conditions); satisfy the normalization constraint: + =1, initially = =0.5;

[0102] Select fusion score The largest type As the fusion classification result of the structured data.

[0103] S5: Adaptive weight update based on credibility difference. If the prediction results of the rule system and the model system are inconsistent, the credibility difference between the rule system and the model system is calculated, and the weight of the fusion score is dynamically adjusted according to the credibility difference.

[0104] Specifically, the weight dynamic update mechanism: when there is label feedback (such as confirming that a certain type is a true type through cross-validation or offline evaluation ), this invention uses manual labeling to obtain the true data type. During implementation, the system displays the classification results to the user, who can manually correct any incorrect classification results. The corrected label is then considered the true type label for the data.

[0105] Based on this feedback mechanism, when the system obtains the true type of a data sample, it updates the weight to enhance the influence of the correct source and weaken the effect of the error source. The weight update adopts a nonlinear dynamic step size strategy:

[0106] ;

[0107] ;

[0108] ;

[0109] in, is the step length adjustment coefficient, and 0.1 is selected in the present invention. If the rule is correct, it is set to 1, otherwise it is set to 0; If the model system judges correctly, it is set to 1, otherwise it is set to 0; Indicates that the weights are constrained within a reasonable range to prevent overfitting; is the credibility difference between the rule system and the model system, For the rule system type credibility.

[0110] S6: Based on the fusion classification results and the confidence interval, the predefined sensitivity level mapping rules are matched and the hierarchical labels of the structured data to be classified are output;

[0111] Specifically, based on the final fusion classification results and combined with a predefined sensitivity level mapping model system, data is accurately classified into sensitivity levels. This supports multi-level and multi-dimensional classification strategies, improving the refined management capabilities of structured data and ensuring the usability and controllability of data classification and classification results in business security scenarios.

[0112] Based on the fusion classification results, the system determines the data type (such as ID card number, bank card number, mobile phone number, etc.), and then performs mapping based on the preset classification model corresponding to that type. For example, the system can pre-set the following mapping relationship based on industry standards, security policies, or risk assessment standards: ID card number: high sensitivity; bank card number: medium sensitivity; email address: low sensitivity.

[0113] The specific mapping relationship can be customized according to the industry and confidentiality level. Sensitivity level mapping can be based on business risk assessment, security compliance standards or industry specifications, and can be dynamically configured and expanded according to business needs. For some data types with fuzzy sensitivity boundaries, the credibility score of the fusion classification results is introduced as a reference:

[0114] a. When the credibility is higher than the set threshold, the sensitivity level is directly assigned according to the mapping table;

[0115] b. When the credibility is in the borderline range, the system will prompt the user to review;

[0116] c. When the credibility is lower than the lower credibility limit, the system does not perform classification by default.

[0117] S7: After receiving user feedback or relying on expert manual annotation, update the rule system rules and model system parameters to optimize the intelligent classification and grading of structured data.

[0118] Specifically, the adaptive feedback optimization mechanism introduces a manual review process to confirm samples with uncertain classification results. The review results are then fed back to the rule engine and model system inference module. Based on this feedback, fusion strategy parameters are dynamically adjusted to automatically optimize rule weights and model system weights, forming a closed-loop learning path that effectively improves the system's adaptability and long-term accuracy.

[0119] Furthermore, this embodiment also provides an intelligent classification and grading system for structured data, including:

[0120] A data input module, used to input structured data to be classified;

[0121] A data preprocessing module is used to preprocess the structured data to be classified;

[0122] A rule recognition module is used to execute the rule set in the rule system and generate the rule system classification output and credibility;

[0123] Model prediction module, used to obtain the prediction classification results of the structured data to be classified and the credibility of the model system;

[0124] The fusion judgment module is used to calculate and output the final fusion classification result based on the credibility of the rule system and the model system;

[0125] Dynamic weight update module, used to dynamically adjust the weight of the fusion score according to the credibility difference between the rule system and the model system;

[0126] A classification strategy module is used to determine the sensitivity level of the structured data to be classified based on the fusion classification results and confidence interval;

[0127] Feedback optimization module, used to update rule system rules and model system parameters based on user feedback.

[0128] This embodiment also provides a computer device, which is suitable for a method for intelligent classification and grading of structured data, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement all or part of the steps of the method described in the embodiment of the present invention as proposed in the above embodiment.

[0129] This embodiment further provides a storage medium having a computer program stored thereon, which, when executed by a processor, executes the method of any optional implementation of the above embodiment. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0130] The storage medium proposed in this embodiment and the data storage method proposed in the above embodiment belong to the same inventive concept. Technical details not fully described in this embodiment can be found in the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0131] In summary, the present invention improves the overall accuracy of classification and grading through the integration of rules and model systems, and is particularly suitable for scenarios with fuzzy boundaries or scarce labels; the rule set can be continuously expanded, and the model system can be continuously trained to support different types of structured data sources and ever-changing business needs; the system supports self-feedback and optimization, with low maintenance costs and moderate human involvement; it can achieve efficient automation while meeting compliance, auditing, and regulatory requirements; it is applicable to a variety of structured data scenarios such as government affairs, finance, energy, manufacturing, and education, and does not rely on the knowledge structure of a specific industry. It breaks through the bottleneck of traditional structured data classification and grading methods, and provides intelligent, efficient, and sustainably evolving data classification and grading technology solutions with broad application prospects and industrial promotion value.

[0132] Example 2

[0133] Reference Figure 1 - Figure 4 This is the second embodiment of the present invention. This example, based on the management of sensitive data assets within a government or enterprise, employs the structured data classification and grading system provided by this invention to process large amounts of heterogeneous structured data from a database. The system integrates expert-defined rules with a data-driven model system to automatically identify data types and assess security levels, offering advantages such as high precision, adaptability, and scalability.

[0134] First, to ensure subsequent classification accuracy and model system stability, the system first standardizes and cleans the original structured data. The system obtains a batch of raw data from the data source, which may include fields such as name, ID number, bank card number, and phone number. The data may contain missing values, duplicate data, or inconsistent formats. Therefore, the system will perform the following processing on the data:

[0135] (1) Cleaning and deduplication: The system checks for duplicates in the data and automatically removes redundant records. It uses a combination of primary key fields (such as ID number + name) to determine uniqueness, and combines this with a hash index to quickly filter out duplicate data. For example, if the same user's name and ID number appear multiple times, the system will retain one record and delete the duplicates.

[0136] (2) Standardization: The system automatically extracts the fields and content of each record based on the database metadata and unifies the format, including but not limited to:

[0137] Date fields: Convert to ISO8601 format (YYYY-MM-DD). For example, convert 2025 / 04 / 22 to 2025-04-22.

[0138] Amount fields: Convert to floating-point type and retain two decimal places to ensure that all amount fields are in a unified number format with two decimal places.

[0139] Text fields: All text is encoded in UTF-8, removing special symbols, spaces, line breaks, and other abnormal characters.

[0140] (3) Missing value handling: For missing fields, the system will select different filling methods based on the field type. For example, for the income field, if there are not many missing values, the median can be used to fill in the missing values; for continuous numeric fields (such as age), the system can use the mean to fill in the missing values.

[0141] (4) Outlier detection and elimination: Use the Z-Score method to check for outliers in each record. For example, if the age field data of a user is 200, the system will use the Z-Score formula to determine whether it is an outlier. If the Z value of the data is greater than the preset threshold, the system will eliminate it.

[0142] In the rule engine, the system defines a set of rules for each data type, such as ID card number, bank card number, mobile phone number, etc. The rules are implemented through regular expressions. Specific examples include:

[0143] ID card number recognition rule: The system defines a regular expression for matching ID card numbers. The rule will match fields that conform to the ID card number format and output the rule matching type and confidence level (for example, 0.95 means that the confidence level that the data is an ID card number is 95%).

[0144] Bank card number rule: A rule defined for bank card numbers. This rule is used to identify the bank card number field. If a match is successful, the corresponding confidence level (for example, 0.9) is returned.

[0145] The system will match each record. If the field value meets a certain rule, the field will be classified as the type corresponding to the rule and the credibility will be given.

[0146] Credibility Assessment and Adjustment Mechanism: The initial credibility is manually assigned (in this example, the rule engine's initial credibility is 0.5), and is subsequently dynamically adjusted based on actual matching results. If the rule's prediction accuracy consistently exceeds the model system's prediction, the system automatically increases its weight.

[0147] In this embodiment, the system uses a Transformer neural network to train a deep learning model system. End-to-end type prediction is performed on field data. The training set is a set of manually annotated structured fields, which include various sensitive information of users and their classification labels. The model system has the advantages of strong generalization ability and context awareness. The training structured data set of the model system consists of manually annotated data.

[0148] Construct a structured dataset: For example, a structured dataset contains fields such as the user's name, ID number, bank card number, and the actual type labels of these fields (such as the ID number type is marked as ID card).

[0149] Model system training: The system uses the cross entropy loss function for supervised learning:

[0150] ;

[0151] in, is the true label, Predict probabilities for the model system. Through multiple rounds of training, the model system will learn the complex semantic relationships between various fields, improving classification accuracy, especially for data with ambiguous boundaries (such as data like ID cards and bank card numbers, which are easily confused).

[0152] Model system output: The model system outputs the multi-classification probability distribution of each record, for example: ID number: 0.82, bank card number: 0.13, mobile phone number: 0.05;

[0153] At this stage, the output of the rule engine and the model system are integrated. Assuming that the matching rules for the ID card number and bank card number of a certain record both return credibility, and the model system also gives the predicted probability of various labels for this data, the system will proceed as follows:

[0154] Rule system credibility calculation: Assume that the rule engine returns a credibility of 0.9 for the ID card number field and a credibility of 0.85 for the bank card number field.

[0155] Model system credibility calculation: Assume that the model system predicts the category of the record as the ID number category and returns a prediction probability of 0.88.

[0156] Fusion score calculation: Initial weight setting: The system initially sets the weights of the rule system and the model system to 0.5, indicating that the two systems have equal weights in the initial prediction. That is, in the example of the present invention, the initial weight of the rule system is set to =0.5, the initial weight of the model system , through the weighted fusion rule and the calculation of the model system fusion score for:

[0157] ;

[0158] in, For the rule system type The current weight of For the model system type The current weight of For rules Pair Type credibility; For the model system type The symbol r represents a rule-based system, and the symbol m represents a model-based system. For all applicable types The sum of the credibility of the rule system; Indicates the current pair type The number of valid rules (i.e., the number of rules that meet the conditions);

[0159] Weight update strategy: Based on the cross-validation results, assuming that the rule system is correct but the model system is wrong, the system will use a dynamic step size to update the weights of the rule and model system. The weight update adopts a nonlinear dynamic step size strategy:

[0160] ;

[0161] ;

[0162] ;

[0163] in, is the step length adjustment coefficient, and 0.1 is selected in the present invention; Indicates that the rule is correct, otherwise it is 0; Indicates that the model system judgment is correct, otherwise it is 0; Indicates that the weights are constrained within a reasonable range to prevent overfitting.

[0164] The weight update process is as follows: if the two systems make the same judgment, the weight remains unchanged; if only one system makes the correct judgment, the weight of that system is increased and the weight of the other system is decreased; the weight value range is [0.1, 0.9];

[0165] Assuming that the rule system is correct and the model system is wrong, the rule system returns a credibility of 0.8 and the model system returns a credibility of 0.5, then we can calculate:

[0166] ,

[0167] ;

[0168] ;

[0169] The rule weights and model system weights can be dynamically updated to new values. Indicates uncertainty or ambiguity. When it approaches 1, it means there is no ambiguity and no competition.

[0170] When the gap between the rule and model systems is small (with similar credibility), any slight difference could affect the outcome, necessitating more drastic adjustments to clearly enhance the influence of the better-performing system. When it approaches 0, it indicates a high degree of competition between the rule and model systems, indicating a clear divergence (e.g., 0.9 for one and 0.1 for the other), and no further drastic adjustments are necessary.

[0171] Assuming that the rule system is correct and the model system is wrong, the rule system returns a credibility of 0.9 and the model system returns a credibility of 0.1, then we can calculate:

[0172] ,

[0173] ;

[0174] ;

[0175] Based on the final classification label, the system uses a predefined level mapping model system to assign a sensitivity level to the data. The specific implementation is as follows:

[0176] Level classification: For data classified as ID card number, the system will assign it a high sensitivity level based on the predefined level classification model; for data classified as bank card number, it will be assigned a medium sensitivity level.

[0177] Multi-dimensional classification: In addition to sensitive information, the classification level can be further refined based on the actual application scenario of the data. For example, certain user data can be classified as high risk or low risk. Some of the levels are as follows:

[0178] When the data type is ID card number, it is a high sensitivity level; when the data type is mobile phone number, it is a medium sensitivity level; when the data type is email address, it is a medium sensitivity level; when the data type is postal code, it is a low sensitivity level.

[0179] The system manually reviews classification results. If a user reports that a record was misclassified, the system will adjust the rule weights and model system weights based on the feedback. For example, if feedback indicates that the model system misclassified a certain type of data, the system will increase the rule system weight and decrease the model system weight, thereby improving the classification results of similar data in the future.

[0180] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for intelligent classification and grading of structured data, characterized by: include, Obtain structured data to be classified and perform data preprocessing; Perform preliminary identification based on the rule system, use manually formulated rules to parse the pre-processed structured data, and output one or more categories of matching results and the credibility of the rule system; Input pre-processed structured data into the model system to obtain the prediction classification results and the credibility of the model system; Conduct credibility fusion and preliminary classification decision-making. For each candidate type, perform weighted fusion of the credibility of the rule system and the model system, calculate the credibility fusion score, and select the candidate type based on the credibility fusion score as the fusion classification result of the structured data. Adaptive weight update based on credibility difference. If the prediction results of the rule system and the model system are inconsistent, the credibility difference between the rule system and the model system is calculated, and the weight of the credibility fusion score is dynamically adjusted according to the credibility difference. According to the fusion classification results and the confidence interval, the predefined sensitivity level mapping rules are matched and the hierarchical labels of the structured data to be classified are output; After receiving user feedback or relying on expert manual annotation, the rule system rules and model system parameters are updated to optimize the intelligent classification and grading of structured data.

2. The intelligent classification and grading method for structured data according to claim 1, characterized in that: The structured data is data with clear field definitions and a fixed format, and is stored in a relational database in a tabular form, including a user information table, a transaction record table, and a business log table.

3. The intelligent classification and grading method for structured data according to claim 2, characterized in that: The data preprocessing includes format standardization, missing value completion, field unification and noise data removal for the structured data to be classified; The manually formulated rules include field name mapping rules, regular expression matching rules, domain dictionary matching rules, threshold rules, and semantic labeling rules coupled with structured fields.

4. The intelligent classification and grading method for structured data according to claim 3, wherein: The model system includes the following: Load a structured dataset containing data samples and their corresponding category labels from the local file system. Each data sample is represented in JSON format. For the loaded structured dataset, the system extracts the label values ​​of all datasets and builds a label-index mapping dictionary; Use the word segmenter corresponding to the pre-trained Chinese language model system to perform word segmentation, encoding, truncation, and padding operations on the original text field, converting the text data into a tensor form acceptable to the model system. At the same time, the original text label is converted into an integer label index through a mapping dictionary; Divide the processed structured dataset into training set and validation set; Based on the labeled data, the deep learning model system is trained to understand the semantics and extract features of the structured data. The model system error is minimized through the loss function, and the cross entropy is used. As the classification loss function, the formula is: ; in, is the true label, is the predicted probability of the model system, i is the classification category, and N is the total number of classification categories; The system loads a general Chinese pre-trained language model system and dynamically adjusts the output dimension of the classification layer according to the number of labels. During the training process, the strategy of saving the model system in each round and loading the optimal model system at the end of training is adopted; Use the advanced training control interface to complete forward propagation, loss function calculation, gradient backpropagation, and optimizer updates during the training process. At the same time, evaluate the model system performance on the validation set after each round of training, and save the optimal parameter configuration based on the evaluation results. After the model system training is completed, save the final model system parameters and word segmentation configuration; The output of the model system is the predicted probability distribution of each classification category , expressed as: ; in, is the activation function, is the encoded context semantic representation, and is the classification layer weight of the model system; The category with the largest prediction probability is selected as the final prediction result, and the probability value of the category is used as the credibility of the model system.

5. The intelligent classification and grading method for structured data according to claim 4, characterized in that: The calculation formula of the credibility fusion score is: ; ; in, is the credibility fusion score, For the rule system type The current weight of For the model system type The current weight of For rules Pair Type credibility; For the model system type credibility; It represents a rule-based system. Represents a model-based system, Indicates the current pair type The number of valid rules, that is, the number of rules that meet the conditions.

6. The intelligent classification and grading method for structured data according to claim 5, characterized in that: The calculation formula for the credibility difference between the rule system and the model system is: ; in, is the credibility difference between the rule system and the model system, Indicates that the rule system is correct; Indicates that the model system is correct; For the rule system type credibility.

7. The intelligent classification and grading method for structured data according to claim 6, characterized in that: The weight update method for dynamically adjusting the weight of the fusion score according to the credibility difference is: ; ; in, represents the limit function used to define the upper and lower bounds, is the step size adjustment coefficient.

8. An intelligent classification and grading system for structured data, based on the intelligent classification and grading method for structured data according to any one of claims 1 to 7, characterized in that: include, A data input module, used to input structured data to be classified; A data preprocessing module is used to preprocess the structured data to be classified; A rule recognition module is used to execute the rule set in the rule system and generate the rule system classification output and credibility; Model prediction module, used to obtain the prediction classification results of the structured data to be classified and the credibility of the model system; The fusion judgment module is used to calculate and output the final fusion classification result based on the credibility of the rule system and the model system; Dynamic weight update module, used to dynamically adjust the weight of the fusion score according to the credibility difference between the rule system and the model system; A classification strategy module is used to determine the sensitivity level of the structured data to be classified based on the fusion classification results and confidence interval; Feedback optimization module, used to update rule system rules and model system parameters based on user feedback.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the processor implements the steps of the method for intelligent classification and grading of structured data described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for intelligent classification and grading of structured data described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Medical image feature mining method based on small-scale data set and related device

    CN114445621A

  • Claim settlement auditing method and system based on knowledge base

    CN119477560A

  • Multi-modal bill processing method based on dynamic knowledge enhancement

    CN120470018A

  • Ensemble machine learning models incorporating a model trust factor

    US20230044102A1

Cited By

  • Intelligent data classification and grading method and device, equipment and medium

    CN122112732A