Financial data security classification and grading system based on multi-strategy learning
Through the financial data security classification and grading system based on multi-strategy learning, the problems of large amount of data, high complexity and high manual review costs faced by financial institutions in the process of data security classification and grading are solved, and automatic, efficient and accurate data security classification and grading are achieved, with the characteristics of self-learning and strong adaptability.
Patent Information
- Application Number
- CN202510061256.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-09
AI Technical Summary
Financial institutions face problems such as large data volume, high complexity, complex label description and high manual review costs in the process of data security classification and classification, which makes it difficult to meet the classification accuracy requirements.
Design a financial data security classification and grading system based on multi-strategy learning, combining the advantages of zero-sample and few-sample learning, combining error-driven incremental learning strategies, and realize automatic, efficient and accurate data security classification and grading through sentence embedding models and multiple classification models.
The system can automatically, efficiently and accurately complete the security classification and grading of financial data under limited resources and costs, reduce the workload of manual review, improve overall efficiency, have self-learning ability, adapt to different data characteristics and classification requirements, and support frequent model updates and upgrades.
Smart Images

Figure CN119961454A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of financial data security management, and in particular to a financial data security classification and grading system based on multi-strategy learning. Background Art
[0002] In the financial industry, data security classification and hierarchical management are crucial. The People's Bank of China issued the "Guidelines for Data Security Classification of Financial Data Security" (JR / T 0197-2020) in September 2020. The "Reference Table of Typical Data Classification Rules for Financial Institutions" provides a detailed list of data security level classifications, including 73 third-level sub-classifications and 303 fourth-level sub-classifications. Financial institutions need to formulate their own data security classification and grading methods based on the guidelines, and accurately classify all their data features into corresponding security levels.
[0003] However, financial institutions face many challenges in the process of data security classification and grading. On the one hand, the data volume is large and complex, with numerous data systems, complicated feature fields, and inconsistent field naming methods, which increases the difficulty of classification. On the other hand, the label description is complex, and the security classification label contains a large amount of detailed description of information, which requires in-depth understanding to accurately match. At the same time, data security involves major risks and requires extremely high classification accuracy. Manual review is usually required to ensure the accuracy of the results, which leads to high manual labeling costs. Manually labeling a large amount of data is time-consuming and laborious, requiring the coordination of multiple relevant personnel, and is costly.
[0004] Regarding the data of financial institutions, the classification and grading methods commonly used by financial institutions each have their own shortcomings:
[0005] (1) Regular expression-based methods: Regular expressions have a good recognition effect on key sensitive data in specific formats (such as ID card numbers and mobile phone numbers) and can quickly match data with fixed patterns. However, their coverage is low, and their recognition ability is insufficient for non-sensitive data with diverse formats and complex structures or new data types, making it difficult to fully cover all data types.
[0006] (2) Methods based on traditional natural language processing: Traditional natural language processing technology has low accuracy in data classification, especially when dealing with a large number of professional terms and complex semantics in the financial field. For non-sensitive data that requires deep understanding of semantics and contextual relationships, traditional methods are difficult to accurately classify, resulting in limited overall classification performance. Summary of the invention
[0007] In order to overcome the above challenges, the present invention hopes to design a system that can automatically, efficiently and accurately complete the security classification and grading of financial data under limited resources and costs, so as to adapt to complex data, be able to handle a large number of feature fields with non-uniform names, and complex label descriptions; and, improve the accuracy of classification, reduce the workload of manual review, improve overall efficiency, have self-learning capabilities, and be able to continuously improve its own classification performance and classification accuracy through error correction. In the absence of pre-labeled data, the accuracy of the model's classification of financial data can be improved to meet high-standard industry requirements. It can further enhance the adaptability and scalability of the model, so that the model can adapt to different data features and classification requirements, and support frequent model updates and upgrades; reduce labor costs: reduce dependence on manual pre-labeling and review, and save human resources.
[0008] In order to achieve the above objectives, the present invention provides a financial data security classification and grading system based on multi-strategy learning, which integrates the advantages of zero-sample and few-sample learning, and combines with error-driven incremental learning strategies, aiming to efficiently and accurately solve the problem of financial data security classification and grading, and meet the strict requirements of the financial industry for data security management.
[0009] The financial data security classification and grading system based on multi-strategy learning described in the present invention includes: a sentence embedding model, multiple classification models,
[0010] The sentence embedding model is used to convert short texts of training data into embedding vectors for use by the classification model;
[0011] The classification model introduces the embedding vector converted by the sentence embedding model and the labels and weights of the corresponding training data to train the classification model;
[0012] Among them, each classification model is constructed as follows:
[0013] S1. Prepare classification data:
[0014] Extract the content in the specification file ("Reference Table of Typical Data Classification Rules for Financial Institutions") and convert it into classified data. The content of the classified data includes the classification name and classification description. All classified data are accumulated to form a classified data set.
[0015] S2. Extract classified data (initial training data) from the classified data set according to preset rules, process the classified data in a unified format, construct a training data set, and set the initial weight of the training data to 1; the training data includes short text, labels, and weights;
[0016] S3, generate sentence pairs based on training data;
[0017] S4, introduce pre-trained sentence embedding model;
[0018] S5. Use the generated sentence pairs to fine-tune the pre-trained sentence embedding model;
[0019] S6. Use the fine-tuned sentence embedding model to convert the short text of the training data into an embedding vector representation and save the sentence embedding model.
[0020] S7, select a classification model;
[0021] S8. Use the encoded sample embedding vector and the labels and weights of the corresponding training data to train the classification model;
[0022] S9. Input the target data set to be predicted and set the initial weight of the sample data to 1. The sample data has the same format as the training data, including short text, label, and weight.
[0023] S10. Use the sentence embedding model to convert the short text of the sample data of the target dataset into an embedded vector representation;
[0024] S11. Use the classification model to predict the sample data of the target data set;
[0025] S12, record the prediction results and prediction probabilities, and match them to the corresponding classification and grading results;
[0026] S13, review the prediction results,
[0027] When the prediction results need to be corrected, correct the wrong prediction results, collect the corrected prediction results and record them in the misclassified samples, and adjust the weight ω according to the prediction probability p using the formula ω=α+β×p; incorporate the misclassified samples into the training data set, and return to S3 to retrain or fine-tune the classification model;
[0028] When the prediction results do not need to be corrected, the fine-tuned embedding model and the trained classification model are stored for use in the system prediction results.
[0029] In the financial data security classification and grading system based on multi-strategy learning described in the present invention, the data labels actually represent the predicted classification and grading results, and the labels are modified multiple times through the data model to make the classification and grading indicated by the labels more accurate.
[0030] Furthermore, the target data set is stored in a database; the database includes multiple tables, and each table includes multiple fields;
[0031] The target data stored in the target data set takes the field as the smallest unit. The interpretation corresponding to the target data includes the description of the content corresponding to the data table and the description of the content corresponding to the field.
[0032] Furthermore, the content of the classification data includes: third-level sub-classification and third-level sub-classification description, fourth-level sub-classification and fourth-level sub-classification description.
[0033] Furthermore, in step S13, when reviewing the prediction results, a step is included to eliminate misclassified samples that are obviously unreasonable or have large noise, so as to ensure data quality.
[0034] Furthermore, the third-level subcategory description is a complete description of each third-level subcategory formed by combining the fourth-level subcategory descriptions below it.
[0035] Furthermore, the classification model includes a three-level sub-classification model, and a plurality of four-level sub-classification models corresponding to each three-level sub-classification model;
[0036] When the three-level sub-classification model is constructed, the content of the classification data includes the three-level sub-classification name and the three-level sub-classification description; after executing the classification model construction process, the retrained or fine-tuned classification model is used as the three-level sub-classification model, and the three-level sub-classification model is used to classify all target data of the target data set to obtain the three-level sub-classification prediction result of the target data;
[0037] In step S12, the classification and grading results include: third-level sub-classifications and their probabilities, fourth-level sub-classifications and their probabilities, and corresponding security levels.
[0038] When the four-level sub-classification model corresponding to each three-level sub-classification model is constructed, the content of the classification data is the four-level sub-classification name and the four-level sub-classification description corresponding to the current three-level sub-classification; after executing the classification model construction process, the retrained or fine-tuned classification model is used as the four-level sub-classification model, and the four-level sub-classification model is used to classify all target data of the target data set to obtain the four-level sub-classification prediction results of the target data.
[0039] Furthermore, the classification description corresponds to the classification data. The classification description of each subclass is a combination of multiple keywords parsed and extracted from the description text of the subclass in the specification file.
[0040] Furthermore, the format of the classification data of the classification data set is consistent with the prediction data format, and the initial weight of each training data is preset to 1.
[0041] Furthermore, the classification model adopts the SetFit model.
[0042] The financial data security classification and grading system based on multi-strategy learning of the present invention meets the strict requirements of the financial industry for data classification management, and can automatically and efficiently complete the security classification and grading tasks of financial data in combination with the data classification rule reference table issued by the People's Bank of China. It has the characteristics of high efficiency, accuracy, flexibility, self-learning and resource-friendly, integrates the optimized zero-sample and few-sample learning classification models, introduces error-driven incremental learning strategies, and constructs a complete set of financial data security classification and grading systems, helping financial institutions to achieve more intelligent data security management. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to better understand the above and other purposes, features, advantages and functions of the present invention, reference may be made to the embodiments shown in the accompanying drawings. The same reference numerals in the accompanying drawings refer to the same components. It should be understood by those skilled in the art that the accompanying drawings are intended to schematically illustrate the preferred embodiments of the present invention and have no limiting effect on the scope of the present invention, and the components in the drawings are not drawn to scale.
[0044] Figure 1 The present invention shows a flowchart of an implementation of a financial data security classification and grading system based on multi-strategy learning.
[0045] Figure 2 Another implementation flow chart of a financial data security classification and grading system based on multi-strategy learning of the present invention is shown
[0046] Figure 3 Part of the content table for the classification data set
[0047] Figure 4 A flow chart is shown when the classification model training data includes four levels of sub-classification.
[0048] Figure 5 This is an example table of the structure and some content of the training dataset.
[0049] Figure 6 This is the training sample table. DETAILED DESCRIPTION
[0050] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0051] As used herein, the term "including" and its variations mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "based at least in part on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0052] In order to at least partially solve one or more of the above problems and other potential problems, an embodiment of the present disclosure proposes a financial data security classification and grading system based on multi-strategy learning.
[0053] like Figure 1 As shown, the financial data security classification and grading system based on multi-strategy learning described in the present invention includes: a sentence embedding model and multiple classification models. The sentence embedding model is used to convert the short text of the training data into an embedding vector for use by the classification model; the classification model introduces the embedding vector converted by the sentence embedding model and the labels and weights of the corresponding training data to train the classification model;
[0054] Among them, each classification model is constructed as follows:
[0055] S1. Prepare classification data:
[0056] Extract the contents of the regulatory document "Reference Table of Typical Data Classification Rules for Financial Institutions" and convert them into classified data. The contents of the classified data include classification name and classification description. All classified data are accumulated to form a classified data set.
[0057] S2. Extract classified data (initial training data) from the classified data set according to preset rules, process the classified data in a unified format, construct a training data set, and set the initial weight of the training data to 1; the training data includes short text, labels, and weights;
[0058] S3, generate sentence pairs based on training data;
[0059] S4, introduce pre-trained sentence embedding model;
[0060] S5. Use the generated sentence pairs to fine-tune the pre-trained sentence embedding model;
[0061] S6. Use the fine-tuned sentence embedding model to convert the short text of the training data into an embedding vector representation and save the sentence embedding model.
[0062] S7, select a classification model;
[0063] S8. Use the encoded sample embedding vector and the labels and weights of the corresponding training data to train the classification model;
[0064] S9. Input the target data set to be predicted and set the initial weight of the sample data to 1. The sample data has the same format as the training data, including short text, label, and weight.
[0065] S10. Use the sentence embedding model to convert the short text of the sample data of the target dataset into an embedded vector representation;
[0066] S11. Use the classification model to predict the sample data of the target data set;
[0067] S12, record the prediction results and prediction probabilities, and match them to the corresponding classification and grading results;
[0068] S13, review the prediction results,
[0069] When the prediction results need to be corrected, correct the wrong prediction results, collect the corrected prediction results and record them in the misclassified samples, and adjust the weight ω according to the prediction probability p using the formula ω=α+β×p; incorporate the misclassified samples into the training data set, and return to S3 to retrain or fine-tune the classification model;
[0070] When the prediction results do not need to be corrected, the fine-tuned embedding model and the trained classification model are stored for use in the system prediction results.
[0071] The security classification and grading system described in the present invention is trained based on the training data generated by the label, and predicts the phrase "table name + field name", using a 1+N hierarchical classification model strategy, that is, a three-level sub-classification model and multiple corresponding four-level sub-classification models. When predicting, the trained classification model is used to predict all fields, and the three-level sub-classification and four-level sub-classification prediction results and corresponding probabilities of each field can be obtained.
[0072] like Figure 2 The table of partial contents of the classification data set shown in the figure takes the standard document "Reference Table of Typical Data Classification Rules for Financial Institutions" as an example, and extracts the classification data set from the standard document. The content of the classification data usually includes the classification name of the third-level sub-classification and its corresponding classification name of the fourth-level sub-classification. The third-level sub-classification "personal natural information" contains 9 fourth-level sub-classifications. Then, in the security classification and grading system described in the present invention, all the third-level sub-classifications (assuming 73 in the reference table) correspond to a third-level sub-classification model, and the predicted labels are these 73 third-level sub-classifications. The third-level sub-classification model is composed of a sentence embedding model and a classification model; each third-level sub-classification, such as Figure 2The personal natural information in corresponds to a four-level sub-classification model, and the predicted labels are these nine four-level sub-classifications, that is, one model predicts nine categories; each four-level sub-classification model is composed of its own embedding model and classification model. Assume that the reference table has 73 three-level sub-classification labels, corresponding to 73 four-level sub-classification models, and the three-level sub-classification of personal natural information corresponds to one of the four-level sub-classification models.
[0073] like Figure 3 Shown is a flowchart of the construction of the above-mentioned financial data security classification and grading system based on multi-strategy learning, the three-level sub-classification model, and the four-level sub-classification model.
[0074] M1. Extract the contents of the specification document "Reference Table of Typical Data Classification Rules for Financial Institutions" and convert them into classified data. The contents of the classified data include the classification name of the third-level sub-classification, the classification description of the third-level sub-classification, the classification name of the fourth-level sub-classification, and the classification description of the fourth-level sub-classification. All classified data are accumulated to form a classified data set.
[0075] The classification description corresponds to the classification data. The classification description of each subclass is a combination of multiple keywords parsed and extracted from the description text of the subclass in the specification file. The classification description of the third-level subclass is a complete classification description of each third-level subclass formed by combining the fourth-level subclass descriptions under it.
[0076] M2. Extract the third-level sub-classification data from the classification data set (as the initial training data for the three-level model), process the third-level sub-classification data in a unified format, and construct a three-level training data set;
[0077] The third-level training data contains short text, label, and weight. The format of each third-level training data is consistent with the prediction data format. For example, the short text of both is a phrase of "table name + field name". The short text is a collection of all the fourth-level sub-classification descriptions under the third-level sub-classification. Here, only one fourth-level sub-classification is used as an example. The label is the label of the third-level sub-classification, for example Figure 2 Personal natural information.
[0078] Short text is a combination of multiple keywords that are parsed and extracted:
[0079] Extract keywords from the category name and category description and combine them into phrases in the form of "table name + field name" to ensure that the format is consistent with the predicted data.
[0080] Extraction steps:
[0081] Identify key entities from the category description: such as "mobile phone", "landline phone", "email address", "WeChat ID", etc.
[0082] Build phrases: Combine key entities with common table names to form phrases:
[0083] "Customer Information Form Mobile Phone"
[0084] "Contact List Landline"
[0085] "User Profile Email Address"
[0086] “Social media account list WeChat ID”
[0087] "Address Book Contact Information"
[0088] Taking the third-level sub-category "personal natural information" in the data security classification and grading standard as an example, collect label descriptions. According to the data security classification and grading standard, collect detailed descriptions of all third-level sub-categories and their corresponding fourth-level sub-categories:
[0089] Category name of the third-level sub-category: Personal natural information
[0090] Classification description of the third-level sub-classification: refers to the natural attribute information of an individual, that is, the attributes that the individual possesses, rather than the attribute information generated in the course of business operations of financial institutions.
[0091] Category name for the fourth-level sub-category: Personal contact information
[0092] Classification description of the fourth-level sub-classification: refers to various types of personal communication contact data, such as mobile phones, landlines, email addresses, WeChat accounts, etc.
[0093] The extracted phrases are paired with the corresponding labels (third-level sub-categories or fourth-level sub-categories) to form the training data in the target dataset. Figure 4 As shown in the figure, the structure of the training dataset and some content examples are shown in the table.
[0094] M3. Set the initial weight of the three-level training data to 1; generate sentence pairs based on the three-level training data; introduce the pre-trained sentence embedding model; use the generated sentence pairs to fine-tune the pre-trained sentence embedding model; use the fine-tuned sentence embedding model to convert the short text of the three-level training data into an embedded vector representation, and save the sentence embedding model;
[0095] M4, select the three-level sub-classification model;
[0096] M5. Use the encoded sample embedding vector and the corresponding three-level training data labels and weights to train the three-level sub-classification model;
[0097] M6. Input the target data set to be predicted and set the initial weight of the sample data to 1. The sample data has the same format as the training data, including short text, labels, and weights. The sample data for prediction stored in the target data set is in fields as the smallest unit.
[0098] M7. Use the sentence embedding model to convert the short text of the sample data of the target dataset into an embedded vector representation;
[0099] M8. Use the three-level sub-classification model to predict the sample data of the target data set;
[0100] M9. Record the three-level prediction results and the three-level prediction probabilities, and match them to the corresponding classification and grading results;
[0101] M10, review the third-level forecast results,
[0102] When the third-level prediction results need to be corrected, correct the wrong third-level prediction results, collect the corrected third-level prediction results and record them in the misclassified samples, and adjust the weight ω according to the prediction probability p using the formula ω=α+β×p; merge the misclassified samples into the training data set, and return to M3 to retrain or fine-tune the third-level sub-classification model;
[0103] For example:
[0104] Input field: "Transaction record table customer age".
[0105] Model prediction results: The third-level sub-classification is "customer behavior information" with a prediction probability of 0.65.
[0106] After proofreading, the third-level sub-category should be "Personal Natural Information".
[0107] During the proofreading process, you can also remove samples that are obviously unreasonable or noisy to ensure the quality of misclassified samples. The screening criteria are: if the model prediction is wrong and the sample is representative, then the sample is retained; if the sample content is vague, incomplete, has appeared before, or contains wrong information, then the sample is removed.
[0108] You can also adjust the weight of each misclassified sample, usually slightly lower than the weight of the initial training sample, to reflect its importance in model training. For example: adjust it to 0.9.
[0109] Or adjust the weights according to the following weighting strategy:
[0110] First, for each misclassified example, record the model’s confidence p (prediction probability) in the wrong prediction.
[0111] Input field: "User Information Form Mobile Number".
[0112] Actual label: Third-level subcategory "Personal contact information".
[0113] Model prediction: The third-level sub-classification is “account information”, and the predicted probability p is 0.85.
[0114] Then, the weight of the misclassified samples is calculated using the formula ω=α+β×p, where ω is the weight of the misclassified samples, p is the predicted probability, α, β are adjustable parameters that control the weight range and sensitivity to the predicted probability, and the initial settings are α=0.5 and β=0.5.
[0115] In the above example, the weight ω=0.5+0.5×0.85=0.925.
[0116] Incorporate misclassified samples into the training data set and return to M3 to retrain or fine-tune the three-level sub-classification model;
[0117] For example, the filtered misclassified samples are incorporated into the training data set, including the short text, the correct label and the set weight, which is actually updating the training data set. For the above misclassified sample example of "customer age in transaction record table", the weight is adjusted to 0.9 and added to the training data set. Figure 5 A sample data table is shown.
[0118] In the model training phase, a weighted strategy can be introduced for all samples (including initial training samples and misclassified samples) to train the model: the loaded training data contains short text, label and weight columns; when training the classification model, the sample weight is passed to the training algorithm (for example, the sample_weight parameter of logistic regression); the encoded samples are trained; Figure 6 The training sample example table shown in the figure shows that the first two rows are initial training samples, and the third row is misclassified samples.
[0119] After weighted training, the model can pay more attention to samples that were previously predicted with high confidence but incorrectly, and improve the classification accuracy of similar samples. By adjusting the values of α and β, the weight range of misclassified samples can be controlled.
[0120] If you want the weight of misclassified samples to be between 0.8 and 1.0, you can set α = 0.8 and β = 0.2. By giving misclassified samples weights related to their confidence in wrong predictions, the model pays more attention to confident but wrong predictions during training. Appropriate weight settings can prevent the model from paying too much attention to misclassified samples, maintain the generalization ability of the model, and effectively improve the model's error correction ability and generalization performance.
[0121] This weighting strategy can be applied to the training process of both the three-level sub-classification model and the four-level sub-classification model.
[0122] After all the prediction results of the current three-level sub-classification model are reviewed, or when all the prediction results do not need to be revised, the trained three-level sub-classification model is stored for use in the system prediction results.
[0123] Each third-level sub-category will definitely correspond to one or more fourth-level sub-categories. For example, personal natural information corresponds to 9 fourth-level sub-categories, and some third-level sub-categories only correspond to 1 fourth-level sub-category.
[0124] When a third-level sub-classification corresponds to only one fourth-level sub-classification, there is no need to train a fourth-level sub-classification model.
[0125] When there are more than or equal to 2 fourth-level subcategories corresponding to the third-level subcategories, it is necessary to train the fourth-level subcategories model. The purpose of the fourth-level subcategories model is to further determine the fourth-level subcategories after the third-level subcategories are determined. The process of training the fourth-level subcategories model for the fourth-level subcategories data corresponding to the current third-level subcategories data is:
[0126] F1. Extract the fourth-level sub-classification data corresponding to the third-level sub-classification data from the classification data set, process the fourth-level sub-classification data in a unified format, and construct a four-level training data set;
[0127] F2. Set the initial weight of the four-level training data to 1; generate sentence pairs based on the four-level training data; introduce the pre-trained sentence embedding model; use the generated sentence pairs to fine-tune the pre-trained sentence embedding model; use the fine-tuned sentence embedding model to convert the short text of the four-level training data into an embedding vector representation, and save the sentence embedding model;
[0128] F3, select the four-level sub-classification model;
[0129] F4, use the encoded sample embedding vector and the corresponding four-level training data labels and weights to train the four-level sub-classification model;
[0130] F5. Enter the target data set to be predicted and set the initial weight of the sample data to 1. The sample data has the same format as the training data, including short text, label, and weight.
[0131] F6. Use the sentence embedding model to convert the short text of the sample data of the target dataset into an embedded vector representation;
[0132] F7. Use the four-level sub-classification model to predict the sample data of the target data set;
[0133] F8. Record the four-level prediction results and four-level prediction probabilities, and match them to the corresponding classification results;
[0134] F9, review the four-level forecast results,
[0135] When the four-level prediction results need to be corrected, correct the wrong four-level prediction results, collect the corrected four-level prediction results and record them in the misclassified samples, and adjust the weight ω according to the prediction probability p using the formula ω=α+β×p; merge the misclassified samples into the four-level training data set, and return F2 to retrain or fine-tune the four-level sub-classification model;
[0136] When the four-level prediction results do not need to be corrected, the three-level sub-classification model and the four-level sub-classification model corresponding to the three-level sub-classification are all trained and used for the system prediction results.
[0137] The training data characteristics of the training data set used in the present invention are:
[0138] Unified format: The phrase form is "table name + field name", such as "customer information table mobile phone".
[0139] Concise and clear: highlight the field name, so that the model can learn the direct relationship between the field and the classification.
[0140] The application effect of constructing classified data:
[0141] Through the data preparation strategy of label description parsing, the classification model can learn the key features of each category in the initial training stage. In the subsequent prediction process, the classification model can more accurately classify the input field into the correct third-level and fourth-level sub-categories.
[0142] For example, for the input field "customer information form mobile phone", the classification model can predict with high probability that its third-level sub-category is "personal natural information" because it has learned similar phrases during training.
[0143] The data preparation strategy for label description parsing is the foundation of the entire system. By parsing and extracting official label descriptions, a high-quality initial training data set is quickly constructed. This strategy is applied in the key steps of model training, prediction, and tuning architecture, ensuring effective model training and laying a solid foundation for achieving high-accuracy financial data security classification and grading.
[0144] Using the updated training data set to retrain the model is essentially: loading the training data set containing misclassified samples, and the initial classification data and previous misclassified samples are still retained in the training data set. Retraining the model according to the model training process can improve the model's classification performance.
[0145] The updated classification model is deployed to the system and applied to subsequent prediction tasks to improve the classification accuracy of new data.
[0146] When a prediction is wrong, record the misclassified samples, set the weights, and add them to the corresponding model training data. By continuously collecting and introducing new misclassified samples, the training data set is gradually enriched and the model performance is continuously improved. When the number of newly added misclassified samples reaches a certain scale, or the model performance needs to be improved, the classification model can be retrained.
[0147] The error-driven incremental learning strategy continuously enriches and updates the model's training data set by collecting, screening, and introducing misclassified samples. By setting appropriate sample weights, the model can incrementally learn samples that were previously mispredicted and improve the classification accuracy of similar samples. This strategy is applied in multiple key steps of the system to ensure that the model is continuously optimized and improves performance in practical applications.
[0148] The classification model can be trained in the above manner. Furthermore, the following manners can be combined arbitrarily to improve the accuracy of the classification results of the classification model.
[0149] For example: the target data set is stored in a database; the database contains multiple tables, and each table contains multiple fields;
[0150] The sample data stored in the target data set is based on fields as the smallest unit. The interpretation of the sample data includes the name of the data table and the description of the corresponding content, the name of the field and the description of the corresponding content.
[0151] For example, the content of classification data includes: third-level sub-classification and third-level sub-classification description, fourth-level sub-classification and fourth-level sub-classification description. The third-level sub-classification description is a complete description of each third-level sub-classification formed by combining the fourth-level sub-classification description below it.
[0152] For example, when reviewing prediction results, you can eliminate samples that are obviously unreasonable or noisy to ensure data quality.
[0153] In the above example, the third-level sub-classification description is preferably a complete description of each third-level sub-classification formed in combination with the fourth-level sub-classification description below it; the classification description corresponds to the classification data, and the classification description of each sub-classification is a combination of multiple keywords parsed and extracted from the description text of the sub-classification in the specification file.
[0154] In the above example, the main model, which is responsible for predicting all three-level sub-categories, belongs to the “1” in the “1+N” hierarchical classification strategy.
[0155] A three-level sub-classification model was trained using a few-shot learning-based model framework.
[0156] Based on the above example, a four-level sub-classification model is constructed, which belongs to the "N" in the "1+N" hierarchical classification method, that is, the sub-model. For each three-level sub-classification, a corresponding four-level sub-classification model is constructed.
[0157] Use as Figure 1 The classification prediction (three-level sub-classification prediction) process shown is:
[0158] Input sample: field information to be predicted, such as "customer name in transaction record table"
[0159] Model prediction: The model predicts the input samples and obtains the prediction results and probabilities of the three-level sub-classifications.
[0160] Prediction result: "Personal natural information"; prediction probability: 0.80.
[0161] Proofread the prediction results of the three-level sub-classification. If the prediction is wrong: modify it to the correct three-level sub-classification; record the misclassified samples, including the input text, correct label, model prediction results and prediction probability, add them to the classification data, set the weight, and add them to the classification data set.
[0162] The above example adopts the "1+N" hierarchical classification model strategy. The four-level sub-classification model actually partially reuses the model training and fine-tuning methods of the three-level sub-classification, reducing the complexity of the model. Through hierarchical model design, the three-level sub-classifications with relatively few categories are first predicted, and then the four-level sub-classifications are predicted under the corresponding three-level sub-classifications, reducing the number of categories that each model needs to process. Hierarchical prediction helps the model capture the features of different levels more accurately and improve the overall classification performance.
[0163] Key points in model training and prediction: You need to ensure that the format of training data and prediction data is consistent, and the model uses the "table name + field name" phrase.
[0164] During the model training and tuning process, the combination of label description parsing, error-driven incremental learning, and flexible weighting for misclassified samples significantly improved the model performance.
[0165] Through the "1+N" hierarchical classification model strategy, the system can efficiently handle a large number of classification tasks, reduce the complexity of a single model, and improve the accuracy and robustness of classification. In the process of model training and prediction, combined with the application of other strategies, the model performance is continuously optimized to meet the high standards of financial data security classification and grading.
[0166] In any of the above examples, the classification model can be a logistic regression model or other classification models. The pre-trained sentence embedding model is preferably an open source, Chinese pre-trained model with high performance and small model size, which can run under limited resources (only CPU is required). The model fine-tuning tool is preferably setfit.
[0167] The financial data security classification and grading system based on multi-strategy learning described in the present invention adopts a data preparation strategy of label description parsing, which has multiple beneficial effects:
[0168] 1. Efficiently obtain high-quality training data: Without a large amount of manual labeling, you can directly use the officially released classification label descriptions to quickly build the initial training data set, reducing the cost and time of data preparation.
[0169] 2. Matching data format with model requirements: According to the input requirements of the model, the label description is parsed, key phrases are extracted, and the training requirements of the model are adapted. Training data suitable for the model is extracted to ensure that the training data and prediction data formats are consistent. High-quality initial training data enables the model to have good classification capabilities from the beginning, reducing the cold start problem.
[0170] 3. Wide applicability: Since the data comes from the "Financial Data Security Data Security Classification Guidelines" issued by the People's Bank of China, each institution generally derives its own security classification and grading specifications based on this guideline. Therefore, this methodology is applicable to most financial institutions.
[0171] 4. Use official authoritative data: Extract classification descriptions from the "Financial Data Security Data Security Classification Guidelines" issued by the People's Bank of China to ensure the accuracy and authority of the data.
[0172] By continuously introducing misclassified samples, the model can gradually correct errors, conduct reinforcement learning for its own weaknesses, adapt to new data characteristics, and reduce the occurrence of misclassification. The model has the ability to self-improve. As the usage time increases, the accuracy of the model is also constantly improving, and it can quickly adapt to the characteristics of different systems and data. Because it can perform incremental learning, when an organization completes the labeling of all systems, the model has learned to a very good level and can be migrated to other organizations for use, reducing the model training costs of new organizations. The model's migration ability enables experience and results to be transferred between different organizations, accelerating the maturity of the model.
[0173] The key points of this strategy are the collection and screening of misclassified samples (in the process of prediction and manual review, misclassified samples are collected in a timely manner and manually screened to ensure data quality), and iterative updating of the model (adding new misclassified samples to training data, retraining or fine-tuning the model to achieve incremental learning of the model).
[0174] By assigning dynamic weights to misclassified samples, the model can focus on learning samples with high confidence but incorrect predictions. Flexible weighting strategies avoid over-focusing on specific samples and maintain the overall performance of the model. During the training process, the model pays more attention to samples that are prone to errors, reduces the recurrence of similar errors, and enhances the model's error correction capabilities. By adjusting the weights, the influence of initial training samples and misclassified samples in model training is balanced.
[0175] The key point of this strategy lies in the design of the weight calculation formula, and by adjusting the values of α and β, the weight range of misclassified samples can be controlled to adapt to different training requirements.
[0176] Through hierarchical model design, each model only needs to process relatively few categories, which reduces the training and prediction burden of the model. The hierarchical prediction method allows the model to focus more on specific features at each level, improving the accuracy of overall classification. The model can quickly and accurately complete the classification task of large amounts of data. When business needs change, sub-models can be flexibly added or adjusted without affecting the entire model system, making the model easier to maintain and expand.
[0177] The key points of this strategy are: establish a main model for all three-level sub-classifications, and a four-level sub-classification model for each three-level sub-classification; first predict the three-level sub-classifications, and then use the corresponding four-level sub-classification model for detailed prediction after review or proofreading.
[0178] In summary, the financial data security classification and grading system based on multi-strategy learning described in the present invention has significantly improved the classification accuracy of the model through the comprehensive application of the above four strategies, meeting the high standards for financial data security classification and grading. The model based on few-sample learning is adopted, combined with a flexible weighting strategy, to achieve efficient training and prediction under limited resources. Since the methodology is based on the grading guidelines of the People's Bank of China, each institution will derive its own security classification and grading specifications based on this, so this method has wide applicability and is suitable for most financial institutions. Through incremental learning and model portability, after an institution completes the labeling of all systems, the model has reached a high level and can be migrated to other institutions for use, reducing the model training costs of new institutions and accelerating the maturity and sharing of the model.
[0179] The key technical points of the financial data security classification and grading system based on multi-strategy learning in the present invention to obtain accurate classification are:
[0180] 1. Organic integration of multiple strategies: The four strategies work together to build a complete and efficient solution from data preparation, model training, prediction to optimization.
[0181] 2. Continuous model optimization: Through error-driven incremental learning and flexible weighting strategies, the model can continuously self-learn and improve performance.
[0182] 3. Effective use of existing resources: Make full use of officially released label descriptions and a small amount of hardware resources to reduce development and maintenance costs.
[0183] 4. Adaptability and scalability: The model system has good adaptability and can be customized and expanded according to the specific needs of different institutions, supporting the continuous optimization and promotion of the model.
[0184] 5. Reduce the workload of manual review: Due to the high accuracy and continuous optimization capabilities of the model, the workload of manual review is significantly reduced, improving overall work efficiency.
[0185] 6. Rapid deployment and implementation: As the solution is resource-friendly and highly applicable, the model can be quickly deployed and implemented in different organizations, shortening the project cycle.
[0186] 7. Improve data security management: This solution helps financial institutions improve data security management, meet regulatory requirements, and reduce data security risks.
[0187] This technical solution successfully builds a set of efficient, accurate, widely applicable and resource-friendly financial data security classification and grading system through the innovative application of four core strategies. The highlight of each strategy is that it solves the key problems in the classification and grading of financial data in a targeted manner, achieving the effect of improving classification accuracy, enhancing model adaptability and portability, and reducing resource consumption. The key point is the organic combination of strategies to form a complete and coordinated technical solution, which provides strong technical support for the data security management of financial institutions and has broad application prospects.
Claims
1. A financial data security classification and grading system based on multi-strategy learning, characterized in that: include: Sentence embedding models, multiple classification models, The sentence embedding model is used to convert short texts of training data into embedding vectors for use by the classification model; The classification model introduces the embedding vector converted by the sentence embedding model and the labels and weights of the corresponding training data to train the classification model; Among them, each classification model is constructed as follows: S1. Prepare classification data: Extract the content in the specification file and convert it into classification data. The content of the classification data includes classification name and classification description. All classification data are accumulated to form a classification data set. S2. Extract classified data from the classified data set as initial training data according to preset rules, process the classified data in a unified format, and construct a training data set. The training data includes short text, labels, and weights. The initial weight of the training data is set to 1. S3, generate sentence pairs based on training data; S4, introduce pre-trained sentence embedding model; S5. Use the generated sentence pairs to fine-tune the pre-trained sentence embedding model; S6. Use the fine-tuned sentence embedding model to convert the short text of the training data into an embedding vector representation and save the sentence embedding model. S7, select a classification model; S8. Use the encoded sample embedding vector and the labels and weights of the corresponding training data to train the classification model; S9. Input the target data set to be predicted. The sample data of the target data set has the same format as the training data, including short text, label, and weight. S10, using the sentence embedding model, converting the short text of the sample data of the target dataset into an embedded vector representation; S11. Use the classification model to predict the sample data of the target data set; S12, record the prediction results and prediction probabilities, and match them to the corresponding classification and grading results; S13, review the prediction results, When the prediction results need to be corrected, correct the wrong prediction results, collect the corrected prediction results and record them in the misclassified samples, and adjust the weight ω according to the prediction probability p using the formula ω=α+β×p; incorporate the misclassified samples into the training data set, and return to S3 to retrain or fine-tune the classification model; When the prediction results do not need to be corrected, the fine-tuned embedding model and the trained classification model are stored for use in the system prediction results.
2. A financial data security classification and grading system based on multi-strategy learning as claimed in claim 1, characterized in that: The target data set is stored in a database; the database contains multiple tables, and each table contains multiple fields; The classified data stored in the classified data set is based on fields as the smallest unit. The corresponding interpretation of the classified data includes the name of the data table and the description of the corresponding content, the name of the field and the description of the corresponding content.
3. A financial data security classification and grading system based on multi-strategy learning as claimed in claim 1, characterized in that: The content of the classification data includes: third-level sub-classification and third-level sub-classification description, fourth-level sub-classification and fourth-level sub-classification description.
4. A financial data security classification and grading system based on multi-strategy learning as claimed in claim 1, characterized in that: In step S13, when reviewing the prediction results, it includes the step of eliminating misclassified samples that are obviously unreasonable or have large noise to ensure data quality.
5. A financial data security classification and grading system based on multi-strategy learning as claimed in claim 3, characterized in that: The third-level sub-category description is a complete description of each third-level sub-category formed by combining the fourth-level sub-category descriptions below it.
6. A financial data security classification and grading system based on multi-strategy learning as claimed in claim 5, characterized in that: The classification model includes a three-level sub-classification model, and a plurality of four-level sub-classification models corresponding to the three-level sub-classification model; When the three-level sub-classification model is constructed, the content of the classification data includes the three-level sub-classification name and the three-level sub-classification description; after executing the classification model construction process, the retrained or fine-tuned classification model is used as the three-level sub-classification model, and the three-level sub-classification model is used to classify all target data of the target data set to obtain the three-level sub-classification prediction result of the target data; In step S12, the classification and grading results include: third-level sub-classifications and their probabilities, fourth-level sub-classifications and their probabilities, and corresponding security levels; When the four-level sub-classification model corresponding to each three-level sub-classification model is constructed, the content of the classification data is the four-level sub-classification name and the four-level sub-classification description corresponding to the current three-level sub-classification; after executing the classification model construction process, the retrained or fine-tuned classification model is used as the four-level sub-classification model, and the four-level sub-classification model is used to classify all target data of the target data set to obtain the four-level sub-classification prediction results of the target data.
7. A financial data security classification and grading system based on multi-strategy learning as claimed in claim 1, characterized in that: The classification description corresponds to the classification data, and the classification description of each subclass is a combination of multiple keywords parsed and extracted from the description text of the subclass in the specification file.
8. The financial data security classification and grading system based on multi-strategy learning as claimed in claim 1, characterized in that: The format of the classification data of the classification data set is consistent with the prediction data format.
9. A financial data security classification and grading system based on multi-strategy learning as claimed in claim 1, characterized in that: The classification model used is the SetFit model.
10. A financial data security classification and grading system based on multi-strategy learning as claimed in claim 1, characterized in that: The initial value of the tag is empty.
Citation Information
Cited By
Surveying and mapping result classification model fusing time dynamic weight and use method
CN121030475A