Data processing method and device, equipment, storage medium and program product
Through a four-stage intelligent agent collaborative process, the problem of low efficiency of manual operation and unreliable automated classification results in financial data classification and grading is solved, achieving efficient and reliable data classification and grading, which is suitable for data security governance in the financial industry.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies for classifying and grading financial data suffer from problems such as low efficiency of manual operations, lack of verification and optimization mechanisms for automated classification results, and insufficient reliability of results, especially in complex semantic or ambiguous scenarios where accurate classification is difficult.
A four-stage intelligent agent collaborative process is adopted. The first intelligent agent performs data standardization processing, the second intelligent agent performs preliminary classification based on the data classification specification knowledge base, the third intelligent agent performs risk assessment and deviation correction, and the fourth intelligent agent resolves conflicts, ultimately generating reliable classification labels.
It has achieved automated processing of data classification and grading, improved the credibility of classification results, reduced the workload of manual review, and ensured the compliance and business adaptability of classification results.
Smart Images

Figure CN121858673A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more particularly to a data processing method, apparatus, device, storage medium, and program product. Background Technology
[0002] Financial institutions need to process massive amounts of financial data every day. The classification and grading of this data is a core part of financial data security governance. It needs to be divided according to the requirements of the regulations, based on dimensions such as data sensitivity and compliance requirements.
[0003] Existing data classification and grading technologies primarily rely on manual review or simple rule-based automated tools. In the manual review model, professionals must analyze the business attributes, content characteristics, and potential risks of each data field according to classification and grading standards, manually labeling the data category and security level. Regarding automated tools, existing technologies mostly employ rule engines based on keyword matching or shallow machine learning models to classify data.
[0004] However, while manual review can guarantee a certain level of accuracy, it suffers from high labor costs, low processing efficiency, and poor result consistency. Automated rule engines cannot handle complex semantics or ambiguous scenarios, leading to biased classification results. Basic models lack the ability to deeply analyze data context and potential risks, making them prone to misjudgments. Furthermore, existing solutions are typically single-point processing flows, lacking multi-stage collaborative mechanisms and failing to dynamically verify and optimize model output, resulting in insufficient reliability of results. Summary of the Invention
[0005] This application provides a data processing method, apparatus, device, storage medium, and program product to solve the problems of low efficiency in manual operation and low reliability of automated classification results when classifying and grading data.
[0006] Firstly, this application provides a data processing method, including:
[0007] Get annotation requests for the target data;
[0008] The first intelligent agent standardizes the metadata, content samples, and context information of the target data to generate a structured data object.
[0009] The second intelligent agent performs a preliminary classification of the structured data object based on the data classification standard knowledge base corresponding to the industry to which the target data belongs, and generates a preliminary classification result.
[0010] A third-party intelligent agent performs risk assessment and bias correction on the preliminary classification results to generate corrected classification results.
[0011] The fourth intelligent agent resolves the conflict between the preliminary classification result and the revised classification result, and generates the classification label of the target data.
[0012] Output the classification labels of the target data.
[0013] Secondly, this application provides a data processing apparatus, comprising:
[0014] The acquisition module is used to acquire annotation requests for target data;
[0015] The first generation module is used to standardize the metadata, content samples and context information of the target data through the first intelligent agent to generate a structured data object;
[0016] The second generation module is used to perform preliminary classification of the structured data object based on the data classification standard knowledge base corresponding to the industry to which the target data belongs, through the second intelligent agent, and generate preliminary classification results;
[0017] The third generation module is used to perform risk assessment and bias correction on the preliminary classification results through a third intelligent agent, and generate corrected classification results.
[0018] The fourth generation module is used to resolve conflicts between the preliminary classification result and the corrected classification result through a fourth intelligent agent, and generate classification labels for the target data.
[0019] The output module is used to output the classification labels of the target data.
[0020] Thirdly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0021] The memory stores computer-executed instructions;
[0022] The processor executes computer execution instructions stored in the memory to implement the data processing method as described in the first aspect and various possible implementations of the first aspect above.
[0023] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions thereon, which, when executed by a processor, are used to implement the data processing method as described in the first aspect and various possible implementations of the first aspect.
[0024] Fifthly, this application provides a program product, including a computer program, which, when executed by a processor, implements the data processing method described above.
[0025] The data processing method, apparatus, equipment, storage medium, and program products provided in this application automate data classification and grading through a four-stage intelligent agent collaborative process: First, the first intelligent agent transforms the raw target data into structured input; second, the second intelligent agent generates preliminary classification results based on rules in the data classification specification knowledge base; subsequently, the third intelligent agent optimizes the preliminary classification results through risk assessment and bias correction; finally, the fourth intelligent agent integrates the outputs of multiple stages and generates the final classification labels. Throughout the process, the intelligent agents form a closed-loop optimization through dynamic interaction, significantly improving the credibility of the classification results while reducing the workload of manual review, ensuring the compliance and business adaptability of the classification results in financial data security governance scenarios. Attached Figure Description
[0026] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0027] Figure 1 A flowchart illustrating a data processing method provided in this application embodiment. Figure 1 ;
[0028] Figure 2 A flowchart illustrating a data processing method provided in this application embodiment. Figure 2 ;
[0029] Figure 3 A flowchart illustrating a data processing method provided in this application embodiment. Figure 3 ;
[0030] Figure 4 A flowchart illustrating a data processing method provided in this application embodiment. Figure 4 ;
[0031] Figure 5 A flowchart illustrating a data processing method provided in this application embodiment. Figure 5 ;
[0032] Figure 6 A schematic diagram of the structure of a data processing device provided in this application;
[0033] Figure 7 This is a schematic diagram of the structure of an electronic device provided in this application.
[0034] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0035] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0036] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, they do not violate public order and good morals, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0037] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0038] It should be noted that the data processing methods, apparatus, devices, storage media, and program products provided in this application can be used in the field of artificial intelligence, or in any field other than artificial intelligence. The application fields of the data processing methods, apparatus, devices, storage media, and program products in this application are not limited.
[0039] First, let me explain the terms used in this application:
[0040] An intelligent agent is a computational entity or system that can function autonomously and continuously in a given environment. It perceives the environment through sensors, acts upon it through effectors, and adjusts its behavior based on environmental feedback to achieve specific goals. It possesses characteristics such as autonomy, responsiveness, sociality, and initiative, and can simulate intelligent human behavior.
[0041] In the financial sector, data security and compliance are directly related to the protection of personal privacy, the safety of institutional operations, and the stability of the financial order. Financial institutions (such as banks) process massive amounts of financial data daily, including customer identity information, transaction records, market data, and internal management data. The classification and grading of this data is a core aspect of financial data security governance, requiring precise categorization based on data sensitivity, leakage risk, and compliance requirements, in accordance with regulatory requirements.
[0042] However, traditional classification and grading processes rely heavily on manual operation, requiring professionals to review data content item by item, match classification rules, and determine security levels based on business scenarios. This model faces pain points such as inefficiency, high labor costs, and poor consistency in results in the context of surging data volumes. For example, a large bank processes hundreds of millions of data records annually, requiring hundreds of person-days for its manual review team. Furthermore, human error or misunderstanding of rules can lead to some data being misclassified, potentially triggering compliance risks or data breaches.
[0043] In terms of automation tools, existing technologies mostly employ rule engines based on keyword matching or shallow machine learning models. For example, they use a pre-defined rule base (such as "ID number" field corresponding to "high sensitivity") or utilize basic classification models to predict labels on data samples.
[0044] However, such methods have significant limitations: First, rule engines cannot handle complex semantics or ambiguous scenarios (such as implicit customer privacy information), leading to biased classification results; second, the basic model lacks the ability to deeply analyze data context and potential risks, making it prone to misjudgment; third, existing solutions are usually single-point processing flows, lacking multi-stage collaborative mechanisms, and cannot dynamically verify and optimize model outputs, resulting in insufficient reliability of results.
[0045] Furthermore, the collaboration model between humans and models is mostly "model output + human review," lacking deep interaction. The decision-making basis of the model is difficult for humans to understand, leading to low collaboration efficiency. In summary, existing technologies struggle to balance the accuracy, efficiency, and compliance of data classification and grading, and fail to meet the high standards of data security governance required by the financial industry.
[0046] To address the aforementioned issues, this application proposes a data processing method that constructs a fully automated processing chain from data input to final classification and grading label output through hierarchical and specialized intelligent agent collaboration. Each intelligent agent undertakes a specific functional module (information processing, preliminary classification, deep verification, and result integration), and through multi-stage collaborative verification and optimization, dynamically corrects and enhances the credibility of the model's output. This method overcomes the limitations of traditional single-point processing models, solving problems such as large discrepancies between AI model output and human review, poor result consistency, and high compliance risks. It provides a structured and interpretable solution for data classification and grading in financial-grade trusted AI scenarios.
[0047] This application can be applied to data security governance scenarios in the financial industry, such as the classification and grading of customer data, transaction data, and credit data of banks and other institutions. In this scenario, data needs to be classified in accordance with standardized requirements. This application, through a multi-agent collaborative framework, can reconstruct the data classification and grading process into a four-stage automated processing chain, and achieve result optimization and human-machine collaboration alignment through dynamic interaction between agents, ultimately outputting classification and grading labels that meet the standardized requirements.
[0048] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0049] Figure 1 A flowchart illustrating a data processing method provided in this application embodiment. Figure 1 .like Figure 1 As shown, the data processing method provided in this embodiment is applied to a data processing system and includes:
[0050] S101: Obtain a labeling request for the target data.
[0051] Understandably, the request to obtain target data is the starting point for the annotation process of target data. The annotation request can be triggered by relevant technical personnel (such as data management personnel) or by an automated system based on preset rules, and is a request to initiate annotation actions on specific target data in accordance with established data classification and grading standards and business needs.
[0052] The annotation request should clearly specify which data should be annotated, what classification and grading standards should be used for the annotation, and so on, so as to provide guidance for the accurate annotation of subsequent data.
[0053] S102: The first intelligent agent standardizes the metadata, content samples, and context information of the target data to generate a structured data object.
[0054] As is understandable, standardization refers to the technical process of transforming raw data into a unified format and structure, including metadata extraction, content sample analysis, and context parsing. For example, the field "customer ID number" is transformed into a structured object containing the field name, data type, content example, and a description of the business scenario.
[0055] The first intelligent agent can be an information processing agent that can standardize the target data. Metadata refers to information describing the characteristics of the data itself, including the data name, data creation time, data format, and the domain to which the data belongs. The first intelligent agent can extract the metadata information of the target data using specific algorithms.
[0056] The generation and construction of information processing intelligent agents takes data standardization as its core objective, requiring the integration of multiple modules such as data cleaning, format conversion, and feature extraction. First, a standardized processing framework is built by defining data standard specifications (such as numerical ranges, text encoding, and missing value imputation rules). Next, a data preprocessing engine is developed, integrating regular expression matching, word segmentation algorithms (for text data), and normalization methods (for numerical data) to ensure that input data can be processed accurately. Finally, an anomaly detection process is embedded, using machine learning models to identify and correct data that deviates from the standards, forming a closed-loop processing flow.
[0057] A content sample refers to a specific value example in the target data. For example, the content sample in the field "Customer ID Number" might be "123456789123456789". The first intelligent agent can use technologies such as natural language processing and machine learning to analyze the semantics, syntactic structure, and the expressed theme and intent in the content sample, thereby understanding the inherent logic of the data content.
[0058] Contextual information refers to the business scenarios or related logic that describe the target data. For example, the table containing the "transaction amount" field may be related to the "bank card transaction records" business scenario. The first intelligent agent can analyze the relationship between a target data and other target data, the business background in which the data was generated, and the position of the data in the business process, thereby comprehensively grasping the contextual information of the data.
[0059] S103: The second intelligent agent performs a preliminary classification of the structured data object based on the data classification standard knowledge base corresponding to the industry to which the target data belongs, and generates a preliminary classification result.
[0060] Understandably, a data classification standard knowledge base refers to a knowledge base that stores the classification and grading rules specified in financial industry standards. For example, the customer's ID number field should be classified as a highly sensitive field.
[0061] The second intelligent agent can be a thinking agent. Based on its own data processing and knowledge application capabilities, the second intelligent agent can classify the target data according to a data classification standard knowledge base built for the industry corresponding to that data. This data classification standard knowledge base includes the industry's long-term accumulated experience, expert knowledge, and industry standards, covering comprehensive knowledge from basic concept definitions to complex business scenario classifications.
[0062] The generation and construction of intelligent agents should revolve around industry knowledge bases and classification algorithms. First, a structured knowledge base is built based on the target industry (e.g., finance), including data classification specifications, contextual association rules, and business logic constraints. Then, natural language processing and machine learning models are integrated to extract features and perform semantic analysis on the standardized data, generating initial classification labels. Finally, a classification decision engine is designed, combining rule trees and statistical models (such as decision trees and support vector machines) from the knowledge base to output preliminary classification results. This intelligent agent needs to be interpretable, meaning it should be able to demonstrate the classification basis through rule tracing functionality and support dynamic updates to the knowledge base.
[0063] The second intelligent agent can perform detailed comparison and precise matching between the generated structured data objects and various classification rules and feature descriptions in the knowledge base, comprehensively analyze the various attributes, features and key information in the structured data objects, and retrieve points that match the classification standards in the knowledge base from different dimensions.
[0064] During the comparison and matching process, the second intelligent agent can use advanced algorithms and logical judgment capabilities to evaluate and judge each possible classification direction, and finally determine a preliminary classification result for the structured data object based on the specifications in the data classification specification knowledge base.
[0065] S104: The preliminary classification results are assessed for risk and biased by a third agent, and a revised classification result is generated.
[0066] Understandably, risk assessment refers to analyzing the compliance or security risks that the classification results may pose. For example, misclassifying the field "customer phone number" as low-sensitivity could potentially lead to privacy breaches. Bias correction refers to correcting unreasonable results in the initial predictions through contextual analysis or rule supplementation. For example, the field "transaction time" is not defined in the rules, but can be inferred to be moderately sensitive based on the context "bank card transaction records".
[0067] The third intelligent agent can be a reflective intelligent agent. It can simulate the critical thinking in human review, and judge whether there are problems such as inapplicability of classification standards or misunderstandings in the preliminary classification results based on the standards and rules on which the preliminary classification results are generated, combined with the results of risk assessment. If such problems are found, the third intelligent agent can reinterpret and apply the classification standards and make corresponding adjustments to the preliminary classification results.
[0068] The construction of a reflective agent requires simulating critical thinking as its core, and building a rule correction engine and feedback mechanism. A rule evaluation module is designed to identify rule flaws in the initial classification (such as misclassification due to over-reliance on a single feature) by comparing historical classification cases with industry expert experience. Then, a rule optimization algorithm is developed, using reinforcement learning or genetic algorithms to generate corrected rules. Finally, the corrected classification results and the basis for the correction are output. This agent can support human intervention, allowing experts to conduct a secondary review of the corrected rules.
[0069] If a problem is found in the preliminary classification result, the third agent can send the judgment process of the preliminary classification result and its own correction suggestions to the second agent, so that the second agent can reclassify according to the standard.
[0070] S105: The fourth agent resolves the conflict between the preliminary classification result and the corrected classification result, and generates the classification label of the target data.
[0071] Understandably, conflict resolution refers to the technical process of integrating classification results from multiple sources through methods such as rule prioritization, confidence weighting, or human feedback. For example, if the initial classification result is "low sensitivity" (confidence 0.7) and the revised suggestion is "medium sensitivity" (confidence 0.8), the revised suggestion should be prioritized. Confidence is the model's assessment of the reliability of the prediction result, typically a floating-point number between 0 and 1.
[0072] The fourth agent can be a summarizing agent. It can review the classification rules and data characteristics on which the preliminary classification results are based. At the same time, the fourth agent can also review the judgment basis and correction rules for correcting the classification results, so as to determine the confidence and reasonableness of the two results.
[0073] The construction of the inductive agent requires the integration of multi-source result review and confidence assessment modules. For preliminary classification results, the classification rules and data features on which they are based are extracted, and their reasonableness score is calculated using a rule matching algorithm. For corrected classification results, the improvement points and basis of the corrected rules are analyzed, and their logical consistency is evaluated. Simultaneously, a confidence calculation model is designed, combining rule rigor, data completeness, and historical accuracy to assign weights to the two types of results. Finally, a report generation module is constructed to generate the final conclusion and attach a detailed review report.
[0074] In this process, the fourth intelligent agent can use algorithmic models and logical reasoning mechanisms to accurately identify and analyze conflicts between the preliminary classification results and the revised classification results, finding the key factors and potential contradictions that lead to the conflicts. It then uses its processing and analytical capabilities to analyze and judge the contradictions, ultimately generating a reasonable, accurate target data classification label that conforms to data characteristics and business needs.
[0075] S106: Output the classification labels of the target data.
[0076] Understandably, after the fourth agent generates the final target data classification label, the data processing system can output this classification label. This classification label can clearly identify the category to which the target data belongs, providing a reliable basis for the subsequent analysis and use of the target data.
[0077] The data processing method provided in this embodiment obtains a labeling request for target data. A first intelligent agent standardizes the metadata, content samples, and contextual information of the target data to generate a structured data object. A second intelligent agent performs a preliminary classification of the structured data object based on a data classification standard knowledge base corresponding to the industry to which the target data belongs, generating a preliminary classification result. A third intelligent agent performs risk assessment and bias correction on the preliminary classification result, generating a revised classification result. A fourth intelligent agent resolves conflicts between the preliminary and revised classification results, generating a classification label for the target data. Finally, the classification label for the target data is output. This method obtains the classification label for the target data through a four-stage collaborative mechanism, improving the credibility of the classification result. Simultaneously, the automated processing chain reduces manual intervention and lowers the cost of manual review.
[0078] Figure 2 A flowchart illustrating a data processing method provided in this application embodiment. Figure 2 .like Figure 2 As shown, in Figure 1 Based on the embodiments, the process of generating classification labels through the fourth intelligent agent is described in detail, including:
[0079] S201: When the confidence levels of the preliminary classification results and the revised classification results are inconsistent, the classification result with higher confidence is determined as the classification label based on the confidence levels of the preliminary classification results and the revised classification results.
[0080] Understandably, for both preliminary and revised classification results, the data processing system can determine the corresponding confidence level based on their respective result generation processes. Confidence level is a quantitative measure of the reliability of the classification result. When the confidence levels of the preliminary and revised classification results are inconsistent, to ensure the accuracy and reliability of the classification labels, the system can select the classification result with the higher confidence level as the final classification label. This is because a higher confidence level indicates that the result was based on more sufficient information, has more rigorous logic, and is more representative of the true category of the data.
[0081] S202: If the confidence levels of the preliminary classification results and the revised classification results are consistent, obtain the historical label data of the target category data provided by human feedback.
[0082] Understandably, if the confidence levels of the preliminary classification results and the revised classification results are at the same level, that is, the reliability of the results output by the two agents is difficult to distinguish directly by confidence level, in order to determine the classification label, the data processing system can obtain the historical label data of the target category data fed back by humans.
[0083] The target category data contains the target data that needs to be classified now. The historical label data for the target category data, provided by human feedback, refers to the category information labeled by humans based on their professional knowledge, experience, and understanding of the data during past data classification processes. By acquiring this historical label data, the system can leverage human experience to provide valuable references for the current classification, thereby more reasonably determining the final classification labels.
[0084] S203: Based on historical label data of target category data from manual feedback, determine the classification label from the preliminary classification results and the revised classification results.
[0085] Understandably, after obtaining historical label data for the target category data from manual feedback, the data processing system can perform data mining and analysis on this historical label data to identify data and their corresponding labels that are similar to the current target data in terms of features and attributes. Through similarity analysis, the system can understand the results of manual classification in similar situations. Then, the preliminary classification results and revised classification results are compared with these analysis results based on historical label data. The most similar result is then used as the final classification label.
[0086] Optionally, classification labels can be determined based on historical label data from manual feedback, or by combining preliminary classification results and revised classification results. Specifically, this includes:
[0087] Based on historical label data of target category data based on human feedback, the matching frequency of preliminary classification results and corrected classification results is calculated, and a matching frequency index is generated.
[0088] Understandably, historical label data can provide a reference for current classification decisions. The data processing system can compare the preliminary classification results and the revised classification results with each label in the historical label data one by one. For the preliminary classification results, the system counts the number of times the same or similar labels appear with those in the historical label data. Similarly, the same statistical operation is performed on the revised classification results.
[0089] After completing the statistical work, the matching frequency of the preliminary classification results and the revised classification results can be calculated based on the number of matches obtained from the statistics and the total amount of historical tag data. The formula for calculating the matching frequency is: Matching Frequency = Number of Matches / Total Amount of Historical Tag Data.
[0090] The two matching probability values calculated using this formula are also two matching frequency indicators. They can intuitively reflect the degree to which the preliminary classification results and the revised classification results fit with historical experience. The higher the matching frequency, the more consistent the classification result is with the classification situation in historical label data, which means that it is more in line with human classification habits and experience judgments to a certain extent.
[0091] Based on the matching frequency index, data classification priority rules, and confidence threshold, the preliminary classification results and the revised classification results are weighted and calculated to generate classification labels.
[0092] Understandably, after considering the matching frequency of two classification results, the system can comprehensively consider the data classification priority rules and confidence thresholds, combining multiple factors to determine the classification label. The data classification priority rules can be formulated according to actual application scenarios and business needs, clarifying the importance and selection order of classification results under different conditions and various specifications.
[0093] The confidence threshold is a threshold used to assess the reliability of the classification results. The confidence threshold can be adjusted appropriately based on the specific classification task and data characteristics.
[0094] The weighted calculation process assigns different weights to the preliminary classification result and the revised classification result based on the matching frequency index, priority rules, and confidence threshold. For example, if the preliminary classification result has a high matching frequency, meets the priority rules, and has a confidence level higher than the set threshold, it may be given a higher weight; conversely, if the revised classification result performs better in these aspects, it will receive a higher weight.
[0095] By using weighted calculations, a score can be obtained, which reflects the overall quality of the preliminary and revised classification results. Finally, the classification result with the higher score can be selected as the final classification label.
[0096] The data processing method provided in this embodiment adopts different strategies under different confidence levels: when the confidence levels are inconsistent, the high-confidence result is selected to ensure that a reliable classification is quickly locked; when the confidence levels are consistent, historical label data with human feedback is introduced to assist decision-making, which improves the fit between the classification results and human classification habits, effectively integrates the efficiency of intelligent classification and the accuracy of human experience, thereby improving the accuracy, reliability and practicality of data classification.
[0097] Figure 3 A flowchart illustrating a data processing method provided in this application embodiment. Figure 3 .like Figure 3 As shown, in Figure 1 Based on the embodiments, the process of generating preliminary classification results through the second intelligent agent is described in detail, including:
[0098] S301: Use a graph neural network model to extract contextual features of structured data objects and generate multidimensional feature vectors.
[0099] Understandably, a graph neural network model refers to a deep learning model that models data relationships through graph structures. Graph structure data consists of nodes and edges; nodes represent data objects, and edges represent relationships between objects. Contextual features refer to the feature representation of the relationships between data fields and table structures, business scenarios, etc. Multidimensional feature vectors refer to vector representations that contain information such as field attributes, business risks, and related fields.
[0100] By inputting structured data objects into a graph neural network model, the model can capture the contextual relationships between nodes and integrate them into the node's feature representation. Through the computation and processing of multiple layers of the graph neural network, the features of each node can be integrated into a multi-dimensional feature vector.
[0101] S302: Based on multi-dimensional feature vectors and a data classification standard knowledge base, a machine learning model is used to perform classification matching and generate preliminary classification results.
[0102] Understandably, a machine learning model is an algorithmic model that can automatically learn patterns and rules from data and use them for classification. Multidimensional feature vectors are used as input to the machine learning model. Based on the input feature vectors and classification information from a data classification knowledge base, the machine learning model learns the mapping relationship between features and categories to classify data objects, thereby generating a preliminary classification result.
[0103] Optionally, the parameters of the machine learning model can be fine-tuned using an incremental learning framework to generate an incremental learning model.
[0104] As is understandable, an incremental learning framework refers to a learning framework that allows the model to update only a subset of its parameters when it receives new data. The parameters of a machine learning model determine how it processes data and its classification ability. Parameter fine-tuning refers to adjusting model parameters through local optimization to improve the model's performance on new data. An incremental learning model is a machine learning model that adapts to changes in new data after parameter fine-tuning within an incremental learning framework.
[0105] When new data arrives, the incremental learning framework can integrate and analyze the new data with the existing training data. Then, based on the characteristics and distribution of the new data, the parameters of the machine learning model are fine-tuned in a targeted manner, enabling the model to maintain good performance in a constantly changing data environment.
[0106] Based on multidimensional feature vectors and a data classification standard knowledge base, an incremental learning model is used for classification matching to generate preliminary classification results.
[0107] Understandably, by inputting multidimensional feature vectors into an incremental learning model, the model can utilize the new knowledge and adjusted parameters learned during the incremental learning process, combined with classification rules and standards in the data classification knowledge base, to analyze and judge the feature vectors. Based on the similarity between the feature vectors and the features of each category, the data objects are assigned to the most likely category, thus obtaining preliminary classification results.
[0108] Preliminary classification results generated using incremental learning models can promptly reflect the impact of new data on classification results, better adapt to dynamic changes in data, and have higher accuracy and reliability.
[0109] The data processing method provided in this embodiment effectively captures complex data relationships through graph neural networks, provides rich information through multi-dimensional feature vectors, and enables the model to adapt to data changes through incremental learning. Ultimately, it improves the accuracy and adaptability of data classification, reduces classification errors caused by dynamic data changes, and enhances the performance and efficiency of the classification system in real-time and variable environments.
[0110] Figure 4A flowchart illustrating a data processing method provided in this application embodiment. Figure 4 .like Figure 4 As shown, in Figure 1 Based on the implementation examples, the update process of the data classification specification knowledge base is described in detail, including:
[0111] S401: Update event of the industry classification standard to which the target data belongs.
[0112] Understandably, an update event refers to a change in the industry classification standards, such as the addition of new industry categories, modification of the definition or scope of existing categories, or deletion of categories that are no longer applicable.
[0113] The data processing system can connect with authoritative channels that publish industry classification standards, such as official websites of relevant departments and information platforms of industry associations. The system can continuously monitor information updates regarding industry classification standards through periodic polling or real-time subscription. Once a new version is released or existing content is changed, the system can confirm that an update to the industry classification standards has occurred.
[0114] S402: Upon detecting an update event, obtain the new clauses of the industry classification standard, parse the new clauses, and generate the machine-executable rule format corresponding to the new clauses.
[0115] Understandably, the new clauses are specific additions or modifications made after the industry classification standards were updated. They detail the new requirements and standards for industry classification.
[0116] Once the data processing system detects the update time, it can obtain the complete new version of the industry classification standard document. Then, using natural language processing technology, it analyzes the new document sentence by sentence to identify new clauses. These new clauses are then parsed to extract key information. Finally, based on a pre-set rule template, this key information is transformed into a machine-executable rule format, such as by writing rule code in a specific programming language or converting it into a defined format supported by the rule engine.
[0117] S403: Add the machine-executable rule format corresponding to the new clause to the data classification specification knowledge base.
[0118] Understandably, after generating the machine-executable rule format corresponding to the new clause, the data processing system can connect to the data classification specification knowledge base. During the supplementation process, the data processing system can first query the knowledge base to see if a rule related to the new clause already exists. If it does, the new rule will be compared and analyzed with the existing rules, and the existing rules will be modified or supplemented accordingly.
[0119] If no relevant rules exist, the new rule will be directly inserted into the appropriate location in the knowledge base. During the update and supplementation process, the system also needs to ensure the consistency and integrity of the knowledge base data to avoid rule conflicts or data errors.
[0120] The data processing method provided in this embodiment monitors industry update events, parses the new clauses after determining the update time, and generates machine-executable rule formats. Finally, the new rules are updated and added to the data classification standard knowledge base. This method can keep pace with changes in industry classification standards, ensuring that data classification criteria remain up-to-date and improving the accuracy and timeliness of data classification.
[0121] Figure 5 A flowchart illustrating a data processing method provided in this application embodiment. Figure 5 .like Figure 5 As shown, in Figure 1 Based on the examples, the process of generating corrected classification results is described in detail, including:
[0122] S501: Assess whether the preliminary classification results conform to industry data classification standards, whether they match the data characteristics of the target data, and whether there are potential classification errors.
[0123] Understandably, data processing systems can employ various evaluation algorithms and models to compare preliminary classification results with industry standards one by one to check whether they meet the requirements. Simultaneously, they can analyze the degree of fit between the classification results and the characteristics of the target data to determine whether the classification accurately reflects the essential characteristics of the data. Furthermore, data processing systems can model different scenarios and conditions to predict the potential risks caused by possible classification errors.
[0124] S502: If at least one of the above conditions is not met, the preliminary classification result is revised based on the evaluation results to obtain a revised classification result.
[0125] Understandably, when the data processing system finds during the evaluation process that the preliminary classification results do not meet at least one of the three requirements of "complying with industry data classification standards", "matching data characteristics with target data", and "not having potential classification error risks", it can initiate a correction process to revise the preliminary classification results.
[0126] The correction process begins by analyzing the specific reasons for non-compliance with the evaluation results. If the data does not conform to industry standards, the industry data classification standards can be understood, and the corrected classification result can be determined based on the understood classification standards. Alternatively, the evaluation results can be pushed to a second agent, allowing the second agent to re-understand the classification standards and determine a new initial classification result.
[0127] If the data characteristics do not match the target data, the data characteristics can be analyzed in depth, and the classification results can be revised based on the data characteristics. If a potential classification error risk is found, the system can analyze the cause and scope of the risk, and take corresponding corrective measures for different types of risks, such as adjusting classification boundaries or adding classification categories. During the correction process, the system can record the correction operations and basis in real time and generate a correction report for subsequent auditing.
[0128] The data processing method provided in this embodiment comprehensively evaluates the preliminary classification results to determine whether they comply with industry standards, match data characteristics, and have any potential classification error risks. If at least one requirement is not met, the preliminary classification results are corrected to obtain corrected classification results. This method can improve the accuracy and compliance of data classification.
[0129] Figure 6 This is a schematic diagram of the structure of a data processing device provided in this application. Figure 6 As shown, this application provides a data processing apparatus, the data processing apparatus 600 including:
[0130] Module 601 is used to obtain annotation requests for target data;
[0131] The first generation module 602 is used to standardize the metadata, content samples and context information of the target data through the first intelligent agent to generate a structured data object;
[0132] The second generation module 603 is used to perform preliminary classification of structured data objects based on the data classification standard knowledge base corresponding to the industry to which the target data belongs, through the second intelligent agent, and generate preliminary classification results;
[0133] The third generation module 604 is used to perform risk assessment and bias correction on the preliminary classification results through a third intelligent agent, and generate corrected classification results.
[0134] The fourth generation module 605 is used to resolve conflicts between the preliminary classification results and the corrected classification results through the fourth intelligent agent, and generate classification labels for the target data.
[0135] Output module 606 is used to output the classification labels of the target data.
[0136] Optionally, the fourth generation module 605 is specifically used to determine the classification result with higher confidence as the classification label when the confidence of the preliminary classification result and the confidence of the corrected classification result are inconsistent.
[0137] The acquisition module 601 is specifically used to acquire historical label data of the target category data fed back by humans, provided that the confidence levels of the preliminary classification results and the corrected classification results are consistent; the target category data includes the target data.
[0138] The fourth generation module 605 is specifically used to determine classification labels from the preliminary classification results and the corrected classification results based on the historical label data of the target category data based on human feedback.
[0139] Optionally, the fourth generation module 605 is specifically used to calculate the matching frequency of the preliminary classification result and the corrected classification result based on the historical label data of the target category data based on human feedback, and generate a matching frequency index.
[0140] The fourth generation module 605 is specifically used to perform weighted calculations on the preliminary classification results and the corrected classification results based on the matching frequency index, the priority rules of data classification and the confidence threshold, and to generate classification labels.
[0141] Optionally, the fourth generation module 605 is specifically used to extract contextual features of structured data objects using a graph neural network model to generate multidimensional feature vectors; based on the multidimensional feature vectors and a data classification standard knowledge base, a machine learning model is used to perform classification matching to generate preliminary classification results.
[0142] Optionally, the fourth generation module 605 is also used to fine-tune the parameters of the machine learning model through the incremental learning framework to generate an incremental learning model; based on the multidimensional feature vector and the data classification specification knowledge base, the incremental learning model is used to perform classification matching to generate preliminary classification results.
[0143] Optionally, the device may also include: a monitoring module 607;
[0144] Monitoring module 607 is used to monitor update events of the industry classification standards to which the target data belongs;
[0145] The first generation module 602 is also used to obtain new clauses of the industry classification standard when an update event is detected, parse the new clauses, generate machine-executable rule formats corresponding to the new clauses, and supplement the machine-executable rule formats corresponding to the new clauses to the data classification standard knowledge base.
[0146] The first generation module 602 is also used to evaluate whether the preliminary classification result conforms to the industry's data classification standards, whether it matches the data characteristics of the target data, and whether there is a potential risk of classification error; if at least one of the above is not met, the preliminary classification result is corrected based on the evaluation result to obtain the corrected classification result.
[0147] The data processing apparatus provided in this application embodiment is similar in principle and technical effect to the implementation of each part of the aforementioned data processing method, and will not be described again here.
[0148] Figure 7 This is a schematic diagram of the structure of an electronic device provided in this application. Figure 7 This application provides an electronic device 700, which includes a receiver 701, a transmitter 702, a processor 703, and a memory 704.
[0149] Receiver 701 is used to receive instructions and data;
[0150] Transmitter 702 is used to send commands and data;
[0151] Memory 704 is used to store instructions executed by the computer;
[0152] The processor 703 is used to execute computer execution instructions stored in the memory 704 to implement the various steps of the data processing method in the above embodiments. For details, please refer to the relevant descriptions in the foregoing data processing method embodiments.
[0153] Optionally, the memory 704 described above can be either standalone or integrated with the processor 703.
[0154] When the memory 704 is set up independently, the electronic device also includes a bus for connecting the memory 704 and the processor 703.
[0155] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the foregoing embodiments, and will not be repeated here.
[0156] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method of any of the foregoing embodiments.
[0157] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method of any of the foregoing embodiments.
[0158] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0159] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0160] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.
[0161] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.
[0162] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.
[0163] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0164] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0165] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0166] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A data processing method, characterized in that, include: Get annotation requests for the target data; The first intelligent agent standardizes the metadata, content samples, and context information of the target data to generate a structured data object. The second intelligent agent performs a preliminary classification of the structured data object based on the data classification standard knowledge base corresponding to the industry to which the target data belongs, and generates a preliminary classification result. A third-party intelligent agent performs risk assessment and bias correction on the preliminary classification results to generate corrected classification results. The fourth intelligent agent resolves the conflict between the preliminary classification result and the revised classification result, and generates the classification label of the target data. Output the classification labels of the target data.
2. The method according to claim 1, characterized in that, The process of resolving conflicts between the preliminary classification result and the revised classification result to generate classification labels for the target data includes: If the confidence levels of the preliminary classification result and the revised classification result are inconsistent, the classification result with the higher confidence level shall be determined as the classification label based on the confidence levels of the preliminary classification result and the revised classification result. If the confidence levels of the preliminary classification results and the revised classification results are consistent, historical label data of the target category data obtained from manual feedback is acquired; the target category data includes the target data. Based on the historical label data of the target category data provided by human feedback, the classification label is determined from the preliminary classification result and the revised classification result.
3. The method according to claim 2, characterized in that, The historical label data based on the target category data from the manual feedback, used to determine the classification label from the preliminary classification result and the revised classification result, includes: Based on historical label data of the target category data based on the manual feedback, the matching frequency of the preliminary classification result and the revised classification result is calculated, and a matching frequency index is generated; Based on the matching frequency index, the priority rules for data classification, and the confidence threshold, the preliminary classification result and the revised classification result are weighted and calculated to generate the classification label.
4. The method according to any one of claims 1-3, characterized in that, The preliminary classification of the structured data objects based on the data classification standard knowledge base, generating preliminary classification results, includes: The contextual features of the structured data object are extracted using a graph neural network model to generate a multidimensional feature vector; Based on the multidimensional feature vector and the data classification standard knowledge base, a machine learning model is used for classification matching to generate the preliminary classification result.
5. The method according to claim 4, characterized in that, Before generating the preliminary classification result by performing classification matching using a machine learning model based on the multidimensional feature vector and the data classification specification knowledge base, the process further includes: The incremental learning model is generated by fine-tuning the parameters of the machine learning model using an incremental learning framework. The preliminary classification result is generated by using a machine learning model for classification matching based on the multidimensional feature vector and the data classification specification knowledge base, including: Based on the multidimensional feature vector and the data classification standard knowledge base, the incremental learning model is used to perform classification matching and generate the preliminary classification result.
6. The method according to any one of claims 1-3, characterized in that, Before the step of standardizing the metadata, content samples, and context information of the target data through the first intelligent agent, the method further includes: Monitor updates to the industry classification standards to which the target data belongs; Upon detecting an update event, the new clauses of the industry classification standard are obtained, the new clauses are parsed, and a machine-executable rule format corresponding to the new clauses is generated. The machine-executable rule format corresponding to the new clauses will be added to the data classification specification knowledge base.
7. The method according to any one of claims 1-3, characterized in that, The step of performing risk assessment and bias correction on the preliminary classification results to generate corrected classification results includes: The preliminary classification results are evaluated to determine whether they conform to the industry's data classification standards, whether they match the data characteristics of the target data, and whether there are any potential classification errors. If at least one of the above conditions is not met, the preliminary classification result is revised based on the evaluation results to obtain the revised classification result.
8. A data processing apparatus, characterized in that, include: The acquisition module is used to acquire annotation requests for target data; The first generation module is used to standardize the metadata, content samples and context information of the target data through the first intelligent agent to generate a structured data object; The second generation module is used to perform preliminary classification of the structured data object based on the data classification standard knowledge base corresponding to the industry to which the target data belongs, through the second intelligent agent, and generate preliminary classification results; The third generation module is used to perform risk assessment and bias correction on the preliminary classification results through a third intelligent agent, and generate corrected classification results. The fourth generation module is used to resolve conflicts between the preliminary classification result and the corrected classification result through a fourth intelligent agent, and generate classification labels for the target data. The output module is used to output the classification labels of the target data.
9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 7.
11. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.