Data classification method and device, computer device and storage medium
By combining the initial classification and deep classification models of metadata in the BI system through edge processing nodes, the problem of low data classification accuracy and efficiency in existing technologies is solved, and efficient and accurate metadata management is achieved.
Patent Information
- Application Number
- CN202610732046.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-08-25
AI Technical Summary
Existing data classification methods in BI systems suffer from low accuracy and efficiency, especially when faced with large amounts of diverse metadata, making it difficult to achieve efficient and accurate identification and management.
Metadata is initially classified by edge processing nodes, category tags are added, and the data is encrypted and transmitted to the data management platform. The metadata to be classified is analyzed using a deep classification model, and the classification results are finally integrated to improve classification accuracy and efficiency.
Distributed data classification was achieved, which improved data processing efficiency. Furthermore, the data volume was reduced and classification accuracy and efficiency were improved through a deep classification model.
Smart Images

Figure CN122634483A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a data classification method, apparatus, computer equipment, and storage medium. Background Technology
[0002] As enterprises deepen their digital transformation, Business Intelligence (BI) systems have become core infrastructure supporting business decision-making, such as telemedicine systems, intelligent triage systems, electronic medical record systems, securities asset management systems, banking systems, insurance systems, and risk management systems. BI platforms generate a large amount of metadata every day. How to efficiently and accurately identify, classify, and manage this metadata directly affects data asset inventory, data governance, data quality monitoring, and subsequent value mining.
[0003] Existing classification methods only perform conceptual-level logical categorization of metadata, such as risk management metadata, insurance claims metadata, and medical record metadata. Furthermore, due to the massive volume and diverse data types of BI metadata, traditional classification models suffer from low accuracy and efficiency. Therefore, improving data classification accuracy and efficiency has become a pressing issue. Summary of the Invention
[0004] This application provides a data classification method, apparatus, computer equipment, and storage medium to improve the accuracy and efficiency of data classification.
[0005] Firstly, this application provides a data classification method, the method comprising: Based on the edge processing nodes corresponding to each data source platform, the metadata of each data source platform is initially classified to obtain a first classification result, and based on the first classification result, classification tags are added to the metadata of each category. The metadata is encrypted and transmitted to the data management platform, and the metadata to be classified is extracted from each of the metadata based on the classification tags; Based on a preset deep classification model, the metadata to be classified is analyzed to obtain a second classification result; The first classification result and the second classification result are integrated to obtain the target classification result.
[0006] Secondly, this application also provides a data classification apparatus, the apparatus comprising: The initial classification module is used to perform initial classification of the metadata of each data source platform based on the edge processing node corresponding to each data source platform, obtain a first classification result, and add classification tags to each type of metadata based on the first classification result; The data extraction module is used to encrypt and transmit the metadata to the data management platform, and extract the metadata to be classified from each of the metadata based on the classification tags; The deep classification module is used to analyze the metadata to be classified based on a preset deep classification model to obtain a second classification result; The result integration module is used to integrate the first classification result and the second classification result to obtain the target classification result.
[0007] Thirdly, this application also provides a computer device, the computer device including a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the data classification method as described above when executing the computer program.
[0008] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the data classification method described above.
[0009] This application discloses a data classification method, apparatus, computer equipment, and storage medium. Based on edge processing nodes corresponding to each data source platform, it performs initial classification of the metadata of each data source platform to obtain a first classification result, and adds classification tags to each type of metadata based on the first classification result. The metadata is then encrypted and transmitted to a data management platform, and metadata to be classified is extracted from each metadata based on the classification tags. Based on a preset deep classification model, the metadata to be classified is analyzed to obtain a second classification result. The first and second classification results are then integrated to obtain a target classification result. This application first performs initial classification of the metadata of each data source platform through edge processing nodes to obtain a first classification result, thus improving data processing efficiency through distributed data classification via edge processing nodes. Then, based on the first classification result, the metadata to be classified that requires deep classification is extracted, and deep classification is performed using a deep classification model, improving classification accuracy and reducing the amount of data that the deep classification model needs to process, thereby improving classification efficiency. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic flowchart of a data classification method provided in the first embodiment of this application; Figure 2 This is a schematic flowchart of a data classification method provided in the second embodiment of this application; Figure 3 This is a schematic flowchart of a data classification method provided in the third embodiment of this application; Figure 4 A schematic block diagram of a data classification device provided for embodiments of this application; Figure 5 A schematic block diagram of the structure of a computer device provided for an embodiment of this application. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0014] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0015] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0016] This application provides a data classification method, apparatus, computer device, and storage medium. The data classification method can be applied to a server, using edge processing nodes to perform initial classification of the metadata of each data source platform, obtaining a first classification result. This distributed data classification via edge processing nodes improves data processing efficiency. Then, based on the first classification result, the metadata to be classified for deep classification is extracted, and deep classification is performed using a deep classification model, improving classification accuracy and reducing the amount of data the deep classification model needs to process, thus improving classification efficiency. The server can be a standalone server or a server cluster.
[0017] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0018] Please see Figure 1 , Figure 1 This is a schematic flowchart illustrating a data classification method provided in an embodiment of this application. This data classification method can be applied to a server to perform initial classification of the metadata of each data source platform through edge processing nodes, obtaining a first classification result. Distributed data classification via edge processing nodes improves data processing efficiency. Then, based on the first classification result, the metadata to be classified for deep classification is extracted, and deep classification is performed using a deep classification model, improving classification accuracy and reducing the amount of data that the deep classification model needs to process, thus improving classification efficiency.
[0019] like Figure 1 As shown, the data classification method specifically includes steps S101 to S104.
[0020] S101. Based on the edge processing nodes corresponding to each data source platform, perform initial classification on the metadata of each data source platform to obtain a first classification result, and add classification tags to each type of metadata based on the first classification result; In one embodiment, each data source platform deploys localized edge processing nodes to perform preliminary processing and classification of metadata according to a rule-based classification model, thereby improving processing efficiency. Data source platforms can be BI tools, databases, and ETL tools, etc. Examples include telemedicine platforms, intelligent triage platforms, electronic medical record platforms, securities asset management platforms, bank deposit platforms, and risk management platforms.
[0021] In one embodiment, the metadata of each data source platform includes field names, field descriptions, data types, field value distribution (such as the number of unique values and the null value rate), and the position of the field in the report. Examples include patient ID, medical record number, symptom description, examination indicator name, examination result value, consultation time, customer ID, account balance, transaction amount, transaction time, risk rating, and product code.
[0022] In one embodiment, the rule-based classification model uses regular expressions, keyword matching, and other methods to perform rule matching on text data, structured data, and business format data to determine the first classification result of the metadata.
[0023] In one embodiment, a classification label is added to each metadata item based on the first classification result. The first classification result includes a primary category, a secondary subcategory, and a confidence score; therefore, the classification label is the label corresponding to the primary category, secondary subcategory, and confidence score. The classification label can be stored in key-value pairs. For example, the metadata "account balance 500,000 yuan" in a bank's core business system has a primary category of "account information," a secondary subcategory of "funds data," and a confidence score of 0.95. The classification label could then be {"primary category": "account information", "secondary subcategory": "funds data", "confidence score": 0.95}.
[0024] S102. The metadata is encrypted and transmitted to the data management platform, and the metadata to be classified is extracted from each of the metadata based on the classification tags; In one embodiment, metadata with category tags is encrypted and transmitted to the data management platform to ensure the security of data transmission. Encryption algorithms such as AES or RSA can be used.
[0025] After receiving the data, the data management platform filters the metadata based on the classification tags of each metadata element and extracts the metadata that needs to be classified in depth.
[0026] Further, the step of extracting the metadata to be classified from each of the metadata based on the classification labels includes: determining the primary category, secondary subcategory, and confidence score corresponding to each of the metadata based on the classification labels; taking the metadata corresponding to the classification labels whose confidence scores are less than a preset scoring threshold, or whose primary category is empty, or whose secondary subcategory is empty as the metadata to be classified, and extracting the metadata to be classified from each of the metadata.
[0027] In one embodiment, each type of metadata corresponds to a category label, which includes the category and confidence score of that type of metadata. The category includes a primary category and a secondary subcategory.
[0028] If the primary category or secondary subcategory is empty, it indicates that the metadata has not been classified or is incompletely classified, and needs to be reclassified. The confidence score reflects the accuracy of the initial classification. When the confidence score is low, it indicates that the current metadata has a low probability of belonging to the current category, and a more accurate deep classification is needed. All metadata that needs to be reclassified is designated as unclassified metadata. The scoring threshold can be freely set by the user according to actual needs. For example, assuming the scoring threshold is 0.9, the metadata "New Wealth Management Product Holding Records" of the securities asset management platform has a confidence score of 0.87 after initial classification, which is lower than the threshold, and is extracted as unclassified metadata.
[0029] In the above embodiments, metadata that needs to be deeply classified is selected based on the classification labels, instead of performing deep classification on all metadata. This reduces the amount of data and improves the efficiency of data classification while ensuring the accuracy of data classification.
[0030] S103. Based on a preset deep classification model, analyze the metadata to be classified to obtain a second classification result; In one embodiment, the metadata that is unclassified and has a low confidence score in the initial classification results is analyzed using a deep classification model to obtain the classification results of the metadata to be classified.
[0031] In one embodiment, the deep classification model includes a structured data processing layer and a text data processing layer, which process the structured data and text data in parallel. The structured data processing layer processes the structured data of the metadata to be classified to obtain structured features, while the text data processing layer processes the text data of the metadata to be classified to obtain text features. For example, the structured data of the metadata to be classified, "Holding Records of New Financial Products," includes the holding quantity and holding amount, while the text data consists of product descriptions and remarks. The structured data processing layer and the text data processing layer of the deep classification model process these data in parallel to quickly extract structured features (such as "Holding Amount: 100,000 yuan") and text features (such as "Non-principal-protected floating return, risk level B").
[0032] In one embodiment, the deep classification model further includes a master classifier. Structured features and textual features are concatenated to obtain a fused feature, which is then transmitted to the master classifier for data classification to obtain the data category of the metadata to be classified, i.e., the second classification result. For example, the structured feature "Holding Amount: 100,000 RMB" is concatenated with the textual feature "Non-principal-protected floating return, risk level B" to form a fused feature. The master classifier then performs classification based on financial risk control rules, outputting an accurate second classification result.
[0033] S104. Integrate the first classification result and the second classification result to obtain the target classification result.
[0034] In one embodiment, the first classification result and the second classification result are combined to obtain the classification result of all metadata, i.e., the target classification result.
[0035] In another embodiment, users can correct the target classification results through a manual verification interface and add the corrected data to the training dataset to retrain the deep classification model periodically.
[0036] The above embodiments provide a data classification method, apparatus, computer equipment, and storage medium. Based on edge processing nodes corresponding to each data source platform, the metadata of each data source platform is initially classified to obtain a first classification result. Based on the first classification result, classification tags are added to each type of metadata. The metadata is encrypted and transmitted to a data management platform, and metadata to be classified is extracted from each metadata based on the classification tags. Based on a preset deep classification model, the metadata to be classified is analyzed to obtain a second classification result. The first and second classification results are integrated to obtain a target classification result. This application first uses edge processing nodes to initially classify the metadata of each data source platform to obtain a first classification result, thus improving data processing efficiency through distributed data classification. Then, based on the first classification result, the metadata to be classified that needs deep classification is extracted, and deep classification is performed using a deep classification model, improving classification accuracy and reducing the amount of data that the deep classification model needs to process, thereby improving classification efficiency.
[0037] Furthermore, after step S104, the method further includes: storing various types of metadata in a preset database according to the target classification results based on the data management platform; analyzing the user query request when a user query request is received to obtain request data information and data purpose; analyzing the data purpose and the request data information to obtain recommended data information; extracting first data corresponding to the request data information and second data corresponding to the recommended data information from the preset database based on the request data information and the recommended data information, and pushing the first data and the second data to the user.
[0038] In one embodiment, a data management platform is a system that centrally manages metadata, responsible for storing, classifying, retrieving, and providing metadata, and may include components such as a database management system, a data warehouse, and a data lake.
[0039] In one embodiment, after processing by a deep classification model, each metadata element has a clear classification result, including a primary category, secondary subcategories, and confidence score. Based on the target classification result, each type of metadata element is stored in a pre-defined database.
[0040] The default database can be a relational database, a NoSQL (Not Only SQL, a general term for a type of non-relational database) database, or a data warehouse.
[0041] Each piece of metadata can be stored as a record, containing information such as field name, data type, description, and classification results (primary category, secondary subcategory, confidence score).
[0042] In one embodiment, a user submits a query request through the data management platform's interface. The query request typically includes the data information the user needs and its intended use. A query request can be a simple field name or a complex query statement containing multiple conditions and filters.
[0043] In one embodiment, the data information (such as field names, data types, etc.) and the purpose of the data (such as market analysis, user profiling, etc.) required by the user are extracted from the query request.
[0044] The system analyzes the user's data usage to determine if they might need other relevant data. For example, if a user requires market analysis data, the system can recommend other fields related to market analysis, generating recommended data information including field names, data types, and descriptions. The recommendation logic can be based on preset rules, historical data usage, or machine learning models. For instance, if a user requests a patient's age, the system can recommend fields such as patient gender and patient health insurance information.
[0045] For example, a data association graph can be pre-built. When a user queries the risk of credit bond default, the graph can be automatically traversed along its edges to recommend related data such as the issuer's financial indicators, guarantee chain information, and industry prosperity index.
[0046] In one embodiment, based on the requested data information, corresponding first data is extracted from a preset database. For example, data on the patient's age field is extracted.
[0047] Based on the recommended data, extract corresponding secondary data from a preset database. For example, extract data related to the patient's gender field.
[0048] In one embodiment, request data and recommendation data are combined into a single response object and pushed to the user.
[0049] For example, when an investment manager requests the historical volatility of a stock (requested data information) to build a quantitative investment strategy (data usage), the system analyzes the data usage and recommends secondary data such as market sentiment indicators and macroeconomic factors; when risk control and compliance personnel request customer credit scores for anti-money laundering suspicious transaction screening, the system recommends relevant data such as counterparty relationship graphs, fund flow paths, and cross-border transaction records.
[0050] In the above embodiments, the data management platform can effectively process user query requests, provide requested data and related recommendation data, centrally manage metadata from various data source platforms, and improve data management efficiency and data value.
[0051] Please see Figure 2 , Figure 2This is a schematic flowchart illustrating a data classification method provided in an embodiment of this application. This data classification method can be applied in a server to efficiently perform preliminary classification of metadata based on edge processing nodes, generating a first classification result. A confidence score can be used to evaluate the reliability of the classification result; a high confidence score indicates a more reliable classification result, while a low confidence score indicates that the classification result may require further verification. Accurate confidence scores help to precisely identify data that needs reclassification, improving the accuracy and reliability of the classification.
[0052] like Figure 2 As shown, the data classification method specifically includes steps S201 to S204.
[0053] S201. Based on the edge processing node, identify the metadata to obtain text data, structured data, and business format data; S202. Based on a preset rule classification model, perform rule matching on the text data, the structured data, and the business format data to obtain the primary category, secondary subcategory, and confidence score of the classification result for each metadata. S203. Based on the primary category, the secondary subcategory, and the confidence score, generate the first classification result.
[0054] In one embodiment, edge processing nodes can be deployed on a data source platform, with a built-in lightweight data parser. The parser identifies metadata and categorizes it into structured data, text data, and business format data. For example, in a medical imaging cloud platform, imaging devices generate millions of image metadata records daily, including: device technical parameters (structured data), image description text (semi-structured reports), and DICOM (Digital Imaging and Communications in Medicine) tags (business format data), etc.
[0055] Specifically, extract textual information from metadata, such as field descriptions and report titles. This textual data typically contains rich semantic information, which helps in understanding the business meaning of the metadata. Extract structured information, such as field names, data types, and field value distribution (e.g., number of unique values, null value rate). Identify business format information in the metadata, such as date format and currency format.
[0056] In one embodiment, the rule-based classification model classifies data based on a pre-defined rule base. The rule base contains a set of rules that define how metadata is categorized into different classes based on its characteristics. For example: fields whose names end with "_id" are dimensions; fields whose names begin with "count" or "sum" are metrics; fields with a low number of unique values in their value distribution are dimensions; fields with a high number of unique values in their value distribution are metrics; and fields whose descriptions contain specific keywords (such as "date" or "time") are time fields.
[0057] Specifically, the rule-based classification model uses regular expressions, keyword matching, and other techniques to perform rule matching on text data, structured data, and business format data to determine the category of metadata.
[0058] In one embodiment, the primary category is the highest-level classification of metadata, such as dimensions, metrics, time fields, etc.
[0059] Second-level subcategories are finer-grained classifications under the first-level categories, such as dimensions being divided into user dimensions, product dimensions, etc.
[0060] For example, edge processing nodes are deployed in the core trading system of a securities firm to perform initial classification of metadata such as transaction flow, position data, and customer assets. For instance, metadata with field names starting with "trade" is automatically marked as the first-level category "trading data" by the rule engine, and the second-level subcategories are further subdivided into "stock trading", "bond trading", or "derivatives trading" based on the field suffix; metadata with field names containing "name_" and data type string is marked as the first-level category "customer information".
[0061] Edge processing nodes are deployed across various hospital business systems (such as outpatient systems, inpatient systems, and imaging systems). Taking the electronic medical record platform as an example, rule matching is performed on the metadata of medical record fields: text data with field names containing "diagnosis," "symptoms," or "chief complaint" are initially classified into the first-level category of "clinical diagnostic information"; metadata with field names starting with "test" and data type being numeric is marked as the first-level category of "test indicators," and the second-level subcategories are further subdivided according to the test item name (e.g., "blood glucose test" corresponds to "blood glucose indicators").
[0062] In one embodiment, the rule matching result is accompanied by a confidence score, indicating the reliability of the classification result. The confidence score is typically calculated based on the degree of rule matching and the rule priority. For example, if a field name matches multiple rules simultaneously, the confidence score will be calculated by combining the rule priority and the degree of matching.
[0063] In specific embodiments, the matching degree of a rule refers to the degree of matching between the metadata and the preset rule. For example, if the metadata fully conforms to the rule, the matching degree is the highest. For instance, if a field name ends with "_id" and the rule definition uses fields ending with "_id" as dimensions, the matching degree is 100%. If the metadata partially conforms to the rule, the matching degree will decrease. For instance, if a field name contains "id" but does not end with "_id," and the rule definition uses fields ending with "_id" as dimensions, the matching degree might be 50%. For some fuzzy rules, such as field descriptions containing specific keywords, the matching degree can be measured based on the frequency and position of the keywords. For example, if the keyword "time" appears in the field description within a certain frequency range, the matching degree is the numerical value corresponding to that frequency range.
[0064] Rule priority refers to which rules are more important when multiple rules match simultaneously. Priority can be explicitly specified in the rule base. For example, rule 1 might have a priority of 1, rule 2 a priority of 2, and so on; the smaller the priority number, the higher the priority. Alternatively, a default priority can be used.
[0065] The confidence score is calculated based on the rule's matching degree and priority. Specifically, for a single rule, the confidence score can be simply equal to the matching degree. For example, if the field name ends with "_id" and the matching degree is 100%, the confidence score is 100%. If the field name contains "id" but does not end with "_id" and the matching degree is 50%, the confidence score is 50%.
[0066] When a field name matches multiple rules simultaneously, the matching degree and priority of each rule need to be considered comprehensively. Specifically, all matching rules, their matching degree, and priority are listed, and the weight of each rule is calculated based on its priority. The weight can be the reciprocal of the priority. The matching degree of each rule is multiplied by its weight to obtain the weighted matching degree. The weights of all rules are summed to obtain the total weight. The weighted matching degrees of all rules are summed and then divided by the total weight to obtain the overall confidence score.
[0067] In one embodiment, the primary category, secondary subcategories, and confidence scores are integrated to generate a first classification result. This first classification result includes not only the classification category of the metadata but also the confidence score for further processing and verification.
[0068] Understandably, during the initial classification process, if there is data in the metadata that does not match the preset rules at all, then the primary category for that data will be empty; if no secondary subcategory is matched, then the secondary subcategory will be empty. If the metadata only matches the primary category, then the confidence score will be the confidence score corresponding to the primary category; if the metadata matches the secondary subcategory, then the confidence score will be the confidence score corresponding to the secondary subcategory.
[0069] In the above embodiments, edge processing nodes can efficiently perform preliminary classification of metadata and generate a first classification result. The confidence score can be used to evaluate the reliability of the classification result. A high confidence score indicates that the classification result is more reliable, while a low confidence score indicates that the classification result may need further verification. An accurate confidence score is helpful in accurately identifying data that needs to be reclassified, thereby improving the accuracy and reliability of the classification.
[0070] Please see Figure 3 , Figure 3 This is a schematic flowchart illustrating a data classification method provided in an embodiment of this application. This data classification method can be applied in a server to process structured data and text data in parallel using a deep classification model, generating structured features and text features respectively, which are then concatenated to generate fused features. This results in high-precision classification prediction, obtaining a second classification result, thereby improving classification efficiency and accuracy.
[0071] like Figure 3 As shown, step S103 of the data classification method specifically includes steps S301 to S304.
[0072] S301. Based on the structured data processing layer and text data processing layer of the deep classification model, the structured data and text data of the metadata to be deeply classified are processed in parallel to obtain structured features and text features. S302. The structured features and the text features are concatenated to obtain fused features; S303. The classification module based on the deep classification model processes the fused features to obtain the second classification result.
[0073] In one embodiment, the structured data processing layer includes a first feature transformation module, used to perform nonlinear transformations on the input structured data to obtain a structured feature vector (containing key information such as data type weights and field name semantic mappings). Specifically, numerical features are normalized or standardized to achieve similar scales; categorical features (such as data types) are encoded and converted into numerical data. The processed data is then subjected to feature transformation to obtain a structured feature vector.
[0074] Specifically, normalization / standardization processing involves normalizing or standardizing numerical features to give them similar scales and prevent certain features from dominating model training.
[0075] Encoding Processing: Categorical features are encoded, such as using one-hot encoding or label encoding, to convert them into numerical data for model processing.
[0076] In one embodiment, the text data processing layer is used to extract text semantic features to obtain a text feature vector (containing information such as keyword weights and contextual semantic relationships). Specifically, the text data is segmented into words, breaking it down into the smallest semantic units (such as words or phrases). The segmented text data is then vectorized into a fixed-length sequence, and feature transformation is performed on the processed data to obtain the text feature vector.
[0077] In one embodiment, the structured data processing layer and the text data processing layer can be computed independently and in parallel, improving processing efficiency and meeting real-time requirements.
[0078] In one embodiment, the structured feature vector and the text feature vector are concatenated into a complete feature vector, i.e., a fused feature vector.
[0079] The classification module based on the deep classification model performs classification prediction on the fused feature vector to obtain a second classification result, including the classification category and confidence score.
[0080] For example, the structured data processing layer transforms the features of medical test indicators, such as normalizing numerical test results like white blood cell count and hematocrit, and labeling categorical features like laboratory departments (e.g., laboratory, pathology, microbiology). The text data processing layer can use a BERT-based pre-trained model for the medical field to process medical record text, extracting semantic features such as chief complaint, present illness, and past medical history, and capturing short keywords and long semantics such as "fever for 3 days" and "cough with sputum" through multi-scale convolution. After fusing the features and inputting them into the main classifier, accurate classification is performed to obtain the secondary classification result.
[0081] In the above embodiments, the deep classification model can process structured data and text data in parallel, generate structured features and text features respectively, and then concatenate them to generate fused features, perform high-precision classification prediction, obtain a second classification result, and improve classification efficiency and accuracy.
[0082] Further, before step S103, the method includes: acquiring a training dataset, and processing the structured data and text data in the training dataset in parallel based on the structured data processing layer and text data processing layer in the pre-trained model to obtain structured features, structured classification loss, structured classification accuracy, text features and text classification loss, and text classification accuracy; fusing the structured features and text features to obtain fused features, and processing the fused features based on the main classifier of the pre-trained model to obtain the predicted classification result and main classification loss of the training dataset; determining the weight coefficients of the structured classification loss, the text classification loss, and the main classification loss based on a preset main classification loss threshold, the main classification loss, the structured classification accuracy, and the text classification accuracy; performing weighted fusion of the structured classification loss, the text classification loss, and the main classification loss based on the weight coefficients to obtain a joint loss; and iteratively training the model based on the joint loss until a preset iteration stopping condition is reached to obtain the deep classification model.
[0083] In one embodiment, the training dataset can be obtained from multiple data sources, including but not limited to BI tools, databases, and ETL tools.
[0084] The training dataset includes structured data (such as field names, data types, field value distributions, etc.) and text data (such as field descriptions, report titles, etc.). The training dataset also includes the data categories corresponding to each data point; these categories can be user-input categories or categories output by the data classification model and manually corrected.
[0085] In one embodiment, the pre-trained model includes a structured data processing layer and a text data processing layer.
[0086] The structured data processing layer processes the structured data, including normalization, standardization, and encoding, to obtain structured features. These structured features are then analyzed to classify the training dataset, yielding the structured classification loss and structured classification accuracy.
[0087] The text data processing layer processes the text data, including word segmentation, breaking it down into its smallest semantic units (such as words or phrases), vectorization, and obtaining text features. These text features are then analyzed to classify the training dataset, yielding text classification loss and accuracy.
[0088] In one embodiment, structured features and textual features are concatenated to obtain fused features. The fused features are then processed using the main classifier of a pre-trained model to obtain the predicted classification results and main classification loss for the training dataset.
[0089] In one embodiment, the weight coefficients of the structured classification loss, text classification loss, and main classification loss are determined based on the main classification loss, structured classification accuracy, and text classification accuracy. For example, if the main classification loss is below a threshold, and the structured classification accuracy and text classification accuracy are high, the weight coefficient of the main classification loss can be appropriately reduced, while the weight coefficients of the structured classification and text classification losses can be increased; conversely, the weight coefficient of the main classification loss can be increased, while the weight coefficients of the structured classification and text classification losses can be decreased. The specific weight coefficients can be calculated using the following formula:
[0090]
[0091]
[0092] in, , , These are the weight coefficients for the structured classification loss, text classification loss, and main classification loss, respectively. , , These are the structured classification accuracy, text classification accuracy, and main classification loss, respectively.
[0093] In one embodiment, the structured classification loss, text classification loss, and main classification loss are weighted and fused based on weight coefficients to obtain the joint loss: ,in, , These are the structured classification loss and the text classification loss, respectively.
[0094] In one embodiment, joint loss is used as the optimization objective, and model parameters are updated through backpropagation for iterative model training. For example, the Adam (Adaptive Moment Estimation, an optimization algorithm used in deep learning to update weight parameters during neural network training to minimize the loss function) optimizer is used to adjust hyperparameters such as learning rate and momentum to optimize the model training process.
[0095] Training stops when the preset iteration stopping conditions are met (such as the joint loss no longer decreasing significantly, or the maximum number of iterations being reached), and a deep classification model is obtained.
[0096] In one embodiment, K-fold cross-validation is used to evaluate the model to reduce the risk of overfitting and improve the model's generalization ability. For example, the dataset is divided into K subsets, and each time K-1 subsets are used for training and 1 subset is used for validation, repeated K times, and the average value is taken as the final evaluation result.
[0097] In the above embodiments, the weight coefficients of the main classification loss, structured classification accuracy, and text classification accuracy are dynamically determined based on the main classification loss, structured classification accuracy, and text classification accuracy. The main classification loss, structured classification loss, and text classification loss are then weighted and fused according to the weight coefficients to obtain the joint classification loss, which improves the accuracy of the joint classification loss value. This allows for joint training of the structured data processing layer and the text data processing layer to obtain a deep classification model, thereby improving the accuracy of model training and enhancing model performance.
[0098] Furthermore, the structured data processing layer and text data processing layer in the pre-trained model process the structured data and text data in the training dataset in parallel to obtain structured features, structured classification loss, structured classification accuracy, text features, text classification loss, and text classification accuracy. This includes: performing feature transformations on the structured data and text data respectively based on the first feature transformation module in the structured data processing layer and the second feature transformation module in the text data processing layer to obtain the structured features and the text features; and processing the structured features and text features respectively based on the first auxiliary classifier in the structured data processing layer and the second auxiliary classifier in the text data processing layer to obtain the structured classification loss, the structured classification accuracy, the text classification loss, and the text classification accuracy.
[0099] In one embodiment, the first feature transformation module processes the preprocessed structured data to generate a structured feature vector.
[0100] In a specific embodiment, the structured data is normalized, standardized, and encoded, and the processed data is then transformed to obtain a structured feature vector.
[0101] In one embodiment, the second feature transformation module is used for text data feature extraction, focusing on local semantic features, filtering redundant information, capturing short keywords and long semantics through multi-scale convolutional kernels, retaining the strongest response value of each convolutional kernel, enhancing feature discriminability, and obtaining text feature vectors.
[0102] In a specific embodiment, the text data is segmented into its smallest semantic units (such as words or phrases). The processed data is then converted into a fixed-length sequence and subjected to feature transformation to obtain a text feature vector. For example, the phrase "The patient has no history of drug allergies, has a 5-year history of hypertension, and is currently taking medication regularly" is segmented to extract key semantics such as no drug allergies, history of hypertension, and regular medication, generating a text feature vector.
[0103] In one embodiment, the first auxiliary classifier is used for independent evaluation of structured features, independently evaluating the classification ability of structured features, and outputting the structured classification loss and structured classification accuracy to provide a basis for dynamic weight calculation.
[0104] Classification is performed using a first auxiliary classifier (such as XGBoost, Random Forest, etc.). Specifically, the received structured feature vectors are used for classification prediction. The cross-entropy loss function is used to calculate the structured classification loss, and the accuracy of structured classification, i.e., the proportion of correctly predicted samples out of the total number of samples, is calculated.
[0105] In one embodiment, the second auxiliary classifier is used for independent classification evaluation of text features, independently evaluating the classification ability of text features, and outputting text classification loss and text classification accuracy.
[0106] A second auxiliary classifier is used to analyze the text data to classify the training dataset. Specifically, classification prediction is performed on the received text feature vectors, and the text classification loss is calculated using the cross-entropy loss function. The accuracy of text classification is calculated, which is the proportion of correctly predicted samples out of the total number of samples.
[0107] Understandably, the structured data processing layer of the pre-trained model includes a first feature transformation module and a first auxiliary classifier, while the text data processing layer includes a second feature transformation module and a second auxiliary classifier. Once the deep classification model is obtained after iterative training, the first and second auxiliary classifiers stop processing the received data. The first and second auxiliary classifiers are only activated during model training and remain dormant during model usage.
[0108] Through the above steps, feature transformation and classification processing of structured data and text data can be effectively performed, obtaining structured classification loss, structured classification accuracy, text classification loss, and text classification accuracy. By using the structured classification loss, structured classification accuracy, text classification loss, and text classification accuracy, the performance of the model in processing structured data and text data can be comprehensively evaluated, providing an important basis for subsequent model optimization and joint training, and improving model training efficiency and training accuracy.
[0109] Please see Figure 4 , Figure 4 This is a schematic block diagram of a data classification apparatus provided in an embodiment of the present application. The data classification apparatus is used to perform the aforementioned data classification method. The data classification apparatus can be configured on a server.
[0110] like Figure 4 As shown, the data classification device 400 includes: The initial classification module 401 is used to perform initial classification of the metadata of each data source platform based on the edge processing node corresponding to each data source platform, obtain a first classification result, and add classification tags to each type of metadata based on the first classification result; The data extraction module 402 is used to encrypt and transmit the metadata to the data management platform, and extract the metadata to be classified from each of the metadata based on the classification tags; The deep classification module 403 is used to analyze the metadata to be classified based on a preset deep classification model to obtain a second classification result; The result integration module 404 is used to integrate the first classification result and the second classification result to obtain the target classification result.
[0111] Furthermore, the initial classification module 401 includes: The data recognition unit is used to recognize the metadata based on the edge processing node to obtain text data, structured data, and business format data; The primary classification unit is used to perform rule matching on the text data, the structured data, and the business format data based on a preset rule classification model, and to obtain the primary category, secondary subcategory, and confidence score of the classification result for each metadata. The first classification result generation unit is used to generate the first classification result based on the first-level category, the second-level sub-category, and the confidence score.
[0112] Furthermore, the data extraction module 402 includes: The information determination unit is used to determine the primary category, secondary subcategory, and confidence score corresponding to each data category of the metadata based on the classification label; The metadata determination unit is used to take the metadata corresponding to the classification label whose confidence score is less than a preset score threshold, or whose first-level category is empty, or whose second-level sub-category is empty as the metadata to be classified, and extract the metadata to be classified from each of the metadata.
[0113] Furthermore, the deep classification module 403 includes: The feature acquisition unit is used to process the structured data and text data of the metadata to be deeply classified in parallel based on the structured data processing layer and text data processing layer of the deep classification model, and to obtain structured features and text features. The feature concatenation unit is used to concatenate the structured features and the text features to obtain fused features; The second classification result acquisition unit is used by the classification module based on the deep classification model to process the fused features and obtain the second classification result.
[0114] Furthermore, the data classification device 400 also includes a model training module, which includes: The branch classification unit is used to acquire the training dataset and, based on the structured data processing layer and text data processing layer in the pre-trained model, process the structured data and text data in the training dataset in parallel to obtain structured features, structured classification loss, structured classification accuracy, text features and text classification loss, and text classification accuracy. The main classification unit is used to fuse the structured features and the text features to obtain fused features, and to process the fused features based on the main classifier of the pre-trained model to obtain the predicted classification result and main classification loss of the training dataset. The weight determination unit is used to determine the weight coefficients of the structured classification loss, the text classification loss, and the main classification loss based on a preset main classification loss threshold, the main classification loss, the structured classification accuracy, and the text classification accuracy. The joint loss acquisition unit is used to perform weighted fusion of the structured classification loss, the text classification loss, and the main classification loss based on the weight coefficients to obtain the joint loss; The deep classification model acquisition unit is used to perform iterative training of the model based on the joint loss until a preset iteration stopping condition is reached, thereby obtaining the deep classification model.
[0115] Further, the branch classification unit includes: The feature conversion unit is used to perform feature conversion on the structured data and the text data respectively based on the first feature conversion module in the structured data processing layer and the second feature conversion module in the text data processing layer to obtain the structured features and the text features; An auxiliary classification unit is used to process the structured features and the text features based on the first auxiliary classifier in the structured data processing layer and the second auxiliary classifier in the text data processing layer, respectively, to obtain the structured classification loss, the structured classification accuracy, the text classification loss, and the text classification accuracy.
[0116] Furthermore, the data classification device 400 also includes a data query module, which includes: A data storage unit is used to store various types of metadata into a preset database based on the data management platform and according to the target classification results; The request analysis unit is used to analyze the user query request when it is received, and to obtain the request data information and the purpose of the data. The data analysis unit is used to analyze the purpose of the data and the requested data information to obtain recommended data information; The data extraction unit is used to extract first data corresponding to the request data information and second data corresponding to the recommendation data information from the preset database based on the request data information and the recommendation data information, and push the first data and the second data to the user.
[0117] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the above-described apparatus and modules can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0118] The aforementioned device can be implemented as a computer program, which can be used in, for example... Figure 5 It runs on the computer device shown.
[0119] Please see Figure 5 , Figure 5 This is a schematic block diagram illustrating the structure of a computer device according to an embodiment of this application. The computer device may be a server.
[0120] See Figure 5 The computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.
[0121] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any data classification method.
[0122] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0123] Internal memory provides an environment for the execution of computer programs stored in non-volatile storage media. When executed by a processor, the computer program enables the processor to perform any data classification method.
[0124] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0125] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0126] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps: Based on the edge processing nodes corresponding to each data source platform, the metadata of each data source platform is initially classified to obtain a first classification result, and based on the first classification result, classification tags are added to the metadata of each category. The metadata is encrypted and transmitted to the data management platform, and the metadata to be classified is extracted from each of the metadata based on the classification tags; Based on a preset deep classification model, the metadata to be classified is analyzed to obtain a second classification result; The first classification result and the second classification result are integrated to obtain the target classification result.
[0127] In one embodiment, when the processor performs initial classification of the metadata of each data source platform based on the edge processing nodes corresponding to each data source platform to obtain a first classification result, it is configured to: Based on the edge processing node, the metadata is identified to obtain text data, structured data, and business format data; Based on a preset rule classification model, rule matching is performed on the text data, the structured data, and the business format data to obtain the primary category, secondary subcategory, and confidence score of the classification results for each metadata. The first classification result is generated based on the primary category, the secondary subcategory, and the confidence score.
[0128] In one embodiment, when the processor extracts the metadata to be classified from each of the metadata based on the classification labels, it is configured to: Based on the classification labels, determine the primary category, secondary subcategory, and confidence score corresponding to each data category for each metadata. The metadata corresponding to the classification labels whose confidence scores are less than a preset scoring threshold, or whose primary category or secondary subcategory is empty, is taken as the metadata to be classified, and the metadata to be classified is extracted from each of the metadata.
[0129] In one embodiment, when the processor analyzes the metadata to be classified based on a preset deep classification model to obtain a second classification result, it is used to: Based on the structured data processing layer and text data processing layer of the deep classification model, the structured data and text data of the metadata to be deeply classified are processed in parallel to obtain structured features and text features. The structured features and the text features are concatenated to obtain fused features; The classification module based on the deep classification model processes the fused features to obtain the second classification result.
[0130] In one embodiment, before the processor analyzes the metadata to be classified based on a preset deep classification model to obtain a second classification result, it is also configured to: Obtain the training dataset, and process the structured data and text data in the training dataset in parallel based on the structured data processing layer and text data processing layer in the pre-trained model to obtain structured features, structured classification loss, structured classification accuracy, text features and text classification loss, and text classification accuracy. The structured features and the text features are fused to obtain fused features. Based on the main classifier of the pre-trained model, the fused features are processed to obtain the predicted classification results and main classification loss of the training dataset. Based on the preset main classification loss threshold, the main classification loss, the structured classification accuracy, and the text classification accuracy, the weight coefficients of the structured classification loss, the text classification loss, and the main classification loss are determined. Based on the weight coefficients, the structured classification loss, the text classification loss, and the main classification loss are weighted and fused to obtain the joint loss; The model is iteratively trained based on the joint loss until a preset iteration stopping condition is met, thus obtaining the deep classification model.
[0131] In one embodiment, when the processor implements a structured data processing layer and a text data processing layer based on a pre-trained model to process structured data and text data in the training dataset in parallel to obtain structured features, structured classification loss, structured classification accuracy, text features, text classification loss, and text classification accuracy, it is used to achieve the following: Based on the first feature conversion module in the structured data processing layer and the second feature conversion module in the text data processing layer, feature conversion is performed on the structured data and the text data respectively to obtain the structured features and the text features; Based on the first auxiliary classifier in the structured data processing layer and the second auxiliary classifier in the text data processing layer, the structured features and the text features are processed respectively to obtain the structured classification loss, the structured classification accuracy, the text classification loss, and the text classification accuracy.
[0132] In one embodiment, after integrating the first classification result and the second classification result to obtain the target classification result, the processor is further configured to: Based on the data management platform, the metadata of each type is stored in a preset database according to the target classification results; Upon receiving a user query request, the user query request is analyzed to obtain the request data information and the purpose of the data; The purpose of the data and the requested data information are analyzed to obtain recommended data information; Based on the requested data information and the recommended data information, the first data corresponding to the requested data information and the second data corresponding to the recommended data information are extracted from the preset database, and the first data and the second data are pushed to the user.
[0133] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the data classification methods provided in the embodiments of this application.
[0134] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.
[0135] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0136] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A data classification method, characterized in that, include: Based on the edge processing nodes corresponding to each data source platform, the metadata of each data source platform is initially classified to obtain a first classification result, and based on the first classification result, classification tags are added to the metadata of each category. The metadata is encrypted and transmitted to the data management platform, and the metadata to be classified is extracted from each of the metadata based on the classification tags; Based on a preset deep classification model, the metadata to be classified is analyzed to obtain a second classification result; The first classification result and the second classification result are integrated to obtain the target classification result.
2. The data classification method according to claim 1, characterized in that, The step of performing preliminary classification of the metadata of each data source platform based on the edge processing nodes corresponding to each data source platform to obtain a first classification result includes: Based on the edge processing node, the metadata is identified to obtain text data, structured data, and business format data; Based on a preset rule classification model, rule matching is performed on the text data, the structured data, and the business format data to obtain the primary category, secondary subcategory, and confidence score of the classification results for each metadata. The first classification result is generated based on the primary category, the secondary subcategory, and the confidence score.
3. The data classification method according to claim 1, characterized in that, The step of extracting the metadata to be classified from each of the metadata based on the classification labels includes: Based on the classification labels, determine the primary category, secondary subcategory, and confidence score corresponding to each data category for each metadata. The metadata corresponding to the classification labels whose confidence scores are less than a preset scoring threshold, or whose primary category or secondary subcategory is empty, is taken as the metadata to be classified, and the metadata to be classified is extracted from each of the metadata.
4. The data classification method according to claim 1, characterized in that, The method of analyzing the metadata to be classified based on a preset deep classification model to obtain a second classification result includes: Based on the structured data processing layer and text data processing layer of the deep classification model, the structured data and text data of the metadata to be deeply classified are processed in parallel to obtain structured features and text features. The structured features and the text features are concatenated to obtain fused features; The classification module based on the deep classification model processes the fused features to obtain the second classification result.
5. The data classification method according to claim 1, characterized in that, Before analyzing the metadata to be classified based on a preset deep classification model to obtain the second classification result, the process further includes: Obtain the training dataset, and process the structured data and text data in the training dataset in parallel based on the structured data processing layer and text data processing layer in the pre-trained model to obtain structured features, structured classification loss, structured classification accuracy, text features and text classification loss, and text classification accuracy. The structured features and the text features are fused to obtain fused features. Based on the main classifier of the pre-trained model, the fused features are processed to obtain the predicted classification results and main classification loss of the training dataset. Based on the preset main classification loss threshold, the main classification loss, the structured classification accuracy, and the text classification accuracy, the weight coefficients of the structured classification loss, the text classification loss, and the main classification loss are determined. Based on the weight coefficients, the structured classification loss, the text classification loss, and the main classification loss are weighted and fused to obtain the joint loss; The model is iteratively trained based on the joint loss until a preset iteration stopping condition is met, thus obtaining the deep classification model.
6. The data classification method according to claim 5, characterized in that, The structured data processing layer and text data processing layer in the pre-trained model process the structured data and text data in the training dataset in parallel to obtain structured features, structured classification loss, structured classification accuracy, text features, text classification loss, and text classification accuracy, including: Based on the first feature conversion module in the structured data processing layer and the second feature conversion module in the text data processing layer, feature conversion is performed on the structured data and the text data respectively to obtain the structured features and the text features; Based on the first auxiliary classifier in the structured data processing layer and the second auxiliary classifier in the text data processing layer, the structured features and the text features are processed respectively to obtain the structured classification loss, the structured classification accuracy, the text classification loss, and the text classification accuracy.
7. The data classification method according to any one of claims 1 to 6, characterized in that, After integrating the first classification result and the second classification result to obtain the target classification result, the method further includes: Based on the data management platform, the metadata of each type is stored in a preset database according to the target classification results; Upon receiving a user query request, the user query request is analyzed to obtain the request data information and the purpose of the data; The purpose of the data and the requested data information are analyzed to obtain recommended data information; Based on the requested data information and the recommended data information, the first data corresponding to the requested data information and the second data corresponding to the recommended data information are extracted from the preset database, and the first data and the second data are pushed to the user.
8. A data classification device, characterized in that, include: The initial classification module is used to perform initial classification of the metadata of each data source platform based on the edge processing node corresponding to each data source platform, obtain a first classification result, and add classification tags to each type of metadata based on the first classification result; The data extraction module is used to encrypt and transmit the metadata to the data management platform, and extract the metadata to be classified from each of the metadata based on the classification tags; The deep classification module is used to analyze the metadata to be classified based on a preset deep classification model to obtain a second classification result; The result integration module is used to integrate the first classification result and the second classification result to obtain the target classification result.
9. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and, in executing the computer program, implement the data classification method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to implement the data classification method as described in any one of claims 1 to 7.