Big data management system and method based on hierarchical label system

The big data governance system with a layered labeling system solves the problem that the labeling system in existing technologies cannot adapt to business changes and cross-industry applications, realizes the accurate classification and management of data, improves data processing efficiency and security, and ensures the stability and adaptability of the labeling system.

CN120596475AActive Publication Date: 2025-09-05BEIJING GUOXINDA DATA TECH CO LTD

Patent Information

Application Number
CN202511086174.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-09-05
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

The existing labeling system cannot flexibly respond to changes in business scenarios and cross-industry application needs. Label parsing and data binding lack comprehensive consideration, resulting in the inability to accurately classify and match data. Traditional data governance methods are difficult to meet the company's strict requirements for data quality and security. The lack of real-time monitoring and dynamic repair mechanisms can easily lead to data management chaos.

Method used

A big data governance system based on a hierarchical label system is adopted, including data access preprocessing, label modeling and hierarchical configuration, label parsing and data binding, label-driven data governance, and label conflict detection and dynamic update units. Through multi-mode matching algorithms, hash comparison and path traversal algorithms, accurate data parsing, binding and real-time monitoring are achieved, and the label system is dynamically adjusted.

Benefits of technology

It achieves accurate classification and management of data, improves data processing efficiency and accuracy, ensures data quality and security, reduces data governance costs, and ensures the stability and adaptability of the labeling system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596475A_ABST
    Figure CN120596475A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of big data governance, and discloses a big data governance system and method based on a hierarchical label system. The system comprises a data access preprocessing unit, a label modeling and hierarchical configuration unit, a label analysis and data binding unit, a label driving data governance unit and a label conflict detection and dynamic updating unit, the label driving data governance unit takes a label binding result as a trigger, a governance rule is loaded through a dynamic rule engine, and the label conflict detection and dynamic updating unit carries out dynamic updating. A multi-dimensional verification strategy and an encryption management and control mechanism are used for cleaning, classifying and performing security protection on the data; the label conflict detection and dynamic updating unit monitors the label system in real time, and adjusts and repairs by means of path traversal and version management technologies after finding conflicts, so as to ensure the stability of the label system; the two functions together to automatically process abnormal data, maintain a label system and guarantee data quality and safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of big data management technology, and specifically relates to a big data management system and method based on a hierarchical label system. Background Art

[0002] Against the backdrop of accelerated digital transformation, the data governance challenges faced by enterprises and organizations are becoming increasingly severe.

[0003] In terms of label system construction, existing label systems mostly use static predefined models, which cannot flexibly respond to changes in business scenarios and cross-industry application needs. When business processes are adjusted or new business forms emerge, fixed label levels and attributes are difficult to adapt quickly, resulting in the inability to accurately classify data, which limits data value mining.

[0004] The label parsing and data binding process lacks comprehensive consideration of multi-label adaptation scenarios. Existing technologies often select binding labels based on a single factor, ignoring key factors such as business sensitivity and permission compliance, resulting in insufficient accuracy in label and data matching, affecting the accuracy and reliability of data analysis.

[0005] In the data governance process, traditional methods for data cleansing and security management are relatively crude, making it difficult to implement refined governance of data tag attributes and failing to meet enterprises' stringent requirements for data quality and security. Furthermore, the tagging system lacks real-time monitoring and dynamic repair mechanisms during operation, making it difficult to promptly detect and address anomalies such as tag conflicts and redundancies, which can easily lead to data management chaos and reduce data governance efficiency. Summary of the Invention

[0006] The purpose of the present invention is to provide a big data governance system and method based on a hierarchical labeling system to solve the problems raised in the above background technology.

[0007] To achieve the above objectives, the present invention provides the following technical solutions: a big data governance system based on a hierarchical tag system, the system comprising: a data access preprocessing unit, a tag modeling and hierarchical configuration unit, a tag parsing and data binding unit, a tag-driven data governance unit, and a tag conflict detection and dynamic update unit; Data access preprocessing unit: This unit is responsible for receiving heterogeneous data from multiple sources in various formats. It uses data format standardization mapping technology to build unified conversion rules and convert the raw data into a standard format to facilitate the subsequent label modeling and hierarchical configuration unit work. Label modeling and hierarchical configuration unit: After receiving pre-processed data, it uses label hierarchical modeling formulas to build a label system based on data types and business scenarios. By setting multi-level labels, label relationships and attributes are clarified, supporting recursive expansion and cross-industry adaptation. Label Parsing and Data Binding Unit: This system, built on the Label Modeling and Hierarchical Configuration Units, uses a multi-mode matching algorithm to parse data and automatically bind data fields and labels based on mapping paths. It uses a priority formula to optimize binding selection and then outputs the binding results to the Label-Driven Data Governance Unit. Tag-driven data governance unit: Based on the results of the tag parsing and data binding units, it dynamically loads governance rules and automatically cleans data with anomalies, redundancy, and other issues. This tag-based governance model ensures data quality, and the processed data is monitored by the tag conflict detection and dynamic update unit. Label conflict detection and dynamic update unit: It monitors the data labels processed by the label-driven data governance unit in real time, uses the detection formula based on hash comparison and path traversal algorithm to find conflicts and make timely adjustments to ensure the stability of the label system.

[0008] Tag binding priority formula

[0009]

[0010] Where: P is the label binding priority score are the weights of business tags, data type tags, and permission tags respectively; x, y, and z are the number of tags hit by the current data respectively; Label conflict detection formula: D=

[0011] Where D is the tag conflict detection result value; Hi is the current mapping hash value of the tag; Hh is the historical mapping hash value of the tag; n is the number of tags; Biaoci drives data governance formula: TiGi Where: R is the result set after cleaning; S is the input data set; T is the first matching anomaly label; Gi is the anomaly cleaning rule; n is the number of anomaly labels.

[0012] Preferably, the data access pre-processing unit includes: (1) Receiving multi-source heterogeneous data: The data access preprocessing unit undertakes the key task of receiving multi-source heterogeneous data, including structured data, semi-structured data, and unstructured data. These data come from a wide range of sources and in a variety of formats, including common data formats such as JSON, CSV, and XML, as well as non-traditional data forms such as images and text. The diversity and complexity of the data highlight the importance of this unit's work; (2) Data format standardization: To ensure the smooth progress of the subsequent tag parsing process, this unit uses data format standardization mapping technology to deeply process the original data. By building unified data conversion rules, data in different formats are accurately converted into a standard data format that supports tag parsing. This process eliminates the parsing barriers caused by data format differences, ensures that the data entering the system is structurally consistent, and effectively improves data processing efficiency and accuracy.

[0013] Data format standardization mapping expression:

[0014] Where: Normalized dataset; The original data set accessed; Preset data standardization parameters (such as encoding format, data type mapping table); Data preprocessing mapping function.

[0015] In existing big data governance, data sources vary greatly. Without preprocessing, subsequent label parsing and data governance rules will become ineffective. This formula effectively improves data adaptability through data standardization mapping functions, providing structural compatibility guarantees for subsequent label parsing.

[0016] Preferably, the label modeling and hierarchical configuration unit includes: (1) Construction of a multi-level label system: After receiving the standardized data processed by the data access pre-processing unit, the core work is carried out based on the data type and business scenario with the help of the label hierarchical modeling formula. By setting different levels of labels such as industry, business, and authority, a multi-level label system is constructed to clarify the parent-child relationship, level, and specific attributes of each label; Label level modeling expression:

[0017] Where: No. Tier tags; Label parent node name; Tag attribute set (such as industry attributes, authority attributes); Traditional tag governance mostly uses flat or static tags, which cannot adapt to complex business structures. This formula introduces tag hierarchy and attribute construction paths to solve the problems of coarse tag granularity and insufficient adaptability in existing technologies, and realizes the construction of a dynamic tag system.

[0018] (2) Flexible adaptation and management support: The label modeling and hierarchical configuration unit has strong flexibility, supports recursive expansion of labels and cross-industry adaptive configuration, and can flexibly adjust the label system to meet the data labeling needs in different industries and business scenarios. This feature not only meets diverse needs, but also provides strong support for refined data management, ensuring that data can be accurately and effectively classified and applied in different scenarios.

[0019] Preferably, the tag parsing and data binding unit includes: (1) Data parsing and binding based on the tag system: The tag parsing and data binding unit works closely with the tag modeling and hierarchical configuration unit to perform accurate data parsing based on the established tag system.

[0020] This unit uses a multi-mode matching algorithm to quickly and efficiently extract data features, and then automatically binds data fields to hierarchical labels according to the label mapping path. In this way, the data is accurately labeled. (2) Priority optimization selection for multi-label adaptation: When there are multiple adaptation label candidates, the label parsing and data binding unit uses the label binding priority formula to comprehensively consider key factors such as the label's sensitivity to the business, the degree of data adaptation, and the rationality of permissions, and scientifically determine the label binding priority; through this priority evaluation mechanism, it ensures that the most appropriate label is selected for binding first, effectively improving the accuracy of label parsing, ensuring accurate and correct matching between data and labels, and enhancing the reliability of data management.

[0021] Label binding priority expression:

[0022] Where: P is the label binding priority score; is the business label weight; is the data type label weight; is the authority label weight; The number of business tags matched by the current data; The number of data type labels hit for the current data; The number of permission tags hit for the current data.

[0023] Preferably, the tag-driven data governance unit includes: (1) Governance rule-driven based on tag binding: Based on the results of tag binding, targeted tag-driven data governance rules are dynamically loaded. The mechanism of flexibly calling rules based on tag binding conditions makes data governance more accurate and adaptable, ensuring that subsequent data governance operations can be carried out in an orderly and scientific manner. (2) Efficient data cleaning and quality and safety assurance: In the data governance process, the data governance rules are dynamically loaded by driving the tags, and the data is automatically screened in depth, abnormal data, redundant data and data with incorrect formats are accurately identified and eliminated, thus achieving efficient data cleaning. This unit deeply integrates data quality control and security management into the tag system, and through the tag-driven governance model, it significantly improves the efficiency and effectiveness of data governance, and comprehensively guarantees the quality and safety of data.

[0024] Tag-driven data governance principle expression:

[0025] Where: Result set after data cleaning Input standardized dataset; For the matched data anomaly labels (such as redundancy, format error); For the label Corresponding data anomaly cleaning rules; The number of abnormal labels.

[0026] Traditional data governance relies on static rules and lacks a dynamic label-driven path. This formula automatically cleans abnormal data based on label indexes, solving the problems of delayed static rule updates and poor data cleaning effects in existing technologies.

[0027] Preferably, the label conflict detection and dynamic update unit includes: (1) Real-time monitoring mechanism for label conflicts: The label conflict detection and dynamic update unit is closely connected with the label-driven data governance unit. As the "guardian" of the data governance process, it relies on label hash comparison and label path traversal algorithms, and uses the label conflict detection formula to conduct real-time monitoring of the current label mapping hash value and the historical label hash value. Through this continuous difference comparison, potential problems in the label system can be keenly captured, so that label conflicts can be handled in a timely manner; Label conflict detection expression:

[0028] Where: Label conflict detection result value; For label The current mapping hash value of ; For label The history mapping hash value of ; : Number of tags; Existing technologies often rely on manual label conflict review, which is inefficient and prone to omissions. This formula uses hash comparison to quickly detect label path changes, enabling real-time detection and automatic adjustment of label conflicts, avoiding label redundancy and label failure.

[0029] (2) Dynamic adjustment ensures system stability: Once a label conflict or abnormal mismatch is detected, the label conflict detection and dynamic update unit will quickly activate the response mechanism. By automatically adjusting the label mapping path and label version, the inconsistencies in the label system are accurately repaired, ensuring the consistency and stability of the entire label system. This dynamic update mechanism effectively avoids data management chaos caused by label conflicts.

[0030] The present invention also provides a big data management method based on a hierarchical tag system, which uses the above-mentioned big data management system based on a hierarchical tag system. The specific steps of the method are as follows: S1. Label modeling and hierarchical configuration: This method uses label modeling and hierarchical configuration units to analyze access data from multiple dimensions such as industry, business, and permissions, and establish a multi-level label system. During the construction process, label attributes are configured in detail, and the label hierarchical relationship and label mapping path are clarified. This provides a clear and orderly structural foundation for subsequent label parsing and data governance work, allowing data to be accurately classified and managed according to different dimensions and needs; S2. Label Parsing and Driven Governance: Based on the label parsing and data binding unit and the label-driven data governance unit, the label parsing and data binding unit first parses the data and automatically binds data fields and labels according to the hierarchical label system. Then, based on label-driven data governance rules, a series of governance operations such as data classification, cleaning, desensitization, and deduplication are dynamically performed at the data level. This label-driven governance model ensures a high degree of matching between data quality and labels, effectively improving data availability and security. S3. Label Conflict Detection and Dynamic Update: This unit monitors conflicts, redundancies, and failures that may arise during label mapping in real time. By utilizing a label path traversal algorithm and label version management mechanism, the label mapping path can be rapidly and dynamically updated upon detection of anomalies, maintaining the continuity and correctness of the labeling system.

[0031] Preferably, the specific steps of label modeling and hierarchical configuration in step S1 are as follows: S11. Multi-dimensional tag system construction: The tag modeling and hierarchical configuration method relies on the tag modeling and hierarchical configuration unit and uses the tag hierarchical modeling formula to deeply analyze the access data from the dimensions of industry, business, and authority. By setting up multi-level tags, a complete tag system framework is built to facilitate the refined management of data; S12. Clear label relationships and paths: When building a label system, configure label attributes in detail, clarify label hierarchical relationships through a hierarchical division algorithm, and determine label mapping paths with the help of path mapping technology. This establishes a clear and orderly structure for subsequent label parsing and data governance, enabling accurate data classification and management.

[0032] Preferably, the specific steps of tag parsing and driver management in step S2 are as follows: S21. Tag parsing and binding: Based on the tag parsing and data binding unit, a multi-mode matching algorithm facilitates in-depth mining of data features for accurate data parsing. Combined with the tag binding priority formula, it comprehensively considers business sensitivity, adaptability, and authority factors, and automatically achieves accurate binding between data fields and tags based on a hierarchical tag system and tag mapping path. S22. Tag-driven dynamic data governance: Dynamically load tag-driven data governance rules, use automated data processing technology to perform data classification, cleaning, desensitization, deduplication and other operations, and achieve accurate matching of data quality and tags through the tag-driven model, thereby improving data availability and security.

[0033] Preferably, the specific steps of tag conflict detection and dynamic update in step S3 are as follows: S31. Real-time monitoring of label anomalies: The label conflict detection and dynamic update method relies on the label conflict detection and dynamic update unit, uses label hash comparison and label path traversal algorithms, and combines label conflict detection formulas to scan the label mapping process in real time. By comparing current and historical label hash values, it accurately locates anomalies such as conflicts, redundancies, and failures. S32. Dynamic update: Once an anomaly is detected, the label path traversal algorithm and label version management mechanism are used to automatically adjust the label mapping path and update the label version to ensure the stability of the label system.

[0034] The beneficial effects of the present invention are as follows: 1. The present invention uses the acquisition architecture and intelligent conversion engine constructed by the data access preprocessing unit to adopt interface adaptation, syntax tree parsing, feature extraction and other technologies for multi-source heterogeneous data, and combines regular expressions and machine learning algorithms to complete format standardization, thereby reducing data processing obstacles; the label parsing and data binding unit uses multi-mode matching algorithms and priority formulas to extract features from structured, semi-structured and unstructured data using corresponding strategies, and determines the binding priority based on multiple factors; through the synergistic effect of the two, data format and label adaptation problems are avoided, thereby greatly improving data processing efficiency and accuracy.

[0035] 2. The present invention uses a label modeling and hierarchical configuration unit with the help of domain knowledge graphs and semantic analysis algorithms, combined with label hierarchical modeling formulas, to build a multi-level label system that supports recursive label expansion; the label conflict detection and dynamic update unit monitors the label system in real time, identifies anomalies by using hash comparison and path traversal algorithms, and dynamically repairs them using version management and path adjustment technology; with the cooperation of the two, the system can flexibly respond to industry and business changes, thereby reducing data governance costs.

[0036] 3. The present invention loads governance rules through a dynamic rule engine, and uses multi-dimensional verification strategies and encryption management mechanisms to achieve data cleaning, classification, and security protection; the label conflict detection and dynamic update unit monitors the label system in real time, and after discovering conflicts, it uses path traversal and version management technology to adjust and repair to ensure the stability of the label system; through the joint action of the two, abnormal data is automatically processed, the label system is maintained, and data quality and security are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a flow chart of the big data management system based on the hierarchical tag system of the present invention; Figure 2 This is a flow chart of the big data management method based on the hierarchical label system of the present invention. DETAILED DESCRIPTION

[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0039] like Figures 1 to 2 As shown, an embodiment of the present invention provides a big data governance system based on a hierarchical label system, which includes: a data access preprocessing unit, a label modeling and hierarchical configuration unit, a label parsing and data binding unit, a label-driven data governance unit, and a label conflict detection and dynamic update unit; Data access preprocessing unit: As the starting point of the system's data processing, it is responsible for receiving multi-source heterogeneous data in various formats. Through data format standardization mapping technology, it builds unified conversion rules to convert raw data into a standard format, facilitating the subsequent label modeling and hierarchical configuration unit work; The label modeling and hierarchical configuration unit receives pre-processed data and, based on data types and business scenarios, uses label hierarchical modeling formulas to build a labeling system. By setting up multiple levels of labels, clarifying label relationships and attributes, and supporting recursive expansion and cross-industry adaptation, the results provide a labeling framework for the label parsing and data binding unit. The label parsing and data binding unit, based on a system built with label modeling and hierarchical configuration units, uses a multi-mode matching algorithm to parse data and automatically binds data fields and labels based on mapping paths. Priority formulas are used to optimize binding selection, and accurate binding results are a prerequisite for the label-driven data governance unit to function. Tag-driven data governance unit: Based on the results of the tag parsing and data binding units, it dynamically loads governance rules and automatically cleans data with anomalies, redundancy, and other issues. This tag-based governance model ensures data quality, and the processed data is monitored by the tag conflict detection and dynamic update unit. Label conflict detection and dynamic update unit: It monitors the data labels processed by the label-driven data governance unit in real time, uses the detection formula based on hash comparison and path traversal algorithm to find conflicts and make timely adjustments to ensure the stability of the label system.

[0040] Application scenarios of multi-mode matching algorithms: Structured data uses regular expressions to extract key fields, semi-structured data uses JSONPath or XPath path language to parse node relationships, and unstructured data uses NLP models based on BERT, GPT, etc. to extract semantic vectors.

[0041] The label-driven data governance unit automatically loads the data governance rules corresponding to the label binding results through the label-rule mapping index library, and dynamically triggers data cleaning, classification, desensitization and security control operations.

[0042] Label conflict detection compares the current label with the historical label in real time through the label hash value, quickly traverses the label path based on the depth-first search (DFS) or breadth-first search (BFS) algorithm through the graph database, locates abnormal label paths in time, and activates the label version management mechanism to dynamically repair conflicts.

[0043] Among them, the data access preprocessing unit realizes the efficient reception and processing of multi-source heterogeneous data such as structured, semi-structured and unstructured data by building a highly compatible and intelligent data processing architecture. For structured data, standard interfaces such as JDBC and ODBC are used to achieve seamless connection with relational databases, supporting real-time incremental extraction and full loading; for semi-structured data such as JSON and XML, syntax tree-based parsing technology is used to accurately identify key-value pairs and hierarchical relationships in the data, and combined with the streaming computing framework, it can efficiently process large amounts of log data. When faced with unstructured data such as pictures and text, the OpenCV library is used to extract image features, and NLP technology is used to convert text into word vectors or semantic units.

[0044] During the data format standardization process, the unit has a built-in industry-standard data template library, which supports users to customize conversion rules through a visual interface and use regular expressions, XSLT style sheets and other tools to achieve format conversion; at the same time, it introduces machine learning algorithms to automatically identify data encoding, delimiters and other features through training on historical data, generate the optimal conversion plan, and adopt a double verification mechanism of hash check and pattern matching to ensure the integrity and consistency of data during the conversion process, and finally output unified format data that meets the label parsing requirements, which is convenient for subsequent label modeling, parsing and data governance.

[0045] Among them, after receiving the standardized data output by the data access preprocessing unit, the label modeling and hierarchical configuration unit uses domain knowledge graph construction technology and semantic analysis algorithm, combines data type characteristics and business scenario requirements, and constructs a multi-level label system through the label hierarchical modeling formula.

[0046] During the system construction process, a tree-like hierarchical structure model is adopted to subdivide business process labels and permission control labels, and the attribute description language is used to accurately define the value range, data type and association rules of each label. The topological sorting algorithm is used to clarify the parent-child hierarchical relationship and dependency path between labels.

[0047] At the same time, based on the metadata-driven architecture design, it supports recursive expansion of tags. When new data objects or business needs emerge, sub-tags can be added to the existing hierarchical structure through a dynamic node insertion algorithm. Through the adaptive parameter adjustment mechanism, according to the data characteristics, business logic and regulatory requirements of different industries, the hierarchical depth, attribute configuration and mapping rules of the tag system are automatically optimized to achieve cross-industry adaptive configuration, which is convenient for subsequent tag parsing, data binding and refined management.

[0048] The label parsing and data binding unit integrates multi-modal matching algorithms with intelligent mapping mechanisms to achieve in-depth data feature mining and precise label binding. In the data feature extraction phase, the unit adopts a hybrid multi-modal matching strategy that includes regular expression matching, semantic similarity calculation, and machine learning model prediction. For structured data, regular expressions are used to quickly locate key data fields. For semi-structured data, node relationships are parsed based on path languages ​​such as JSONPath and XPath, combined with semantic similarity algorithms to identify data semantics. For unstructured text data, semantic vector features are extracted using pre-trained natural language processing models such as BERT and GPT.

[0049] During the automated binding process, the system uses the label topology stored in the graph database, based on the hierarchical relationships and attribute constraints defined by the label mapping path, to efficiently match data fields and label nodes through a path search algorithm. Data lineage tracking technology is also introduced to record the binding process, ensuring that each data label association is traceable. When a multi-label adaptation scenario occurs, the label parsing and data binding unit uses a multi-objective decision-making algorithm to sort candidate labels based on the label binding priority formula, comprehensively evaluating the label business weight, data compatibility score, and permission compliance verification results. Labels that meet business requirements, have high data adaptability, and meet permission requirements are prioritized for binding.

[0050] Among them, the tag-driven data governance unit has a built-in dynamic rule engine. By establishing a tag-rule mapping index library, different types of tags are pre-associated with corresponding governance rules. When tag-bound data is received, it can achieve millisecond-level response and load targeted data governance rule sets based on the rule priority algorithm and event triggering mechanism.

[0051] During the data cleansing phase, a multi-dimensional data validation strategy is employed: For structured data, compliance checks are performed using pre-set field format regular expressions and value range constraints. For semi-structured data, structural verification is performed using JSON Schema or XML DTD. For unstructured data, anomaly detection models in natural language processing are used to identify issues such as sensitive information leaks and semantic inconsistencies. Furthermore, data lineage analysis techniques are incorporated to trace the data's source and processing history, and association analysis algorithms are used to locate the root causes of anomalies.

[0052] At the data security management level, a hierarchical encryption strategy is adopted based on the permission attributes carried by the tags, AES-256 encryption is implemented for highly sensitive data, and lightweight hash summary processing is used for ordinary data. A permission management mechanism that combines access control lists (ACLs) and mandatory access control (MAC) is used to ensure that the data governance process and the tag system are deeply integrated and promoted in a coordinated manner.

[0053] Among them, the label conflict detection and dynamic update unit adopts a dual-track parallel detection strategy in terms of the real-time monitoring mechanism of label conflicts: on the one hand, the label hash comparison algorithm is used to perform block hashing on the label mapping data, and the hash fingerprint of the current label mapping is compared bit by bit with the hash record of the historical version by calculating the SHA-256 or MD5 hash value. Even slight changes in the data structure or modifications to the attribute value can be accurately identified; on the other hand, with the help of the label path traversal algorithm, the label topology structure built based on the graph database adopts the depth-first search (DFS) or breadth-first search (BFS) algorithm to perform a systematic scan along the parent-child hierarchical relationship and associated paths of the label, which can not only detect directly conflicting label nodes, but also identify indirect logical contradictions caused by changes in hierarchical relationships.

[0054] When a conflict or abnormal mismatch is detected, the dynamic adjustment module is immediately started. Through the label version management mechanism, the historical version snapshot of the conflicting label is extracted for difference analysis. The adjustment priority is calculated based on the conflict detection formula. The automatic mapping rerouting algorithm is used to re-plan the label mapping path. At the same time, the version number, modification time and other information in the label metadata are updated to ensure that the entire label system maintains structural integrity and logical consistency during the dynamic adjustment process.

[0055] The embodiment of the present invention further provides a big data management method based on a hierarchical tag system, which uses the above-mentioned big data management system based on a hierarchical tag system. The specific steps of the method are as follows: S1. Label Modeling and Hierarchical Configuration: Using the label modeling and hierarchical configuration unit, we analyze access data from multiple dimensions, such as industry, business, and permissions, to establish a multi-level labeling system. During the construction process, we configure label attributes in detail, clarify the label hierarchical relationship and label mapping path. This provides a clear and orderly structural foundation for subsequent label parsing and data governance work. S2. Label Parsing and Driven Governance: First, the label parsing and data binding unit parses the data and automatically binds data fields and labels based on the hierarchical labeling system. Then, based on label-driven data governance rules, a series of governance operations such as data classification, cleansing, desensitization, and deduplication are dynamically executed at the data level. This label-driven governance model ensures a high degree of consistency between data quality and labels. S3. Label Conflict Detection and Dynamic Update: The label conflict detection and dynamic update unit monitors potential conflicts, redundancies, and failures during the label mapping process in real time. By utilizing a label path traversal algorithm and label version management mechanism, the label mapping path can be dynamically updated if an anomaly is detected.

[0056] Among them, the label modeling and hierarchical configuration in step S1 is based on the label modeling and hierarchical configuration unit. When building the label system, the label hierarchical modeling formula is first applied to conduct in-depth analysis of the data from multiple dimensions such as industry, business, and authority. For the industry dimension, refer to the industry standard classification and the company's own business characteristics to build an industry label tree. For example, in the financial industry, top-level industry labels such as "credit", "insurance", and "securities" are set; In terms of business, we categorize sales, production, and warehousing into business tags based on the company's actual business processes, and further refine sub-tags for each business segment. Regarding permissions, we assign permissions such as public, internal, and confidential based on data sensitivity and access control requirements. We also employ domain knowledge graph technology to assist with labeling hierarchies, using semantic analysis algorithms to uncover potential connections between data and refine the labeling architecture.

[0057] In the process of clarifying the tag relationship and path, the tag attributes are configured in detail through the attribute description language, including the tag name, data type, value range, description information, etc. For example, the numeric type, range limit and other attributes are set for the "transaction amount" tag.

[0058] Through a hierarchical partitioning algorithm, labels are organized in a tree structure, the parent-child hierarchical relationship of each label is determined, and a topological sorting algorithm is used to ensure the logical consistency of the hierarchical relationship. Utilizing path mapping technology, directed edge relationships between label nodes are established in the graph database. By setting node attributes and edge weights, the label mapping path is accurately described, achieving an ordered and structured label system.

[0059] Among them, label parsing and driving governance in step S2 refers to the multi-modal matching algorithm adopting a layered processing strategy in the label parsing and binding link: for structured data, by establishing a field feature library, using regular expressions and pattern recognition technology, quickly locate and extract formatted fields such as dates and numbers; for semi-structured data, based on JSONPath and XPath syntax parsing nodes, combined with semantic similarity calculation, identify the business meaning of the data; for unstructured text, use pre-trained NLP models (such as BERT, RoBERTa) to perform word vector conversion and semantic analysis to extract key information.

[0060] During the binding process, the tag binding priority formula comprehensively considers business weight (determining the importance of data through business process diagrams), data adaptability (calculating the cosine similarity between data features and tag attributes) and permission compliance (verified according to the enterprise data access policy), and uses a multi-objective decision-making algorithm to sort candidate tags to ensure binding accuracy.

[0061] Entering the label-driven dynamic data governance stage, dynamic loading is achieved through the rule engine. The rule engine has a built-in governance rule library based on the event-condition-action (ECA) model, covering data classification rules (such as dividing data sensitivity according to the label "customer level"), cleaning rules (processing logic for null values ​​and abnormal values), desensitization rules (application of algorithms such as masking and generalization) and deduplication rules (identification of duplicate data based on hash value comparison).

[0062] The system uses an automated data processing pipeline to send data to each processing module in sequence. During the processing, the data transformation trajectory is recorded through data lineage tracking technology to ensure the traceability of the governance process, and ultimately achieve deep integration and collaborative governance of data and labeling systems.

[0063] Among them, the label conflict detection and dynamic update in step S3 refers to the use of a dual-track parallel detection strategy in the real-time monitoring of label anomalies: on the one hand, the label hash comparison algorithm is used to split the label mapping data into independent blocks according to semantic units, and the SHA-256 hash value is calculated for each block. The sliding window technology is used to continuously compare the current block hash value with the historical version hash record, which can accurately capture subtle data structure changes or attribute value modifications; on the other hand, with the help of the label path traversal algorithm, based on the label topology network constructed by the graph database, a combination of depth-first search (DFS) and breadth-first search (BFS) is used to perform a systematic scan along the parent-child hierarchical relationship and associated dependency path of the label. It can not only identify directly conflicting label nodes, but also discover indirect logical contradictions caused by changes in hierarchical relationships through the path backtracking algorithm.

[0064] At the same time, combined with the label conflict detection formula, multi-dimensional factors such as the frequency of label attribute changes and changes in the number of associated nodes are comprehensively considered, and the monitored differences are weighted to accurately locate abnormal situations such as conflicts, redundancies and failures.

[0065] Once an anomaly is detected, the affected label association paths are first re-sorted through the label path traversal algorithm, and the shortest path algorithm and topological sorting algorithm in graph theory are used to generate the optimal path adjustment plan; at the same time, relying on the label version management mechanism, historical version snapshots of conflicting labels are retrieved for difference analysis, and the version merge algorithm is used to integrate the effective changes into the current version.

[0066] During the adjustment process, data consistency is ensured through the transaction management mechanism, modification operations on the label mapping path are atomically processed, and information such as the version number, modification timestamp, and operation log in the label metadata are synchronously updated, ultimately achieving dynamic repair and stable maintenance of the label system.

[0067] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0068] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A big data management system based on a hierarchical labeling system, characterized by: The system includes: data access preprocessing unit, label modeling and hierarchical configuration unit, label parsing and data binding unit, label driven data governance unit and label conflict detection and dynamic update unit; Data access preprocessing unit: responsible for receiving multi-source heterogeneous data, building unified conversion rules through data format standardization mapping technology, and converting raw data into a standard format; Label modeling and hierarchical configuration unit: Receives data preprocessed by the data access preprocessing unit, uses the label hierarchical modeling formula to build a label system, and clarifies label relationships and attributes by setting multi-level labels; Label parsing and data binding unit: Based on the system built by label modeling and hierarchical configuration units, it uses a multi-mode matching algorithm to parse data, automatically binds data fields and labels based on mapping paths, and optimizes binding selection using the label binding priority formula; Tag-driven data governance unit: After receiving the results of the tag parsing and data binding units, it dynamically loads tag-driven data governance rules to automatically clean abnormal and redundant problem data; Label conflict detection and dynamic update unit: It monitors the data labels processed by the label-driven data governance unit in real time, and uses the label conflict detection formula based on hash comparison and path traversal algorithm to detect conflicts and make timely adjustments.

2. The big data management system based on a hierarchical tag system according to claim 1, characterized in that: The data access pre-processing unit includes: (1) Multi-source heterogeneous data reception: Receive structured data, semi-structured data, and unstructured data from multiple sources, including common data formats such as JSON, CSV, and XML, as well as non-traditional data formats such as images and text; (2) Data format standardization processing: Using data format standardization mapping technology, by building unified data conversion rules, data in different formats can be accurately converted into a standard data format that supports label parsing.

3. The big data management system based on a hierarchical tag system according to claim 2, characterized in that: The label modeling and hierarchical configuration unit includes: (1) Construction of a multi-level label system: After receiving the standardized data processed by the data access preprocessing unit, a multi-level label system is constructed with the help of the label hierarchy modeling formula to clarify the parent-child relationship, hierarchy and specific attributes of each label; (2) Flexible adaptation and management support: It supports recursive expansion of labels and cross-industry adaptive configuration. It can flexibly adjust the labeling system to meet the data labeling needs in different industries and business scenarios.

4. The big data management system based on a hierarchical tag system according to claim 3 is characterized by: The tag parsing and data binding unit includes: (1) Data parsing and binding: Based on the established labeling system, the data is accurately parsed, and multi-modal matching algorithms are used to quickly and efficiently extract data features. Then, data fields are automatically bound to hierarchical labels according to the label mapping path. (2) Priority optimization selection for multi-tag adaptation: With the help of the tag binding priority formula, the tag’s sensitivity to the business, the degree of data adaptation, and the key factors of authority rationality are comprehensively considered to scientifically determine the tag binding priority.

5. The big data management system based on a hierarchical tag system according to claim 4 is characterized by: The tag-driven data governance unit includes: (1) Governance rule-driven based on tag binding: Based on the results of tag binding, targeted tag-driven data governance rules are dynamically loaded; (2) Efficient data cleaning and quality assurance: Through the dynamic loading of data governance rules driven by tags, the data is automatically screened in depth to accurately identify and eliminate abnormal data, redundant data, and data with incorrect formats.

6. The big data management system based on a hierarchical tag system according to claim 5, characterized in that: The label conflict detection and dynamic update unit includes: (1) Real-time label conflict monitoring mechanism: Relying on label hash comparison and label path traversal algorithm, with the help of label conflict detection formula, the current label mapping hash value and historical label hash value are monitored in real time, and the difference is compared continuously; (2) Dynamic adjustment ensures system stability: When label conflicts or abnormal mismatches are detected, the label mapping path and label version are automatically adjusted to accurately repair the contradictions in the label system.

7. A big data management method based on a hierarchical tag system, using the big data management system based on a hierarchical tag system according to claim 6, characterized in that: The specific steps of this method are as follows: S1. Label modeling and hierarchical configuration: Use the label modeling and hierarchical configuration unit to establish a multi-level label system, configure label attributes in detail, and clarify the label hierarchical relationship and label mapping path; S2. Tag Parsing and Driven Governance: This unit uses tag parsing and data binding to parse data and automatically bind data fields and tags based on a hierarchical tagging system. It then dynamically performs data classification, cleansing, desensitization, and deduplication operations based on tag-driven data governance rules. S3. Label conflict detection and dynamic update: The label conflict detection and dynamic update unit monitors the possible conflicts, redundancies and failures that may occur during the label mapping process in real time. The label path traversal algorithm and label version management mechanism are used to quickly and dynamically update the label mapping path.

8. The big data management method based on a hierarchical label system according to claim 7 is characterized by: The specific steps of label modeling and hierarchical configuration in step S1 are as follows: S11. Multi-dimensional tag system construction: The tag modeling and hierarchical configuration method relies on the tag modeling and hierarchical configuration unit and uses the tag hierarchical modeling formula to analyze access data from the industry, business, and permission dimensions. By setting up multi-level tags, a complete tag system framework is built. S12. Label relationships and paths are clear: When building a label system, the label hierarchy relationship is clarified through a hierarchical division algorithm, and the label mapping path is determined with the help of path mapping technology.

9. The big data management method based on a hierarchical label system according to claim 8, characterized in that: The specific steps of tag parsing and driver management in step S2 are as follows: S21. Label parsing and binding: The label parsing and driving governance method relies on the label parsing and data binding unit, adopts a multi-modal matching algorithm, and deeply mines data features; Combined with the tag binding priority formula, it comprehensively considers business sensitivity, adaptability, and authority factors, and automatically achieves accurate binding between data fields and tags based on the hierarchical tag system and tag mapping path; S22. Tag-driven dynamic data governance: With the help of tag-driven data governance units, tag-driven data governance rules are dynamically loaded, and automated data processing technology is used to perform data classification, cleaning, desensitization, and deduplication operations.

10. The big data management method based on a hierarchical label system according to claim 9, characterized in that: The specific steps of tag conflict detection and dynamic update in step S3 are as follows: S31. Real-time monitoring of label anomalies: The label conflict detection and dynamic update method relies on the label conflict detection and dynamic update unit, uses label hash comparison and label path traversal algorithms, combined with the label conflict detection formula, to scan the label mapping process in real time. By comparing the current and historical label hash values, it can accurately locate conflicts, redundancies, and failure anomalies. S32. Dynamic update: When an anomaly is detected, this method uses the label path traversal algorithm and the label version management mechanism to automatically adjust the label mapping path and update the label version.

Citation Information

Patent Citations

  • Large-scale sample storage management method and system based on multi-dimensional label system

    CN119474493A

  • Label management method and device for multi-modal data, equipment and storage medium

    CN119829735A

  • Media content intelligent label generation and target classification method

    CN119939314A

  • Intelligent recommendation method and system based on scholar academic background and user tag

    CN119961443A

  • Classification tagging management method and platform for full life cycle data

    CN120217054A

Cited By

  • Data weaving method for integration and treatment of multi-source heterogeneous data

    CN120910144A

  • Processing method for preventing data leakage

    CN121118115A