Real-time Adaptive Standardization System for Metadata in Multi-domain Data Sharing

Through metadata information layering, intelligent information collection and real-time adaptive standardization modules, combined with dynamic rule database and visual management, the real-time adaptability and cross-domain sharing problems of metadata standardization in the existing technology are solved, and efficient and semantically consistent data processing and management are achieved.

CN120144549BActive Publication Date: 2025-08-01ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510632109.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-01
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The existing metadata standardization methods lack real-time adaptability and cannot adapt to the frequent changes in data source formats and requirements. There are missing or inconsistent semantic information in cross-domain data sharing. Heterogeneity leads to data comprehension bias. Manual definition and maintenance templates are costly, difficult to expand to multi-domain scenarios, and low processing efficiency.

Method used

The metadata information layering module, intelligent information collection module, real-time adaptive standardization module and visual management module are adopted. Through the dynamic rule library and reinforcement learning mechanism, real-time adaptive standardization of metadata is realized, supporting unified collection and semantic mapping of different types of data, and providing an intuitive visual management interface.

Benefits of technology

Real-time adaptive standardization of cross-domain data has been realized, semantic interoperability capabilities have been improved, the threshold for use has been lowered, the accuracy and efficiency of processing unknown data have been improved, and it has adapted to changing application needs, and is highly scalable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144549B_ABST
    Figure CN120144549B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time adaptive standardization system for multi-domain data sharing, which includes a metadata information layering module, an intelligent information acquisition module, a real-time adaptive standardization module, and a visualization management module; the metadata information layering module identifies the sources of various metadata and performs preprocessing of cleaning, classification, and layering on the metadata in sequence; the intelligent information acquisition module uniformly acquires metadata information and converts the acquired data into a standardized intermediate representation; the real-time adaptive standardization module is used to realize the dynamic adaptation and real-time optimization of rules in the metadata standardization process, and its built-in dynamic rule library is used to store dynamically generated metadata mapping rules, supporting the automatic loading, updating, and semantic extension of rules; the visualization management module realizes the management, analysis, and dynamic adjustment of metadata through an intuitive interface. The present invention has high accuracy in processing unknown data, can adapt to different data scenarios in real time, and has strong semantic interoperability capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of metadata standardization, and more particularly to a real-time adaptive metadata standardization system for multi-domain data sharing. Background Art

[0002] With the rapid development of big data technology, the amount of data from all walks of life has grown explosively. At the same time, the demand for data sharing between different industries and fields is also continuously increasing. Whether in academia, industry or government departments, data sharing has become an important means to promote scientific research, business innovation and social development. However, due to the large diversity of data between different industries and fields in terms of type, format, semantics and metadata standards, cross-domain data sharing and interoperability face major challenges.

[0003] In actual operation, due to the lack of unified standards and specifications, data from different sources are often stored in different formats and contain different semantic information. This not only hinders efficient data exchange and integration, but also increases the cost and complexity of data analysis and utilization. Metadata, as the core bridge for data sharing, is used to describe the structure, content and semantics of data, and its standardization and normalization are crucial for the success of data sharing. Currently, most technical solutions for multi-source metadata standardization still have limitations and are difficult to fully meet the needs of multi-domain data sharing. Some platforms attempt to solve this problem by predefined templates or manual mapping, but this method is inefficient and difficult to cope with the dynamically changing data environment.

[0004] With the development of artificial intelligence technology, the current metadata standardization mainly adopts the following means: First, predefined multiple standard metadata templates and frameworks to apply to different industries and fields, such as natural science, finance, healthcare, retail, etc.; after the data source is accessed, extract structured metadata information, and select appropriate metadata standards based on the industry characteristics of the extracted information and the specific needs of users; align the collected metadata with the existing standards. Based on semantic web technology, use RDF (Resource Description Framework) and OWL (Web Ontology Language) to define the semantic relationships of different metadata standards and realize the semantic mapping between each metadata standard. That is, for each metadata standard (such as Dublin Core, DataCite, ISO 19115), establish the semantic mapping between standards through automatic and manual methods.

[0005] However, the existing methods for metadata standardization still have obvious deficiencies. Specifically, they are manifested as follows:

[0006] (1) Lack of real-time adaptability, unable to adjust standardization rules in real time according to user needs and data context, and difficult to adapt to the frequent changes in data source formats and requirements;

[0007] (2) In cross-domain data sharing, the lack or inconsistency of semantic information will lead to data understanding deviation and sharing obstacles; the semantic mapping between different standards cannot meet unknown metadata patterns or abnormal situations, and the effect is limited when dealing with data in unknown fields, and manual intervention in rule updates or troubleshooting is still required;

[0008] (3) The current metadata standards adopted in different fields (such as Dublin Core, DataCite, ISO 19115) have significant heterogeneity, and it is impossible to effectively unify these standards to achieve seamless sharing of cross-domain data;

[0009] (4) The cost of manually defining and maintaining templates is relatively high, and it is difficult to expand to multi-domain scenarios. When dealing with large-scale or real-time data, it shows high computational overhead and latency, and it is difficult to meet the performance requirements of actual applications;

[0010] (5) The existing metadata standardization tools have low processing efficiency. Summary of the Invention

[0011] In view of the deficiencies of the prior art, the present invention proposes a real-time adaptive standardization system for multi-domain data sharing, and the specific technical solutions are as follows:

[0012] A real-time adaptive standardization system for multi-domain data sharing includes a metadata information layering module, an intelligent information acquisition module, a real-time adaptive standardization module, and a visualization management module;

[0013] The metadata information layering module is used to preprocess the input metadata from various sources, including identifying the data source, and sequentially cleaning, classifying, and layering the metadata of different types and sources;

[0014] The intelligent information acquisition module supports the unified acquisition of metadata information, and adopts different quality control strategies and paths for metadata information at different layers, and finally converts the acquired data into a standardized intermediate representation;

[0015] The real-time adaptive standardization module is used to achieve dynamic adaptation and real-time optimization of rules in the metadata standardization process. It has a built-in dynamic rule library for storing dynamically generated metadata mapping rules, and supports automatic loading, updating, and semantic extension of rules;

[0016] The visualization management module realizes the management, analysis, and dynamic adjustment of metadata through an intuitive interface.

[0017] Furthermore, the metadata information layering module classifies metadata of different types and sources into the following three categories:

[0018] (1) Structured data with clear field definitions and data types, which is easy to parse;

[0019] (2) Semi-structured data that has a certain structure but is not as strict as structured data;

[0020] (3) Unstructured data that lacks a predefined schema or structure and requires complex analysis techniques to extract valuable information.

[0021] Furthermore, the metadata information layering module layers structured, semi-structured, and unstructured data as follows:

[0022] Uniformly classify the information recording the basic attributes of the record file into the basic layer, and based on the above information, verify whether the file is complete or damaged for subsequent tracking management;

[0023] Analyze the internal structures of structured and semi-structured data, extract the structure information, and map the extracted structure information into a unified data model to ensure a consistent representation form; and classify these structure information into the structure layer;

[0024] Identify the unstructured parts in semi-structured data, and uniformly classify this part of the data and the unstructured data into the semantic layer.

[0025] Furthermore, the intelligent information acquisition module includes a semantic layer preprocessing sub-module, a key information feature extraction sub-module, and a standardization conversion sub-module; among them,

[0026] The semantic layer preprocessing sub-module is used to clean the data in the semantic layer to remove noise elements that may interfere with subsequent analysis;

[0027] The key information feature extraction sub-module combines regular expression matching and natural language processing technologies to optimize the parsing of data content and identify the key terms, entities, and their relationships therein;

[0028] The standardization conversion sub-module is used to convert the information in the basic layer, the structure layer, and the processed semantic layer into a standardized intermediate representation.

[0029] Furthermore, the standardized intermediate representation is in the form of key-value pairs, and the key-value pair form includes basic data content and carries context information, enabling the dynamic rule library to better understand and process this data.

[0030] Furthermore, the key information feature extraction sub-module combines regular expression matching and natural language processing technologies to optimize the parsing of data content, identify key terms, entities, and the relationships between them, specifically including:

[0031] The key information feature extraction sub-module extracts key information in important dimensions including research methods, time range, and geographical location from the data at the semantic level based on a pre-constructed term vocabulary; for unconventional or unknown entity words not in the term vocabulary, the following two-step strategy is adopted for processing:

[0032] ① Sentence segmentation: Use regular expressions to segment the text according to rules including punctuation marks and abbreviations to ensure that each sentence is independent and complete;

[0033] ② Entity recognition: Apply a sequence labeling model to extract and identify entities in the text.

[0034] Furthermore, the key information feature extraction sub-module adopts an advanced sequence labeling model that combines a bidirectional long short-term memory network and a conditional random field, where the bidirectional long short-term memory network is used to capture the context information of the text, process the output sequence at one time, and automatically extract lexical features and their context environment; the output of the bidirectional long short-term memory network is a sequence of feature vectors, and the vectors at each position of the sequence of feature vectors are context-aware feature vectors that capture the context information of the current position and its surroundings;

[0035] The conditional random field receives the sequence of feature vectors as input, considers the dependencies between the labels in the sequence, converts the labels into a joint probability distribution, and outputs the optimal label sequence.

[0036] Furthermore, the real-time adaptive normalization module includes a pattern discrimination sub-module, a rule verification sub-module, a reinforcement learning sub-module, and an automated quality control sub-module;

[0037] The pattern discrimination sub-module is used to receive the standardized intermediate representation processed by the intelligent information acquisition module, retrieve the dynamic rule library, and determine whether the input data is in a known format. If it is in a known format, the dynamic rule library automatically loads the corresponding rules for mapping; if it is in an unknown format, an incremental learning algorithm is used to gradually build new rules;

[0038] The rule verification sub-module uses the built-in test data set to verify the constructed new rules, including verifying the integrity, consistency, and accuracy of the mapping results using the test set data, and storing the verified rules in the dynamic database in a standardized format; at the same time, generating a detailed report of the test results for developers to reference;

[0039] The reinforcement learning sub-module is used to collect user feedback and optimize the rules in real time after each data conversion task is completed;

[0040] The automated quality control sub-module performs quality checks on the mapped standardized metadata from three dimensions: integrity, logic, and consistency.

[0041] Further, the reinforcement learning sub-module records each successful data conversion as a reference, continuously optimizes the processing logic of the dynamic rule library itself, and reorders the priorities of the rules.

[0042] Further, the visualization management module includes a metadata browsing and retrieval sub-module, a dynamic rule configuration sub-module, a real-time monitoring and alerting sub-module, and an interactive learning sub-module;

[0043] The metadata browsing and retrieval sub-module is used for visual browsing and searching of metadata structures, contents, and semantics;

[0044] The dynamic rule configuration sub-module allows users to dynamically create, modify, and delete metadata mapping rules through a visual interface;

[0045] The real-time monitoring and alerting sub-module provides a real-time monitoring view of metadata processing, including data flow, rule execution status, and anomaly detection;

[0046] The interactive learning sub-module supports users to check and repair the generated metadata, and the front end supports operations such as batch editing of fields and data completion; at the same time, a quality control report is generated and provided to the user, and the received user repair actions and feedback are submitted to the reinforcement learning sub-module.

[0047] The beneficial effects of the present invention are as follows:

[0048] (1) Adaptive optimization ability: The present invention can optimize mapping rules according to user feedback and historical data to achieve continuous improvement. Compared with the insufficient accuracy of existing automated tools when facing data in unknown fields, the adaptive optimization ability of the present invention greatly improves the accuracy of processing unknown data.

[0049] (2) Enhanced dynamic adaptability: Through the real-time adaptive standardization module, the present invention can dynamically adjust metadata mapping rules in combination with context information to ensure semantic consistency of cross-domain data. Compared with traditional static standard mapping tools, the present invention can adapt to different data scenarios in real time to meet changing application requirements.

[0050] (3) Improved semantic interoperability ability: The present invention uses semantic network technology to enhance the semantic relevance between metadata in different fields and realizes deep integration at the semantic level. Compared with existing technologies that rely on a single semantic model, the present invention has more advantages in processing heterogeneous data.

[0051] (4) Strong user - friendliness: Provide an intuitive visual management module, enabling non - professional users to conveniently adjust standardized rules and feedback optimization results. Compared with template tools that require professional knowledge, the present invention reduces the usage threshold and enhances the universality of the tool.

[0052] (5) Modular design, easy to expand: Through a modular architecture, each functional component of the present invention is independent and can be expanded or integrated into an existing data management system according to actual needs. Compared with existing solutions for specific fields, the present invention has a wider application scope in cross - domain and cross - platform data sharing. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a schematic diagram of the composition of the real - time adaptive standardization system for multi - domain data sharing according to an embodiment of the present invention.

[0054] Figure 2 It is a schematic diagram of the metadata information layering module for layering structured, semi - structured, and unstructured metadata information.

[0055] Figure 3 It is a schematic diagram of the intelligent information acquisition module.

[0056] Figure 4 It is a specific architecture diagram of the intelligent information acquisition module.

[0057] Figure 5 It is an architecture diagram of the real - time adaptive standardization module.

[0058] Figure 6 It is a flowchart of the implementation of the real - time adaptive standardization module.

[0059] Figure 7 It is a schematic diagram of the architecture of the visual management module and its interaction with the backend. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] The present invention will be described in detail below according to the drawings and preferred embodiments. The purpose and effects of the present invention will become clearer. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0061] As Figure 1 shown, the real - time adaptive standardization system for multi - domain data sharing in this embodiment includes a metadata information layering module, an intelligent information acquisition module, a real - time adaptive standardization module, and a visual management module. The function implementation of each module will be introduced in detail below.

[0062] I. Metadata Information Layering Module

[0063] The metadata information layering module is used to preprocess the metadata from various sources. The preprocessing specifically includes identifying the data sources (such as files, databases, APIs, etc.), and sequentially cleaning, classifying, and layering the metadata of different types and sources to ensure that appropriate measures can be taken for different types of data sources in subsequent processing steps.

[0064] The metadata information layering module first performs preprocessing such as noise cleaning and data cleaning on the raw data collected from various sources. Then, according to the degree of structuring of the metadata information, it can be divided into three categories:

[0065] (1) Structured data with clear field definitions and data types, which is easy to parse;

[0066] (2) Semi-structured data that has a certain structure but is not as strict as structured data;

[0067] (3) Unstructured data that lacks a predefined schema or structure and requires complex analysis techniques to extract valuable information.

[0068] If the metadata is stored in a relational database (RDBMS, Relational Database Management System), or in the format of a CSV file, it is usually structured data; semi-structured data is usually in formats such as JSON, XML, YAML, etc.; unstructured data is usually stored in formats such as text files, images, audio, and video.

[0069] Finally, as Figure 2 shown, the metadata information layering module also layers the structured, semi-structured, and unstructured metadata information respectively:

[0070] (1) Uniformly classify the information recording the basic attributes of the file (including file name, path, extension, size, creation time, and modification time, etc.) into the basic layer. At the same time, based on the above information, verify whether the file is complete or damaged for subsequent tracking management.

[0071] (2) For structured and semi-structured data, parse its internal structure, extract information such as field names, data types, and table structures; identify the data pattern, such as the row and column structure in a table and the key-value pair relationship in JSON. Map the extracted structure information into a unified data model to ensure a consistent representation form. And classify this part of the structure information into the structure layer.

[0072] (3) Identify the unstructured part in the semi-structured data, and this part of the data is uniformly classified into the semantic layer together with the unstructured data.

[0073] II. Intelligent Information Acquisition Module

[0074] The intelligent information collection module supports the unified collection of data at the basic layer, structural layer, and semantic layer. Different quality control strategies and paths are adopted for the metadata information at different layers. The semantic layer information needs to be preprocessed such as data cleaning, and then key information feature extraction is carried out to form metadata information in a structured form. After completing the above steps, the metadata information is integrated with the structural layer information and the basic layer information, and is output in the form of a standardized intermediate representation, such as Figure 3 shown.

[0075] such as Figure 4 shown, the intelligent information collection module specifically includes a semantic layer preprocessing sub-module, a key information feature extraction sub-module, and a standardization conversion sub-module.

[0076] (2.1) Semantic layer preprocessing sub-module

[0077] The semantic layer preprocessing sub-module is used to clean the data in the semantic layer to remove noise elements that may interfere with subsequent analysis. This includes but is not limited to deleting irrelevant characters, HTML tags, and special symbols, so as to provide a clean data set for subsequent semantic analysis.

[0078] (2.2) Key information feature extraction sub-module

[0079] The key information feature extraction sub-module combines regular expression matching and natural language processing (NLP, Natural Language Processing) technologies, aiming to optimize the parsing of data content, identify key terms, entities and the relationships between them, specifically including:

[0080] The key information feature extraction sub-module extracts key information covering important dimensions such as research methods, time range, geographical location, etc. from the data in the semantic layer based on a pre-constructed term vocabulary; for unconventional or unknown entity words not in the term vocabulary, the following two-step strategy is adopted for processing:

[0081] ① Sentence segmentation: Use regular expressions to segment the text according to rules such as punctuation marks and abbreviations to ensure that each sentence is independent and complete.

[0082] ② Entity recognition: Apply a sequence labeling model to extract and identify entities in the text.

[0083] In this embodiment, the key information feature extraction sub-module adopts an advanced sequence labeling model BILSTM-CRF that combines a bidirectional long short-term memory network (BILSTM) and a conditional random field (CRF). Among them, the bidirectional long short-term memory network is used to capture the context information of the text, process the output sequence at one time, and automatically extract lexical features and their context environment without manually defining features. Each vector at each position in the feature vector sequence output by BILSTM is a context-aware feature vector that captures the context information of the current position and its surroundings; CRF is used to receive the feature vector sequence output by BILSTM as input, consider the dependencies between the labels in the sequence, convert the labels into a joint probability distribution, and output the optimal label sequence. CRF can adjust the prediction of the current label during the labeling process, taking into account the influence of the previous and subsequent labels, and finally determine the optimal label sequence. Each label can represent information such as entity category and syntactic structure. At each position in the input sequence, a label needs to be assigned to it. The joint distribution function defines the probability of the label at each position. Specifically, the label prediction at each position on the label sequence can be regarded as a random variable, and the label prediction situations at all positions are used as the joint probability function of the joint distribution. In this way, only by modeling and optimizing the joint probability distribution of the label sequence can the optimal prediction result of the entire label sequence be obtained. The advantage of doing this is not only to improve the accuracy of entity recognition but also to enhance the model's ability to understand complex contexts. The output results of the BILSTM-CRF model are as follows:

[0084]

[0085] The key information feature extraction sub-module can effectively achieve the automated analysis of semantic layer information, ensure that the data content is fully understood and utilized, and at the same time avoid the problem of repeated expressions.

[0086] (2.3) Standardization conversion sub-module

[0087] The standardization conversion sub-module is used to convert the basic layer information, structure layer information, and processed semantic layer information into a standardized intermediate representation, so as to maintain the consistency of information transmission between different modules and provide a clear and definite data basis for subsequent rule applications. In this embodiment, the standardized intermediate representation is in the form of key-value pairs. In addition to the basic data content, the standardized intermediate representation will carry context information, etc., enabling the dynamic rule library to better understand and process this data.

[0088] III. Real-time adaptive standardization module

[0089] The real-time adaptive standardization module aims to achieve dynamic adaptation and real-time optimization of rules in the metadata standardization process. Through incremental learning, reinforcement learning, and real-time feedback mechanisms, it meets the standardization requirements of multi-domain and multi-type metadata. The real-time adaptive standardization module has a built-in dynamic rule library, which is a collection of updatable rules used to guide the data standardization process. Different from traditional static configuration methods, the dynamic rule library is not fixed but can automatically expand and optimize according to newly encountered data patterns. The dynamic rule library is used to store dynamically generated metadata mapping rules and supports automatic loading, updating, and semantic extension of rules.

[0090] As Figure 5 shown, the real-time adaptive standardization module includes a pattern discrimination sub-module, a rule verification sub-module, a reinforcement learning sub-module, and an automated quality control sub-module.

[0091] (3.1)Pattern Discrimination Sub-module

[0092] As Figure 6 shown, the pattern discrimination sub-module is used to receive the standardized intermediate representation processed by the intelligent information collection module and retrieve the dynamic rule library to determine whether the input data is in a known format. If it is in a known format, the dynamic rule library automatically loads the corresponding rules for mapping. The criteria for judging a known format (meeting one of the following):

[0093] Condition 1: There are rules in the dynamic rule library that are exactly the same as the field semantics, structure, or name.

[0094] Condition 2: It conforms to the predefined standard field name (such as dc.title in Dublin Core).

[0095] When the pattern discrimination sub-module determines that the input data is in an unknown format, an incremental learning algorithm is used to gradually construct new rules. Specifically, based on the context information carried by the standardized intermediate representation, the similarity between the key information and the terms in the metadata standard is evaluated for semantic similarity matching recommendation, which specifically includes the following steps:

[0096] ① Calculate the cosine similarity between two word vectors. This step aims to quantify the direct similarity between keywords and terms:

[0097]

[0098] Among them, a is the extracted keyword, b is the term in the metadata standard, a , i ,

[0098] ,

[0099] , i represents the word vector corresponding to a, and b i represents the word vector corresponding to b.

[0099] ②Next, to more comprehensively understand the context environment, first calculate the semantic similarity of the vocabulary sets on both sides of the keyword, and then calculate the semantic similarity of keywords a and b. When the calculated semantic similarity is higher than the preset threshold, it is preliminarily determined as a relevant field, and a new mapping rule is formed:

[0100]

[0101]

[0102]

[0103] Among them, represents the vocabulary set on the left side of keyword a, represents the vocabulary set on the right side of keyword a, represents the vocabulary set on the left side of keyword b, represents the vocabulary set on the right side of keyword b, represents the number of words in the vocabulary set, represents the number of words in the vocabulary set, represents and the semantic similarity of, represents and the semantic similarity of, sim(a,b) represents the semantic similarity of keywords a and b; w1, w2, w3 are preset weight parameters, and w1 + w2 + w3 = 1. In this embodiment, w1 = 0.6, w2 = w3 = 0.2. In this embodiment, the preset threshold is 0.85.

[0104] (3.2) Rule Verification Sub-module

[0105] To further ensure that the new rule meets the requirements for entering the dynamic rule library, the rule verification sub-module uses the built-in test data set to verify the rule, including but not limited to aspects such as the integrity, consistency, and accuracy of the test set data mapping results. The rule verification sub-module stores the rules that pass the verification in the dynamic rule library in a standardized format. At the same time, it generates a detailed report of the test results for developers to refer to.

[0106] (3.3) Automatic Quality Control Sub-module

[0107] The automatic quality control sub-module conducts quality inspections on the mapped standardized metadata from three dimensions: integrity, logic, and consistency. The specific rule definitions are as follows:

[0108] (a) Integrity: Ensure that the data in each field of the mapping result is complete, and avoid situations of missing or non-compliant data.

[0109] ① Field non - null check: Check whether there are null values in the key fields (such as title, data source, subject field, etc.) in the mapping result.

[0110] ② Field length check: It is agreed that the minimum length threshold of some fields (such as keywords and authors) is not less than 4 characters.

[0111] (b)Logic: Verify whether the logical relationships in the mapping result are reasonable to ensure that the associations between data meet the expectations.

[0112] ① Temporal logic check, such as: Observation time ≤ Data release time ≤ Metadata creation time;

[0113] ② Set a reasonable time range and statistically identify the fields that exceed the time threshold;

[0114] ③ Standardized term matching: It is agreed that the value range of some metadata fields must come from a controlled vocabulary. For example, language = ‘eng’ will be identified and should be replaced with ‘en’.

[0115] (c)Consistency

[0116] ① Dynamic consistency: It is required that the real - time data stream processing delay consistency ≤ 50ms (P99 index); the tolerance for the difference between batch processing and real - time processing ≤ 0.1%;

[0117] ② Field semantic consistency: It is required that the value of the mapped meta - field is consistent with the original semantics.

[0118] (3.4)Reinforcement learning sub - module

[0119] The reinforcement learning sub - module selects the Apache Flink framework that supports real - time data stream processing. The reinforcement learning sub - module is used to collect user feedback after each data conversion task is completed and optimize the rules in real time. The reinforcement learning sub - module records each successful data conversion as a reference, continuously optimizes the processing logic of the dynamic rule library itself, and re - ranks the priorities of the rules. With the accumulation of converted data, the dynamic rule library will not only be able to better process the current type of data, but also accumulate experience in processing other potential data types.

[0120] The reinforcement learning sub - module adopts stream processing technology, can efficiently process large - scale and real - time data streams, and significantly reduces latency. Compared with existing semantic mapping tools and automated matching tools, the present invention has obvious advantages in processing performance, with high real - time processing efficiency and is suitable for scenarios with high real - time requirements.

[0121] IV. Visualization management module

[0122] The Visualization Management Module realizes the management, analysis, and dynamic adjustment of metadata through an intuitive interface. This module is designed based on an architecture that separates the front end from the back end, as Figure 7 shown. The Visualization Management Module includes a Metadata Browsing and Retrieval Sub-module, a Dynamic Rule Configuration Sub-module, a Real-time Monitoring and Alerting Sub-module, and an Interactive Learning Sub-module.

[0123] The Metadata Browsing and Retrieval Sub-module is used for the visual browsing and searching of metadata structure, content, and semantics. This sub-module displays the metadata hierarchical relationship through a tree structure and supports highlighting search results, fuzzy matching, and advanced queries.

[0124] The Dynamic Rule Configuration Sub-module supports users to dynamically create, modify, and delete metadata mapping rules through a visual interface. The front-end component adopts a drag-and-drop interface and supports modular grouping of rule conditions; the back-end engine is based on Drools to implement the logical processing and execution of rules. The front-end component receives the source and target mapping relationships of fields selected by the user in the interface, provides a syntax checking function to prevent rule conflicts, and dynamically displays the execution process of the rules in the form of a flowchart.

[0125] The Real-time Monitoring and Alerting Sub-module provides a real-time monitoring view of metadata processing, including data flow, rule execution status, anomaly detection, etc. Echarts is used to draw the data flow diagram to show the distribution of metadata at different processing stages. Real-time monitoring and threshold alerting are implemented based on the Prometheus tool.

[0126] The Interactive Learning Sub-module supports users to check and repair the generated metadata. The front end supports operations such as batch editing of fields and data completion. At the same time, users can view the quality control reports generated by the automated quality control sub-module, which contain detailed descriptions of the problem sources. The Interactive Learning Sub-module receives the actions and feedback repaired by the user and submits them to the reinforcement learning sub-module.

[0127] Those of ordinary skill in the art can understand that the above are only preferred examples of the invention and are not used to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, for those skilled in the art, they can still modify the technical solutions described in the foregoing examples, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, etc. made within the spirit and principle of the invention shall be included within the protection scope of the invention.

Claims

1. A real-time adaptive standardization system for multi-domain data sharing, characterized in that It includes a metadata information layering module, an intelligent information collection module, a real-time adaptive standardization module, and a visualization management module; The metadata information layering module is used to preprocess the metadata from various sources of input, including identifying the data source, and sequentially cleaning, classifying, and layering the metadata of different types and sources; The intelligent information collection module supports the unified collection of metadata information, and adopts different quality control strategies and paths for the metadata information of different layers, and finally converts the collected data into a standardized intermediate representation; The real-time adaptive standardization module is used to achieve dynamic adaptation and real-time optimization of rules in the metadata standardization process. It has a built-in dynamic rule library for storing dynamically generated metadata mapping rules, and supports automatic loading, updating, and semantic extension of rules; The real-time adaptive standardization module includes a pattern discrimination sub-module, a rule verification sub-module, a reinforcement learning sub-module, and an automated quality control sub-module; The pattern discrimination sub-module is used to receive the standardized intermediate representation processed by the intelligent information collection module, and retrieve the dynamic rule library to determine whether the input data is in a known format. If it is in a known format, the dynamic rule library automatically loads the corresponding rules for mapping; If it is in an unknown format, an incremental learning algorithm is used to gradually construct new rules; The rule verification sub-module uses the built-in test data set to verify the constructed new rules, including verifying the integrity, consistency, and accuracy of the mapping results using the test set data, and storing the verified rules in the dynamic database in a standardized format; At the same time, generate a detailed report of the test results for developers to refer to; The reinforcement learning sub-module is used to collect user feedback and optimize the rules in real time after each data conversion task is completed; The automated quality control sub-module performs quality checks on the mapped standardized metadata from three dimensions: integrity, logic, and consistency; The visualization management module realizes the management, analysis, and dynamic adjustment of metadata through an intuitive interface.

2. The real-time adaptive standardization system for multi-domain data sharing metadata according to claim 1, characterized in that, The metadata information layering module classifies the metadata of different types and sources into the following three categories: (1) Structured data with clear field definitions and data types, which is easy to parse; (2) Semi-structured data with a certain structure but not as strict as structured data; (3) Unstructured data that lacks a predefined schema or structure and requires complex analysis techniques to extract valuable information.

3. The real-time adaptive standardization system for multi-domain data sharing metadata according to claim 2, characterized in that, The metadata information layering module layers the structured, semi-structured, and unstructured data as follows: Uniformly classify the information recording the basic attributes of the record file into the basic layer, and based on the above information, verify whether the file is complete or damaged for subsequent tracking management; Analyze the internal structure of structured and semi-structured data, extract the structure information, and map the extracted structure information to a unified data model to ensure a consistent representation form; And classify this structure information into the structure layer; Identify the unstructured part in the semi-structured data, and uniformly classify this part of the data and the unstructured data into the semantic layer.

4. The real-time adaptive standardization system for multi-domain data sharing according to claim 3, characterized in that, The intelligent information collection module includes a semantic layer preprocessing sub-module, a key information feature extraction sub-module, and a standardization conversion sub-module; among them, The semantic layer preprocessing sub-module is used to clean the data of the semantic layer to remove noise elements that may interfere with subsequent analysis; The key information feature extraction sub-module combines regular expression matching and natural language processing techniques to optimize the parsing of data content, and identify key terms, entities, and the relationships between them; The standardization conversion sub-module is used to convert the basic layer information, structure layer information, and processed semantic layer information into a standardized intermediate representation.

5. The real-time adaptive standardization system for metadata for multi-domain data sharing according to claim 4, wherein, The standardized intermediate representation is in the form of key-value pairs, and the key-value pair form includes basic data content and carries context information, enabling the dynamic rule library to better understand and process this data.

6. The real-time adaptive standardization system for multi-domain data sharing according to claim 4, characterized in that The key information feature extraction sub-module combines regular expression matching and natural language processing techniques to optimize the parsing of data content, and identify key terms, entities, and the relationships between them, specifically including: The key information feature extraction sub-module extracts key information of important dimensions including research methods, time range, and geographical location from the data of the semantic layer based on a pre-constructed term vocabulary; for unconventional or unknown entity words not in the term vocabulary, the following two-step strategy is adopted for processing: ① Sentence segmentation: Use regular expressions to segment the text according to rules including punctuation marks and abbreviations to ensure that each sentence is independent and complete; ② Entity recognition: Apply a sequence labeling model to extract and identify entities in the text.

7. The real-time adaptive standardization system for metadata for multi-domain data sharing according to claim 6, characterized in that, The key information feature extraction sub-module adopts an advanced sequence labeling model that combines a bidirectional long short-term memory network and a conditional random field. Among them, the bidirectional long short-term memory network is used to capture the context information of the text, process the output sequence at one time, and automatically extract lexical features and their context environment; the output of the bidirectional long short-term memory network is a sequence of feature vectors, and the vector at each position of the sequence of feature vectors is a context-aware feature vector that captures the context information of the current position and its surroundings; The conditional random field receives the sequence of feature vectors as input, considers the dependencies between labels in the sequence, converts the labels into a joint probability distribution, and outputs the optimal label sequence.

8. The real-time adaptive standardization system for metadata for multi-domain data sharing according to claim 1, characterized in that, The reinforcement learning sub-module records each successful data conversion as a reference, continuously optimizes the processing logic of the dynamic rule library itself, and reorders the priorities of the rules.

9. The real-time adaptive standardization system for multi-domain data sharing according to claim 1, wherein The visualization management module includes a metadata browsing and retrieval sub-module, a dynamic rule configuration sub-module, a real-time monitoring and alerting sub-module, and an interactive learning sub-module; The metadata browsing and retrieval sub-module is used for visual browsing and searching of metadata structure, content, and semantics; The dynamic rule configuration sub-module allows users to dynamically create, modify, and delete metadata mapping rules through a visual interface; The real-time monitoring and alerting sub-module provides a real-time monitoring view of metadata processing, including data flow, rule execution status, and anomaly detection; The interactive learning sub-module supports the user to check and repair the generated metadata, and the front end supports operations such as batch editing of fields and data completion; at the same time, a quality control report is generated and provided to the user, and the actions and feedback of the user's repair received are submitted to the reinforcement learning sub-module.

Citation Information

Patent Citations

  • Multi-source heterogeneous ecological environment big data processing method and system based on data lake

    CN111459908A

  • Multi-platform metadata standardization processing method and device

    CN117851387A