Metadata real-time adaptive standardization system for multi-field data sharing

By designing a real-time adaptive standardization system for metadata, using dynamic rule bases and reinforcement learning algorithms to achieve real-time adaptation and optimization of metadata standardization, the problem of difficult real-time adaptation and cross-domain data sharing barriers in the existing technology is solved, and efficient cross-domain data sharing and interoperability are achieved.

CN120144549AActive Publication Date: 2025-06-13ZHEJIANG LAB

Patent Information

Application Number
CN202510632109.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

Existing metadata standardization technologies are difficult to achieve real-time adaptation, and cannot effectively handle the missing or inconsistency of semantic information of cross-domain data. The heterogeneity of metadata standards in different fields leads to data sharing obstacles, and manual definition and maintenance of templates are expensive, making it difficult to expand to multi-domain scenarios.

Method used

A real-time adaptive standardization system for metadata is designed, including metadata information layering module, intelligent information collection module, real-time adaptive standardization module and visual management module. The system realizes dynamic adaptation and real-time optimization of rules in the process of metadata standardization through dynamic rule bases and reinforcement learning algorithms, supporting semantic integration of cross-domain data.

Benefits of technology

Real-time adaptation and optimization of the metadata standardization process is realized, the semantic consistency and interoperability of cross-domain data are improved, the cost of manually defining and maintaining templates is reduced, and the dynamic adaptability and scalability of the system is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144549A_ABST
    Figure CN120144549A_ABST
Patent Text Reader

Abstract

The invention discloses a metadata real-time self-adaptive standardization system for multi-field data sharing. The metadata real-time self-adaptive standardization system comprises a metadata information layering module, an intelligent information acquisition module, a real-time self-adaptive standardization module and a visual management module, the metadata information layering module identifies sources of various metadata, and sequentially performs cleaning, classification and layering preprocessing on the metadata; the intelligent information acquisition module uniformly acquires metadata information and converts the acquired data into standardized intermediate representation; the real-time adaptive standardization module is used for realizing rule dynamic adaptation and real-time optimization in a metadata standardization process, is internally provided with a dynamic rule library, is used for storing dynamically generated metadata mapping rules and supports automatic loading, updating and semantic extension of the rules; and the visual management module realizes management, analysis and dynamic adjustment of metadata through a visual interface. The method is high in accuracy of processing unknown data, can adapt to different data scenes in real time, and is high in semantic interoperation capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of metadata standardization, and particularly to a real-time adaptive metadata standardization system for multi-domain data sharing. Background Art

[0002] With the rapid development of big data technology, the amount of data from all walks of life has grown explosively. At the same time, the demand for data sharing between different industries and fields is also continuously increasing. Whether in academia, industry or government departments, data sharing has become an important means to promote scientific research, business innovation and social development. However, due to the large diversity of data between industries and fields in terms of type, format, semantics and metadata standards, cross-domain data sharing and interoperability face major challenges.

[0003] In actual operation, due to the lack of unified standards and specifications, data from different sources are often stored in different formats and contain different semantic information. This not only hinders efficient data exchange and integration, but also increases the cost and complexity of data analysis and utilization. Metadata, as the core bridge for data sharing, is used to describe the structure, content and semantics of data. Its standardization and normalization are crucial for the success of data sharing. Currently, most technical solutions for multi-source metadata standardization still have limitations and are difficult to fully meet the needs of multi-domain data sharing. Some platforms attempt to solve this problem by predefined templates or manual mapping, but this method is inefficient and difficult to cope with the dynamically changing data environment.

[0004] With the development of artificial intelligence technology, the current metadata standardization mainly adopts the following means: First, predefined multiple standard metadata templates and frameworks to apply to different industries and fields, such as natural science, finance, medical, retail, etc.; after the data source is accessed, extract the structured type of metadata information, and select the appropriate metadata standard based on the industry characteristics of the extracted information and the specific needs of users; align the collected metadata with the existing standards. Based on semantic web technology, use RDF (Resource Description Framework) and OWL (Web Ontology Language) to define the semantic relationships of different metadata standards and achieve the semantic mapping between each metadata standard. That is, for each metadata standard (such as Dublin Core, DataCite, ISO 19115), establish the semantic mapping between standards through automatic and manual methods.

[0005] However, the existing methods for metadata standardization still have obvious deficiencies. Specifically manifested as:

[0006] (1) Lack of real-time self-adaptability, unable to adjust the standardization rules in real time according to user needs and data context, and it is difficult to adapt to the frequent changes in data source formats and requirements;

[0007] (2) In cross-domain data sharing, the lack or inconsistency of semantic information will lead to data understanding deviation and sharing obstacles; the semantic mapping between different standards cannot meet unknown metadata patterns or abnormal situations, and the effect is limited when dealing with data in unknown fields, and manual intervention in rule updates or troubleshooting is still required;

[0008] (3) The current metadata standards adopted in different fields (such as Dublin Core, DataCite, ISO 19115) have significant heterogeneity and cannot effectively unify these standards to achieve seamless sharing of cross-domain data;

[0009] (4) The cost of manually defining and maintaining templates is relatively high, and it is difficult to expand to multi-domain scenarios. It shows high computational overhead and latency when dealing with large-scale or real-time data, and it is difficult to meet the performance requirements of actual applications;

[0010] (5) The existing metadata standardization tools have low processing efficiency. Summary of the Invention

[0011] In view of the deficiencies of the prior art, the present invention proposes a real-time adaptive standardization system for multi-domain data sharing, and the specific technical solutions are as follows:

[0012] A real-time adaptive standardization system for multi-domain data sharing, including a metadata information layering module, an intelligent information acquisition module, a real-time adaptive standardization module, and a visualization management module;

[0013] The metadata information layering module is used to preprocess the input metadata from various sources, including identifying the data source, and sequentially cleaning, classifying, and layering the metadata of different types and sources;

[0014] The intelligent information acquisition module supports the unified acquisition of metadata information, and adopts different quality control strategies and paths for metadata information at different layers, and finally converts the acquired data into a standardized intermediate representation;

[0015] The real-time adaptive standardization module is used to realize the dynamic adaptation and real-time optimization of rules in the metadata standardization process. It has a built-in dynamic rule library for storing dynamically generated metadata mapping rules, and supports automatic loading, updating, and semantic extension of rules;

[0016] The visualization management module realizes the management, analysis, and dynamic adjustment of metadata through an intuitive interface.

[0017] Further, the metadata information layering module classifies metadata of different types and sources into the following three categories:

[0018] (1) Structured data with clear field definitions and data types, which is easy to parse;

[0019] (2) Semi-structured data with a certain structure but not as strict as structured data;

[0020] (3) Unstructured data lacking predefined schemas or structures, which requires complex analysis techniques to extract valuable information.

[0021] Further, the metadata information layering module layers structured, semi-structured, and unstructured data as follows:

[0022] Uniformly classify the information recording the basic attributes of the record file into the basic layer, and based on the above information, verify whether the file is complete or damaged for subsequent tracking management;

[0023] Analyze the internal structures of structured and semi-structured data, extract the structure information, and map the extracted structure information into a unified data model to ensure a consistent representation form; and classify these structure information into the structure layer;

[0024] Identify the unstructured parts in the semi-structured data, and uniformly classify this part of the data and the unstructured data into the semantic layer.

[0025] Further, the intelligent information acquisition module includes a semantic layer preprocessing sub-module, a key information feature extraction sub-module, and a standardization conversion sub-module; among them,

[0026] The semantic layer preprocessing sub-module is used to clean the data in the semantic layer to remove noise elements that may interfere with subsequent analysis;

[0027] The key information feature extraction sub-module combines regular matching and natural language processing technologies to optimize the parsing of data content and identify the key terms, entities, and their relationships therein;

[0028] The standardization conversion sub-module is used to convert the information in the basic layer, the structure layer, and the processed semantic layer into a standardized intermediate representation.

[0029] Further, the standardized intermediate representation is in the form of key-value pairs, and the key-value pair form includes basic data content and carries context information, enabling the dynamic rule library to better understand and process this data.

[0030] Furthermore, the key information feature extraction sub-module combines regular expression matching and natural language processing techniques to optimize the parsing of data content, identify key terms, entities, and the relationships between them, specifically including:

[0031] The key information feature extraction sub-module extracts key information in important dimensions including research methods, time range, and geographical location from the data at the semantic level based on a pre-constructed term vocabulary; for unconventional or unknown entity words not in the term vocabulary, the following two-step strategy is adopted for processing:

[0032] ① Sentence segmentation: Use regular expressions to segment the text according to rules including punctuation marks and abbreviations to ensure that each sentence is independent and complete;

[0033] ② Entity recognition: Apply a sequence annotation model to extract and identify entities in the text.

[0034] Furthermore, the key information feature extraction sub-module adopts an advanced sequence annotation model that combines a bidirectional long short-term memory network and a conditional random field. The bidirectional long short-term memory network is used to capture the context information of the text, process the output sequence at one time, and automatically extract lexical features and their context environment; the output of the bidirectional long short-term memory network is a sequence of feature vectors, and the vectors at each position of the sequence of feature vectors are context-aware feature vectors that capture the context information of the current position and its surroundings;

[0035] The conditional random field receives the sequence of feature vectors as input, considers the dependencies between the labels in the sequence, converts the labels into a joint probability distribution, and outputs the optimal label sequence.

[0036] Furthermore, the real-time adaptive normalization module includes a pattern discrimination sub-module, a rule verification sub-module, a reinforcement learning sub-module, and an automated quality control sub-module;

[0037] The pattern discrimination sub-module is used to receive the standardized intermediate representation processed by the intelligent information collection module, retrieve the dynamic rule library, and determine whether the input data is in a known format. If it is in a known format, the dynamic rule library automatically loads the corresponding rules for mapping; if it is in an unknown format, an incremental learning algorithm is used to gradually construct new rules;

[0038] The rule verification sub-module uses the built-in test data set to verify the constructed new rules, including verifying the integrity, consistency, and accuracy of the mapping results using the test set data, and storing the verified rules in the dynamic database in a standardized format; at the same time, generating a detailed report of the test results for developers to reference;

[0039] The reinforcement learning sub-module is used to collect user feedback and optimize the rules in real time after each data conversion task is completed;

[0040] The automated quality control sub-module performs quality checks on the mapped standardized metadata from three dimensions: integrity, logic, and consistency.

[0041] Furthermore, the reinforcement learning sub-module records each successful data conversion as a reference, continuously optimizes the processing logic of the dynamic rule library itself, and reorders the priorities of the rules.

[0042] Furthermore, the visualization management module includes a metadata browsing and retrieval sub-module, a dynamic rule configuration sub-module, a real-time monitoring and alerting sub-module, and an interactive learning sub-module;

[0043] The metadata browsing and retrieval sub-module is used for visual browsing and searching of metadata structures, contents, and semantics;

[0044] The dynamic rule configuration sub-module allows users to dynamically create, modify, and delete metadata mapping rules through a visual interface;

[0045] The real-time monitoring and alerting sub-module provides a real-time monitoring view of metadata processing, including data flow, rule execution status, and anomaly detection;

[0046] The interactive learning sub-module supports users to check and repair the generated metadata. The front end supports operations such as batch editing of fields and data completion; at the same time, it generates a quality control report for the user and submits the received user repair actions and feedback to the reinforcement learning sub-module.

[0047] The beneficial effects of the present invention are as follows:

[0048] (1) Adaptive optimization ability: The present invention can optimize mapping rules according to user feedback and historical data to achieve continuous improvement. Compared with the insufficient accuracy of existing automated tools when facing data in unknown fields, the adaptive optimization ability of the present invention greatly improves the accuracy of processing unknown data.

[0049] (2) Enhanced dynamic adaptability: The present invention can dynamically adjust metadata mapping rules in combination with context information through a real-time adaptive standardization module to ensure semantic consistency of cross-domain data. Compared with traditional static standard mapping tools, the present invention can adapt to different data scenarios in real time to meet changing application requirements.

[0050] (3) Improved semantic interoperability ability: The present invention uses semantic network technology to enhance the semantic correlation between metadata in different fields and realizes deep integration at the semantic level. Compared with existing technologies that rely on a single semantic model, the present invention has more advantages in processing heterogeneous data.

[0051] (4) Strong user - friendliness: Provide an intuitive visual management module, enabling non - professional users to easily adjust standardized rules and feedback optimization results. Compared with template tools that require professional knowledge, the present invention reduces the usage threshold and enhances the universality of the tool.

[0052] (5) Modular design, easy to expand: Through a modular architecture, each functional component of the present invention is independent and can be expanded or integrated into an existing data management system according to actual needs. Compared with existing solutions for specific fields, the present invention has a wider application scope in cross - domain and cross - platform data sharing. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a schematic diagram of the composition of the real - time adaptive standardization system for multi - domain data sharing according to the embodiment of the present invention.

[0054] Figure 2 It is a schematic diagram of the metadata information layering module for layering structured, semi - structured, and unstructured metadata information.

[0055] Figure 3 It is a schematic diagram of the intelligent information acquisition module.

[0056] Figure 4 It is a specific architecture diagram of the intelligent information acquisition module.

[0057] Figure 5 It is an architecture diagram of the real - time adaptive standardization module.

[0058] Figure 6 It is a flowchart of the implementation of the real - time adaptive standardization module.

[0059] Figure 7 It is a schematic diagram of the architecture of the visual management module and its interaction with the backend. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] The present invention will be described in detail below according to the accompanying drawings and preferred embodiments. The purpose and effects of the present invention will become more apparent. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0061] As Figure 1 shown, the real - time adaptive standardization system for multi - domain data sharing in this embodiment includes a metadata information layering module, an intelligent information acquisition module, a real - time adaptive standardization module, and a visual management module. The function implementation of each module will be introduced in detail below.

[0062] I. Metadata Information Layering Module

[0063] The metadata information layering module is used to preprocess the metadata from various sources. The preprocessing specifically includes identifying the data sources (such as files, databases, APIs, etc.), and sequentially cleaning, classifying, and layering the metadata of different types and sources to ensure that appropriate measures can be taken for different types of data sources in subsequent processing steps.

[0064] The metadata information layering module first performs preprocessing such as noise cleaning and data cleaning on the raw data collected from various sources. Then, according to the degree of structuring of the metadata information, it can be divided into three categories:

[0065] (1)Structured data with clear field definitions and data types, which is easy to parse;

[0066] (2)Semi-structured data that has a certain structure but is not as strict as structured data;

[0067] (3)Unstructured data that lacks a predefined schema or structure and requires complex analysis techniques to extract valuable information.

[0068] If the metadata is stored in a relational database (RDBMS, Relational Database Management System), or in the format of a CSV file, it is usually structured data; semi-structured data is usually in formats such as JSON, XML, YAML, etc.; unstructured data is usually stored in formats such as text files, images, audio, and video.

[0069] Finally, as Figure 2 shown, the metadata information layering module also layers the structured, semi-structured, and unstructured metadata information respectively:

[0070] (1)Uniformly classify the information recording the basic attributes of the file (including file name, path, extension, size, creation time, and modification time, etc.) into the basic layer. At the same time, based on the above information, verify whether the file is complete or damaged for subsequent tracking management.

[0071] (2)For structured and semi-structured data, parse its internal structure, extract information such as field names, data types, and table structures; identify the data pattern, such as the row and column structures in a table and the key-value pair relationships in JSON. Map the extracted structure information into a unified data model to ensure a consistent representation form. And classify this part of the structure information into the structure layer.

[0072] (3)Identify the unstructured part in the semi-structured data, and this part of the data is uniformly classified into the semantic layer together with the unstructured data.

[0073] II. Intelligent Information Acquisition Module

[0074] The intelligent information collection module supports the unified collection of data at the basic layer, structural layer, and semantic layer. For the metadata information at different layers, different quality control strategies and paths are adopted. The semantic layer information needs to undergo preprocessing such as data cleaning, and then key information feature extraction is performed to form metadata information in a structured form. After completing the above steps, the metadata information, together with the structural layer information and basic layer information, is uniformly integrated and output in the form of a standardized intermediate representation, such as Figure 3 shown.

[0075] such as Figure 4 shown, the intelligent information collection module specifically includes a semantic layer preprocessing sub-module, a key information feature extraction sub-module, and a standardization conversion sub-module.

[0076] (2.1) Semantic layer preprocessing sub-module

[0077] The semantic layer preprocessing sub-module is used to clean the data at the semantic layer to remove noise elements that may interfere with subsequent analysis. This includes but is not limited to deleting irrelevant characters, HTML tags, and special symbols, so as to provide a clean data set for subsequent semantic analysis.

[0078] (2.2) Key information feature extraction sub-module

[0079] The key information feature extraction sub-module combines regular expression matching and natural language processing (NLP) techniques, aiming to optimize the parsing of data content, identify key terms, entities, and the relationships between them, specifically including:

[0080] The key information feature extraction sub-module extracts key information covering important dimensions such as research methods, time range, and geographical location from the data at the semantic layer based on a pre-constructed term vocabulary; for unconventional or unknown entity words not in the term vocabulary, the following two-step strategy is adopted for processing:

[0081] ① Sentence segmentation: Use regular expressions to segment the text according to rules such as punctuation marks and abbreviations to ensure that each sentence is independent and complete.

[0082] ② Entity recognition: Apply a sequence labeling model to extract and identify entities in the text.

[0083] In this embodiment, the key information feature extraction sub-module adopts an advanced sequence labeling model BILSTM-CRF that combines a bidirectional long short-term memory network (BILSTM) and a conditional random field (CRF). Among them, the bidirectional long short-term memory network is used to capture the context information of the text, process the output sequence at one time, and automatically extract lexical features and their context environment without manually defining features. Each vector at each position in the feature vector sequence output by BILSTM is a context-aware feature vector that captures the context information of the current position and its surroundings; CRF is used to receive the feature vector sequence output by BILSTM as input, consider the dependencies between labels in the sequence, convert the labels into a joint probability distribution, and output the optimal label sequence. CRF can adjust the prediction of the current label during the labeling process, taking into account the influence of the previous and subsequent labels, and finally determine the optimal label sequence. Each label can represent information such as entity category and syntactic structure. At each position in the input sequence, a label needs to be assigned to it. The joint distribution function defines the probability of the label at each position. Specifically, the label prediction at each position on the label sequence can be regarded as a random variable, and the label prediction situations at all positions are used as the joint probability function of the joint distribution. In this way, only by modeling and optimizing the joint probability distribution of the label sequence can the optimal prediction result of the entire label sequence be obtained. The advantage of doing this is not only to improve the accuracy of entity recognition but also to enhance the model's ability to understand complex contexts. The output results of the BILSTM-CRF model are as follows:

[0084]

[0085] The key information feature extraction sub-module can effectively realize the automated analysis of semantic layer information, ensure that the data content is fully understood and utilized, and at the same time avoid the problem of repeated expressions.

[0086] (2.3) Standardization conversion sub-module

[0087] The standardization conversion sub-module is used to convert the basic layer information, structure layer information, and processed semantic layer information into a standardized intermediate representation, so as to maintain the consistency of information transmission between different modules and provide a clear and definite data basis for subsequent rule applications. In this embodiment, the standardized intermediate representation is in the form of key-value pairs. In addition to the basic data content, the standardized intermediate representation will carry context information, etc., enabling the dynamic rule library to better understand and process this data.

[0088] III. Real-time adaptive standardization module

[0089] The real-time adaptive standardization module aims to achieve dynamic adaptation and real-time optimization of rules in the metadata standardization process. Through incremental learning, reinforcement learning, and real-time feedback mechanisms, it meets the standardization requirements of multi-domain and multi-type metadata. The real-time adaptive standardization module has a built-in dynamic rule library, which is a collection of updatable rules used to guide the data standardization process. Different from traditional static configuration methods, the dynamic rule library is not fixed but can automatically expand and optimize according to newly encountered data patterns. The dynamic rule library is used to store dynamically generated metadata mapping rules and supports automatic loading, updating, and semantic extension of rules.

[0090] As Figure 5 shown, the real-time adaptive standardization module includes a pattern discrimination sub-module, a rule verification sub-module, a reinforcement learning sub-module, and an automated quality control sub-module.

[0091] (3.1) Pattern Discrimination Sub-module

[0092] As Figure 6 shown, the pattern discrimination sub-module is used to receive the standardized intermediate representation processed by the intelligent information collection module and retrieve the dynamic rule library to determine whether the input data is in a known format. If it is in a known format, the dynamic rule library automatically loads the corresponding rules for mapping. The criteria for judging the known format (meeting one of the following):

[0093] Condition 1: There are rules in the dynamic rule library that are exactly the same as the field semantics, structure, or name.

[0094] Condition 2: It conforms to the predefined standard field name (such as dc.title in Dublin Core).

[0095] When the pattern discrimination sub-module determines that the input data is in an unknown format, an incremental learning algorithm is used to gradually construct new rules. Specifically, based on the context information carried by the standardized intermediate representation, the similarity between the key information and the terms in the metadata standard is evaluated for semantic similarity matching recommendation, which specifically includes the following steps:

[0096] ① Calculate the cosine similarity between two word vectors. This step aims to quantify the direct similarity between keywords and terms:

[0097]

[0098] Among them, a is the extracted keyword, b is the term in the metadata standard, a i represents the word vector corresponding to a, and b i represents the word vector corresponding to b.

[0099] ②Next, to more comprehensively understand the context environment, first calculate the semantic similarity of the vocabulary sets on both sides of the keyword, and then calculate the semantic similarity of keywords a and b. When the calculated semantic similarity is higher than the preset threshold, it is initially determined as a relevant field, and a new mapping rule is formed:

[0100]

[0101]

[0102]

[0103] Among them, represents the vocabulary set on the left side of keyword a, represents the vocabulary set on the right side of keyword a, represents the vocabulary set on the left side of keyword b, represents the vocabulary set on the right side of keyword b, represents the number of words in the vocabulary set, represents the number of words in the vocabulary set, represents and the semantic similarity of, represents and the semantic similarity of, sim(a,b) represents the semantic similarity of keywords a and b; w 1 、w 2 、w 3 are preset weight parameters, w 1 +w 2 +w 3 =1. In this embodiment, w 1 =0.6, w 2 =w 3 =0.2. In this embodiment, the preset threshold is 0.85.

[0104] (3.2) Rule Verification Sub-module

[0105] To further ensure that the new rule meets the requirements for entering the dynamic rule library, the rule verification sub-module uses the built-in test data set for rule verification, including but not limited to aspects such as the integrity, consistency, and accuracy of the test set data mapping results. The rule verification sub-module stores the rules that pass the verification in the dynamic rule library in a standardized format. At the same time, it generates a detailed report of the test results for developers to refer to.

[0106] (3.3) Automatic Quality Control Sub-module

[0107] The automated quality control sub-module conducts quality checks on the mapped standardized metadata from three dimensions: integrity, logic, and consistency. The specific rule definitions are as follows:

[0108] (a)Integrity: Ensure the data in each field of the mapping result is complete, avoiding missing or non-compliant situations.

[0109] ①Field non-empty check: Check whether there are null values in the key fields (such as title, data source, subject area, etc.) in the mapping result.

[0110] ②Field length check: It is agreed that the minimum length threshold for some fields (such as keywords and authors) is not less than 4 characters.

[0111] (b)Logic: Verify whether the logical relationships in the mapping result are reasonable, ensuring that the associations between data meet expectations.

[0112] ①Temporal logic check, such as: Observation time ≤ Data release time ≤ Metadata creation time;

[0113] ②Set a reasonable time range and statistically identify fields that exceed the time threshold;

[0114] ③Standardized term matching: It is agreed that the value range of some metadata fields must come from a controlled vocabulary. For example, language=‘eng’ will be identified and should be replaced with ‘en’.

[0115] (c)Consistency

[0116] ①Dynamic consistency: It is required that the real-time data stream processing latency consistency ≤ 50ms (P99 metric); the tolerance for the difference between batch processing and real-time processing ≤ 0.1%;

[0117] ②Field semantic consistency: It is required that the value of the mapped meta-field is consistent with the original semantics.

[0118] (3.4)Reinforcement learning sub-module

[0119] The reinforcement learning sub-module selects the Apache Flink framework that supports real-time data stream processing. The reinforcement learning sub-module is used to collect user feedback after each data conversion task is completed, and perform real-time optimization on the rules. The reinforcement learning sub-module records each successful data conversion as a reference, continuously optimizes the processing logic of the dynamic rule library itself, and reorders the priorities of the rules. With the accumulation of converted data, the dynamic rule library will not only be able to better process the current type of data, but also accumulate experience in processing other potential data types.

[0120] The reinforcement learning sub-module adopts stream processing technology, which can efficiently process large-scale and real-time data streams and significantly reduce latency. Compared with existing semantic mapping tools and automated matching tools, the present invention has obvious advantages in processing performance, with high real-time processing efficiency and is suitable for scenarios with high real-time requirements.

[0121] IV. Visualization Management Module

[0122] The visualization management module realizes the management, analysis and dynamic adjustment of metadata through an intuitive interface. This module is designed based on the architecture of separating the front end from the back end. As Figure 7 shown, the visualization management module includes a metadata browsing and retrieval sub-module, a dynamic rule configuration sub-module, a real-time monitoring and warning sub-module, and an interactive learning sub-module.

[0123] The metadata browsing and retrieval sub-module is used for the visual browsing and searching of metadata structures, contents and semantics. This sub-module displays the hierarchical relationship of metadata through a tree structure and supports highlighting search results, fuzzy matching and advanced queries.

[0124] The dynamic rule configuration sub-module supports users to dynamically create, modify and delete metadata mapping rules through a visual interface. The front-end component adopts a drag-and-drop interface and supports modular grouping of rule conditions; the back-end engine is based on Drools to implement the logical processing and execution of rules. The front-end component receives the source and target mapping relationships of fields selected by the user in the interface, provides a syntax checking function to prevent rule conflicts, and dynamically displays the execution flow of rules in the form of a flowchart.

[0125] The real-time monitoring and warning sub-module provides a real-time monitoring view of metadata processing, including data streams, rule execution status, anomaly detection, etc. Use Echarts to draw a data flow diagram to show the distribution of metadata in different processing stages. Realize real-time monitoring and threshold warning based on the Prometheus tool.

[0126] The interactive learning sub-module supports users to check and repair the generated metadata. The front end supports operations such as batch editing of fields and data completion. At the same time, users can view the quality control reports generated by the automated quality control sub-module, which contain detailed descriptions of the problem sources. The interactive learning sub-module receives the actions and feedback repaired by the user and submits them to the reinforcement learning sub-module.

[0127] Those of ordinary skill in the art can understand that the above are only preferred examples of the invention and are not used to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art can still modify the technical solutions described in the foregoing examples or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, etc. made within the spirit and principles of the invention shall be included within the protection scope of the invention.

Claims

1. A metadata real-time adaptive standardization system for multi-domain data sharing, characterized in that: It includes metadata information stratification module, intelligent information collection module, real-time adaptive standardization module and visual management module; The metadata information stratification module is used to pre-process the metadata input from various sources, including identifying the data source, and sequentially cleaning, classifying and stratifying metadata of different types and sources; The intelligent information collection module supports the unified collection of metadata information, and adopts different quality control strategies and paths for metadata information at different layers, and finally converts the collected data into a standardized intermediate representation; The real-time adaptive standardization module is used to realize dynamic adaptation and real-time optimization of rules in the metadata standardization process. It has a built-in dynamic rule library for storing dynamically generated metadata mapping rules and supports automatic loading, updating and semantic extension of rules. The visual management module implements management, analysis and dynamic adjustment of metadata through an intuitive interface.

2. The metadata real-time adaptive standardization system for multi-domain data sharing according to claim 1, characterized in that: The metadata information layering module divides metadata of different types and sources into the following three categories: (1) Structured data with clear field definitions and data types that are easy to parse; (2) Semi-structured data, which has a certain structure but is not as strict as structured data; (3) Unstructured data that lacks predefined patterns or structures and requires complex analytical techniques to extract valuable information.

3. The metadata real-time adaptive standardization system for multi-domain data sharing according to claim 2, characterized in that: The metadata information stratification module stratifies structured, semi-structured and unstructured data as follows: The information of the basic attributes of the recorded files is uniformly classified into the basic layer, and based on the above information, it is verified whether the files are complete or damaged for subsequent tracking and management; Parse the internal structure of structured and semi-structured data, extract structural information, and map the extracted structural information into a unified data model to ensure consistent representation; And these structural information are classified into the structural layer; Identify the unstructured part of the semi-structured data, and classify this part of data and the unstructured data into the semantic layer.

4. The metadata real-time adaptive standardization system for multi-domain data sharing according to claim 3, characterized in that: The intelligent information collection module includes a semantic layer preprocessing submodule, a key information feature extraction submodule and a standardization conversion submodule; wherein, The semantic layer preprocessing submodule is used to clean the data of the semantic layer to remove noise elements that may interfere with subsequent analysis; The key information feature extraction submodule combines regular matching and natural language processing technology to optimize the analysis of data content and identify key terms, entities and the relationships between them; The standardization conversion submodule is used to convert the basic layer information, the structural layer information and the processed semantic layer information into a standardized intermediate representation.

5. The metadata real-time adaptive standardization system for multi-domain data sharing according to claim 4, characterized in that: The standardized intermediate representation is in the form of a key-value pair, which includes basic data content and carries context information, so that the dynamic rule base can better understand and process the data.

6. The metadata real-time adaptive standardization system for multi-domain data sharing according to claim 4, characterized in that: The key information feature extraction submodule combines regular matching and natural language processing technology to optimize the analysis of data content and identify key terms, entities and the relationships between them, including: The key information feature extraction submodule extracts key information covering important dimensions such as research methods, time range, and geographic location from the semantic layer data based on a pre-built terminology vocabulary; for unconventional or unknown entity words that are not in the terminology vocabulary, the following two-step strategy is used for processing: ① Sentence segmentation: Use regular expressions to segment text according to rules including punctuation marks and abbreviations to ensure that each sentence is independent and complete; ②Entity recognition: Apply sequence labeling models to extract and identify entities in text.

7. The metadata real-time adaptive standardization system for multi-domain data sharing according to claim 6, characterized in that: The key information feature extraction submodule adopts an advanced sequence labeling model that combines a bidirectional long short-term memory network and a conditional random field, wherein the bidirectional long short-term memory network is used to capture the context information of the text, process the output sequence at one time, and automatically extract vocabulary features and their context environment; the output of the bidirectional long short-term memory network is a feature vector sequence, and the vector at each position of the feature vector sequence is a context-aware feature vector that captures the context information of the current position and its surroundings; The conditional random field receives the feature vector sequence as input, considers the dependencies between labels in the sequence, converts the labels into joint probability distribution, and outputs an optimal label sequence.

8. The metadata real-time adaptive standardization system for multi-domain data sharing according to claim 1, characterized in that: The real-time adaptive standardization module includes a pattern discrimination submodule, a rule verification submodule, a reinforcement learning submodule and an automated quality control submodule; The pattern discrimination submodule is used to receive the standardized intermediate representation obtained by the intelligent information acquisition module, and retrieve the dynamic rule library to determine whether the input data is in a known format. If it is in a known format, the dynamic rule library automatically loads the corresponding rules for mapping; if it is in an unknown format, an incremental learning algorithm is used to gradually construct new rules; The rule verification submodule uses the built-in test data set to verify the constructed new rules, including using the test set data to verify the completeness, consistency and accuracy of the mapping results, and stores the verified rules in a standardized format in the dynamic database; at the same time, the test results are generated into a detailed report for developers' reference; The reinforcement learning submodule is used to collect user feedback and optimize the rules in real time after each data conversion task is completed; The automated quality control submodule performs quality checks on the mapped standardized metadata from three dimensions: completeness, logic, and consistency.

9. The metadata real-time adaptive standardization system for multi-domain data sharing according to claim 8, characterized in that: The reinforcement learning submodule records each successful data conversion as a reference, continuously optimizes the processing logic of the dynamic rule base itself, and reorders the priority of the rules.

10. The metadata real-time adaptive standardization system for multi-domain data sharing according to claim 8, characterized in that: The visual management module includes a metadata browsing and retrieval submodule, a dynamic rule configuration submodule, a real-time monitoring and alarm submodule, and an interactive learning submodule; The metadata browsing and retrieval submodule is used for visual browsing and searching of metadata structure, content and semantics; The dynamic rule configuration submodule user dynamically creates, modifies and deletes metadata mapping rules through a visual interface; The real-time monitoring and alarm submodule provides a real-time monitoring view of metadata processing, including data flow, rule execution status, and anomaly detection; The interactive learning submodule supports users to check and repair the generated metadata, and the front end supports batch editing of fields and data completion operations; at the same time, a quality control report is generated and provided to the user, and the received user repair actions and feedback are submitted to the reinforcement learning submodule.

Citation Information

Patent Citations

  • Data-source-irrelevant data full-life-cycle management platform and method

    CN108717456A

  • Multi-source heterogeneous ecological environment big data processing method and system based on data lake

    CN111459908A

  • Multi-platform metadata standardization processing method and device

    CN117851387A

  • Data security enhanced metadata integration method of data management and control platform

    CN118673194A

  • Adaptive metadata acquisition and change tracking system

    CN118689863A

Cited By

  • Energy efficiency analysis method and device based on multi-source data

    CN120355532A

  • Energy efficiency analysis method and device based on multi-source data

    CN120355532B

  • Exhibition information synchronization and real-time updating management method and platform based on cloud computing

    CN120670439A

  • Generation method of standardized detection rule base of fully mechanized coal mining face and related device

    CN120743921A

  • Unified data metadata conversion system

    CN121078139A