Data component encapsulation construction method based on standardized description model
By building a data component encapsulation method based on a standardized description model, the problem of tight coupling in data management is solved, seamless interoperability and efficient flow of data between different platforms are achieved, and the data reuse efficiency and governance flexibility are improved.
Patent Information
- Application Number
- CN202510736860.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-16
AI Technical Summary
In existing data management methods, data is deeply bound to specific systems and application processes, resulting in diverse structural formats, difficulty in unifying semantics, and inconsistent interface standards, making it difficult for data to be effectively circulated and used collaboratively between different entities and platforms.
Build a data component encapsulation method based on a standardized description model, use data component construction tools to perform structured encapsulation of raw data, give it a unique identifier, metadata description and access interface, introduce interoperability protocols such as DOIP, achieve seamless interoperability of data between different platforms, and realize dynamic addressing and calling through the identity resolution system.
It improves the addressability, operability and reusability of data, supports efficient circulation and governance between cross-domain systems, and promotes the high-quality development of the data element market.
Smart Images

Figure CN120653631A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to technologies such as artificial intelligence, blockchain, data element governance, and metadata management. Specifically, it relates to a multi-source heterogeneous data circulation and intelligent processing method based on standardized data components and dynamic assembly. Background Art
[0002] In the intelligent and digital age, data has been endowed with two new attributes: resource element and value processing. Data, similar to traditional production factors, involves processes such as acquisition, storage, circulation, transaction, security, and assetization. Data requires computation, analysis, mining, and modeling to unlock its value. Data infrastructure primarily addresses the large-scale aggregation and flow of data, including the construction of data hubs and data circulation and transaction facilities. Data processing technology has been continuously developing and evolving, from early database management to information retrieval, to data analysis and business intelligence, and now to deep data processing and data intelligence. The continuous "decoupling" of data is the primary driver of this evolution, bringing with it changes in the fundamental abstraction of data. The first decoupling involved decoupling data from applications, aiming to mask the complexity of data access and lower the barrier to entry for application system development. The fundamental abstraction for data is the entity-relationship model (ER model), with databases and data warehouses as core systems. The second decoupling is the decoupling of data from business systems. Its goal is to mitigate the complexity of data aggregation and analysis and lower the barrier to entry for enterprise-level system development. The basic abstraction for data is a key-value model, and the core system includes a data lake. The third decoupling is the decoupling of data producers and consumers. Its goal is to mitigate the complexity of data flow and use and lower the barrier to entry for the social supply, circulation, and application of data elements.
[0003] The transformation of data into essential elements has driven the rapid development of the data industry, giving rise to emerging sectors such as data collection, data governance, and data quality management. However, existing data management methods suffer from severe coupling issues. Data is highly dependent on specific systems and applications, resulting in heterogeneous structures, inconsistent semantics, and restricted access. This hinders the circulation, sharing, and integration of data, and severely restricts the availability and computing efficiency of data resources.
[0004] In response to the above problems, this patent, facing the third decoupling trend of data processing and utilization, proposes a data component encapsulation construction method based on a standardized description model. As the basic unit for the circulation and use of data elements, data components are standardized encapsulation entities constructed under the concept of data decoupling. Its core idea is to liberate data from the tightly coupled state of dependence on specific applications, platforms and tasks, and to make it independent, portable and composable by abstracting a unified description structure and behavioral interface. Data components are designed based on unified standards, protocols and mechanisms, and have the characteristics of addressability, exchangeability, operability and controllability, supporting efficient discovery, transparent access and trusted sharing of data among different subjects and different systems.
[0005] Each data component contains a globally unique identifier, structured metadata, a self-describing interface, and data routing information, supporting the encapsulation of structured, semi-structured, and unstructured data. Through an identity resolution mechanism, logical identifiers can be mapped to actual storage and service locations, enabling cross-platform and cross-system data location and access. Through unified interoperability protocols such as DOIP, transparent flow across multiple environments (on-premises, cloud, and edge) is guaranteed.
[0006] The construction process of data components relies on a complete set of functional systems. Its basic functions are mainly to standardize the semantics, structure and basic operations of heterogeneous multi-source data. The operational functions involve the interactive interfaces provided by data components to external applications or other components. These interfaces support multiple operation modes such as query, call, update, and combination, giving data components flexible operational capabilities in cross-system and cross-domain applications. Through standardized APIs and metadata descriptions, data components can interoperate with different platforms, application systems or other components to ensure the effective flow and use of data in different scenarios. Extended functions can cover the embedding of data processing logic, the integration of quality verification mechanisms, access permission control, etc.; derivative functions support the dynamic assembly, on-demand reuse, cross-domain circulation and compliance supervision of components. This multi-level functional system supports the dynamic loading, combination and reuse of components, can dispatch data resources on demand, and supports the integrated utilization of multi-source heterogeneous data.
[0007] Through standardized encapsulation and loosely coupled deployment, data components achieve decoupling of data and subjects, data and applications, structure and semantics, effectively improving data reuse efficiency and governance flexibility, promoting the transformation of data from static resources to dynamic services, building infrastructure capabilities to support trusted circulation and market-based transactions, enabling intelligent data applications in multiple fields such as government affairs, medical care, and finance, and promoting high-quality development of the data factor market. Summary of the Invention
[0008] The technical problem solved by the present invention is: to address the common tight coupling problem in existing data management methods, that is, data is deeply bound to specific systems and application processes, resulting in diverse structural formats, difficult to unify semantics, inconsistent interface standards, and difficulty in achieving effective circulation and collaborative use of data between different subjects and platforms. In essence, data still exists in the form of "appendages" and lacks abstract and decoupled expression as an independent element. This limitation not only restricts the social supply and reuse capabilities of data, but also makes it difficult to embed it into the cross-domain fusion process dominated by information flow, and it cannot be deeply integrated and interoperable with other digital elements such as computing, algorithms, and models. To this end, a standardized encapsulation method for data element governance needs is proposed, namely a data component. The data component achieves decoupling of data at the description, storage, and interaction levels by encapsulating the original data into an independent data unit with a unique identifier, metadata description, access interface, and routing information.
[0009] The technical solution of this invention is as follows: This research develops a data component encapsulation and construction method based on a standardized description model. This method first receives raw data through a data component construction tool and structurally encapsulates the data resources according to a standardized description model (covering four dimensions: identification, metadata, data routing, and interface). This method rapidly constructs data components with unique identification and operational capabilities. The encapsulated data components are then connected to a data component repository system for logical storage and service-oriented deployment. Simultaneously, the system binds the component's identity to its access point and metadata description and submits it to an identity resolution system through a registration service. The identity resolution system uses a resolution mechanism similar to the Handle system to establish a mapping relationship between the identity and the resource location, supporting the unique positioning and dynamic addressing of components in cyberspace. To ensure seamless interoperability between data components across different platforms and systems, this method introduces an interoperability protocol (such as DOIP) to standardize the encapsulation of component interfaces, supporting the full lifecycle interaction process, including read, query, and execute operations. Furthermore, an identity resolution protocol (such as DOIRP) is responsible for resolving and mapping the identity to specific component resources, enabling transparent and accurate access to components in a distributed environment. This method achieves a standardized processing path from data encapsulation, identification management to access and call through a unified data modeling method, registration tools, identification mechanism and interface protocol, significantly improving the addressability, operability and reusability of data elements across cross-domain systems, and providing key support for building a loosely coupled, highly autonomous data circulation and governance system. The specific steps are as follows: (1) Perform basic data processing on the original data by parsing, cleaning, filtering, completing, and deduplicating. The steps include: a. Data parsing. Read raw data from various data sources (such as CSV, Excel, JSON, databases, APIs, etc.) and convert it into a processable format. Understand the data's hierarchical structure, identifying data fields, data types, nested relationships (such as the nested structure of JSON), and associations between tables for subsequent processing. Convert data to appropriate formats, such as converting strings to dates, numbers to floating-point numbers, and parsing Booleans, to ensure consistent data formats. Check the data for format mismatches, missing values, duplicate rows, and other issues to provide a basis for data cleaning.
[0010] b. Data cleaning. Data cleaning is a key step in improving data quality, including addressing missing values, standardizing data formats, correcting outliers, and removing duplicate data. Missing values can be filled in through deletion, filling, or model prediction to ensure data integrity. Data formats from different sources must be standardized, such as date formats and numeric type conversions. Outliers must be detected and corrected to eliminate unreasonable data, such as ages or prices outside the normal range. Furthermore, deduplication can avoid data redundancy and ensure data accuracy. Data cleaning can improve the reliability of data analysis and provide a solid foundation for subsequent modeling and decision-making.
[0011] c. Data deduplication. This removes duplicate records from a dataset to ensure data uniqueness and accuracy. Deduplication can be performed based on uniqueness rules for a single column or multiple columns. Common methods include complete deduplication, deduplication based on specific fields, and fuzzy deduplication.
[0012] (2) Construct a standardized description model for data components. The data component construction tool enables the rapid construction and assembly of data components. The specific steps include: a. The standardized description model for data components includes four core dimensions: identification, description, data routing, and interface. It combines key features such as assemblability, measurability, and controllability to create a unified, operational, and scalable standardized description framework. The "identification" dimension is used to assign a globally unique semantic identifier to each data component. This identifier is stable and resolvable, supporting data addressing, identity resolution, and lifecycle management in cross-domain environments, ensuring the locatability and uniqueness of components during circulation, invocation, and traceability. The "description" dimension is used to define the metadata attributes of data components, including their data source, type structure, generation time, ownership information, version number, access rights, and other content. It has self-describing capabilities and supports component discovery, indexing, classification, and evolution management. The "data routing" dimension is used to configure the storage location, access path, scheduling rules, and security policies of data components, supporting dynamic addressing and strategic transmission in distributed systems, improving the access efficiency and flexibility of data components across network environments. The "interface" dimension defines the set of standardized operations that can be performed on data components, including but not limited to query, update, execute, delete, and other functional operations. The interface encapsulates operational semantics through unified protocol specifications (such as REST API, DOIP), ensuring that different systems or platforms have consistent interaction capabilities with components. Figure 2 Standardized description model for data components and Figure 3 The data component structure is shown in Figure 1.
[0013] (3) Design and implement a data component-based registration, publishing, and discovery mechanism to support the registration and publishing of data components. The data component registration and publishing mechanism aims to provide a unified management method for registering, storing, querying, publishing, and subscribing to data components to improve the reusability, manageability, and sharing of data. It is usually composed of a registration center, a storage system, and a publishing mechanism. Typical application scenarios include data governance, data sharing platforms, and data management in microservice architectures. The specific steps include: a. Component Registration Submission: Component developers submit component data information through the registration service, including the component name, data format, metadata description, interface definition, and ownership information. Each component is assigned a globally unique identifier (e.g., based on the Handle / DOI format) upon registration, ensuring addressability and uniqueness during distribution and invocation. The registration service is built on the Flask framework, storing component metadata and registration information in a structured database system (e.g., MySQL or PostgreSQL), providing efficient data read / write support and transaction consistency.
[0014] b. Information storage and unified management: The physical data of data components can be stored in distributed file systems (such as HDFS), object storage systems (such as MinIO), or integrated API management platforms (such as APISIX and Kong). The registration system establishes a mapping relationship between the data component identifier and its storage location and interface path to achieve decoupled management of logic and physical, supporting subsequent identity resolution and dynamic access.
[0015] c. Registration, Release, and Discovery Mechanism: The registration system integrates a semantic search mechanism, supporting data component discovery based on multiple criteria, including keywords, tags, semantic entities, and metadata fields. The system utilizes a composite search architecture based on an inverted index and a graph database (knowledge graph), enabling users or service systems to quickly locate target components by name, version, and functional type, improving component reuse and search precision.
[0016] d. Registration, Release, and Service Interfaces: After a component is successfully registered and passes review, it is centrally released by the registration system. External systems can access component data and metadata through a variety of standardized interfaces, including RESTful APIs, GraphQL, message queues, and asynchronous file-based sharing. All external interfaces are centrally mounted under the service gateway, and access control and traceability are implemented through permission authentication and call logging.
[0017] (4) Construct a data component protocol cluster. The data component interconnection protocol realizes the transparency and standardization of data component addressing and transmission. The specific steps include: a. Data Usage (Application Interface): A standardized interface model based on the DOIP protocol or RESTful API is used to achieve unified encapsulation and service-oriented invocation of data components. The interface definition supports multiple operations, including read, query, update, and execute, offering excellent compatibility and scalability, significantly improving the compatibility and reuse efficiency of data components across heterogeneous systems.
[0018] b. Data location (naming protocol). Globally unique identification systems such as Handle / DOI are introduced to assign semantic identifiers to each data component, ensuring uniqueness and resolvability during cross-domain circulation. Identification supports multi-dimensional semantic mapping, facilitating the organization, search, and association of data components by industry, space, time, and other dimensions.
[0019] c. Data resolution (routing resolution). This mechanism utilizes a decentralized semantic routing (DSR) mechanism, implementing distributed, hierarchical path mapping through identity resolution services. This combines batch address resolution with a location-aware proximity addressing strategy to improve routing efficiency and system fault tolerance, optimizing the dynamic acquisition path for data components.
[0020] d. Data Acquisition (Data Exchange). Based on the on-demand access mechanism supported by the DOIP protocol, an asynchronous incremental transmission (AIT) capability is designed to enable segmented retrieval of data components, simultaneous processing and transmission, and status update notifications, reducing data transmission latency and improving interactive performance and real-time performance.
[0021] (5) Build a dynamic integration and call system based on the data component identification resolution mechanism. This system aims to use the globally unique identifier (Handle format) to achieve precise positioning, permission control and automatic assembly of data components in a distributed environment, and build a component-level data service capability that is programmable, reusable and traceable. Its core is to support dynamic component discovery, dependency assembly and call path mapping through identification resolution. The specific steps include: a. Each data component is assigned a semantic, globally unique identifier upon registration (e.g., a "prefix / suffix" structure based on the Handle system). This identifier is mapped to the component's metadata, interface path, and storage address. The identity resolution system converts the identifier into the actual service entry address and access metadata upon invocation.
[0022] b. Components declare their dependencies in metadata (including dependent component IDs, version constraints, required interface types, etc.). When called, the identity resolution system automatically resolves and reverse-resolves the dependent component's identity. Based on the resolution results, the system calls the registry and repository system to retrieve dependent components on demand, enabling dynamic injection and loading of component-level resources.
[0023] c. The parsing system has a built-in double-layer mapping structure and access permission verification module, which supports real-time judgment and access authorization of conditions such as the caller's identity, permission policy, and call frequency, ensuring data security and behavior controllability of components during distributed operation.
[0024] d. Leveraging the identity system and integrating with workflow engines (e.g., embeddable Camunda and Apache Airflow) to achieve logical orchestration of data components. Describing the dependency structure and execution order between components using formats such as JSON or BPMN, the system dynamically routes data flows and schedules component services based on parsed paths.
[0025] (6) Provide data applications and data services based on the basic mechanism of data components to support the circulation, sharing and utilization of data.
[0026] The advantages of the present invention compared with the prior art are: 1. Decoupling data from applications improves flexibility. Traditional data management (such as databases and data warehouses) often tightly binds data to application logic, resulting in strong data dependencies. Adjusting data structure or logic requires modifying application code. Data components encapsulate data and its processing logic (such as cleansing and analysis) as independent units, decoupling them from applications. Through standardized interfaces, components can be seamlessly replaced or upgraded without rebuilding the entire system.
[0027] 2. Standardized interfaces and descriptions enhance data reusability. Traditional data sharing relies on customized interfaces or file transfers (such as CSV / Excel), resulting in high integration costs due to heterogeneous interfaces. Standardized interfaces such as RESTful APIs and GraphQL reduce integration complexity. Descriptions based on knowledge graphs or ontologies improve data understandability and discovery efficiency. 3. Promote data circulation, sharing, and utilization. Data silos are a serious problem, and cross-organizational data sharing relies on cumbersome protocols and manual operations. The data component registration center supports global search and on-demand access.
[0028] 4. Dynamic integration and complex process orchestration. Traditional ETL tools or scripts lack flexibility and are unable to quickly respond to business changes. Define component execution order through visualization or code. Dynamically discover and load dependent components, reducing manual configuration.
[0029] In the process of building a government data sharing platform, there are problems such as difficulty in intercommunication of cross-departmental data, inconsistent data item definitions, and unstable interface calls. The data component encapsulation construction method proposed in the present invention is used to uniformly model and encapsulate government data (such as legal person information, credit records, administrative licensing matters, etc.) selected from the business databases of various commissions and bureaus. By building a unified data component identification system and interface specifications, the platform realizes the registration, release and on-demand calling of data components. For example, the enterprise credit information component has a unique identifier, metadata description, RESTful interface and routing information, and can be dynamically referenced by multiple systems such as social security, taxation, industry and commerce, etc. Compared with traditional methods, this method improves the efficiency of data reuse, supports on-demand retrieval and asynchronous access across systems, significantly shortens the system integration cycle, and reduces maintenance costs.
[0030] The regional medical collaboration platform faces the problems of heterogeneous data structures and inconsistent analysis processes among hospitals. The platform introduces the data component construction mechanism of the present invention, which encapsulates basic patient information, test records, imaging data, etc. into data components with clear structures and consistent interfaces.
[0031] Various intelligent analysis algorithms (such as chronic disease prediction and medication risk analysis) are connected as independent services. Through the identification resolution mechanism, the required data components can be automatically assembled to realize the analysis process of "on-demand retrieval and nearest routing". The system loads the patient's basic information and blood sugar test components in the past six months on demand, and combines them with risk analysis model components to realize the automatic orchestration of the entire predictive analysis process. The prediction accuracy is improved by 12.4% compared with the traditional ETL method, and the system response speed is improved by about 30%. The invention of data components and their mechanisms has significantly improved the efficiency and flexibility of data management through standardization, modularization, automation and security enhancement, providing a better technical path for data factorization, intelligent application and digital transformation. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 It is a flow chart of the data component encapsulation structure of the present invention.
[0032] Figure 2 It is a standardized description model diagram of data components.
[0033] Figure 3 It is a data component structure diagram.
[0034] Figure 4 It is a data component protocol cluster diagram. DETAILED DESCRIPTION
[0035] The present invention comprises the following steps: (1) Perform basic data processing on the original data by parsing, cleaning, filtering, completing, and deduplicating. The steps include: a. Data parsing. Read raw data from various data sources (such as CSV, Excel, JSON, databases, APIs, etc.) and convert it into a processable format. Understand the data's hierarchical structure, identifying data fields, data types, nested relationships (such as the nested structure of JSON), and associations between tables for subsequent processing. Convert data to appropriate formats, such as converting strings to dates, numbers to floating-point numbers, and parsing Booleans, to ensure consistent data formats. Check the data for format mismatches, missing values, duplicate rows, and other issues to provide a basis for data cleaning.
[0036] b. Data cleaning. Data cleaning is a key step in improving data quality, including addressing missing values, standardizing data formats, correcting outliers, and removing duplicate data. Missing values can be filled in through deletion, filling, or model prediction to ensure data integrity. Data formats from different sources must be standardized, such as date formats and numeric type conversions. Outliers must be detected and corrected to eliminate unreasonable data, such as ages or prices outside the normal range. Furthermore, deduplication can avoid data redundancy and ensure data accuracy. Data cleaning can improve the reliability of data analysis and provide a solid foundation for subsequent modeling and decision-making.
[0037] c. Data deduplication. This removes duplicate records from a dataset to ensure data uniqueness and accuracy. Deduplication can be performed based on uniqueness rules for a single column or multiple columns. Common methods include complete deduplication, deduplication based on specific fields, and fuzzy deduplication.
[0038] (2) Construct a standardized description model for data components. The data component construction tool enables the rapid construction and assembly of data components. The specific steps include: The standardized description model for data components includes four core dimensions: identification, description, data routing, and interface. It combines key features such as assemblability, measurability, and controllability to build a unified, operational, and scalable standardized description framework. The "identification" dimension is used to assign a globally unique semantic identifier to each data component. The identifier is stable and resolvable, supporting data addressing, identity resolution, and lifecycle management in cross-domain environments, ensuring the locatability and uniqueness of components during circulation, invocation, and traceability. The "description" dimension is used to define the metadata attributes of data components, including their data source, type structure, generation time, ownership information, version number, access rights, and other content. It has self-describing capabilities and supports component discovery, indexing, classification, and evolution management. The "data routing" dimension is used to configure the storage location, access path, scheduling rules, and security policies of data components, supporting dynamic addressing and strategic transmission in distributed systems, and improving the access efficiency and flexibility of data components in cross-network environments. The "interface" dimension defines the set of standardized operations that can be performed on data components, including but not limited to query, update, execute, delete, and other functional operations. The interface encapsulates operational semantics through unified protocol specifications (such as REST API, DOIP), ensuring that different systems or platforms have consistent interaction capabilities with components.
[0039] (3) Design and implement a data component-based registration, publishing, and discovery mechanism to support the registration and publishing of data components. The data component registration and publishing mechanism aims to provide a unified management method for registering, storing, querying, publishing, and subscribing to data components to improve the reusability, manageability, and sharing of data. It is usually composed of a registration center, a storage system, and a publishing mechanism. Typical application scenarios include data governance, data sharing platforms, and data management in microservice architectures. The specific steps include: a. Component Registration Submission: Component developers submit component data information through the registration service, including the component name, data format, metadata description, interface definition, and ownership information. Each component is assigned a globally unique identifier (e.g., based on the Handle / DOI format) upon registration, ensuring addressability and uniqueness during distribution and invocation. The registration service is built on the Flask framework, storing component metadata and registration information in a structured database system (e.g., MySQL or PostgreSQL), providing efficient data read / write support and transaction consistency.
[0040] b. Information storage and unified management: The physical data of data components can be stored in distributed file systems (such as HDFS), object storage systems (such as MinIO), or integrated API management platforms (such as APISIX and Kong). The registration system establishes a mapping relationship between the data component identifier and its storage location and interface path to achieve decoupled management of logic and physical, supporting subsequent identity resolution and dynamic access.
[0041] c. Registration, Release, and Discovery Mechanism: The registration system integrates a semantic search mechanism, supporting data component discovery based on multiple criteria, including keywords, tags, semantic entities, and metadata fields. The system utilizes a composite search architecture based on an inverted index and a graph database (knowledge graph), enabling users or service systems to quickly locate target components by name, version, and functional type, improving component reuse and search precision.
[0042] d. Registration, Release, and Service Interfaces: After a component is successfully registered and passes review, it is centrally released by the registration system. External systems can access component data and metadata through a variety of standardized interfaces, including RESTful APIs, GraphQL, message queues, and asynchronous file-based sharing. All external interfaces are centrally mounted under the service gateway, and access control and traceability are implemented through permission authentication and call logging.
[0043] (4) Construct a data component protocol cluster. The data component interconnection protocol realizes the transparency and standardization of data component addressing and transmission. The specific steps include: a. Data Usage (Application Interface): A standardized interface model based on the DOIP protocol or RESTful API is used to achieve unified encapsulation and service-oriented invocation of data components. The interface definition supports multiple operations, including read, query, update, and execute, offering excellent compatibility and scalability, significantly improving the compatibility and reuse efficiency of data components across heterogeneous systems.
[0044] b. Data location (naming protocol). Globally unique identification systems such as Handle / DOI are introduced to assign semantic identifiers to each data component, ensuring uniqueness and resolvability during cross-domain circulation. Identification supports multi-dimensional semantic mapping, facilitating the organization, search, and association of data components by industry, space, time, and other dimensions.
[0045] c. Data resolution (routing resolution). This mechanism utilizes a decentralized semantic routing (DSR) mechanism, implementing distributed, hierarchical path mapping through identity resolution services. This combines batch address resolution with a location-aware proximity addressing strategy to improve routing efficiency and system fault tolerance, optimizing the dynamic acquisition path for data components.
[0046] d. Data Acquisition (Data Exchange). Based on the on-demand access mechanism supported by the DOIP protocol, an asynchronous incremental transmission (AIT) capability is designed to enable segmented retrieval of data components, simultaneous processing and transmission, and status update notifications, reducing data transmission latency and improving interactive performance and real-time performance.
[0047] (5) Build a dynamic integration and call system based on the data component identification resolution mechanism. This system aims to use the globally unique identifier (Handle format) to achieve precise positioning, permission control and automatic assembly of data components in a distributed environment, and build a component-level data service capability that is programmable, reusable and traceable. Its core is to support dynamic component discovery, dependency assembly and call path mapping through identification resolution. The specific steps include: a. Each data component is assigned a semantic, globally unique identifier upon registration (e.g., a "prefix / suffix" structure based on the Handle system). This identifier is mapped to the component's metadata, interface path, and storage address. The identity resolution system converts the identifier into the actual service entry address and access metadata upon invocation.
[0048] b. Components declare their dependencies in metadata (including dependent component IDs, version constraints, required interface types, etc.). When called, the identity resolution system automatically resolves and reverse-resolves the dependent component's identity. Based on the resolution results, the system calls the registry and repository system to retrieve dependent components on demand, enabling dynamic injection and loading of component-level resources.
[0049] c. The parsing system has a built-in double-layer mapping structure and access permission verification module, which supports real-time judgment and access authorization of conditions such as the caller's identity, permission policy, and call frequency, ensuring data security and behavior controllability of components during distributed operation.
[0050] d. Leveraging the identity system and integrating with workflow engines (e.g., embeddable Camunda and Apache Airflow) to achieve logical orchestration of data components. Describing the dependency structure and execution order between components using formats such as JSON or BPMN, the system dynamically routes data flows and schedules component services based on parsed paths.
[0051] (6) Provide data applications and data services based on the basic mechanism of data components to support the circulation, sharing and utilization of data.
Claims
1. A data component encapsulation construction method based on a standardized description model, characterized in that Here are the steps: (1) The original data is input into the data component encapsulation construction system. After pre-processing such as parsing, cleaning, and standardization, it is encapsulated according to the standardized description model of the data component to generate a data component with unique identification, self-description of semantics, and unified structure; The standardized description model includes identification, metadata, data routing, and interfaces; (2) Data component registration, metadata index generation and permission control are realized through the data component registration and release mechanism. Handle or DID identification system is used to complete identification registration, and data components are registered to the warehouse system based on the DOIP protocol to achieve unified addressing and access; (3) Build a data component identification resolution mechanism to map the unique identifier of each component with its actual storage location, access interface, and metadata information in the warehouse system, support a multi-level resolution architecture and permission control strategy, and ensure that the identifier is traceable, verifiable, and operable; (4) Connect data components to the data interconnection protocol system and build a protocol cluster based on unified addressing rules and transmission protocols to support standardized interaction and transparent calls of data components between different platforms and systems; (5) Based on the interoperability mechanism of data components, provide standardized data service interfaces to support cross-domain access, sharing, combination and trusted use of data components; (6) Provide data applications and data services based on the basic mechanism of data components to support the circulation, sharing and utilization of data.
2. The data component encapsulation construction method based on the standardized description model according to claim 1 is characterized in that: In step (1), the construction of the data component includes: (1) Analyze, clean, complete, and remove duplicates from the original data to form a standardized structural input; (2) Design a standardized data model for data components, support the encapsulation of multiple data types including structured, semi-structured and unstructured data, and use the lightweight data format JSON to encapsulate data.
3. The data component encapsulation construction method based on the standardized description model according to claim 1 is characterized in that: The registration, publishing and discovery mechanism in step (2) includes: Assign a globally unique identifier to each data component, use a decentralized identification system or Handle system to generate the ID, and bind the component's metadata, access interface, and permission information; (2) Provide RESTful API to support uploading, registering, updating and deleting data components, and perform structured indexing of metadata; (3) Construct a data component indexing mechanism that supports semantic retrieval, combining tags, keywords, and entity relationships to achieve intelligent discovery. Furthermore, the system introduces a semantically enhanced retrieval method that integrates ontology structure and domain knowledge graph to achieve conceptual hierarchical understanding and fuzzy matching capabilities. Component registration information supports multimodal semantic tags, improving the semantic coverage and context matching of the retrieval process. At the same time, using a graph-based indexing algorithm, it supports component aggregation and navigational search based on entity associations, upstream and downstream data dependency paths, and application task scenarios, significantly improving discovery efficiency and reuse rate.
4. The data component encapsulation construction method based on the standardized description model according to claim 1 is characterized in that: The identity resolution mechanism in step (3) includes: Build a multi-level identity resolution system consisting of root nodes, domain nodes, and end nodes, supporting the mapping management of identities to resource locations, interface paths, and metadata; Supports identity resolution and access rights binding, performs access verification based on the caller's identity and policy, and ensures the legality and security of data component use; Provides a unified interface service based on HTTP or Handle protocol to achieve automatic parsing and dynamic routing of identifiers.
5. The data component encapsulation construction method based on the standardized description model according to claim 1 is characterized in that: The data interconnection protocol in step (4) includes: Build a unified data component protocol cluster, covering mainstream protocols such as DOIP, DOIRP, REST API, and GraphQL, to achieve compatible access on different platforms; Supports an addressing mechanism based on identity resolution to achieve distributed access, transparent forwarding and permission control of data components.
6. The data component encapsulation construction method based on the standardized description model according to claim 1 is characterized in that: The core mechanisms of the data component in step (5) include: The standardized description model defines the structural boundaries and behavioral specifications of data components; The registration and publishing mechanism implements metadata registration, index publishing, and permission management of data components; The identification resolution mechanism ensures the unique addressing, dynamic positioning and secure calling of data components; The access acquisition mechanism supports the loading, parsing, combination and reuse of data components in a multi-system environment, realizing unified services and trusted governance of cross-domain data.
Citation Information
Cited By
Data weaving method for integration and treatment of multi-source heterogeneous data
CN120910144A
Data processing method and system driven by data piece, medium and program product
CN121542342A
A data-driven data processing method, system, medium and program product
CN121542342B
Data processing method and device, electronic equipment, storage medium and program product
CN122132466A