Data lake construction method and system based on multi-source distributed data

By deploying data acquisition agents on each node of a distributed data source and establishing a distributed collaboration mechanism, the problem of large-scale distributed data processing performance bottlenecks under centralized processing is solved, and efficient processing and unified management of distributed data is realized.

CN120144562AInactive Publication Date: 2025-06-13BEIJING LINGDING LANHAI TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510306867.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-15
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the construction of data lakes usually adopts a centralized data processing method, resulting in performance bottlenecks when processing large-scale distributed data.

Method used

By deploying data acquisition agents at each node of a distributed data source and establishing a distributed collaboration mechanism, data processing tasks can be carried out at the node where the data source is located, avoiding large-scale data transmission. The data acquisition agent performs structured analysis of the data in the distributed data source, obtains data patterns, field types and constraints, and establishes the primary and foreign key relationship, reference relationship and business relationship between fields, generates RDF triples and stores them on the source node.

Benefits of technology

It realizes efficient processing and unified management of distributed data, avoids performance bottlenecks under centralized processing methods, and improves data processing efficiency and data utilization value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144562A_ABST
    Figure CN120144562A_ABST
Patent Text Reader

Abstract

The invention provides a data lake construction method and system based on multi-source distributed data, and relates to the field of data lake construction. The method comprises the following steps: firstly, deploying a data acquisition agent at each node of a distributed data source and establishing a distributed collaborative mechanism, and performing structured analysis on data through the data acquisition agent to obtain data characteristics and establish a relationship between fields; and then semantic description information is constructed, and an RDF triple is generated based on the information and stored in a source end node. And then inputting the RDF triple into an ontology registration service for processing to generate a distributed ontology, and analyzing a corresponding relationship between the RDF triple and concepts in the distributed ontology to obtain a semantic equivalence rule and a mapping rule. And finally, organizing the distributed data sources into a unified data view according to the rules, establishing a distributed index, configuring an access permission, and finally constructing a data lake containing the unified data view, the distributed index and the access permission. According to the method, the performance bottleneck problem caused by centralized processing is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data lake construction based on multi-source distributed data, and particularly to a data lake construction method and system. Background Art

[0002] With the rapid development of big data technology, enterprises and organizations are facing huge challenges in processing and managing multi-source heterogeneous data. As a new type of data management architecture, a data lake can store and manage raw data from different sources, providing support for data analysis and value mining. In the prior art, the construction of a data lake usually adopts a centralized data processing method, that is, data in data sources distributed at different locations is uniformly extracted to a centralized storage system for processing and storage. This method requires all data to be transmitted to the central node, and then semantic analysis and association construction of the data are carried out. However, with the continuous growth of data scale and the increasing dispersion of data distribution, the centralized processing method faces serious performance bottlenecks in processing large-scale distributed data. Summary of the Invention

[0003] This application provides a data lake construction method and system based on multi-source distributed data, which avoids the performance bottleneck problem caused by centralized processing.

[0004] In a first aspect, this application provides a data lake construction method based on multi-source distributed data, including: Deploy data collection agents at each node of the distributed data source, and establish a distributed cooperation mechanism through the data collection agents; Based on the distributed cooperation mechanism, perform structured analysis on the data in the distributed data source through the data collection agents to obtain a data schema, field types, and constraint conditions, and establish primary-foreign key relationships, reference relationships, and business relationships between fields based on the data schema, the field types, and the constraint conditions; Construct semantic description information including data structure features, field attribute features, and relationship features according to the primary-foreign key relationships, the reference relationships, and the business relationships; Under the control of the distributed cooperation mechanism, generate an RDF subject according to the data structure features, generate an RDF predicate according to the field attribute features, generate an RDF object according to the relationship features, thereby form an RDF triple according to the RDF subject, the RDF predicate, and the RDF object, and store the RDF triple in the corresponding source-end node; Input the RDF triples into the ontology registration service under the scheduling of the distributed collaboration mechanism. The ontology registration service classifies the RDF triples according to the data domain division rules, determines the ontology attributes and relationship definitions, and establishes the inheritance relationship between ontologies according to the ontology attributes and the relationship definitions to generate a distributed ontology. Analyze the correspondence between the concepts in the RDF triples and the concepts defined in the distributed ontology based on the distributed collaboration mechanism to obtain semantic equivalence rules and mapping rules. According to the semantic equivalence rules and the mapping rules, organize the distributed data sources into a unified data structure, generate a unified data view, and establish a distributed index based on the unified data view. Configure access permissions for the unified data view according to the preset permission rules and the distributed index, so as to obtain a data lake containing the unified data view, the distributed index, and the access permissions.

[0005] In the above technical solution, the present invention deploys data collection agents at each node of the distributed data source and establishes a distributed collaboration mechanism, enabling data processing tasks to be carried out at the nodes where the data sources are located, avoiding large-scale data transmission. At the same time, the distributed collaboration mechanism ensures the synchronization of data processing among nodes. On this basis, the data collection agents perform structured analysis on the data in the distributed data source, obtain data patterns, field types, and constraint conditions, and establish primary-foreign key relationships, reference relationships, and business relationships among fields, realizing the accurate extraction of data features. By converting the primary-foreign key relationships, reference relationships, and business relationships into semantic description information including data structure features, field attribute features, and relationship features, it provides a basis for subsequent semantic modeling. Under the control of the distributed collaboration mechanism, the system generates RDF subjects, RDF predicates, and RDF objects respectively according to the data structure features, field attribute features, and relationship features, forms RDF triples and stores them in the source-side nodes, ensuring the integrity and consistency of semantic descriptions. The ontology registration service classifies the RDF triples based on the data domain division rules, constructs the inheritance relationship between ontologies by establishing ontology attribute and relationship definitions, and generates a distributed ontology, realizing the semantic unification of data. The system analyzes the correspondence between the concepts in the RDF triples and the concepts defined in the distributed ontology to obtain semantic equivalence rules and mapping rules, establishing semantic associations between different data sources. Finally, based on the semantic equivalence rules and the mapping rules, the distributed data sources are organized into a unified data structure, a unified data view is generated, and a distributed index is established, realizing the unified management and efficient access of data. By configuring access permissions for the unified data view, the security of data access is ensured. This solution not only solves the performance bottleneck problem under the traditional centralized processing method, but also realizes the semantic association and unified management of distributed data, improving the data processing efficiency and data utilization value.

[0006] In a second aspect of the present application, a data lake construction system based on multi-source distributed data is provided. The system includes: a distributed collaboration module, a structured analysis module, a semantic description module, an RDF generation module, an ontology generation module, a semantic mapping module, a data view generation module, and a permission control module; The distributed collaboration module is used to deploy data collection agents at each node of the distributed data source and establish a distributed collaboration mechanism through the data collection agents; The structured analysis module is used to perform structured analysis on the data in the distributed data source through the data collection agents based on the distributed collaboration mechanism, obtain data patterns, field types, and constraint conditions, and establish primary-foreign key relationships, reference relationships, and business relationships between fields based on the data patterns, the field types, and the constraint conditions; The semantic description module is used to construct semantic description information including data structure characteristics, field attribute characteristics, and relationship characteristics according to the primary-foreign key relationships, the reference relationships, and the business relationships; The RDF generation module is used to generate an RDF subject according to the data structure characteristics, generate an RDF predicate according to the field attribute characteristics, and generate an RDF object according to the relationship characteristics under the control of the distributed collaboration mechanism, so as to form an RDF triple according to the RDF subject, the RDF predicate, and the RDF object, and store the RDF triple in the corresponding source node; The ontology generation module is used to input the RDF triple into an ontology registration service under the scheduling of the distributed collaboration mechanism. The ontology registration service classifies the RDF triple according to data domain division rules, determines ontology attribute and relationship definitions, and establishes inheritance relationships between ontologies according to the ontology attributes and the relationship definitions to generate a distributed ontology; The semantic mapping module is used to analyze the correspondence between the concepts in the RDF triple and the concepts defined in the distributed ontology based on the distributed collaboration mechanism to obtain semantic equivalence rules and mapping rules; The data view generation module is used to organize the distributed data source into a unified data structure according to the semantic equivalence rules and the mapping rules, generate a unified data view, and establish a distributed index based on the unified data view; The permission control module is used to configure access permissions for the unified data view according to preset permission rules and the distributed index, so as to obtain a data lake including the unified data view, the distributed index, and the access permissions.

[0007] In a third aspect of the present application, a computer storage medium is provided. The computer storage medium stores multiple instructions, and the instructions are adapted to be loaded and executed by a processor to perform the above method steps.

[0008] In a fourth aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs the above method.

[0009] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. By deploying data collection agents at each node of the distributed data source and establishing a distributed collaboration mechanism in the present invention, data processing tasks can be performed at the nodes where the data sources are located, avoiding large-scale data transmission. At the same time, the distributed collaboration mechanism ensures the synchronization of data processing among nodes. On this basis, the data collection agent performs structured analysis on the data in the distributed data source, obtains data patterns, field types, and constraint conditions, and establishes primary-foreign key relationships, reference relationships, and business relationships between fields, realizing the accurate extraction of data features. By converting the primary-foreign key relationships, reference relationships, and business relationships into semantic description information including data structure features, field attribute features, and relationship features, it provides a basis for subsequent semantic modeling. Under the control of the distributed collaboration mechanism, the system generates RDF subjects, RDF predicates, and RDF objects respectively according to the data structure features, field attribute features, and relationship features, forms RDF triples and stores them in the source-side nodes, ensuring the integrity and consistency of semantic descriptions. The ontology registration service classifies the RDF triples based on the data domain division rules, constructs the inheritance relationship between ontologies by establishing ontology attribute and relationship definitions, and generates a distributed ontology, realizing the semantic unification of data.

[0010] 2. By analyzing the correspondence relationship between the concepts in the RDF triples and the concepts defined in the distributed ontology in the present application, semantic equivalence rules and mapping rules are obtained, and semantic associations between different data sources are established. Finally, based on the semantic equivalence rules and mapping rules, the distributed data sources are organized into a unified data structure, a unified data view is generated, and a distributed index is established, realizing the unified management and efficient access of data. By configuring access permissions for the unified data view, the security of data access is ensured. This solution not only solves the performance bottleneck problem in the traditional centralized processing mode, but also realizes the semantic association and unified management of distributed data, improving the data processing efficiency and data utilization value. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1Schematic flowchart of a method for constructing a data lake based on multi-source distributed data provided by an embodiment of the present application; Figure 2 Architecture diagram of a system for constructing a data lake based on multi-source distributed data provided by an embodiment of the present application; Figure 3 Schematic structural diagram of an electronic device provided by the present application. Detailed implementation manners

[0012] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0013] In the description of the embodiments of the present application, words such as "for example" or "for illustration" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "for example" or "for illustration" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "for example" or "for illustration" is intended to present relevant concepts in a specific manner.

[0014] In the description of the embodiments of the present application, the meaning of the term "a plurality of" refers to two or more. For example, a plurality of systems refers to two or more systems, and a plurality of screen terminals refers to two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0015] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0016] On the basis of the above background technology, further, please refer to Figure 1 , Figure 1 Schematic flowchart of a method for constructing a data lake based on multi-source distributed data provided by an embodiment of the present application. This system can be implemented depending on a computer program or run as an independent tool-like application. Specifically, in the embodiments of the present application, this method can be applied to a server, but can also be applied to an electronic device such as a server. A method for constructing a data lake based on multi-source distributed data includes the following steps: S101. Deploy data collection agents at each node of the distributed data source, and establish a distributed collaboration mechanism through the data collection agents; Specifically, the present invention first deploys data collection agents on each storage node of the distributed data source. The data collection agent is a data processing program used to implement access to the data source and extraction of data. A distributed collaboration mechanism is established through the data collection agents deployed on each node. The distributed collaboration mechanism refers to a collaborative working mechanism established between each node, which is specifically implemented in the following manner: First, communication connections are established between the data collection agents of each node to form a collaborative network. Second, a unified data processing protocol is implemented in the collaborative network. This protocol defines the data exchange format and processing rules between nodes. Third, each data collection agent maintains the stable operation of the collaborative network by regularly sending status information. When it is detected that a data collection agent of a certain node is abnormal, the data collection agents of other nodes can promptly perceive and perform corresponding processing. Through this distributed collaboration mechanism, the data collection agents on each node can work collaboratively, enabling data processing tasks to be carried out at the nodes where the data sources are located, avoiding the performance bottleneck problem of centrally transmitting data to a central node for processing. At the same time, the distributed collaboration mechanism also realizes task allocation and load balancing between nodes, improving the processing efficiency of the entire system. In addition, the data collection agent also realizes real-time monitoring of the data source, can promptly detect changes in the data source and perform corresponding processing, ensuring the real-time and accuracy of data processing.

[0017] S102. Based on the distributed collaboration mechanism, perform structured analysis on the data in the distributed data source through the data collection agents to obtain data patterns, field types, and constraint conditions, and establish primary-foreign key relationships, reference relationships, and business relationships between fields based on the data patterns, the field types, and the constraint conditions; Specifically, after establishing the distributed collaboration mechanism, the system performs structured analysis on the data in the distributed data sources through the data collection agent. The structured analysis first scans the data tables in the data sources, extracts the definition information of the tables to obtain the data schema, which includes the basic structure definitions of the tables, such as table names, field compositions, etc. Then, it analyzes the specific attributes of the fields in each data table to obtain the field types, which record the basic characteristics of each field, such as data type, length limit, whether null values are allowed, default values, etc. Next, it extracts various defined rules in the data tables to obtain the constraint conditions, which include constraint rules such as field value range limits and dependencies between fields. After obtaining the data schema, field types, and constraint conditions, the system further analyzes the associations between them and establishes the primary-foreign key relationships, reference relationships, and business relationships between fields. Among them, the primary-foreign key relationship represents the association relationship between data tables, which is determined by the corresponding relationship between the primary key and the foreign key; the reference relationship represents the reference dependency relationship between fields, indicating the association method of data between different tables; the business relationship represents the field association relationship established based on business rules, reflecting the association relationship of data at the business level. In this way, the system realizes in-depth analysis of the distributed data sources, laying a foundation for subsequent construction of semantic description information.

[0018] Based on the above embodiments, as an alternative embodiment, the performing structured analysis on the data in the distributed data sources through the data collection agent to obtain the data schema, field types, and constraint conditions includes: S201, parsing the data tables in the distributed data sources to obtain the table structure information, and parsing the table name, table description, and primary key definition in the table structure information to obtain the data schema; Specifically, during the process of performing structured analysis, the system first accesses the data tables in the distributed data sources through the data collection agent. The data collection agent reads the definition files of each data table, extracts the structure definitions of the tables to obtain the table structure information, which includes the basic organization form of the data tables. Then, the data collection agent deeply parses the table structure information, extracts the table name, which is used to uniquely identify each data table; extracts the table description, which is used to describe the purpose and content of the data table; extracts the primary key definition, which is used to determine the unique identifier field of the data table. Through the extraction and organization of these information, the system obtains the complete data schema, which reflects the basic structure characteristics of the data tables. This way enables the system to accurately understand the organization method of the data tables, providing a foundation for subsequent field analysis and relationship establishment.

[0019] S202, parsing the name, data type, length limit, whether null, and default value information of the fields in the table structure information to obtain the field types; Specifically, after obtaining the table structure information, the system starts to parse the field information in the table. First, the system extracts the names of the fields from the table structure information. The field names are used to identify each data column in the data table. Then, the system extracts the data types of each field. The data types define the storage formats of the data in the fields, such as numeric type, character type, etc. Next, the system extracts the length limits of the fields. The length limits specify the range of data lengths that the fields can store. The system also extracts the null value constraints of the fields, that is, the settings of whether they can be null. This setting specifies whether the fields allow storing null values. Finally, the system extracts the default value information of the fields. The default value information specifies the values of the fields when no explicit assignment is made. By parsing and organizing these field attributes, the system obtains the complete field type information. These information accurately describe the characteristics of each field in the data table and provide detailed field attribute information for subsequent semantic descriptions.

[0020] S203. Parse the constraint rules in the table structure information to obtain the constraint conditions.

[0021] Specifically, after completing the extraction of the field types, the system starts to parse the constraint rules in the table structure information. First, the system obtains the field value range constraints from the table structure information. The value range constraints specify the range limits of the field values. Then, the system extracts the value constraints between fields. The value constraints between fields define the value rules between different fields in the same data table. Next, the system extracts the cross-table association constraints. The cross-table association constraints describe the association rules between different data tables. By parsing these constraint rules, the system obtains the complete constraint conditions. The constraint conditions ensure the integrity and consistency of the data and provide a constraint basis for subsequent establishment of the association relationships between the data.

[0022] S103. Construct semantic description information including data structure characteristics, field attribute characteristics, and relationship characteristics according to the primary-foreign key relationship, the reference relationship, and the business relationship. Specifically, after obtaining the primary foreign key relationship, reference relationship, and business relationship, the system begins to construct semantic description information. First, the system combines the primary key table name, primary key fields, and foreign key table information that references the primary key in the primary foreign key relationship to form a hierarchical structure, obtaining the data structure feature, which reflects the hierarchical inclusion relationship between data tables. Then, the system combines the reference field information in the reference relationship, including attribute information such as the data type, length limit, nullability, and default value of the field, to form an attribute set, obtaining the field attribute feature, which describes the specific value-taking rules of the field. Next, the system combines the field information in the business relationship, including the business meaning of the field, data flow rules, and business constraints between fields, to form an association rule, obtaining the relationship feature, which reflects the business logic relationship between data. Finally, the system organizes and associates the data structure feature, field attribute feature, and relationship feature according to a predefined semantic description format to generate complete semantic description information. In this way, the system transforms the structural relationship, attribute features, and business relationship of data into standardized semantic descriptions, providing a semantic basis for the subsequent generation of RDF triples.

[0023] Based on the above embodiments, as an alternative embodiment, the constructing semantic description information including data structure features, field attribute features, and relationship features according to the primary foreign key relationship, the reference relationship, and the business relationship includes: S301, combining the primary key table name, primary key fields, and foreign key table information that references the primary key in the primary foreign key relationship to form a hierarchical structure, obtaining the data structure feature, where the hierarchical structure is used to represent the inclusion relationship between data tables; Specifically, after obtaining the primary foreign key relationship, the system begins to construct the data structure feature. First, the system extracts the primary key table name from the primary foreign key relationship, and the primary key table name identifies the main entity of the data table. Then, the system extracts the primary key field information, and the primary key field is used to uniquely identify the data records in the primary key table. Next, the system extracts the foreign key table information that references the primary key, and the foreign key table information records other data tables associated with the primary key table. The system combines this information into a hierarchical structure according to the reference relationship. In the hierarchical structure, the primary key table is located at the upper node, and the foreign key table that references the primary key is located at the lower node. Through this hierarchical relationship, the inclusion relationship between data tables is represented. The finally obtained data structure feature completely describes the hierarchical organizational structure between data tables, providing a structured expression way for subsequent semantic descriptions.

[0024] S302, combining the attribute information of the reference fields in the reference relationship to form an attribute set, obtaining the field attribute feature, where the attribute set is used to describe the value-taking rules of the fields; Specifically, after obtaining the reference relationship, the system begins to construct the field attribute features. First, the system extracts the attribute information of the reference fields from the reference relationship, including basic features such as the data type, length limit, nullability, default value, and value range of the fields. Then, the system organizes this attribute information into an attribute set, and each attribute in the attribute set contains a complete definition of the value-taking rules. Among them, the data type specifies the data storage format of the field, the length limit determines the storage length of the data, the null value constraint specifies whether null values are allowed, the default value sets the initial value of the field, and the value range limits the valid range of the data. By combining this attribute information, the system obtains the complete field attribute features, which clearly define the value-taking rules of the fields and provide a basis for ensuring the validity and consistency of the data.

[0025] S303. Combine the associated field information in the business relationship to form an association rule to obtain the relationship feature. Specifically, after obtaining the business relationship, the system begins to construct the relationship feature. First, the system extracts the associated field information from the business relationship, including the business meaning of the fields, the data flow rules, and the business constraints between the fields. Then, the system combines this associated field information into an association rule. The business meaning in the association rule describes the role of the field in the business, the data flow rule defines the way of data transfer in different business processes, and the business constraint stipulates the business restriction conditions between the fields. By combining this information, the system obtains the complete relationship feature, which reflects the association relationship of the data at the business level and provides business-level support for subsequent semantic descriptions.

[0026] S304. Organize and associate the data structure feature, the field attribute feature, and the relationship feature according to a preset semantic description format to generate the semantic description information.

[0027] Specifically, after obtaining the data structure feature, the field attribute feature, and the relationship feature, the system begins to generate the semantic description information. First, the system reads the data structure feature according to the preset semantic description format and converts the hierarchical relationship between the data tables into a standard semantic expression. Then, the system converts the value-taking rules in the field attribute feature according to the semantic description format to form the semantic representation of the field attributes. Next, the system converts the business association rules in the relationship feature into semantic relationship descriptions. Finally, the system organizes and associates the converted data structure semantic expression, the field attribute semantic representation, and the business relationship semantic description to generate the complete semantic description information. Through this standardized semantic description method, the system realizes the unified semantic expression of the data structure, attributes, and relationships, providing a standardized semantic basis for subsequent RDF generation.

[0028] S104. Under the control of the distributed collaboration mechanism, generate an RDF subject according to the data structure feature, generate an RDF predicate according to the field attribute feature, and generate an RDF object according to the relationship feature, so as to form an RDF triple according to the RDF subject, the RDF predicate, and the RDF object, and store the RDF triple in the corresponding source node; Specifically, after obtaining the semantic description information, the system starts to generate RDF triples under the control of the distributed collaboration mechanism. The system first generates an RDF subject according to the data structure feature. The RDF subject refers to the part in the Resource Description Framework used to identify data entities, which is specifically realized by converting the hierarchical structure of the data table into a resource identifier. Then, the system generates an RDF predicate according to the field attribute feature. The RDF predicate is the part used to describe the relationship between the subject and the object, and is expressed by converting the attribute information of the field into a descriptive predicate. Then, the system generates an RDF object according to the relationship feature. The RDF object is the part used to represent the attribute value or the associated object, and is expressed by converting the business relationship into a specific attribute value or an associated entity. After generating the RDF subject, the RDF predicate, and the RDF object, the system combines them to form an RDF triple, and each RDF triple contains complete semantic description information. Finally, the system stores the generated RDF triples in the corresponding source nodes. This storage method maintains the distributed characteristics of the data and avoids the performance problems caused by centralized data storage. In this way, the system converts the semantic description information into a standard RDF representation form, providing a basis for subsequent ontology construction and semantic mapping.

[0029] S105. Input the RDF triple into the ontology registration service under the scheduling of the distributed collaboration mechanism. The ontology registration service classifies the RDF triple according to the data domain division rule, determines the ontology attributes and relationship definitions, and establishes the inheritance relationship between ontologies according to the ontology attributes and the relationship definitions to generate a distributed ontology; Specifically, after the generation of RDF triples is completed, the system inputs the RDF triples into the ontology registration service under the scheduling of the distributed collaboration mechanism. The ontology registration service first classifies the input RDF triples according to the data domain division rules. The data domain division rules define the classification criteria for different business domains. Through this classification method, the relevant RDF triples are classified into the corresponding data domains. After the classification is completed, the system analyzes the RDF triples in each data domain, extracts the common attribute features and relationship features, so as to determine the ontology attributes and relationship definitions. Among them, the ontology attributes describe the basic features of data entities, and the relationship definitions describe the association methods between entities. After obtaining the ontology attributes and relationship definitions, the system establishes the inheritance relationship between ontologies according to their semantic associations. The inheritance relationship represents the hierarchical inclusion relationship between different ontologies, and realizes the transfer of attributes and relationships through the inheritance mechanism. Finally, the system organizes the ontologies with inheritance relationships into a distributed ontology. The distributed ontology maintains the distributed characteristics of data and provides a unified semantic description framework. In this way, the system realizes the semantic unification of distributed data and lays a foundation for subsequent semantic mapping.

[0030] Based on the above embodiments, as an alternative embodiment, the establishing the inheritance relationship between ontologies according to the ontology attributes and the relationship definitions to generate a distributed ontology includes: S401, analyzing the feature distribution of the ontology attributes, clustering the ontologies with the same attribute features, and obtaining a set of ontology categories; Specifically, after obtaining the ontology attributes, the system starts to establish the inheritance relationship between ontologies. First, the system analyzes the feature distribution of the ontology attributes, extracts the attribute features of each ontology, including the type, value range and constraint conditions of the attributes. Then, the system identifies the ontologies with the same attribute features by comparing the attribute features of different ontologies, and classifies these ontologies into the same category. Through this clustering analysis of attribute features, the system organizes the ontologies with similar attribute features together to form a set of ontology categories. This clustering method enables the system to establish the association between ontologies according to the similarity of attribute features, laying a foundation for establishing the inheritance relationship subsequently.

[0031] S402, determining the association method between different ontologies in the set of ontology categories according to the relationship definition, and establishing the inheritance relationship between the ontologies; Specifically, after obtaining the ontology category set, the system begins to establish the inheritance relationship between ontologies. First, the system analyzes each ontology in the ontology category set according to the relationship definition to determine the association method between different ontologies. Among them, the association methods include attribute inheritance relationship, relationship inheritance relationship, and constraint inheritance relationship. The attribute inheritance relationship means that the sub-ontology inherits the attribute characteristics of the parent ontology. The relationship inheritance relationship means that the sub-ontology inherits the association rules of the parent ontology. The constraint inheritance relationship means that the sub-ontology inherits the constraint conditions of the parent ontology. Then, the system establishes the inheritance relationship between ontologies according to these association methods, and realizes the transfer of attributes, relationships, and constraints through the inheritance relationship. This inheritance mechanism enables the system to establish a hierarchical relationship between ontologies, reduces duplicate definitions, and improves the reusability of ontologies.

[0032] S403. Construct the hierarchical structure of the ontology based on the inheritance relationship, and determine the parent-child relationship, parallel relationship, and cross relationship between ontologies. Specifically, after establishing the ontology inheritance relationship, the system begins to construct the hierarchical structure of the ontology. First, the system analyzes the relationship type between ontologies based on the inheritance relationship. Among them, the inheritance relationship directly reflects the parent-child relationship between ontologies, and the parent-child relationship means that the upper ontology is a generalization of the lower ontology. Then, the system identifies the ontologies with the same parent ontology and determines the parallel relationship between them. The parallel relationship means that these ontologies have the same abstract level. Next, the system analyzes the ontologies with partial attribute overlap and determines the cross relationship between them. The cross relationship means that these ontologies have something in common in some features. By constructing this hierarchical structure, the system realizes the complete expression of the relationship between ontologies and provides a structured organization method for the subsequent generation of distributed ontologies.

[0033] S404. Combine the ontology category set, the inheritance relationship, and the hierarchical structure to generate the distributed ontology.

[0034] Specifically, after completing the construction of the ontology hierarchical structure, the system begins to generate the distributed ontology. First, the system reads the ontology definitions in the ontology category set, and the ontology category set contains ontology information classified according to attribute characteristics. Then, the system combines the attribute inheritance, relationship inheritance, and constraint inheritance rules defined in the inheritance relationship with the ontology definitions to form an ontology description with an inheritance mechanism. Next, the system integrates the parent-child relationship, parallel relationship, and cross relationship in the hierarchical structure into the ontology description to establish the organizational relationship between ontologies. Finally, the system constructs these combined information into a distributed ontology. The distributed ontology maintains the distributed characteristics of the data and provides a unified semantic description framework at the same time. This combination method enables the system to realize the unified management and semantic expression of ontologies while maintaining the data distribution.

[0035] S106. Analyze the correspondence between the concepts in the RDF triples and the concepts defined in the distributed ontology based on the distributed collaboration mechanism, and obtain semantic equivalence rules and mapping rules. Specifically, after generating the distributed ontology, the system starts to analyze the correspondence between the concepts in the RDF triples and the concepts defined in the distributed ontology based on the distributed collaboration mechanism. First, the system extracts concepts from the RDF triples, performs semantic feature analysis on each extracted concept, and obtains the attribute feature set and relationship feature set of the concept. The attribute feature set contains the basic attribute information of the concept, and the relationship feature set contains the association information between the concept and other concepts. Then, the system matches the attribute feature set with the concept attributes defined in the distributed ontology, calculates the semantic similarity between them, and establishes semantic equivalence rules based on the calculated semantic similarity. The semantic equivalence rules are used to identify concepts with the same semantic meaning. Next, the system matches the relationship feature set with the concept relationships defined in the distributed ontology, calculates the structural similarity between them, and establishes mapping rules based on the calculated structural similarity. The mapping rules are used to define the conversion methods between different concepts. In this way, the system realizes the semantic association between the concepts in the RDF triples and the concepts in the distributed ontology, providing a basis for constructing a unified data view in the future.

[0036] Based on the above embodiments, as an alternative embodiment, the step of analyzing the correspondence between the concepts in the RDF triples and the concepts defined in the distributed ontology based on the distributed collaboration mechanism, and obtaining semantic equivalence rules and mapping rules includes: S501. Extract the concepts in the RDF triples, and perform semantic feature analysis on each concept to obtain the attribute feature set and relationship feature set of the concept. Specifically, after obtaining the attribute feature set, the system starts to establish semantic equivalence rules. First, the system compares each attribute feature in the attribute feature set with the concept attributes defined in the distributed ontology. The comparison process includes data type matching, value range matching, and constraint condition matching of the attributes. Then, the system calculates the matching degree of each pair of attributes in each dimension to obtain the semantic similarity between the attributes. The semantic similarity numerically represents the corresponding degree of the attributes at the semantic level. Next, the system sets a threshold according to the calculated semantic similarity, determines the attribute pairs with similarity exceeding the threshold as semantic equivalent attributes, and forms these equivalent relationships into semantic equivalence rules. In this way, the system establishes the semantic mapping relationship of the attributes in different data sources, providing a rule basis for realizing semantic unification of data.

[0037] S502. Match the set of attribute features with the concept attributes defined in the distributed ontology to obtain the semantic similarity between the attributes, and establish the semantic equivalence rules based on the semantic similarity. Specifically, after obtaining the set of attribute features, the system starts to establish semantic equivalence rules. First, the system compares each attribute feature in the set of attribute features with the concept attributes defined in the distributed ontology. The comparison process includes data type matching, value range matching, and constraint condition matching of the attributes. Then, the system calculates the matching degree of each pair of attributes in each dimension to obtain the semantic similarity between the attributes. The semantic similarity numerically represents the corresponding degree of the attributes at the semantic level. Next, the system sets a threshold based on the calculated semantic similarity, determines the attribute pairs with similarity exceeding the threshold as semantically equivalent attributes, and forms these equivalent relationships into semantic equivalence rules. In this way, the system establishes the semantic mapping relationship of the attributes in different data sources, providing a rule basis for realizing the semantic unification of data.

[0038] S503. Match the set of relationship features of the concept with the concept relationships defined in the distributed ontology to obtain the structural similarity between the relationships, and establish the mapping rules based on the structural similarity.

[0039] Specifically, after obtaining the set of relationship features, the system starts to establish mapping rules. First, the system compares each relationship feature in the set of relationship features with the concept relationships defined in the distributed ontology. The comparison process includes relationship type matching, association method matching, and constraint rule matching. Then, the system calculates the corresponding degree of each pair of relationships in terms of structural composition to obtain the structural similarity between the relationships. The structural similarity reflects the matching degree of the relationships at the structural level. Next, the system sets a threshold based on the calculated structural similarity, determines the relationship pairs with similarity exceeding the threshold as structurally equivalent relationships, and forms these equivalent relationships into mapping rules. In this way, the system establishes the mapping rules of the concept relationships in different data sources, providing an operation basis for realizing the structural transformation of data.

[0040] S107. According to the semantic equivalence rules and the mapping rules, organize the distributed data source into a unified data structure, generate a unified data view, and establish a distributed index based on the unified data view. Specifically, after obtaining the semantic equivalence rules and mapping rules, the system begins to organize the distributed data sources into a unified data structure. First, the system identifies the data with the same semantics in the distributed data sources according to the semantic equivalence rules, classifies the data items with the same semantics into the same data category, and obtains a set of semantically associated data categories. Then, the system determines the conversion methods between different data categories in the data category set according to the mapping rules, and converts the data in the data category set into a unified representation form based on these conversion methods, obtaining a standardized data category set. Next, the system constructs a hierarchical data structure for the semantically related data categories in the standardized data category set according to the preset organization rules, obtains a unified data organization structure, and establishes a data access interface and data operation methods based on this unified data organization structure to generate a unified data view. After generating the unified data view, the system scans the structure definition of the unified data view, obtains the classification hierarchy of the data categories, and identifies the storage node information and storage path information of each data category according to the classification hierarchy to establish a distributed index. In this way, the system realizes the unified organization and efficient access of the distributed data sources, laying a foundation for the final construction of the data lake.

[0041] Based on the above embodiments, as an alternative embodiment, the organizing the distributed data sources into a unified data structure and generating a unified data view according to the semantic equivalence rules and the mapping rules includes: S601, identifying the data with the same semantics in the distributed data sources according to the semantic equivalence rules, classifying the data items with the same semantics into the same data category, and obtaining a set of semantically associated data categories; Specifically, after obtaining the semantic equivalence rules, the system begins to perform semantic classification on the data in the distributed data sources. First, the system scans each data item in the distributed data sources according to the semantic equivalence rules to identify the semantic features of the data item. Then, the system matches the semantic features with the semantic equivalence rules to identify the data items with the same semantics. Next, the system organizes the data items with the same semantic identification into the same data category to form a semantically associated data category. Finally, the system organizes all data categories into a data category set, and each category in this data category set contains data items with the same semantics. Through this semantic classification method, the system realizes the unified organization of the semantically related data in the distributed data sources, providing a foundation for subsequent data structure conversion.

[0042] S602, determining the conversion methods between different data categories in the data category set according to the mapping rules, and converting the data in the data category set into a unified representation form based on the conversion methods, obtaining a standardized data category set; Specifically, after obtaining the data category set, the system starts data conversion and standardization processing. First, the system analyzes the relationships between different data categories in the data category set according to the mapping rules to determine the conversion methods between data categories. Then, the system uses these conversion methods to process the data in the data category set, converting data in different forms into a unified representation form. The conversion process includes data format conversion, data unit conversion, and data encoding conversion. Among them, data format conversion unifies data in different formats into a standard format, data unit conversion converts data in different units into a unified unit, and data encoding conversion converts data in different encoding methods into a unified encoding. Through this conversion process, the system obtains a standardized data category set, and all data in the standardized data category set adopt a unified representation form, providing a standardized data foundation for subsequent construction of a unified data view.

[0043] S603. Construct a hierarchical data structure for the semantically related data categories in the standardized data category set according to the preset organization rules to obtain a unified data organization structure, and establish a data access interface and data operation methods based on the unified data organization structure to generate the unified data view.

[0044] Specifically, after obtaining the standardized data category set, the system starts to construct a unified data view. First, the system analyzes the semantic associations between data categories in the standardized data category set according to the preset organization rules. The preset organization rules define the hierarchical relationships and organization methods between data categories. Then, the system organizes the semantically related data categories in a hierarchical manner to construct a unified data organization structure, which reflects the hierarchical inclusion relationships and reference relationships between data categories. Next, the system constructs a data access interface based on the unified data organization structure. The data access interface defines the query methods and access paths of the data. At the same time, the system also establishes data operation methods, and the data operation methods stipulate the operation rules for data addition, deletion, modification, and query. Finally, the system integrates the data organization structure, data access interface, and data operation methods to generate a unified data view, which provides a unified access and operation method for distributed data sources.

[0045] Based on the above embodiments, as an alternative embodiment, establishing a distributed index based on the unified data view includes: S701. Scan the structure definition of the unified data view to obtain the classification hierarchy of data categories, and identify the storage node information and storage path information of each data category according to the classification hierarchy to generate a location index table; Specifically, when building a distributed index, the system first scans the unified data view, analyzes the structure definitions in the unified data view, which include the organization method and association relationship of the data. Then, the system extracts the classification hierarchy of data categories from the structure definitions, and the classification hierarchy reflects the hierarchical relationship between data categories. Next, the system analyzes each data category according to the classification hierarchy to identify the storage locations of the data categories in the distributed environment, including storage node information and storage path information. The storage node information indicates the physical location where the data is stored, and the storage path information indicates the access path of the data. Finally, the system organizes the identified storage node information and storage path information into a location index table, which records the complete location information of the data in the distributed environment and provides location navigation for subsequent data access.

[0046] S702. Based on the location index table, count the access frequencies and access time distributions of each data category in the unified data view, and establish a data heat index table, where the data heat index table records the access statistical information and cache allocation policies of the data categories; Specifically, after generating the location index table, the system starts to build the data heat index table. First, the system counts the access situations of each data category in the unified data view based on the storage location information recorded in the location index table. Through statistical analysis, the access frequency of the data category is obtained, and the access frequency reflects the number of times the data is accessed; at the same time, the access time distribution of the data category is counted, and the access time distribution reflects the time pattern of data access. Then, the system records the statistically obtained access frequency and access time distribution information as access statistical information. Next, the system formulates cache allocation policies for different data categories according to the access statistical information, and the cache allocation policies determine the caching method of the data in the distributed environment. Finally, the system organizes the access statistical information and cache allocation policies into a data heat index table, and this index structure enables the system to optimize data storage and access efficiency according to data access characteristics.

[0047] S703. According to the semantic association relationships recorded in the unified data view, combined with the location index table and the data heat index table, establish a semantic index table, where the semantic index table contains the semantic mapping relationships between data categories and cross-node access paths; Specifically, after establishing the data heat index table, the system begins to construct the semantic index table. First, the system reads the semantic association relationships of the records from the unified data view, and the semantic association relationships describe the semantic connections between different data categories. Then, the system combines the semantic association relationships with the storage location information in the location index table to determine the physical storage locations of the semantically related data. Next, the system combines the semantic association relationships with the access feature information in the data heat index table to analyze the access patterns of the semantically related data. On this basis, the system establishes the semantic mapping relationships between data categories, and the semantic mapping relationships represent the semantic correspondence relationships between different data categories; at the same time, it establishes cross-node access paths, and the cross-node access paths specify the ways to access the semantically related data distributed on different nodes. Finally, the system organizes the semantic mapping relationships and cross-node access paths into a semantic index table, and this index table supports semantic-based distributed data access.

[0048] S704, associate and organize the storage location information in the location index table, the caching policy in the data heat index table, and the semantic mapping relationships in the semantic index table to generate the distributed index.

[0049] Specifically, after completing the establishment of various index tables, the system begins to generate the distributed index. First, the system reads the storage location information from the location index table, and the storage location information records the physical storage locations and access paths of the data. Then, the system reads the caching policy from the data heat index table, and the caching policy defines the caching methods and update rules of the data. Next, the system reads the semantic mapping relationships from the semantic index table, and the semantic mapping relationships describe the semantic connections between the data. The system associates and organizes these three types of information to establish the corresponding relationships between the storage locations, caching policies, and semantic mappings, forming a complete distributed index structure. Through this multi-dimensional index organization method, the system realizes the unified management and efficient access of distributed data, while ensuring the semantic consistency of data access.

[0050] S108, configure access permissions for the unified data view according to the preset permission rules and the distributed index, so as to obtain a data lake including the unified data view, the distributed index, and the access permissions.

[0051] Specifically, after the establishment of the unified data view and the distributed index, the system begins to configure access permissions. First, the system defines the control policy for data access according to the preset permission rules, and the preset permission rules include the access levels and operation restrictions of different user roles to the data. Then, the system combines the data location information and access paths recorded in the distributed index to set corresponding access control policies for each data category in the unified data view. Through permission configuration, the system realizes the refined control of data access and ensures the security of data access. Finally, the system combines the unified data view with configured access permissions, the distributed index for supporting efficient data retrieval, and the permission rules for controlling data access to form a complete data lake. This data lake not only realizes the unified management and efficient access of distributed data, but also ensures the security of data use through access permission control, thus meeting the requirements of large-scale distributed data management.

[0052] On the other hand, the present invention also provides a data lake construction system based on multi-source distributed data, such as Figure 2 , the system includes: a distributed collaboration module, a structured analysis module, a semantic description module, an RDF generation module, an ontology generation module, a semantic mapping module, a data view generation module, and a permission control module; The distributed collaboration module 1 is used to deploy data collection agents at each node of the distributed data source and establish a distributed collaboration mechanism through the data collection agents; The structured analysis module 2 is used to perform structured analysis on the data in the distributed data source through the data collection agents based on the distributed collaboration mechanism, obtain data patterns, field types, and constraint conditions, and establish primary-foreign key relationships, reference relationships, and business relationships between fields based on the data patterns, the field types, and the constraint conditions; The semantic description module 3 is used to construct semantic description information including data structure features, field attribute features, and relationship features according to the primary-foreign key relationships, the reference relationships, and the business relationships; The RDF generation module 4 is used to generate an RDF subject according to the data structure features, generate an RDF predicate according to the field attribute features, and generate an RDF object according to the relationship features under the control of the distributed collaboration mechanism, so as to form RDF triples according to the RDF subject, the RDF predicate, and the RDF object, and store the RDF triples in the corresponding source node; The ontology generation module 5 is configured to input the RDF triples into an ontology registration service under the scheduling of the distributed collaboration mechanism. The ontology registration service classifies the RDF triples according to data domain division rules, determines ontology attributes and relationship definitions, and establishes inheritance relationships between ontologies based on the ontology attributes and the relationship definitions to generate a distributed ontology. The semantic mapping module 6 is configured to analyze the correspondence between the concepts in the RDF triples and the concepts defined in the distributed ontology based on the distributed collaboration mechanism to obtain semantic equivalence rules and mapping rules. The data view generation module 7 is configured to organize the distributed data sources into a unified data structure according to the semantic equivalence rules and the mapping rules, generate a unified data view, and establish a distributed index based on the unified data view. The permission control module 8 is configured to configure access permissions for the unified data view according to preset permission rules and the distributed index, so as to obtain a data lake including the unified data view, the distributed index, and access permissions.

[0053] Please refer to Figure 3 This application also discloses an electronic device. Figure 3 FIG. is a schematic structural diagram of an electronic device disclosed in an embodiment of this application. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.

[0054] Among them, the communication bus 302 is used to implement connection communication between these components.

[0055] Among them, the user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.

[0056] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0057] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling the data stored in the memory 305, it executes various functions of the server and processes data. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately by a single chip.

[0058] Among them, the memory 305 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 305 includes a non-transitory computer-readable medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store the data involved in the above-mentioned method embodiments. Optionally, the memory 305 may also be at least one storage system located far from the aforementioned processor 301. Refer to Figure 3 , as a computer storage medium, the memory 305 may include an operating system, a network communication module, a user interface module, and an application program for a method of constructing a data lake based on multi-source distributed data.

[0059] In Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an interface for the user to input and obtain the data input by the user; while the processor 301 can be used to call the application program for the road evaluation method stored in the memory 305. When executed by one or more processors 301, the electronic device 300 is caused to execute the method as described in one or more of the above embodiments. It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0060] In several implementation manners provided by this application, it should be understood that the disclosed system can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some service interfaces. The indirect couplings or communication connections of the system or unit can be in electrical or other forms.

[0061] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0062] This application embodiment also provides a computer storage medium. The computer storage medium can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor as described in the above Figure 1 shown embodiment of the data lake construction method based on multi-source distributed data. The specific execution process can be referred to the Figure 1 specific description of the shown embodiment, and details are not described herein again.

[0063] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0064] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.

[0065] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will readily think of other implementation manners of the present disclosure after considering the specification and the disclosure of the practical truth.

[0066] The present application aims to cover any variations, uses, or adaptive changes of the present disclosure that follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A method for constructing a data lake based on multi-source distributed data, characterized in that: The following steps are involved: Deploy a data collection agent at each node of a distributed data source, and establish a distributed collaboration mechanism through the data collection agent; Based on the distributed collaborative mechanism, the data acquisition agent performs structured analysis on the data in the distributed data source to obtain data mode, field type and constraint conditions, and establishes primary and foreign key relationships, reference relationships and business relationships between fields based on the data mode, field type and constraint conditions; Constructing semantic description information including data structure characteristics, field attribute characteristics and relationship characteristics according to the primary and foreign key relationships, the reference relationships and the business relationships; Under the control of the distributed collaboration mechanism, an RDF subject is generated according to the data structure feature, an RDF predicate is generated according to the field attribute feature, and an RDF object is generated according to the relationship feature, so that an RDF triple is formed according to the RDF subject, the RDF predicate and the RDF object, and the RDF triple is stored in a corresponding source node; The RDF triples are input into the ontology registration service under the scheduling of the distributed collaborative mechanism. The ontology registration service classifies the RDF triples according to the data domain division rule, determines the ontology attributes and relationship definitions, and establishes inheritance relationships between ontologies according to the ontology attributes and the relationship definitions to generate a distributed ontology; Analyze the correspondence between the concepts in the RDF triples and the concepts defined in the distributed ontology based on the distributed collaboration mechanism to obtain semantic equivalence rules and mapping rules; According to the semantic equivalence rule and the mapping rule, the distributed data sources are organized into a unified data structure, a unified data view is generated, and a distributed index is established based on the unified data view; Access permissions are configured for the unified data view according to preset permission rules and the distributed index, thereby obtaining a data lake including the unified data view, the distributed index and the access permissions.

2. The data lake construction method according to claim 1, characterized in that: The data acquisition agent performs structured analysis on the data in the distributed data source to obtain data mode, field type and constraint conditions, including: Parsing the data table in the distributed data source to obtain table structure information, and parsing the table name, table description and primary key definition in the table structure information to obtain the data model; Parse the name, data type, length limit, whether it is empty and default value information of the field in the table structure information to obtain the field type; The constraint rules in the table structure information are parsed to obtain the constraint conditions.

3. The data lake construction method according to claim 1, characterized in that: The step of constructing semantic description information including data structure features, field attribute features and relationship features according to the primary and foreign key relationships, the reference relationships and the business relationships includes: Combining the primary key table name, primary key field and foreign key table information referencing the primary key in the primary-foreign key relationship to form a hierarchical structure to obtain the data structure feature, wherein the hierarchical structure is used to represent the inclusion relationship between data tables; Combining the attribute information of the referenced fields in the reference relationship to form an attribute set to obtain the field attribute feature, wherein the attribute set is used to describe the value rule of the field; Combining the associated field information in the business relationship to form an association rule to obtain the relationship feature; The data structure features, the field attribute features and the relationship features are organized and associated according to a preset semantic description format to generate the semantic description information.

4. The data lake construction method according to claim 1, characterized in that: The step of establishing inheritance relationships between ontologies according to the ontology attributes and the relationship definitions to generate a distributed ontology includes: Analyze the characteristic distribution of the ontology attributes, cluster ontologies with the same attribute characteristics, and obtain an ontology category set; Determine the association mode between different ontologies in the ontology category set according to the relationship definition, and establish the inheritance relationship between the ontologies; Building a hierarchical structure of the ontology based on the inheritance relationship, and determining the parent-child relationship, parallel relationship and cross relationship between the ontologies; The ontology category set, the inheritance relationship and the hierarchical structure are combined to generate the distributed ontology.

5. The data lake construction method according to claim 1, characterized in that: The step of analyzing the correspondence between the concepts in the RDF triples and the concepts defined in the distributed ontology based on the distributed collaborative mechanism to obtain semantic equivalence rules and mapping rules includes: Extracting concepts from the RDF triples, and performing semantic feature analysis on each of the concepts to obtain an attribute feature set and a relationship feature set of the concepts; Matching the attribute feature set with the concept attributes defined in the distributed ontology to obtain the semantic similarity between the attributes, and establishing the semantic equivalence rule based on the semantic similarity; The relationship feature set of the concept is matched with the concept relationship defined in the distributed ontology to obtain the structural similarity between the relationships, and the mapping rule is established based on the structural similarity.

6. The data lake construction method according to claim 1, characterized in that: The step of organizing the distributed data sources into a unified data structure and generating a unified data view according to the semantic equivalence rule and the mapping rule includes: Identifying data with the same semantics in the distributed data source according to the semantic equivalence rule, classifying data items with the same semantics into the same data category, and obtaining a set of semantically associated data categories; Determining a conversion method between different data categories in the data category set according to the mapping rule, and converting the data in the data category set into a unified representation form based on the conversion method to obtain a standardized data category set; The semantically related data categories in the standardized data category set are constructed into a hierarchical data structure according to preset organizational rules to obtain a unified data organizational structure, and a data access interface and data operation method are established based on the unified data organizational structure to generate the unified data view.

7. The data lake construction method according to claim 1, characterized in that: The establishing of a distributed index based on the unified data view includes: Scan the structural definition of the unified data view to obtain a classification hierarchy of data categories, and identify storage node information and storage path information of each data category according to the classification hierarchy to generate a location index table; Based on the location index table, statistics are collected on the access frequency and access time distribution of each data category in the unified data view to establish a data heat index table, wherein the data heat index table records the access statistics information and cache allocation strategy of the data category; According to the semantic association relationship recorded in the unified data view, in combination with the location index table and the data heat index table, a semantic index table is established, wherein the semantic index table contains semantic mapping relationships between data categories and cross-node access paths; The storage location information in the location index table, the cache strategy in the data heat index table, and the semantic mapping relationship in the semantic index table are associated and organized to generate the distributed index.

8. A data lake construction system based on multi-source distributed data, characterized in that: include: Distributed collaboration module, structured analysis module, semantic description module, RDF generation module, ontology generation module, semantic mapping module, data view generation module and authority control module; The distributed collaboration module is used to deploy data acquisition agents at each node of the distributed data source and establish a distributed collaboration mechanism through the data acquisition agents; The structured analysis module is used to perform structured analysis on the data in the distributed data source through the data acquisition agent based on the distributed collaborative mechanism to obtain data mode, field type and constraint conditions, and establish primary and foreign key relationships, reference relationships and business relationships between fields based on the data mode, field type and constraint conditions; The semantic description module is used to construct semantic description information including data structure characteristics, field attribute characteristics and relationship characteristics according to the primary and foreign key relationships, the reference relationships and the business relationships; The RDF generation module is used to generate an RDF subject according to the data structure characteristics, generate an RDF predicate according to the field attribute characteristics, and generate an RDF object according to the relationship characteristics under the control of the distributed collaboration mechanism, so as to form an RDF triple according to the RDF subject, the RDF predicate and the RDF object, and store the RDF triple in a corresponding source node; The ontology generation module is used to input the RDF triples into the ontology registration service under the scheduling of the distributed collaboration mechanism, and the ontology registration service classifies the RDF triples according to the data domain division rules, determines the ontology attributes and relationship definitions, and establishes the inheritance relationship between the ontologies according to the ontology attributes and the relationship definitions to generate a distributed ontology; The semantic mapping module is used to analyze the correspondence between the concepts in the RDF triples and the concepts defined in the distributed ontology based on the distributed collaboration mechanism to obtain semantic equivalence rules and mapping rules; The data view generation module is used to organize the distributed data sources into a unified data structure according to the semantic equivalence rule and the mapping rule, generate a unified data view, and establish a distributed index based on the unified data view; The permission control module is used to configure access permissions for the unified data view according to preset permission rules and the distributed index, so as to obtain a data lake including the unified data view, the distributed index and the access permissions.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: It includes a processor, a memory and a transceiver, the memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Data lake and warehouse integrated management method and system based on port traffic

    CN120832388A

  • Semiconductor level data merging method, device and equipment based on MergeMap structure

    CN121560894A

  • Method, apparatus and device for semiconductor hierarchical data merging based on merge map structure

    CN121560894B