Automated Method and Apparatus for Converting Relational Data into Property Graphs

By inputting relational data into large language models to generate initial ontology sets and using semantic transformation grammar to construct attribute diagrams, the problem of difficulty in capturing implicit association relationships in biological entities in the prior art is solved, and a more comprehensive data relationship description and personalized query are achieved.

CN119782405BActive Publication Date: 2025-07-22TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510272237.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-22
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture potential implicit associations between biological entities, and attribute graphs rely on expert definitions and are difficult to fully manage complex information.

Method used

By inputting relational data into large language models, the initial ontology collection is generated using weighted scores and semantic transformation grammar, the target ontology collection is constructed and mapped, the relational algebra is generated, and the attribute graph is finally constructed.

Benefits of technology

It realizes a comprehensive and accurate description of biological entity relationships, meets users' personalized query needs, and improves the comprehensiveness and accuracy of data relationships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119782405B_ABST
    Figure CN119782405B_ABST
Patent Text Reader

Abstract

The present invention provides an automated method and apparatus for converting relational data into an attribute graph, which can be applied to the fields of general relational data and artificial intelligence technologies. The method includes: inputting the relational schema of the data in the relational data and the data instances corresponding to the relational schema into a large language model to obtain an initial ontology set, where the initial ontology set includes multiple initial entities and the association relationships between the multiple initial entities, and the initial entities include at least one initial attribute data extracted from the attribute sampling data; updating the initial ontology set according to the weighted scores of the respective initial attribute data in the initial entities to obtain a target ontology set, where the weighted scores are determined according to the user query frequency and information entropy of the initial attribute data; using semantic transformation grammar to map the target ontology set to obtain the relationships of the target ontology set; generating a set of relational algebras based on the mapping relationships between the target ontology set and the relational schema, and constructing an attribute graph for the relational data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of general relational data and artificial intelligence, and particularly to an automated method and device for converting relational data into an attribute graph. Background Art

[0002] An attribute graph is a graphical structure based on entities and relationships. General relational data is represented and managed by an attribute graph to represent biological entities (such as viruses, bacteria, etc.) and the relationships between biological entities (such as the relationship between a virus and a disease), so as to solve complex information problems in the field of biosafety.

[0003] Attribute graphs usually rely on expert definitions or simple constraints and are difficult to capture all relationships in the data. For example, in relational data, entity relationships are usually defined by foreign key constraints. This method can handle explicit relationships, but it is difficult to discover potential implicit associations that may exist between data. Summary of the Invention

[0004] In view of the above problems, the present invention provides an automated method and device for converting relational data into an attribute graph.

[0005] According to a first aspect of the present invention, there is provided an automated method for converting relational data into an attribute graph, including: inputting the relationship schema of the data in the relational data and the data instances corresponding to the relationship schema into a large language model to obtain an initial ontology set, the initial ontology set including a plurality of initial entities and the association relationships between the plurality of initial entities, the initial entities including at least one initial attribute data extracted from the attribute sampling data, and the attribute sampling data being sampled and extracted from the data instances; updating the initial ontology set according to the weighted scores of the respective initial attribute data in the initial entities to obtain a target ontology set, the weighted scores being determined according to the user query frequency and information entropy of the initial attribute data; mapping the target ontology set using a semantic transformation grammar to obtain the relational algebra of the target ontology set, the relational algebra being used to retrieve the relational data; and constructing an attribute graph based on the target ontology set and the relational algebra.

[0006] According to an embodiment of the present invention, updating the initial ontology set according to the weighted scores of the respective initial attribute data in the entities to obtain a target ontology set includes: for each initial entity, calculating the weighted scores of the initial attribute data in the initial entity; in the case where the weighted scores meet a preset threshold, determining the initial attribute data as a target entity; and updating the initial ontology set according to the target entity to obtain a target ontology set.

[0007] According to an embodiment of the present invention, calculating the weighted score of the initial attribute data in each initial entity includes: for each initial entity, calculating the query score of the initial attribute data according to the user query frequency; calculating the information score of the initial attribute data according to the information entropy of the initial attribute data; performing a weighted sum of the query score and the information score to obtain the weighted score.

[0008] According to an embodiment of the present invention, updating the initial ontology set according to the target entity to obtain the target ontology set includes: extracting the target attribute data of the target entity from the initial attribute data in the initial entity; using a large language model to determine the target association relationship between the target entity and multiple initial entities in the initial ontology set; adding the target attribute data and the target association relationship of the target entity to the initial ontology set to obtain the target ontology set.

[0009] According to an embodiment of the present invention, the relational algebra includes projection relational algebra, and mapping the target ontology set using the semantic transformation grammar to obtain the relational algebra of the target ontology set includes: for the target attribute data in the target entity coming from the same relational schema, performing a projection operation on the target attribute data to obtain the projection relational algebra; for the target attribute data in the target entity coming from different relational schemas, establishing query mappings for different relational schemas.

[0010] According to an embodiment of the present invention, the relational algebra further includes full outer join relational algebra and inner join relational algebra. Establishing query mappings for different relational schemas for the target attribute data in the target entity coming from different relational schemas includes: for the target attribute data in the target entity coming from different relational schemas, performing a full outer join operation on the target attribute data to obtain the full outer join relational algebra; performing an inner join between different relational schemas to obtain the inner join relational algebra.

[0011] According to an embodiment of the present invention, constructing an attribute graph based on the target ontology set and the relational algebra includes: constructing an attribute graph based on the target ontology set, the projection relational algebra, the full outer join relational algebra, and the inner join relational algebra, where the entity and the target entity are the nodes of the attribute graph, and the connection edges between the nodes represent the association relationships between multiple entities or between an entity and the target entity.

[0012] According to an embodiment of the present invention, the above method further includes: performing data sampling from the relational data according to a data sampling strategy to obtain attribute sampling data under multiple relational schemas, and the data collection strategy is determined according to the data value range in the relational data.

[0013] According to an embodiment of the present invention, data sampling is performed on relational data according to a data sampling strategy to obtain attribute sampling data under multiple relational patterns, including: dividing a data set in the relational data into multiple subsets according to a data value range, where the data value ranges represented by the multiple subsets are different; performing data sampling on the multiple subsets of the relational data respectively to obtain attribute sampling data under the relational pattern.

[0014] The second aspect of the present invention provides an automated device for converting relational data into an attribute graph, including: an input module, configured to input the relational pattern of the data in the relational data and the data instances corresponding to the relational pattern into a large language model to obtain an initial ontology set. The initial ontology set includes multiple initial entities and the association relationships between the multiple initial entities. The initial entities include at least one initial attribute data extracted from the attribute sampling data, and the attribute sampling data is sampled and extracted from the data instances; an update module, configured to update the initial ontology set according to the weighted scores of the respective initial attribute data in the initial entities to obtain a target ontology set, where the weighted scores are determined according to the user query frequency and information entropy of the initial attribute data; a mapping module, configured to map the target ontology set using semantic transformation grammar to obtain the relational algebra of the target ontology set, and the relational algebra is used to retrieve the relational data; a construction module, configured to construct an attribute graph based on the target ontology set and the relational algebra.

[0015] The third aspect of the present invention provides an electronic device, including: one or more processors; a memory, configured to store one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above method.

[0016] The fourth aspect of the present invention further provides a computer-readable storage medium, on which an executable computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0017] The fifth aspect of the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0018] According to the automated method and device for converting relational data into an attribute graph provided by the present invention, by inputting the relational pattern of the data in the relational data and the data instances corresponding to the relational pattern into a large language model to obtain an initial ontology set, the potential associations existing between the data in the data instances can be analyzed using the large language model, making the initial text set more accurate. The initial ontology set is updated according to the weighted scores of the respective initial attribute data in the initial entities to obtain a target ontology set. Since the weighted scores are determined according to the user query frequency and information entropy of the attribute data, the personalized query needs of users can be satisfied.

[0019] Map the target ontology set using semantic transformation grammar to obtain the relational algebra of the target ontology set; construct an attribute graph based on the target ontology set and the relational algebra, so that the data relationships presented by the constructed attribute graph are more comprehensive and accurate, and can also meet the needs of users' personalized queries. Brief Description of the Drawings

[0020] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0021] Figure 1 Shows an application scenario diagram of an automated method for converting relational data into an attribute graph according to an embodiment of the present invention;

[0022] Figure 2 Shows a flowchart of an automated method for converting relational data into an attribute graph according to an embodiment of the present invention;

[0023] Figure 3 Shows an output schematic diagram of a large language model according to an embodiment of the present invention;

[0024] Figure 4 Shows a schematic diagram of updating the initial ontology set according to the weighted scores of each initial attribute data in the initial entity to obtain the target ontology set;

[0025] Figure 5 Shows a schematic diagram of calculating the weighted scores of the initial attribute data in the initial entity for each initial entity according to an embodiment of the present invention;

[0026] Figure 6 Shows a schematic diagram of an automated method for converting relational data into an attribute graph according to an embodiment of the present invention;

[0027] Figure 7 Shows a structural block diagram of an automated device for converting relational data into an attribute graph according to an embodiment of the present invention;

[0028] Figure 8 Shows a block diagram of an electronic device suitable for implementing an automated method for converting relational data into an attribute graph according to an embodiment of the present invention. Detailed Description of the Embodiments

[0029] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present invention. However, evidently, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.

[0030] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "comprising", "including" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0031] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having meanings consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0032] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0033] In the technical solution of the present invention, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, invention, and application, all comply with relevant laws, regulations and standards, adopt necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0034] In the scenario of making automated decisions using personal information, the method, device, and system provided by the embodiments of the present invention all provide corresponding operation entrances for users to choose to agree or reject the results of automated decisions; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision" here refers to the activity of automatically analyzing and evaluating an individual's behavior habits, hobbies, or economic, health, credit status, etc. through a computer program and making decisions. The expression "expert decision" here refers to the activity of making decisions by personnel who are engaged in work in a specific field, have specialized experience, knowledge, and skills, and have reached a certain professional level.

[0035] Embodiments of the present invention provide an automated method for converting relational data into an attribute graph, including: inputting the relationship pattern of the data in the relational data and the data instances corresponding to the relationship pattern into a large language model to obtain an initial ontology set. The initial ontology set includes multiple initial entities and the association relationships between the multiple initial entities. The initial entities include at least one initial attribute data extracted from the attribute sampling data, and the attribute sampling data is sampled and extracted from the data instances; updating the initial ontology set according to the weighted scores of the respective initial attribute data in the initial entities to obtain a target ontology set, where the weighted scores are determined according to the user query frequency and information entropy of the initial attribute data; using semantic transformation grammar to map the target ontology set to obtain the relational algebra of the target ontology set, and the relational algebra is used to retrieve the relational data; constructing an attribute graph based on the target ontology set and the relational algebra.

[0036] Figure 1 The application scenario diagram of the automated method for converting relational data into an attribute graph according to the embodiments of the present invention is shown.

[0037] As Figure 1 shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0038] Users can use at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for examples).

[0039] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0040] The server 105 may be a server that provides various services, such as a background management server (only as an example) that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0041] It should be noted that the automated method for converting relational data into an attribute graph provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the automated device for converting relational data into an attribute graph provided by the embodiments of the present invention can generally be set in the server 105. The automated method for converting relational data into an attribute graph provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the automated device for converting relational data into an attribute graph provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0042] It should be understood, Figure 1 the number of terminal devices, networks, and servers in

[0043] Figure 2 shows a flowchart of an automated method for converting relational data into an attribute graph according to an embodiment of the present invention.

[0044] As Figure 2 shown, the automated method for converting relational data into an attribute graph in this embodiment includes operation S210 to operation S240.

[0045] In operation S210, the relationship pattern of the data in the relational data and the data instances corresponding to the relationship pattern are input into the large language model to obtain an initial ontology set.

[0046] According to an embodiment of the present invention, a relational schema can be a basic concept for describing a data structure in relational data. The relational schema can define the name of a relational table, the sampled attribute data, and its attribute type. The name of the relational table is usually the name of the data being described.

[0047] According to an embodiment of the present invention, the relational schema includes at least one attribute data for describing data, and the sampled attribute data is sampled and extracted from data instances.

[0048] Define the relational schema R of the data in the relational data j as ( ). are different sampled attribute data respectively. For example, the data can be virus transmission information in biosecurity. At this time, the relational schema of the virus transmission information involves (pathogen, host, transmission route, prevention and control measures).

[0049] The data instances can be biosecurity-related literature, reports, etc. stored in the relational data. Biological experts construct the types of entities and the attributes contained under the entities according to experience. The sampled attribute data can include expert experience attribute data and also the attribute data obtained by extracting attribute features from the data instances.

[0050] For example, applying a large language model as a classifier to an ontology generation tool can support and improve the process of obtaining and representing biological knowledge in data instances.

[0051] Figure 3 Shows a schematic diagram of the output of the large language model according to an embodiment of the present invention.

[0052] The large language model can automatically extract ontology information from the relational schema and data instances. This process includes identifying the entity types in the relational schema and understanding the semantics of the relationships connecting multiple entities. The data instances can be text summaries of the corresponding relational schema. The output of the large language model is to identify the relevant entity types and their attributes, the relationships between entity types, and the relationships between entity instances in the data instances and the sampled attribute data. The entity type can characterize an entity.

[0053] According to an embodiment of the present invention, the initial ontology set includes multiple initial entities and the association relationships between multiple initial entities. The initial entities include at least one initial attribute data extracted from the sampled attribute data.

[0054] For example, for entity type extraction in a large language model, the process is described as follows: For each relational schema R j and the attribute A i in it, the large language model will output a relevant entity . Where A ijis a specific attribute name. Therefore, in the relational schema R j for the large language model, all attributes of the entity V oi in this schema . When the attributes of entity V oi` are scattered in multiple schemas (such as R j and R k ), the large language model uses foreign key dependencies to merge attributes to form the same entity. Specifically, according to the foreign key definition in the relational data, for a given foreign key in the attributes of different relational tables in the database, it represents that the entity associated with A i is the same as the entity associated with A j , that is, A i , A j describe the same entity information. Therefore, the later traversed attribute A j can be merged into the entity to which A i belongs. In addition, construct an attribute tuple, denoted as , and update the relevant attributes and entities in the attribute tuple. As more attributes of entity 𝑣 are identified through the entity recognition process, continuously update the attribute tuple to include the newly identified attributes. In addition, the large language model will also standardize the naming format of the data.

[0055] For entity types in the same relational schema R j , the large language model identifies the relationships between pairs of entities, denoted as v i and v j . The relationship between entities is represented as an edge connecting v i and v j , and is denoted as e=(v i, v j ). The large language model can interpret the semantics of the relationships between entities. Entities are presented as nodes in the attribute graph, and the relationships between entities are presented as edges between nodes. The direction of the edges between nodes can be determined by the large language model or based on the specific requirements of downstream tasks.

[0056] The prompts of the large language model are organized into three parts: Instruction, Input, and Output. Instruction elaborates the goal, Input includes attribute sampling data, data instances, and related entities, and Output details the relationships among the above three. For the training of the large language model, hundreds of ontologies are selectively extracted for training, specifically designed for extracting ontologies, including attribute sampling data related to features, entity names, relationships between entities, and specific data instances. The structure of the prompt words is as Figure 3As shown in the input, the prompt helps to maintain the consistency label of the relationship data according to the trained ontology, improving the overall quality and uniformity of the extracted data.

[0057] The graph construction method of the embodiments of the present invention is not limited to the biosafety direction, and can also be graph construction in other directions. Figure 3 The input prompt words and output templates are provided. After inputting the relationship pattern and its corresponding data instances, the large language model will output entity types and attribute information. For example: import other relationship data into the output template, including the relationship pattern (student name, student ID, instructor name) and related data instances into the large language model, and then output the belonging entities and the relationships among them, which can output: Entity 1: student, with attributes including student name and student ID; Entity 2: instructor, with attributes including instructor name; There is a teacher-student relationship between the student and the instructor, and the semantics of this relationship can be automatically generated in the large language model.

[0058] In operation S220, update the initial ontology set according to the weighted scores of the respective initial attribute data in the initial entity to obtain the target ontology set.

[0059] According to an embodiment of the present invention, the weighted score is determined according to the user query frequency and information entropy of the initial attribute data.

[0060] According to an embodiment of the present invention, the user query frequency represents the number of times the user queries the initial attribute data in the relationship data within a certain time period.

[0061] According to an embodiment of the present invention, information entropy is used to evaluate the frequency of occurrence of attribute values. In the case of less diversity of attribute values, information entropy can be used to evaluate the initial attribute data with more query value.

[0062] For example, the initial entity is bacteria, and the initial attribute data of bacteria includes decomposition, symbiotic relationship, diseases caused, and type. The weighted score of the symbiotic relationship in the initial attribute data meets the threshold for promoting from attribute to entity, and the symbiotic relationship in the initial attribute data can be determined as the target entity in the target ontology set. Then analyze the semantic relationship between the symbiotic relationship and other target entities to determine the relationship between the target entities.

[0063] Figure 4 Shows a schematic diagram of updating the initial ontology set according to the weighted scores of the respective initial attribute data in the initial entity to obtain the target ontology set.

[0064] Although the initial attribute data can be regarded as the attributes of the initial entity in the relationship, the initial attribute data can also be regarded as entities themselves.

[0065] For example, in Figure 4Among them, (a) the database fragment includes data such as "name", "student ID", and "address". (b) The result before optimization only presents the corresponding data.

[0066] The "address" attribute can be regarded as either an attribute of the initial entity or an entity. It is worth noting that regarding certain attributes as entities can improve performance and shorten query response time in some query-intensive applications.

[0067] For example, consider Figure 4 In (c) where the user query extracts information about Location A, aiming to query the database log to return all customers with the address of "Location A". If the "address" attribute is regarded as an entity, then the (d) optimized result can be found by checking the "Location A" entity, that is, presenting all data related to "Location A", and not querying the data related to "Location B".

[0068] However, if "Location A" is not regarded as an entity, then all customer entities need to be searched and their attributes checked to find the required result. The target ontology set can include the entity "Location A", and the target attribute data includes "name" and "student ID".

[0069] To decide which initial attribute data should be promoted to entities, a scoring mechanism called "benefit gain" is introduced. This scoring combines two aspects: user query frequency and information entropy. Initial attribute data that appears infrequently but frequently in historical queries will obtain a higher weighted score and is more likely to be promoted to independent entities.

[0070] In operation S230, the target ontology set is mapped using semantic transformation grammar to obtain the relational algebra of the target ontology set.

[0071] According to an embodiment of the present invention, relational algebra is used to retrieve relational data.

[0072] According to an embodiment of the present invention, the semantic transformation grammar can be STG (Semantic Tree Grammar), and STG can map the target ontology set to a set of algebraic queries without losing any information, and can map the target ontology set to a set of relational algebras (or algebraic queries) without losing any information. Once a set of relational algebras is generated, mature relational query optimization and execution techniques are used to execute these queries. Then the query results are used to generate an attribute graph.

[0073] In operation S240, an attribute graph is constructed based on the target ontology set and relational algebra.

[0074] According to an embodiment of the present invention, nodes in the property graph represent target entities in the target ontology set, and the edges between the nodes represent the relationships between the target entities.

[0075] According to an embodiment of the present invention, by inputting the relationship patterns of the data in the relationship data and the data instances corresponding to the relationship patterns into a large language model, an initial ontology set is obtained. Therefore, the large language model can analyze the potential associations existing between the data in the data instances, making the initial text set more accurate. The initial ontology set is updated according to the weighted scores of the respective initial attribute data in the initial entities to obtain a target ontology set. Since the weighted scores are determined based on the user query frequency and information entropy of the attribute data, the personalized query needs of users can be satisfied. The target ontology set is mapped using semantic transformation grammar to obtain the relational algebra of the target ontology set; an attribute graph is constructed based on the target ontology set and the relational algebra. Therefore, the data relationships presented by the constructed attribute graph are more comprehensive and accurate, and can also meet the needs of users' personalized queries.

[0076] The construction of the attribute graph usually relies on expert definitions or simple constraints, which may not be able to fully capture all the relationships in the data. The large language model is trained using a wide range of data sources, covering multiple fields and topics, and can understand the semantics of the data and the complex relationships between data points. The large language model has a more comprehensive and accurate understanding ability, can automatically identify relevant points in the data, and capture the potential relationships between entities, including implicit and non-explicit associations. This makes it possible to describe entities and their relationships in the data more comprehensively, thereby generating a more accurate and complete ontology.

[0077] An intuitive method for generating an ontology from relationship data is to input the relationship pattern into a large language model, and then the large language model can infer the ontology existing in the relationship pattern. However, this method is limited due to the lack of specific data support for the relationship pattern itself, making it difficult for the large language model to fully understand the complex relationships and changes in real-world data. In addition, without a detailed description of specific data, it may be difficult for the language model to understand various situations and anomalies that may occur in the actual data. If there are errors or incompleteness in the relationship pattern itself, it may also have a negative impact on the accuracy of the output of the large language model.

[0078] To address these limitations, the relational schema and data instances can be input into a large language model. However, processing a large number of data instances may result in high computational costs, and much of the data may be redundant for ontology generation, thereby reducing the efficiency and accuracy of the ontology construction process. To overcome this challenge, representative data samples are selected or key features are extracted to create a compressed data summary. The data summary represents the key information extracted from the original data, capturing the key data features and relational patterns, making it possible to better understand the relationships in real-world data. By analyzing and processing this data summary, a balance can be achieved between computational resource utilization and ontology construction accuracy. Therefore, selecting an appropriate data summary becomes a key issue to be addressed in this context. The present invention can first perform text processing such as data instance cleaning and speech parsing to obtain a key data summary. Then, the data summary and its corresponding relational schema are input into the large language model to obtain an initial ontology set.

[0079] According to an embodiment of the present invention, the above method further includes: performing data sampling from the relational data according to a data sampling strategy to obtain attribute sampling data under multiple relational schemas, and the data sampling strategy is determined according to the data value range in the relational data.

[0080] The attribute sampling data under multiple relational schemas can cover all data value ranges and can clearly show the structural characteristics of the data values. The attribute sampling data can cover information about data formats, data ranges, and data dependencies to enhance the understanding of the attribute sampling data.

[0081] For example, being between 16 and 18 years old may be a student. Being small in volume and non-cellular in structure may be a virus. However, due to the input limitations of the large language model, it is not feasible to directly input all the data. Therefore, the most representative attribute sampling data is input after sampling to optimize the use of the available input space.

[0082] To aggregate the data, a data sampling strategy is designed, and the data sampling strategy is determined according to the data value range in the relational data.

[0083] According to an embodiment of the present invention, performing data sampling from the relational data according to a data sampling strategy to obtain attribute sampling data under multiple relational schemas includes: dividing the data set in the relational data into multiple subsets according to the data value range, and the data value ranges represented by the multiple subsets are different; performing data sampling from the multiple subsets of the relational data respectively to obtain the attribute sampling data under the relational schema.

[0084] For example, the number of instances is 100, but the range of attribute data values can be from 1 to 100. The dataset can be divided into multiple subsets with different ranges of data values according to the data value range. That is, the range of attribute data values of the first subset is from 1 to 20, the range of attribute data values of the second subset is from 20 to 40... The range of attribute data values of the fifth subset is from 80 to 100, ensuring that the dataset used to train the large language model is both diverse and representative.

[0085] For example, if the number of instances is 100, the dataset can be divided into 10 subsets. Then, data is randomly sampled from each subset. The effectiveness of our sampling method is evaluated by examining the characteristics of the sampled data. This is achieved by verifying whether the value range of the sampled data comprehensively covers the entire data value range. If the value range is incomplete, it indicates a lack of characteristics and adjustments are made based on the initial sampling results. These adjustments may include dynamically changing the sampling rate or adjusting the sampling area, taking into account the characteristics of the obtained samples. Emphasis is placed on enhancing data diversity while eliminating redundancy. This is achieved by removing duplicate or redundant data points from the sampled data. Subsequently, resampling is performed from the refined data pool to supplement the initial samples, ensuring that the dataset used to train the large language model is both diverse and representative.

[0086] Although large language models can understand the relationships in data to a certain extent, they still have limitations. For example, when dealing with professional knowledge or complex relationships in a specific domain, large models may encounter biases or incompleteness. To overcome these limitations, multiple attribute graph ontologies constructed by domain experts can be utilized. These ontologies contain rich entity and relationship information. By combining these ontologies with the large language model, the model's domain knowledge can be further enhanced, achieving more accurate ontology construction.

[0087] The enhanced use of the ontology integration strategy for large language models enriches the training data together with multiple ontologies, mainly focusing on the DBpedia ontology (a structured framework for describing and organizing knowledge), while also integrating other ontology resources. More specifically, to integrate multiple ontology data sources, an ontology merging method is adopted. This process is based on the ontology knowledge base, with the DBpedia ontology as the underlying structure. The result is a new ontology structure , for the entity type information and relationship information between entities in the new ontology set. During the merging process, the knowledge of multiple ontologies is transferred to the newly created ontology structure. To achieve this, different information needs to be aligned according to the specified common format . After alignment, the knowledge content of each individual ontology is merged and integrated into the newly established ontology structure. The output of this merging process is a unified ontology , which is subsequently used to enhance and improve the functionality of the large language model.

[0088] According to an embodiment of the present invention, updating the initial ontology set according to the weighted scores of the respective initial attribute data in the entity to obtain a target ontology set includes: for each initial entity, calculating the weighted score of the initial attribute data in the initial entity; in the case where the weighted score meets a preset threshold, determining the initial attribute data as a target entity; and updating the initial ontology set according to the target entity to obtain the target ontology set.

[0089] According to an embodiment of the present invention, sorting the weighted scores of the initial attribute data in the initial entity, and setting a preset threshold with the weighted score at a preset position. That is, determining the initial attribute data before the preset position in the ranking of the weighted scores as the target entity. For example, taking the first k initial attribute data with the highest weighted scores as the target entity, where k is an integer greater than 1.

[0090] The preset threshold can be adjusted according to the actual situation and is not limited herein.

[0091] According to an embodiment of the present invention, calculating the weighted score of the initial attribute data in the initial entity for each initial entity includes: for each initial entity, calculating the query score of the initial attribute data according to the user query frequency; calculating the information score of the initial attribute data according to the information entropy of the initial attribute data; and performing a weighted sum on the query score and the information score to obtain the weighted score.

[0092] Given an attribute A i , the weighted score is defined as follows:

[0093] (1);

[0094] where 𝛼 is a weight parameter in the range of [0, 1], is defined as the information score of A i and is calculated as:

[0095] (2);

[0096] p(x i ) represents the probability of A i when the value is x i . Query(A i ) is the query score of A i , and n represents the number of initial attribute data and is determined by normalizing the frequency of its appearance in the historical query log.

[0097] Figure 5 FIG. shows a schematic diagram of calculating the weighted score of the initial attribute data in the initial entity for each initial entity according to an embodiment of the present invention.

[0098] Given a relational table as Figure 5 shown, the attribute sampling data included in the relational table are "name", "address", and "other content".

[0099] The large language model (such as the classifier in Figure 5 ) outputs an initial ontology set. For example, the initial ontology set includes the entity "customer", and the initial attribute data are "name", "address", and "other content".

[0100] By querying the log, the query scores corresponding to the three initial attribute data are obtained. Through sampling, the information scores of the initial attribute data can be calculated according to the information entropy of the three initial attribute data. The weights of the query score and the information score are both 0.5, and the weighted score of the initial attribute data "address" (the total score in Figure 5 ) can be calculated to be the highest among the three weighted scores, and the initial attribute data "address" can be selected as the target entity. Then, the large language model is used to determine the target association relationship between the target entity and multiple initial entities in the initial ontology set.

[0101] According to an embodiment of the present invention, the initial ontology set is updated according to the target entity, and the obtained target ontology set includes: extracting the target attribute data of the target entity from the initial attribute data in the initial entity; using the large language model to determine the target association relationship between the target entity and multiple initial entities in the initial ontology set; adding the target attribute data and the target association relationship of the target entity to the initial ontology set to obtain the target ontology set.

[0102] According to an embodiment of the present invention, the large language model can semantically analyze the initial attribute data related to the target entity in the data instance, and determine the initial attribute data related to the target entity as the target attribute data.

[0103] For example, starting from the initialization Set := , then using the large language model to perform ontology extraction on the relationship pattern of the relational data and the instance text corresponding to the relationship pattern, and updating the output entities and the relationships between the entities to Set.

[0104] More specifically, for each relationship pattern R i and its data instance, the ontology generation tool applies the large language model to generate entities. Then, the entities and attributes are updated to the initial ontology Set to obtain the initial ontology set. If the entity is already part of the ontology, the ontology generation tool will update the ontology with the attribute A; if the entity is not yet in the ontology (i.e., a new entity), the entity will be added to the ontology 𝑂, and the initial ontology set will be updated and the corresponding attributes will be inserted.

[0105] Next, the relationships generated by the large language model are updated into the new initial ontology set to obtain the target ontology set. More specifically, for entities generated under the same R i (i.e., belonging to the same relationship pattern), the ontology generation tool uses the large language model to generate a new relationship E. Next, the gain score of each attribute A is calculated. Then, the attribute A with a high weighted score is extracted j , and A j is regarded as the target entity V Aj . Finally, V Aj and its new relationship are updated into the initial ontology set Set.

[0106] The specific steps include: initializing the ontology O as Set = (empty set); for each relationship pattern R i , using the large language model to perform ontology extraction on the relationship pattern to obtain the initial entity Entity i , and integrating to generate a set of relationship entity pairs {<A j, Entity i >}.

[0107] For each attribute entity pair <A j, Entity i >, if the initial entity Entity i already exists in the ontology O, use Entity i to update Entity i , and update and add A j to the attribute group of Entity i in O; otherwise, add Entity i to the ontology O. At the same time, for each entity pair (Entity i , Entity i ) generated under R j , use the large language model to generate the relationship between the entity pair (Entity i , Entity j ) to obtain the target ontology set.

[0108] According to the embodiments of the present invention, the relational algebra includes projection relational algebra. Mapping the target ontology set using the semantic transformation grammar, the relational algebra of the target ontology set includes: performing a projection operation on the target attribute data for the target attribute data of the target entity from the same relationship pattern to obtain the projection relational algebra; establishing a query mapping of different relationship patterns for the target attribute data of the target entity from different relationship patterns.

[0109] According to an embodiment of the present invention, the relational algebra further includes a full outer join relational algebra and an inner join relational algebra. For the target attribute data in the target entity coming from different relational schemas, establishing query mappings for different relational schemas includes: for the target attribute data in the target entity coming from different relational schemas, performing a full outer join operation on the target attribute data to obtain a full outer join relational algebra; performing an inner join between different relational schemas to obtain an inner join relational algebra.

[0110] Given a relational schema and an ontology , Vo represents an entity, and E O represents an entity relationship. Define a set of mappings from R to O, and this mapping includes:

[0111] Attribute tuples: For each entity v in the ontology, there exists a set of associated attribute tuples x,..., y, and this tuple is defined as the mapped attribute of entity v, denoted as $v, and the mapped attributes can be expressed as $v.x,...,$v.y.

[0112] Entity generation rule: For each entity in the ontology, there exists a mapping rule from an attribute A in R to the attribute of entity v, denoted as .

[0113] Relationship generation rule. For each relationship between entities in the ontology, for the given entity generation rule and , there exists a mapping rule from an attribute pair (A, B) in R to the entity relationship (u, v).

[0114] In terms of implementation, a set of algebraic queries is used to interpret the STG. For each entity generation rule, there exists an algebraic query Q v to specify how to generate an instance of the entity, where Q is defined as follows: For attributes A1,..., An in the relational schema R i , at this time , for those from different relational schemas .

[0115] (3);

[0116] where represents the full outer join symbol.

[0117] For each relationship generation rule (representing the mapping rule from the attribute pair (A, B) to the entity relationship (u, v)) there exists an algebraic query , representing the inner join of attributes A and B in .

[0118] According to an embodiment of the present invention, constructing an attribute graph based on a target ontology set and relational algebra includes: constructing an attribute graph based on the target ontology set, projection relational algebra, full outer join relational algebra, and inner join relational algebra, where entities and target entities are nodes of the attribute graph, and the connection edges between the nodes represent the association relationships between multiple entities or between an entity and a target entity.

[0119] STG maps the ontology to a set of algebraic queries. The generated relational algebra queries can generate an attribute graph G, where nodes create attributes through queries defined by rules for entities, and edges are generated through queries defined by rules for relationships. The rules of STG can be obtained during the ontology generation process, and the algebraic queries can be obtained according to the semantic definition of STG.

[0120] In STG, the relationship generation rule determines the relationship between entities based on the relationships between subsets of attributes of different entities, but only when these subsets exist in a single relationship. If the attributes of two entities do not co-occur, it means there is no direct relationship between the two entities. Indirect relationships are not considered in the current stage, but can be incorporated into the generated graph through optimization techniques.

[0121] The STG transformation process does not result in any information loss. This is proven by checking whether the original database can be reconstructed from the generated graph through reverse transformation.

[0122] If a transformation from a relational schema R to an attribute graph G satisfies the following conditions, then it is considered lossless: First, the transformation is injective, meaning that each element in the original database can be uniquely mapped to an element in the generated graph. Second, the transformation is surjective, meaning that each element in the generated graph has a corresponding element in the original database.

[0123] By satisfying these conditions, we can ensure that the STG transformation process retains all the information present in the original database, thus allowing the database to be fully reconstructed from the generated graph.

[0124] The STG transformation has the property of being lossless. The lossless property is proven as follows: (1) Injectivity. Injectivity is obvious. During the ontology construction process, each attribute of the relational schema R corresponds to a unique entity type in O. This ensures that each element (attribute) in the relational instance is mapped to a single attribute of the corresponding entity instance in the attribute graph G. (2) Surjectivity. To prove surjectivity, show that the relational instance I' obtained from a set of mappings from R to O is equivalent to the original relational instance I. Here, equivalence means that the set of tuples in I and I' is the same, regardless of the order of the tuples. The construction process of I' is as follows. For each entity generation rule and the attribute tuples to which the entity belongs, create a table by inductively grouping all entity groups of the same entity. Then, combine or create new relational tables for each entity table according to the relational generation rules. At this time, for the newly constructed I', it is necessary to prove that I can be transformed into I' through a finite number of conversions. Define an inverse mapping from O to R and reconstruct it based on the id information in the relational data. It can be seen that essentially through the number of rows in the table and the id information, the reconstruction of I to I' can be achieved (filling in the number of rows in the original table and the id information hidden in the entities in I' into I). Essentially, since the relational table is bounded, I can be transformed into I' in a finite number of steps. This proves that the STG transformation is lossless, ensuring that the original database can be fully reconstructed from the generated graph.

[0125] Therefore, it is proven that STG provides a powerful method to maintain data integrity and consistency without losing information during the process of converting relational data into graph data. This is particularly important for application scenarios that require ensuring data quality and reliability.

[0126] In summary, STG provides an effective way to convert complex relational data models into graph structures by defining clear syntax and semantic rules, as well as through the guarantee of lossless transformation properties, enabling richer and more intuitive expressions for data analysis and processing in the graph domain.

[0127] In practical applications, STG can not only ensure the losslessness of data conversion but also simplify and automate the graph construction process through automatically generated algebraic queries. Once STG is defined, the query optimization and execution engines of existing relational data management systems can be used to execute these queries, effectively constructing the attribute graph.

[0128] In addition, the STG framework also provides a certain degree of flexibility, allowing users to adjust and optimize the generated graph structure according to specific requirements. For example, by modifying the entity and relationship generation rules, new node types or edge types can be easily added, or the existing structure can be adjusted to better reflect the semantics of the data.

[0129] After the construction of the graph is completed, some hidden relationships that are not directly captured may emerge. Although these relationships are not directly reflected in the original relationship data, they can be inferred through the logical relationships between the data. To complement these potential relationships, graph inference techniques such as rule-based inference or machine learning methods can be further applied to discover and add these missing edges. This step helps to enrich the semantics of the graph, enhance the connectivity between data, and make the graph a more comprehensive and accurate model reflecting real-world entities and their relationships.

[0130] Finally, it is worth noting that although STG provides an effective data transformation method, its performance and efficiency highly depend on the query execution ability and optimization strategies of the underlying database management system. Therefore, when applying STG to large-scale datasets, it becomes particularly important to select an appropriate database system and appropriately adjust query optimization parameters. In addition, for specific application scenarios, further customization development may be required to ensure the efficiency of the transformation process and the quality of the generated graph.

[0131] In summary, STG provides a powerful and flexible framework for converting relational data into graph data while maintaining data integrity and richness. Through further optimization and extension, STG can play an important role in various application fields, providing support for data analysis, knowledge discovery, and intelligent applications, etc.

[0132] For example, for the customer entity, the associated attributes are all included in the relation schema of the "Customer" schema. Therefore, only projection operations are used to extract the associated instances. For the address entity, its associated attributes are included in multiple schemas (for example, two schemas of "Customer: custkey, address, comment, mstsegment, nationKey" and "Nation: nationKey, name, regionkey"). The ontology recognized by the large language model is {Customer: custkey, address, comment, mstsegment; Nation: nationKey, name,}. Therefore, full outer join operations are used to aggregate the attributes of the "Customer" and "Nation" schemas.

[0133] The relationship between "customer" and "nation" is mainly stored in the "Customer" schema. To achieve this, simple identifiers are created using the table ID and row ID. Inner join operations are applied and these identifiers are introduced to ensure that each customer instance corresponds to a unique relationship with each address instance.

[0134] The representation of the transformation grammar is: (nationKey, name) →V nation ; Algebra: V nation : Π((nationKey)) Π((name, nationKey)) (Nation) to achieve a full outer join operation.

[0135] (custkey) →V customer ; Algebra: V customer : Π((custkey)) (Customer) to achieve a projection operation.

[0136] E(nation - customer) = (V nation , V customer ); Algebra: E(nation - customer):(Π(tableId, rowId, nationKey) (Customer)) (Π(tableId, rowId, nationKey)(Nation)) (Π(tableId, rowId, custkey) (Customer)) to achieve an inner join operation.

[0137] Through these steps, a set of relational algebras are used to transform the original relational data into an attribute graph. The relational algebras can find the corresponding SQL statements and output specific entities and relationships, which are regarded as the nodes and edges of the graph.

[0138] Given at least one set of relational tables in relational data, the relational tables and their instances are extracted to form an equivalent attribute graph , where V represents entities, E represents entity relationships, L represents entity labels, represents an attribute function to identify and represent the entities and relationships existing in the relational data.

[0139] Thus, an automated attribute graph construction method for relational data is proposed, which includes two stages: one is the generation of the target ontology set, and the other is the construction of the attribute graph.

[0140] Figure 6 The figure shows a schematic diagram of the automated method for transforming relational data into an attribute graph according to an embodiment of the present invention.

[0141] As Figure 6As shown, attribute sampling is performed from a relational database (or database) to obtain a relational schema and data instances. In the generation stage of the target ontology set, the attribute sampling data under the relational schema and the data instances corresponding to the relational schema are input into a large language model to extract and generate an initial ontology set. The initial ontology set includes multiple initial entities and the relationships between the initial entities. According to the log records of the initial attribute data under the initial entities, the weighted scores of the initial attribute data can be calculated. The initial ontology set is updated according to the weighted scores of the initial attribute data to obtain the target ontology set.

[0142] In the automation stage of converting relational data into an attribute graph, the relational algebra in the semantic transformation grammar is used to map the target ontology set in the relational data to obtain an attribute graph. The nodes in the attribute graph represent the target entities in the target ontology set, and the edges between the nodes represent the relationships between the target entities.

[0143] The large language model can be a large model classifier. The relational schema obtained by attribute sampling and its corresponding data instance data are input into the large language model to extract the initial ontology set. A mapping translation grammar between ontology and relational data (such as semantic transformation grammar) is defined. Using this grammar, the mapping relationship between entity relationships and data can be transformed into a relational algebra, which can be used to extract examples from relational data for outputting an attribute graph.

[0144] Taking the automated construction of a biological risk attribute graph as an example, according to the embodiments of the present invention, the automated method for converting relational data into an attribute graph may generally include: Ontology definition step: Adopting a top-down method, biological risk factor entities and relationships are defined based on a pre-determined ontology set, where the ontology set at least includes entities such as pathogens, hosts, transmission routes, and prevention and control measures. Extraction step: According to the defined entities and relationships, information entities and their attributes in the field of biosafety risk are automatically extracted from structured data sources, including but not limited to researchers, research results, research institutions, and research events, and the relationships between these entities are defined. Instance extraction step: The specific instances in the data are automatically identified and integrated into the attribute graph. Construction step: Based on the instance data, an attribute graph for the field of biosafety risk is integrated and generated, involving the generation, integration, linking, and output of graph data. Data matching step: An advanced string matching algorithm and semantic understanding technology are used to accurately match the data attribute names and instances to ensure that all data describes the same entity, thereby ensuring data consistency and accuracy; Graph data connection step: For multi-relational table structured data, according to the predefined entity relationship pattern, an efficient connection algorithm is designed and applied to link the same instances in different relational tables to ensure high efficiency and high accuracy in practical applications, including but not limited to low time complexity, low space complexity, and high connection accuracy.

[0145] Read and save the data information in the relational data with the CSV (Comma-Separated Values) or TSV (Tab-Separated Values) format, and at the same time receive the predefined entity information of the biosafety risk area as input; identify and match the biosafety risk instance information in the data according to the predefined entities and relationships; based on the matched biosafety risk instance information, generate entities and relationships through aggregation and connection operations to form an attribute graph; save the constructed graph data entities and relationships as node files and edge files for the storage, visualization display and analysis of the graph database. Clean, standardize and initially classify the imported data. Understand and interpret the entities and relationships in the data based on the ontology to enhance the accuracy and robustness of data matching.

[0146] In the ontology generation stage of constructing the attribute graph, the key lies in accurately identifying and classifying the entity types, entity attributes and the association relationships between entities in the relational data. This process starts with using a large language model (LLM) pre-trained with domain knowledge as a classifier, and is optimized specifically for the complexity and professionalism of domain knowledge. For example, by training and fine-tuning the LLM with biosafety domain knowledge, it makes it pay special attention to the key concepts within the biosafety domain, such as papers, patents, research institutions, researchers, etc.

[0147] After the LLM classifier generates the initial ontology set, combined with domain knowledge, it optimizes and refines the output ontology, including the log records of the database and the attribute distribution of relationship instances, analyzes which biosafety attribute information the user is most concerned about, and extracts valuable attributes as new entities. The generated ontology can be used to guide the mapping syntax defined later, so that the nodes and relationships in the attribute graph can generate instances through the query statements of the relational data, and then the attribute graph is constructed.

[0148] Figure 7 The structural block diagram of the automated device for converting relational data into an attribute graph according to an embodiment of the present invention is shown.

[0149] As Figure 7 shown, the automated device 700 for converting the relational data of this embodiment into an attribute graph includes an input module 710, an update module 720, a mapping module 730 and a construction module 740.

[0150] The input module 710 is used to input the relationship schema of the data in the relational data and the data instances corresponding to the relationship schema into the large language model to obtain an initial ontology set. In one embodiment, the input module 710 can be used to execute the operation S210 described above, which will not be elaborated here.

[0151] The update module 720 is used to update the initial ontology set according to the weighted scores of the respective initial attribute data in the initial entities, obtaining a target ontology set. The weighted scores are determined based on the user query frequency and information entropy of the initial attribute data. In one embodiment, the update module 720 can be used to perform the operation S220 described above, which will not be elaborated here.

[0152] The mapping module 730 is used to map the target ontology set using semantic transformation grammar, obtaining the relational algebra of the target ontology set. The relational algebra is used to retrieve relational data. In one embodiment, the mapping module 730 can be used to perform the operation S230 described above, which will not be elaborated here.

[0153] The construction module 740 is used to construct an attribute graph based on the target ontology set and the relational algebra. In one embodiment, the construction module 740 can be used to perform the operation S240 described above, which will not be elaborated here.

[0154] According to an embodiment of the present invention, the update module 720 includes a first calculation sub-module, a first determination sub-module, and an update sub-module. The first calculation sub-module is used to calculate the weighted score of the initial attribute data in each initial entity for each initial entity. The first determination sub-module is used to determine the initial attribute data as the target entity when the weighted score meets a preset threshold. The update sub-module is used to update the initial ontology set according to the target entity, obtaining the target ontology set.

[0155] According to an embodiment of the present invention, the first calculation sub-module includes a first calculation unit, a second calculation unit, and a weighted summation unit. The first calculation unit is used to calculate the query score of the initial attribute data according to the user query frequency for each initial entity. The second calculation unit is used to calculate the information score of the initial attribute data according to the information entropy of the initial attribute data. The weighted summation unit is used to perform weighted summation on the query score and the information score to obtain the weighted score.

[0156] According to an embodiment of the present invention, the update module 720 includes an extraction unit, a determination unit, and an addition unit. The extraction unit is used to extract the target attribute data of the target entity from the initial attribute data in the initial entity. The determination unit is used to determine the target association relationship between the target entity and multiple initial entities in the initial ontology set using a large language model. The addition unit is used to add the target attribute data of the target entity and the target association relationship to the initial ontology set, obtaining the target ontology set.

[0157] According to an embodiment of the present invention, the relational algebra includes projection relational algebra, and the mapping module 730 includes a projection sub-module and a query mapping sub-module. The projection sub-module is used to perform a projection operation on the target attribute data in the target entity where the target attribute data comes from the same relational schema, so as to obtain the projection relational algebra. The query mapping sub-module is used to establish query mappings for different relational schemas for the target attribute data in the target entity that comes from different relational schemas.

[0158] According to an embodiment of the present invention, the relational algebra further includes full outer join relational algebra and inner join relational algebra, and the query mapping sub-module includes a full outer join unit and an inner join unit. The full outer join unit is used to perform a full outer join operation on the target attribute data in the target entity where the target attribute data comes from different relational schemas, so as to obtain the full outer join relational algebra. The inner join unit is used to perform an inner join between different relational schemas to obtain the inner join relational algebra.

[0159] According to an embodiment of the present invention, the construction module 740 includes a construction sub-module, and the construction sub-module is used to construct an attribute graph based on the target ontology set, projection relational algebra, full outer join relational algebra, and inner join relational algebra. Entities and the target entity are nodes of the attribute graph, and the connection edges between the nodes represent the association relationships between multiple entities or between an entity and the target entity.

[0160] According to an embodiment of the present invention, the above device further includes a data sampling module, and the data sampling module is used to perform data sampling from the relational data according to a data sampling strategy to obtain attribute sampling data under multiple relational schemas. The data collection strategy is determined according to the data value range in the relational data.

[0161] According to an embodiment of the present invention, the data sampling module includes a partitioning sub-module and a data sampling sub-module. The partitioning sub-module is used to partition the data set in the relational data into multiple subsets according to the data value range, and the data value ranges represented by the multiple subsets are different. The data sampling sub-module is used to perform data sampling from multiple subsets of the relational data respectively to obtain the attribute sampling data under the relational schema.

[0162] It should be noted that the part of the automated device for converting relational data into an attribute graph in the embodiment of the present invention corresponds to the part of the automated method for converting relational data into an attribute graph in the embodiment of the present invention. For the description of the part of the automated device for converting relational data into an attribute graph, please refer to the part of the automated method for converting relational data into an attribute graph specifically, and it will not be elaborated here.

[0163] According to an embodiment of the present invention, any plurality of modules among the input module 710, the update module 720, the mapping module 730, and the construction module 740 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the input module 710, the update module 720, the mapping module 730, and the construction module 740 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as integrating or packaging circuits, etc., in hardware or firmware, or in any one or a suitable combination of the three implementation manners of software, hardware, and firmware. Alternatively, at least one of the input module 710, the update module 720, the mapping module 730, and the construction module 740 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.

[0164] Figure 8 A block diagram of an electronic device suitable for implementing an automated method for converting relational data into an attribute graph according to an embodiment of the present invention is shown.

[0165] As Figure 8 shown, the electronic device 800 according to an embodiment of the present invention includes a processor 801, which may perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 801 may also include on-board memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0166] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 802 and / or the RAM 803. It should be noted that the programs can also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 can also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in the one or more memories.

[0167] According to an embodiment of the present invention, the electronic device 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the input / output (I / O) interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 810 as needed so that a computer program read from it can be installed into the storage portion 808 as needed.

[0168] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.

[0169] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 802 and / or RAM 803 and / or one or more memories other than ROM 802 and RAM 803.

[0170] An embodiment of the present invention also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the automated method for converting relational data into property graphs provided by the embodiments of the present invention for relational data.

[0171] When the computer program is executed by the processor 801, it executes the above functions defined in the system / apparatus of the embodiments of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0172] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 809, and / or be installed from the removable medium 811. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0173] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 809, and / or be installed from the removable medium 811. When the computer program is executed by the processor 801, it executes the above functions defined in the system of the embodiments of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0174] According to embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0175] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0176] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present invention can be combined or / and combined in various ways, even if such combinations or combinations are not explicitly recited in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recited in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0177] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should fall within the scope of the present invention.

Claims

1. An automated method for converting relational data into property graphs, characterized in that, The method includes: Inputting the relationship schema of the data in the relational data and the data instances corresponding to the relationship schema into a large language model to obtain an initial ontology set, the initial ontology set including a plurality of initial entities and the association relationships between the plurality of initial entities, the initial entities including at least one initial attribute data extracted from the attribute sampling data, and the attribute sampling data being sampled and extracted from the data instances; Updating the initial ontology set according to the weighted scores of the respective initial attribute data in the initial entities to obtain a target ontology set, the weighted scores being determined according to the user query frequency and information entropy of the initial attribute data; Mapping the target ontology set using a semantic transformation grammar to obtain the relational algebra of the target ontology set, the relational algebra being used to retrieve the relational data; Constructing an attribute graph based on the target ontology set and the relational algebra.

2. The method according to claim 1, characterized in that, The updating the initial ontology set according to the weighted scores of the respective initial attribute data in the initial entities to obtain a target ontology set includes: For each initial entity, calculating the weighted score of the initial attribute data in the initial entity; When the weighted score meets a preset threshold, determining the initial attribute data as a target entity; Updating the initial ontology set according to the target entity to obtain the target ontology set.

3. The method according to claim 2, wherein The calculating the weighted score of the initial attribute data in the initial entity for each initial entity includes: For each initial entity, calculating the query score of the initial attribute data according to the user query frequency; Calculating the information score of the initial attribute data according to the information entropy of the initial attribute data; Performing weighted summation on the query score and the information score to obtain the weighted score.

4. The method according to claim 2, characterized in that, The updating the initial ontology set according to the target entity to obtain the target ontology set includes: Extracting the target attribute data of the target entity from the initial attribute data in the initial entity; Using the large language model to determine the target association relationship between the target entity and the plurality of initial entities in the initial ontology set; Adding the target attribute data and the target association relationship of the target entity to the initial ontology set to obtain the target ontology set.

5. The method according to claim 1, wherein The relational algebra includes projection relational algebra, The mapping the target ontology set using a semantic transformation grammar to obtain the relational algebra of the target ontology set includes: For the target attribute data in the target entity coming from the same relationship schema, performing a projection operation on the target attribute data to obtain the projection relational algebra; For the target attribute data in the target entity coming from different relationship schemas, establishing a query mapping for the different relationship schemas.

6. The method according to claim 5, characterized in that, The relational algebra further includes full outer join relational algebra and inner join relational algebra, The establishing a query mapping for the different relationship schemas for the target attribute data in the target entity coming from different relationship schemas includes: For the target attribute data in the target entity coming from the different relational schemas, perform a full outer join operation on the target attribute data to obtain the full outer join relational algebra; Perform an inner join between the different relational schemas to obtain the inner join relational algebra.

7. The method according to claim 6, wherein The constructing an attribute graph based on the target ontology set and the relational algebra includes: Construct the attribute graph based on the target ontology set, the projection relational algebra, the full outer join relational algebra, and the inner join relational algebra. Entities and the target entity are nodes of the attribute graph, and the connection edges between the nodes represent the association relationships among multiple entities or between an entity and the target entity.

8. The method according to claim 1, wherein The method further includes: Perform data sampling from the relational data according to a data sampling strategy to obtain multiple sets of attribute sampling data under the relational schemas, where the data sampling strategy is determined according to the data value range in the relational data.

9. The method according to claim 8, wherein The performing data sampling from the relational data according to a data sampling strategy to obtain multiple sets of attribute sampling data under the relational schemas includes: Divide the data set in the relational data into multiple subsets according to the data value range, and the data value ranges represented by the multiple subsets are different; Perform data sampling from the multiple subsets of the relational data respectively to obtain the attribute sampling data under the relational schema.

10. An automated device for converting relational data into a property graph, characterized in that, The apparatus includes: An input module, configured to input the relational schema of the data in the relational data and the relational instances corresponding to the relational schema into a large language model to obtain an initial ontology set, where the initial ontology set includes multiple initial entities and the association relationships between the multiple initial entities, and the initial entities include at least one initial attribute data extracted from the attribute sampling data, and the attribute sampling data is sampled and extracted from data instances; An update module, configured to update the initial ontology set according to the weighted scores of the respective initial attribute data in the initial entities to obtain a target ontology set, where the weighted scores are determined according to the user query frequency and information entropy of the initial attribute data; A mapping module, configured to map the target ontology set by using semantic transformation grammar to obtain the relational algebra of the target ontology set, and the relational algebra is used to retrieve the relational data; A construction module, configured to construct an attribute graph based on the target ontology set and the relational algebra.

Citation Information

Patent Citations

  • Mapping method and equipment for converting relational database into graph database and medium

    CN118820376A

  • Ensemble learning enhanced prompting for open relation extraction

    WO2024233222A1