Hypergraph database construction method and device for high-order correlation
Patent Information
- Application Number
- CN202310257006.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-08
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-03-08
AI Technical Summary
[0005]本申请提供一种面向高阶关联的超图数据库构建方法及装置,以解决相关技术中在面对复杂的高阶关联结构时,无法直接存储和分析高阶关联结构,降低了数据库系统的建模和信息挖掘能力,并且降低了图数据库的适用性的问题
[0020]本申请实施例可以将原始数据归为有序属性数据、无序属性数据和高阶关联结构数据三类基础数据格式,对无序属性数据进行分类存储,并构建相应哈希函数以将无序属性数据映射成满足规整条件的数据进行索引,对有序属性数据进行分类存储,并分别构建B+树进行索引,并且构建交叉双向链表对高阶关联结构数据进行存储,以联合超图神经网络算法对高阶关联结构数据进行统计分析,从而有效的提升了数据库系统的建模和信息挖掘能力,并且提升了图数据库的适用性。由此,解决了相关技术中在面对复杂的高阶关联结构时,无法直接存储和分析高阶关联结构,降低了数据库系统的建模和信息挖掘能力,并且降低了图数据库的适用性的问题。
Smart Images

Figure CN116340577B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of database technology, and in particular to a method and apparatus for constructing a hypergraph database oriented towards high-order associations. Background Technology
[0002] Graph databases have gradually become a mainstay of the new generation of databases due to their superior performance in storing and processing unstructured relational data. Compared with traditional relational database systems, graph databases have a higher response speed for relational expansion and are more suitable for relational analysis and predictive analysis.
[0003] In related technologies, current graph databases can store and analyze low-order associations, such as simple graphs and directed graphs. Furthermore, when faced with complex high-order association structures, graph data can be stored in an approximate manner.
[0004] However, when faced with complex high-order association structures, related technologies cannot directly store and analyze these structures, which reduces the modeling and information mining capabilities of database systems and diminishes the applicability of graph databases, and urgently needs to be addressed. Summary of the Invention
[0005] This application provides a method and apparatus for constructing a hypergraph database oriented towards higher-order associations, in order to solve the problem in related technologies that when faced with complex higher-order association structures, it is impossible to directly store and analyze higher-order association structures, which reduces the modeling and information mining capabilities of the database system and reduces the applicability of the graph database.
[0006] The first aspect of this application provides a method for constructing a hypergraph database oriented towards higher-order associations, comprising the following steps: classifying raw data into three basic data formats, wherein the three basic data formats include ordered attribute data, unordered attribute data, and higher-order association structure data; classifying and storing the unordered attribute data, and constructing corresponding hash functions to map the unordered attribute data into data that meets preset regularization conditions for indexing; classifying and storing the ordered attribute data, and constructing B+ trees for indexing respectively; constructing a cross-linked doubly linked list to store the higher-order association structure data, and performing statistical analysis on the higher-order association structure data in conjunction with a preset hypergraph neural network algorithm.
[0007] Optionally, in one embodiment of this application, classifying the original data into three basic data formats, wherein the three basic data formats include ordered attribute data, unordered attribute data, and higher-order relational structure data, includes: cleaning the original data to obtain cleaned data; and extracting the ordered attribute data, the unordered attribute data, and the higher-order relational structure data from the cleaned data.
[0008] Optionally, in one embodiment of this application, the step of classifying and storing the unordered attribute data and constructing a corresponding hash function to map the unordered attribute data into data that meets a preset regularization condition for indexing includes: constructing the hash function according to the type of the unordered attribute data; and using the hash function to map the unordered attribute data.
[0009] Optionally, in one embodiment of this application, the step of classifying and storing the ordered attribute data and constructing B+ trees for indexing includes: selecting the order of the B+ tree according to the type of the ordered attribute data; and constructing B+ trees for indexing different types of ordered attribute data.
[0010] Optionally, in one embodiment of this application, the step of constructing a cross-linked doubly linked list to store the higher-order association structure data, and then using a preset hypergraph neural network algorithm to perform statistical analysis on the higher-order association structure data, includes: constructing a node list and an edge list; generating multiple higher-order association storage units based on the connection relationship between nodes and hyperedges, and connecting them with the corresponding nodes to generate a doubly linked list of nodes; and bidirectionally linking each hyperedge with the corresponding higher-order association storage unit based on the connection relationship between the hyperedge and the node to obtain a doubly linked list of higher-order associations.
[0011] Optionally, in one embodiment of this application, the step of constructing a cross-linked doubly linked list to store the higher-order association structure data, and then using a preset hypergraph neural network algorithm to perform statistical analysis on the higher-order association structure data, further includes: selecting nodes and hyperedges of interest from the hypergraph data storage structure; generating a corresponding hypergraph association matrix from the doubly linked list of the higher-order associations; generating node features from the selected node attributes and hyperedge features from the corresponding hyperedge attributes, and generating a corresponding feature matrix; constructing a corresponding hypergraph neural network model according to the downstream task, and inputting the hypergraph association matrix and the feature matrices of the nodes and hyperedges into the hypergraph neural network model to obtain the final prediction score.
[0012] A second aspect of this application provides a hypergraph database construction apparatus for higher-order associations, comprising: a classification module for classifying raw data into three basic data formats, wherein the three basic data formats include ordered attribute data, unordered attribute data, and higher-order association structure data; a first classification storage module for classifying and storing the unordered attribute data, and constructing corresponding hash functions to map the unordered attribute data into data that meets preset regularization conditions for indexing; a second classification storage module for classifying and storing the ordered attribute data, and constructing B+ trees for indexing each; and a construction module for constructing cross-linked doubly linked lists to store the higher-order association structure data, and performing statistical analysis on the higher-order association structure data in conjunction with a preset hypergraph neural network algorithm.
[0013] Optionally, in one embodiment of this application, the classification module includes: a cleaning unit for cleaning the original data to obtain cleaned data; and an extraction unit for extracting the ordered attribute data, the unordered attribute data, and the higher-order association structure data from the cleaned data.
[0014] Optionally, in one embodiment of this application, the first classification storage module includes: a first construction unit, configured to construct the hash function according to the type of the unordered attribute data; and a mapping unit, configured to map the unordered attribute data using the hash function.
[0015] Optionally, in one embodiment of this application, the second classification storage module includes: a selection unit, configured to select the order of the B+ tree according to the type of the ordered attribute data; and a second construction unit, configured to construct B+ trees for indexing different ordered attribute data respectively.
[0016] Optionally, in one embodiment of this application, the construction module includes: a third construction unit for constructing a node list and an edge list; a generation unit for generating multiple higher-order association storage units based on the connection relationship between nodes and hyperedges, and connecting them with the corresponding nodes to generate a doubly linked list of nodes; and a linking unit for bidirectionally linking each hyperedge with the corresponding higher-order association storage unit based on the connection relationship between the hyperedge and the node to obtain a doubly linked list of higher-order associations.
[0017] Optionally, in one embodiment of this application, the construction module further includes: a selection unit, configured to select nodes and hyperedges of interest from the hypergraph data storage structure; a first generation unit, configured to generate a corresponding hypergraph association matrix from the doubly linked list of the higher-order associations; a second generation unit, configured to generate node features from the selected node attributes and hyperedge features from the corresponding hyperedge attributes, and generate a corresponding feature matrix; and an acquisition unit, configured to construct a corresponding hypergraph neural network model according to the downstream task, and input the hypergraph association matrix and the feature matrices of nodes and hyperedges into the hypergraph neural network model to obtain the final prediction score.
[0018] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the hypergraph database construction method for higher-order associations as described in the above embodiments.
[0019] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for constructing a hypergraph database oriented towards higher-order associations.
[0020] This application's embodiments categorize raw data into three basic data formats: ordered attribute data, unordered attribute data, and high-order association structure data. Unordered attribute data is stored in a categorized manner, and a corresponding hash function is constructed to map it to data that meets regularity conditions for indexing. Ordered attribute data is also stored in a categorized manner, and B+ trees are constructed for indexing. Furthermore, a cross-linked doubly linked list is constructed to store high-order association structure data. A hypergraph neural network algorithm is then used to perform statistical analysis on this high-order association structure data, effectively improving the modeling and information mining capabilities of the database system and enhancing the applicability of the graph database. This solves the problem in related technologies where, when faced with complex high-order association structures, the inability to directly store and analyze these structures reduces the modeling and information mining capabilities of the database system and diminishes the applicability of the graph database.
[0021] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0022] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0023] Figure 1 This is a flowchart illustrating a method for constructing a hypergraph database oriented towards higher-order associations, according to an embodiment of this application.
[0024] Figure 2 This is a schematic diagram of the storage structure of a hypergraph database system oriented towards higher-order associations according to a specific embodiment of this application;
[0025] Figure 3 This is a schematic diagram of the structure of a hypergraph database construction apparatus for high-order associations provided in an embodiment of this application;
[0026] Figure 4 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0027] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0028] The following description, with reference to the accompanying drawings, describes a method and apparatus for constructing a hypergraph database oriented towards higher-order associations according to embodiments of this application. Addressing the problem mentioned in the background art that the inability to directly store and analyze complex higher-order association structures reduces the modeling and information mining capabilities of database systems and decreases the applicability of graph databases, this application provides a method for constructing a hypergraph database oriented towards higher-order associations. In this method, raw data can be categorized into three basic data formats: ordered attribute data, unordered attribute data, and higher-order association structure data. Unordered attribute data is classified and stored, and a corresponding hash function is constructed to map the unordered attribute data into data that meets regularity conditions for indexing. Ordered attribute data is classified and stored, and B+ trees are constructed for indexing. A cross-linked doubly linked list is constructed to store the higher-order association structure data. A hypergraph neural network algorithm is then used to perform statistical analysis on the higher-order association structure data, thereby effectively improving the modeling and information mining capabilities of the database system and enhancing the applicability of the graph database. This solves the problem in related technologies that, when faced with complex high-order association structures, cannot directly store and analyze them, thus reducing the modeling and information mining capabilities of database systems and decreasing the applicability of graph databases.
[0029] Specifically, Figure 1 This is a flowchart illustrating a method for constructing a hypergraph database oriented towards higher-order associations, as provided in an embodiment of this application.
[0030] like Figure 1 As shown, this method for constructing a hypergraph database oriented towards higher-order associations includes the following steps:
[0031] In step S101, the original data is classified into three basic data formats, which include ordered attribute data, unordered attribute data, and high-order relational structure data.
[0032] It is understood that the embodiments of this application can classify the raw data in the following steps into three basic data formats: ordered attribute data, unordered attribute data, and high-order relational structure data, thereby effectively improving the executability of data systems oriented towards high-order relations.
[0033] Optionally, in one embodiment of this application, the original data is classified into three basic data formats, which include ordered attribute data, unordered attribute data, and high-order relational structure data. This includes: cleaning the original data to obtain cleaned data; and extracting ordered attribute data, unordered attribute data, and high-order relational structure data from the cleaned data.
[0034] In actual implementation, the embodiments of this application can clean the original data by escaping special characters and converting the data to a suitable data type, thereby obtaining cleaned data. Ordered attribute data, unordered attribute data, and high-order relational structure data can be extracted from the cleaned data, thereby improving the storage capacity of the data system oriented towards high-order relations.
[0035] For example, embodiments of this application can determine whether the data contains ordered attribute data, wherein ordered attribute data is a type of sortable data, such as integers, floating-point numbers, strings, etc., wherein given ordered data data ordered ={o1,o2,…,o n According to the rules defined by ordered data, each element supports comparison of size; therefore, we have o1≤o2≤o3≤…≤o n .
[0036] For example, embodiments of this application can also determine whether the data contains unordered attribute data, wherein unordered attribute data is a type of data that is not sortable but is enumerable, such as integers, floating-point numbers, strings, etc., wherein, given unordered data data unorder ={u1,u2,…,u m According to the definition of unordered data, each element can be compared to determine whether they are equivalent.
[0037] For example, embodiments of this application can also determine whether the data contains higher-order relational data, wherein the higher-order relational data is a hypergraph structure constructed from several nodes and their associated groups. The higher-order relational structure (hypergraph) can be defined as follows:
[0038]
[0039] in, Let ε denote the set of nodes, and let ε denote the set of hyperedges.
[0040] in, ε={e1,e2,…,e m}
[0041] Furthermore, edges in a higher-order association structure can connect any number of nodes. For example, e1 = {v1, v3, v8, v9} means that e1 connects four nodes v1, v3, v8, and v9. Thus, a higher-order association structure can be represented by an association matrix H, which is defined as follows:
[0042]
[0043] Where v represents a node and e represents a hyperedge.
[0044] Therefore, the embodiments of this application can simultaneously store unordered attribute data, ordered attribute data, and high-order relational data, greatly improving the data storage capacity of existing database systems.
[0045] In step S102, the unordered attribute data is classified and stored, and a corresponding hash function is constructed to map the unordered attribute data into data that meets the preset regularization conditions for indexing.
[0046] It is understood that the embodiments of this application can classify and store the unordered attribute data in the following steps, and construct corresponding hash functions to map the unordered attribute data into data that meets the regularity conditions for indexing, thereby improving the applicability of the data system.
[0047] Optionally, in one embodiment of this application, unordered attribute data is classified and stored, and a corresponding hash function is constructed to map the unordered attribute data into data that meets preset regularization conditions for indexing, including: constructing a hash function according to the type of unordered attribute data; and using the hash function to map the unordered attribute data.
[0048] For example, such as Figure 2 As shown, in this embodiment of the application, a hash function can be constructed based on the type of unordered attribute data extracted from the cleaned data in the above steps, and the constructed hash function can be used to map the unordered attribute data to ensure that the data can be fully utilized.
[0049] In step S103, the ordered attribute data is classified and stored, and B+ trees are constructed for indexing.
[0050] It is understood that the embodiments of this application can classify and store the ordered attribute data in the following steps, and construct B+ trees for indexing, thereby improving the applicability of the data system.
[0051] Optionally, in one embodiment of this application, the ordered attribute data is classified and stored, and B+ trees are constructed for indexing, including: selecting the order of the B+ tree according to the type of ordered attribute data; and constructing B+ trees for indexing different ordered attribute data.
[0052] For example, such as Figure 2 As shown, in this embodiment of the application, the order m of the B+ tree can be selected according to the type of ordered attribute data, and B+ trees can be constructed for indexing according to different ordered attribute data, thereby effectively improving the full utilization of data.
[0053] In step S104, a cross-linked doubly linked list is constructed to store the high-order association structure data, and a preset hypergraph neural network algorithm is used to perform statistical analysis on the high-order association structure data.
[0054] It is understood that the embodiments of this application can construct the cross doubly linked list in the following steps to store high-order association structure data, so as to perform statistical analysis on high-order association structure data in conjunction with the hypergraph neural network algorithm, thereby effectively improving the data storage and modeling capabilities of the database system, and improving the training speed and inference speed of the model for high-order association data.
[0055] Optionally, in one embodiment of this application, a cross-linked doubly linked list is constructed to store the higher-order association structure data, so as to perform statistical analysis on the higher-order association structure data in conjunction with a preset hypergraph neural network algorithm, including: constructing a node list and an edge list; generating multiple higher-order association storage units according to the connection relationship between nodes and hyperedges, and connecting them with the corresponding nodes to generate a doubly linked list of nodes; based on the connection relationship between hyperedges and nodes, bidirectionally linking each hyperedge with the corresponding higher-order association storage unit to obtain a doubly linked list of higher-order associations.
[0056] For example, embodiments of this application can construct a node list and an edge list. After preprocessing the original data, the node list and the edge list are defined as lists respectively. v ={v1,v2,…,v n} and list e ={e1,e2,…,e m}
[0057] Furthermore, in this embodiment, several higher-order associative storage units can be generated based on the connection relationship between nodes and hyperedges, and connected to the corresponding nodes to generate a doubly linked list of nodes. The required number of higher-order associative storage units is calculated based on the constructed node list, hyperedge list, and hypergraph association matrix, and this number is consistent with the number of non-zero elements in the hypergraph association matrix. Figure 2As shown, this is the design of a high-order associative storage unit, which consists of four pointers and one data bit. The four pointers point to the previous node, the next node, the previous hyperedge, and the next hyperedge, respectively.
[0058] Furthermore, such as Figure 2 As shown, in this embodiment of the application, each hyperedge can be bidirectionally linked to its corresponding higher-order association storage unit based on the connection relationship between the hyperedge and the node. Each node and hyperedge directly point to its adjacent hyperedge and node. By storing higher-order associations in this structure, constant-level retrieval of nearest neighbor data can be achieved.
[0059] Optionally, in one embodiment of this application, constructing a cross-linked doubly linked list to store high-order association structure data, and using a pre-defined hypergraph neural network algorithm to perform statistical analysis on the high-order association structure data, further includes: selecting nodes and hyperedges of interest from the hypergraph data storage structure; generating a corresponding hypergraph association matrix from the doubly linked list of high-order associations; generating node features from the selected node attributes, and simultaneously generating hyperedge features from the corresponding hyperedge attributes, and generating a corresponding feature matrix; constructing a corresponding hypergraph neural network model according to the downstream task, and inputting the hypergraph association matrix and the feature matrices of nodes and hyperedges into the hypergraph neural network model to obtain the final prediction score.
[0060] As one possible implementation, embodiments of this application can select nodes and hyperedges of interest from the hypergraph data storage structure, generate the corresponding hypergraph association matrix H from the doubly linked list of higher-order associations, generate node features from the selected node attributes, generate hyperedge features from the corresponding hyperedge attributes, and generate corresponding feature matrices X and Y. Based on the downstream task, a corresponding hypergraph neural network model is constructed, and the constructed hypergraph association matrix and the feature matrices of nodes and hyperedges are input into the hypergraph neural network model to obtain the final prediction score.
[0061] The hypergraph neural convolutional layer is defined as follows:
[0062]
[0063] In summary, as Figure 2 As shown, the embodiments of this application can simultaneously store unordered attribute data, ordered attribute data, low-order association data, and high-order association data. Furthermore, by integrating the storage and analysis of high-order associations and designing a hypergraph neural network model oriented towards specific tasks, the data system can achieve more efficient modeling and analysis capabilities for high-order association data.
[0064] The hypergraph database construction method for high-order association proposed in this application can classify raw data into three basic data formats: ordered attribute data, unordered attribute data, and high-order association structure data. Unordered attribute data is categorized and stored, and a corresponding hash function is constructed to map it to data that meets regularity conditions for indexing. Ordered attribute data is categorized and stored, and B+ trees are constructed for indexing. Furthermore, a cross-linked doubly linked list is constructed to store high-order association structure data. A hypergraph neural network algorithm is then used to perform statistical analysis on the high-order association structure data, thereby effectively improving the modeling and information mining capabilities of the database system and enhancing the applicability of the graph database. This solves the problem in related technologies where, when faced with complex high-order association structures, the inability to directly store and analyze these structures reduces the modeling and information mining capabilities of the database system and decreases the applicability of the graph database.
[0065] Next, referring to the accompanying drawings, a hypergraph database construction apparatus for high-order associations proposed according to an embodiment of this application is described.
[0066] Figure 3 This is a block diagram of a hypergraph database construction apparatus for high-order associations according to an embodiment of this application.
[0067] like Figure 3 As shown, the hypergraph database construction device 10 for high-order association includes: a classification module 100, a first classification storage module 200, a second classification storage module 300, and a construction module 400.
[0068] Specifically, the classification module 100 is used to classify the raw data into three basic data formats, which include ordered attribute data, unordered attribute data, and high-order relational structure data.
[0069] The first classification storage module 200 is used to classify and store unordered attribute data, and to construct corresponding hash functions to map unordered attribute data into data that meets preset regularization conditions for indexing.
[0070] The second category storage module 300 is used to classify and store ordered attribute data, and to build B+ trees for indexing.
[0071] Module 400 is used to construct a cross-linked doubly linked list to store high-order association structure data, and to perform statistical analysis on the high-order association structure data in conjunction with a preset hypergraph neural network algorithm.
[0072] Optionally, in one embodiment of this application, the classification module 100 includes a cleaning unit and an extraction unit.
[0073] The cleaning unit is used to clean the raw data to obtain cleaned data.
[0074] The extraction unit is used to extract ordered attribute data, unordered attribute data, and high-order relational structure data from the cleaned data.
[0075] Optionally, in one embodiment of this application, the first classification storage module 200 includes: a first construction unit and a mapping unit.
[0076] The first building unit is used to construct a hash function based on the type of unordered attribute data.
[0077] Mapping unit, used to map unordered attribute data using a hash function.
[0078] Optionally, in one embodiment of this application, the second classification storage module 300 includes: a selection unit and a second construction unit.
[0079] The selection unit is used to select the order of the B+ tree based on the type of ordered attribute data.
[0080] The second building unit is used to construct B+ trees for indexing different ordered attribute data.
[0081] Optionally, in one embodiment of this application, the construction module 400 includes: a third construction unit, a generation unit, and a linking unit.
[0082] The third building unit is used to build the node list and the edge list.
[0083] The generation unit is used to generate multiple higher-order associative storage units based on the connection relationship between nodes and hyperedges, and connect them with the corresponding nodes to generate a doubly linked list of nodes.
[0084] Linking units are used to bidirectionally link each hyperedge to its corresponding higher-order associative storage unit based on the connection relationship between the hyperedge and the node, thus obtaining a doubly linked list of higher-order associations.
[0085] Optionally, in one embodiment of this application, the construction module 400 further includes: a selection unit, a first generation unit, a second generation unit, and an acquisition unit.
[0086] The selection unit is used to select the nodes and hyperedges of interest from the hypergraph data storage structure.
[0087] The first generation unit is used to generate the corresponding hypergraph association matrix from the doubly linked list of higher-order associations.
[0088] The second generation unit is used to generate node features from the selected node attributes, generate hyperedge features from the corresponding hyperedge attributes, and generate the corresponding feature matrix.
[0089] The acquisition unit is used to construct the corresponding hypergraph neural network model based on the downstream task, and input the hypergraph association matrix and the feature matrices of nodes and hyperedges into the hypergraph neural network model to obtain the final prediction score.
[0090] It should be noted that the foregoing explanation of the embodiment of the hypergraph database construction method for higher-order associations also applies to the hypergraph database construction apparatus for higher-order associations in this embodiment, and will not be repeated here.
[0091] The hypergraph database construction apparatus for high-order associations proposed in this application can classify raw data into three basic data formats: ordered attribute data, unordered attribute data, and high-order association structure data. Unordered attribute data is categorized and stored, and a corresponding hash function is constructed to map it to data that meets regularity conditions for indexing. Ordered attribute data is categorized and stored, and B+ trees are constructed for indexing. Furthermore, a cross-linked doubly linked list is constructed to store high-order association structure data. A hypergraph neural network algorithm is then used to perform statistical analysis on the high-order association structure data, thereby effectively improving the modeling and information mining capabilities of the database system and enhancing the applicability of the graph database. This solves the problem in related technologies where, when faced with complex high-order association structures, the inability to directly store and analyze these structures reduces the modeling and information mining capabilities of the database system and decreases the applicability of the graph database.
[0092] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include:
[0093] The memory 401, the processor 402, and the computer program stored on the memory 401 and capable of running on the processor 402.
[0094] When the processor 402 executes the program, it implements the hypergraph database construction method for high-order associations provided in the above embodiments.
[0095] Furthermore, electronic devices also include:
[0096] Communication interface 403 is used for communication between memory 401 and processor 402.
[0097] The memory 401 is used to store computer programs that can run on the processor 402.
[0098] The memory 401 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0099] If the memory 401, processor 402, and communication interface 403 are implemented independently, then the communication interface 403, memory 401, and processor 402 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0100] Optionally, in a specific implementation, if the memory 401, processor 402, and communication interface 403 are integrated on a single chip, then the memory 401, processor 402, and communication interface 403 can communicate with each other through an internal interface.
[0101] Processor 402 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0102] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for constructing a hypergraph database oriented towards higher-order associations.
[0103] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0104] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0105] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0106] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0107] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0108] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it includes one or a combination of the steps of the method embodiments.
[0109] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0110] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A method for constructing a hypergraph database oriented towards higher-order associations, characterized in that, Includes the following steps: The raw data is categorized into three basic data formats, namely, ordered attribute data, unordered attribute data, and high-order relational structure data. The unordered attribute data is classified and stored, and a corresponding hash function is constructed to map the unordered attribute data into data that meets a preset regularization condition for indexing; The ordered attribute data is categorized and stored, and B+ trees are constructed for indexing each category; and A cross-linked doubly linked list is constructed to store the higher-order association structure data, and a preset hypergraph neural network algorithm is used to perform statistical analysis on the higher-order association structure data; The construction of a cross-linked doubly linked list to store the higher-order association structure data, and the joint use of a preset hypergraph neural network algorithm to perform statistical analysis on the higher-order association structure data, includes: Construct a list of nodes and a list of edges; Multiple higher-order associative storage units are generated based on the connection relationship between nodes and hyperedges, and then connected to the corresponding nodes to generate a doubly linked list of nodes. Based on the connection relationship between the hyperedge and the node, each hyperedge is bidirectionally linked to the corresponding higher-order associative storage unit to obtain a higher-order associative doubly linked list.
2. The method according to claim 1, characterized in that, The original data is categorized into three basic data formats, which include ordered attribute data, unordered attribute data, and higher-order relational structure data, including: The original data is cleaned to obtain cleaned data; Extract the ordered attribute data, the unordered attribute data, and the higher-order relational structure data from the cleaned data.
3. The method according to claim 1, characterized in that, The step of classifying and storing the unordered attribute data, and constructing a corresponding hash function to map the unordered attribute data into data that meets preset regularization conditions for indexing, includes: Construct the hash function based on the type of the unordered attribute data; The hash function is used to map the unordered attribute data.
4. The method according to claim 1, characterized in that, The step of classifying and storing the ordered attribute data, and constructing B+ trees for indexing each, includes: Select the order of the B+ tree based on the type of the ordered attribute data; For the different ordered attribute data, B+ trees are constructed for indexing.
5. The method according to claim 1, characterized in that, The step of constructing a cross-linked doubly linked list to store the higher-order association structure data, and then using a pre-defined hypergraph neural network algorithm to perform statistical analysis on the higher-order association structure data, further includes: Select the nodes and hyperedges of interest from the hypergraph data storage structure; Generate the corresponding hypergraph association matrix from the doubly linked list of the higher-order association; Node features are generated from the selected node attributes, and hyperedge features are generated from the corresponding hyperedge attributes, and the corresponding feature matrix is generated. A corresponding hypergraph neural network model is constructed based on the downstream task, and the hypergraph association matrix and the feature matrices of nodes and hyperedges are input into the hypergraph neural network model to obtain the final prediction score.
6. A hypergraph database construction apparatus for high-order associations, characterized in that, include: The classification module is used to classify the raw data into three basic data formats, namely, ordered attribute data, unordered attribute data, and high-order relational structure data. The first classification storage module is used to classify and store the unordered attribute data, and construct a corresponding hash function to map the unordered attribute data into data that meets preset regularization conditions for indexing; The second classification storage module is used to classify and store the ordered attribute data, and to construct B+ trees for indexing each category; and The construction module is used to construct a cross-linked doubly linked list to store the higher-order association structure data, and to perform statistical analysis on the higher-order association structure data in conjunction with a preset hypergraph neural network algorithm; The construction module includes: a third construction unit, used to construct a node list and an edge list; The generation unit is used to generate multiple higher-order associative storage units based on the connection relationship between nodes and hyperedges, and connect them with the corresponding nodes to generate a doubly linked list of nodes. The linking unit is used to bidirectionally link each hyperedge with its corresponding higher-order association storage unit based on the connection relationship between the hyperedge and the node, thereby obtaining a doubly linked list of higher-order associations.
7. The apparatus according to claim 6, characterized in that, The classification module includes: A cleaning unit is used to clean the original data to obtain cleaned data. The extraction unit is used to extract the ordered attribute data, the unordered attribute data, and the higher-order association structure data from the cleaned data.
8. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the hypergraph database construction method for higher-order associations as described in any one of claims 1-5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the hypergraph database construction method for higher-order associations as described in any one of claims 1-5.
Citation Information
Patent Citations
Internal memory database system and method and device for implementing internal memory data base
CN101315628A
Partitioning and parallel distribution processing method of super-large scale RDF graph data
CN104809168A