Graph data enhancement method and device, equipment and medium

By aligning the graph data and fine-tuning the processing using large language models, the problem of low accuracy in graph data enhancement is solved, and the enrichment of node features and the improvement of data enhancement effect is achieved.

CN120337985APending Publication Date: 2025-07-18SHENZHEN INST OF COMPUTING SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510667138.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the process of graph data enhancement, the accuracy of graph data is low, especially in the face of missing node features, sparse labels and noise interference, the traditional methods are inefficient and cannot dynamically screen key features, and the heterogeneous data is insufficient in adaptability.

Method used

By aligning the graph data to be enhanced with the preset graph data, determining the matching node and writing the features of the preset graph data, fine-tuning processing is used for large language model, a collection of candidate features is obtained, and data augmentation is used for fine-tuning large language model to enrich the semantic information of the node.

Benefits of technology

It improves the accuracy of graph data, enhances the comprehensiveness and depth of node characteristics, and improves the effect of data enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337985A_ABST
    Figure CN120337985A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a graph data enhancement method and device, equipment and a medium. The to-be-enhanced graph data is aligned with preset graph data, matched nodes in the to-be-enhanced graph data and the preset graph data are determined, features corresponding to the preset graph data in the matched nodes are written into the corresponding nodes of the to-be-enhanced graph data, candidate features of the corresponding nodes of the to-be-enhanced graph data are obtained, and a preset large language model is obtained; and according to the candidate feature set, performing fine tuning processing on the large language model to obtain a fine-tuned large language model, and performing data enhancement processing on the to-be-enhanced graph data by using the fine-tuned large language model to obtain enhanced graph data. The cross-domain heterogeneous data is used for graph data enhancement, additional semantic information is provided for the nodes, semantic features of the nodes in the to-be-enhanced graph data are enriched, the nodes are represented more comprehensively and deeply, the accuracy of the corresponding node features is improved, and therefore the accuracy of the enhanced graph data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a graph data enhancement method, device, equipment and medium. Background Art

[0002] With the wide application of graph neural networks, the quality of training data has become a key factor affecting the performance of the model. In practical applications, graph data often faces problems such as missing node features, sparse labels, and noise interference. Traditional data enhancement methods mostly rely on manual annotation or simple feature extension, which have problems such as low efficiency, feature redundancy, and the risk of introducing noise. Existing technologies such as knowledge graph-based feature alignment methods (such as JedAI, MAGNN) can partially solve the problem of missing features, but cannot dynamically screen key features and have insufficient adaptability to heterogeneous data, resulting in low accuracy of graph data. Therefore, in the process of enhancing graph data, how to improve the accuracy of graph data has become an urgent problem to be solved. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a graph data enhancement method, device, equipment and medium to solve the problem of low accuracy of graph data in the process of enhancing graph data. In a first aspect, an embodiment of the present invention provides a graph data enhancement method, and the graph data enhancement method includes: Obtain the graph data to be enhanced and preset graph data, where the graph data to be enhanced includes a first node set, and the preset graph data includes a second node set; For any first target node in the first node set, determine a second target node in the second node set that matches the first target node; According to the preset graph data, determine the features of the second target node, write the features of the second target node into the first target node to obtain a candidate feature set of the first target node, and traverse all the first target nodes to obtain a candidate feature set of each first target node; Obtain a preset large language model, fine-tune the model parameters of the large language model according to the candidate feature set of each first target node to obtain a fine-tuned large language model, and use the fine-tuned large language model to perform data enhancement processing on the graph data to be enhanced to obtain the enhanced graph data of the graph data to be enhanced.

[0004] In a second aspect, an embodiment of the present invention provides a graph data enhancement device, and the graph data enhancement device includes: An acquisition module, configured to obtain the graph data to be enhanced and preset graph data, where the graph data to be enhanced includes a first node set, and the preset graph data includes a second node set; A matching module, configured to determine, for any first target node in the first node set, a second target node that matches the first target node in the second node set; A writing module, configured to determine the features of the second target node according to the preset graph data, write the features of the second target node into the first target node to obtain a candidate feature set of the first target node, and traverse all the first target nodes to obtain the candidate feature set of each first target node; An enhancement module, configured to obtain a preset large language model, fine-tune the model parameters of the large language model according to the candidate feature set of each first target node to obtain a fine-tuned large language model, and use the fine-tuned large language model to perform data enhancement processing on the graph data to be enhanced to obtain the enhanced graph data of the graph data to be enhanced.

[0005] In a third aspect, an embodiment of the present invention provides a computer device, where the computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the graph data enhancement method described in the first aspect is implemented.

[0006] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the graph data enhancement method described in the first aspect is implemented.

[0007] The beneficial effects of the present invention compared with the prior art are as follows: In this application, the graph data to be enhanced is aligned with the preset graph data, the matching nodes in the graph data to be enhanced and the preset graph data are determined, the features corresponding to the preset graph data in the matching nodes are written into the corresponding nodes of the graph data to be enhanced to obtain the candidate features of the corresponding nodes of the graph data to be enhanced, a preset large language model is obtained, the large language model is fine-tuned according to the candidate feature set to obtain a fine-tuned large language model, and the fine-tuned large language model is used to perform data enhancement processing on the graph data to be enhanced to obtain the enhanced graph data of the graph data to be enhanced. Cross-domain heterogeneous data is used for graph data enhancement, providing additional semantic information for the nodes, enriching the semantic features of the nodes in the graph data to be enhanced, making the node representation more comprehensive and in-depth, improving the accuracy of the features of the corresponding nodes, and thus improving the accuracy of the enhanced graph data. Description of the Drawings

[0008] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0009] Figure 1 It is a schematic diagram of the application environment of a graph data enhancement method provided by an embodiment of the present invention; Figure 2 It is a schematic flowchart of a graph data enhancement method provided by an embodiment of the present invention; Figure 3 It is a graph data enhancement device provided by an embodiment of the present invention; Figure 4 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present invention. Detailed implementation manners

[0010] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0011] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are proposed to thoroughly understand the embodiments of the present invention. However, those skilled in the art should understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.

[0012] It should be understood that when used in the specification and claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0013] It should also be understood that the term "and / or" used in the specification and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0014] As used in the specification of the present invention and the appended claims, the term "if" may be construed, depending on the context, as "when", or "once", or "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [a described condition or event] is detected" may be construed, depending on the context, as meaning "once determined", or "in response to determining", or "once [a described condition or event] is detected", or "in response to detecting [a described condition or event]".

[0015] In addition, in the description of the specification of the present invention and the appended claims, the terms "first", "second", "third", etc. are used only for distinguishing descriptions, and cannot be construed as indicating or implying relative importance.

[0016] Reference to "one embodiment" or "some embodiments" or the like described in the specification of the present invention means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0017] Embodiments of the present invention may acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results.

[0018] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, mechatronics, etc. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0019] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution is prior or subsequent. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0020] In order to illustrate the technical solutions of the present invention, specific embodiments will be used for illustration below.

[0021] A method for graph data enhancement provided by an embodiment of the present invention can be applied in an application environment such as Figure 1 , where the client communicates with the server. The client includes, but is not limited to, computer devices such as a palmtop computer, a desktop computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a cloud computer device, a personal digital assistant (PDA), etc. The server can be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0022] See Figure 2 , which is a schematic flowchart of a method for graph data enhancement provided by an embodiment of the present invention. The above graph data enhancement method can be applied to the server in Figure 1 , as shown in Figure 2 , and the graph data enhancement method can include the following steps.

[0023] S201: Obtain the graph data to be enhanced and the preset graph data. The graph data to be enhanced includes a first node set, and the preset graph data includes a second node set.

[0024] In step S201, obtain the graph data to be enhanced and the preset graph data. Among them, the preset graph data is known graph data. The graph data to be enhanced includes a first node set, and the preset graph data includes a second node set. The first node set is the set of all nodes in the graph data to be enhanced, and the second node set is the set of all nodes in the preset graph data.

[0025] In this embodiment, obtain the graph data to be enhanced and the preset graph data. Among them, the preset graph data is a large-scale and multi-domain knowledge external knowledge graph. The graph data to be enhanced and the preset graph data are in graph structure. The representation of the graph structure is G = (V, E, T, FA), where V is the node set, and each node is a corresponding entity, such as a commodity, a user, etc.; E is the edge set, and each edge is the relationship between the corresponding nodes, such as purchase, evaluation, belonging to a category, etc.; T is the label set, such as the category to which the commodity belongs, etc. FA is the attribute set, such as the price of the commodity node, the age range of the user node, etc.

[0026] It should be noted that the preset graph data is heterogeneous graph data and is used for data enhancement of the graph data to be enhanced.

[0027] In this embodiment, the to-be-enhanced graph data and the preset graph data are obtained to update the nodes in the to-be-enhanced graph data using the preset graph data.

[0028] S202: For any first target node in the first node set, determine a second target node in the second node set that matches the first target node.

[0029] In step S202, the to-be-enhanced graph data and the preset graph data are aligned to determine the matching nodes in the to-be-enhanced graph data and the preset graph data, that is, the second target node that matches the first target node.

[0030] In this embodiment, for any first target node in the first node set, the preset graph neural network is used to extract the features of the first target node to obtain the first feature of the first target node, and the preset graph neural network is used to extract the features of each node in the second node set to obtain the second feature of each node in the second node set. When determining the second target node in the second node set that matches the first target node, calculate the similarity between the first feature and the second feature, and the node corresponding to the maximum similarity is the matching node, obtaining the second target node.

[0031] In this embodiment, by extracting the features of the nodes, according to the corresponding features, the node with the maximum feature similarity is used as the matching point, and the second target node that matches the first target node is determined in the second node set, so as to use the features of the second target node as the candidate features of the first target node. Due to the matching between the first target node and the second target node, the accuracy of the candidate features of the first target node is improved.

[0032] Optionally, determining the second target node in the second node set that matches the first target node includes: Determine the set of nodes to be matched in the second node set that are similar to the first target node; According to the to-be-enhanced graph data, determine the optimal first node path starting from the first target node, and according to the preset graph data, determine the optimal node path to be matched starting from each node to be matched in the set of nodes to be matched; Calculate the relevance between the optimal first node path and each optimal node path to be matched, and according to the relevance, determine that the optimal node path to be matched corresponding to the maximum value of the relevance is the target node path to be matched that matches the optimal first node path; Determine the node to be matched in the target node path to be matched as the second target node.

[0033] In this embodiment, a set of nodes to be matched in the second node set similar to the first target node is determined, where the set of nodes to be matched is a set of nodes with a relatively large similarity to the first target node. When calculating the similarity between the nodes in the second node set and the first target node, the similarity between the text features of the labels of the nodes can be calculated, that is, using a text feature extraction model to extract the features of the label of the first target node to obtain the first text feature, and extract the text features of the labels of the nodes in the second node set to obtain the second text feature. Among them, the text feature extraction model is a trained neural network model, such as the Sentence-BERT model. The similarity between the first text feature and the second text feature is calculated, and the calculation formula is as follows: Where is the first target node, is the node in the second node set, is the similarity between the first target node and the node in the second node set, is the first text feature, is the second text feature.

[0034] Traverse all the nodes in the second node set to obtain the similarity between the first target node and each node in the second node set, and determine the nodes in the second node set with a similarity greater than the preset threshold as the set of nodes to be matched. Among them, the size of the preset threshold can be determined according to the actual situation and is not limited in this embodiment.

[0035] According to the data of the graph to be enhanced, determine the optimal first node path starting from the first target node, and according to the preset graph data, determine the optimal node path to be matched starting from each node to be matched in the set of nodes to be matched. Among them, the optimal first node path has the best semantic association among the node paths starting from the first target node, that is, the strongest semantic association between nodes, and the optimal node path to be matched has the best semantic association among the node paths starting from the node to be matched, that is, the strongest semantic association between nodes.

[0036] After determining the optimal first node path starting from the first target node and the optimal node path to be matched starting from each node to be matched in the set of nodes to be matched, calculate the value of the association between the optimal first node path and each optimal node path to be matched, and according to the value of the association, determine the optimal node path to be matched corresponding to the maximum value of the association value as the target node path to be matched that matches the optimal first node path.

[0037] Among them, when calculating the relevance between the optimal first node path and each optimal node path to be matched, it can be determined according to the similarity between the corresponding nodes, that is, perform node matching on the nodes between the optimal first node path and the optimal node path to be matched. For example, the start node of the optimal first node path is matched with the start node of the optimal node path to be matched, and the second node of the optimal first node path is matched with the second node of the optimal node path to be matched, and so on. Calculate the similarity between the matched nodes, and determine the average value of the similarities of all the matched nodes as the magnitude of the relevance between the optimal first node path and each optimal node path to be matched.

[0038] It should be noted that when calculating the similarity between the matched nodes, the similarity of the labels of the corresponding nodes can be calculated, or other methods can be used to calculate the similarity, which is not limited in this embodiment.

[0039] In this embodiment, according to the relevance, the optimal node path to be matched corresponding to the maximum value of the relevance value is determined as the target node path to be matched that matches the optimal first node path; Determine the node to be matched in the target node path to be matched as the second target node, that is, the second target node of the second node set that matches the first target node.

[0040] In this embodiment, through the node path, the first target node in the matched first node set and the second target node of the second node set are determined, comprehensively considering the similarity between the first target node and the second target node, and improving the accuracy of determining the second target node that matches the first target node.

[0041] Optionally, according to the graph data to be enhanced, determining the optimal first node path starting from the first target node includes: According to the graph data to be enhanced, determine multiple initial first node paths starting from the first target node; For any initial first node path, calculate the first semantic relevance in the initial first node path; Take the initial first node path corresponding to the maximum value of the first semantic relevance value as the optimal first node path.

[0042] In this embodiment, according to the graph data to be enhanced, determine multiple initial first node paths starting from the first target node. Among them, when determining each initial first node path, search and match all connected and strongly semantically related descendant nodes starting from the first target node. The search termination condition is that if the node contains too many descendant nodes and adding these nodes will weaken the semantic connection between the node and the first target node, then stop the search. It should be noted that the initial first node path is obtained by searching through a preset language model. The preset language model is a neural network model used to search for descendant nodes that start from a node and match all connected and strongly semantically related nodes. One or two or more node paths can be generated.

[0043] For any initial first node path, calculate the first semantic relevance in the initial first node path. The formula for calculating the first semantic relevance is as follows: Where is the value of the first semantic relevance in the initial first node path ; is the set of descendant nodes of the node in the initial first node path, that is, the set of nodes connected to the node starting from the node and connected to the node ; is the number of nodes in the initial first node path.

[0044] Calculate the first semantic relevance in each initial first node path, and take the initial first node path corresponding to the maximum value of the first semantic relevance as the optimal first node path.

[0045] In this embodiment, quantifying the semantic relevance in the initial first node path facilitates determining the optimal first node path according to the quantified value of the semantic relevance.

[0046] Optionally, according to the preset graph data, determining the optimal path of the nodes to be matched connected to each node to be matched in the set of nodes to be matched includes: For any node to be matched in the set of nodes to be matched, according to the preset graph data, determine multiple initial paths of the nodes to be matched connected to the node to be matched as the start node; For any initial path of the nodes to be matched, calculate the second semantic relevance in the initial path of the nodes to be matched; Take the initial path of the nodes to be matched corresponding to the maximum value of the second semantic relevance as the optimal path of the nodes to be matched corresponding to the node to be matched; Traverse each node to be matched to obtain the optimal path of the nodes to be matched connected to each node to be matched in the set of nodes to be matched as the start node.

[0047] In this embodiment, according to the preset graph data, multiple initial paths of nodes to be matched are determined with the node to be matched as the starting node. When determining each initial path of nodes to be matched, starting from the node to be matched, all connected descendant nodes with strong semantic relationships are searched and matched. The search termination condition is that if the node contains too many descendant nodes and adding these nodes will weaken the semantic connection between the node and the node to be matched, the search stops. It should be noted that the initial paths of nodes to be matched are obtained by searching with a preset language model. The preset language model is a neural network model, which is used to search for all connected descendant nodes with strong semantic relationships starting from a node. One or two or multiple node paths can be generated.

[0048] For any initial path of nodes to be matched, the second semantic relevance in the initial path of nodes to be matched is calculated. The formula for calculating the second semantic relevance is as follows: Where is the value of the second semantic relevance in the initial path of nodes to be matched ; is the set of descendant nodes of the node in the initial path of nodes to be matched, that is, the set of nodes connected to the node starting from the node ; is the number of nodes in the initial path of nodes to be matched.

[0049] The second semantic relevance in each initial path of nodes to be matched is calculated, and the initial path of nodes to be matched corresponding to the maximum value of the first semantic relevance is taken as the optimal path of nodes to be matched.

[0050] Each node to be matched is traversed to obtain the optimal path of nodes to be matched with each node to be matched in the set of nodes to be matched as the starting node.

[0051] In this embodiment, the semantic relevance in the initial path of nodes to be matched is quantified, which is convenient for determining the optimal path of nodes to be matched according to the quantified value of the semantic relevance.

[0052] Optionally, calculating the relevance between the optimal first node path and each optimal path of nodes to be matched includes: Obtaining the mapping relationship between the node path relevance and the node paths; According to the mapping relationship, calculating the relevance between the optimal first node path and each optimal path of nodes to be matched.

[0053] In this embodiment, the mapping relationship between the node path relevance and the node path is obtained, and the formula of the mapping relationship is as follows: Wherein, is the node path and the node path the relevance between them, is the encoding of the label sequence of the node path , is the encoding of the label sequence of the node path , is a neural network for calculating and the similarity between them, and the output is limited to [0,1], is the length of the node path , that is, the number of nodes included in the node path , is the length of the node path , that is, the number of nodes included in the node path .

[0054] It should be noted that in this embodiment, the BERT model can be used to encode the label sequence of the node path , and encode the label sequence of the node path . The neural network includes 3 fully connected layers, and the output layer uses the Sigmoid function.

[0055] Substitute the optimal first node path and the optimal node path to be matched into the formula of the mapping relationship, and calculate the relevance between the optimal first node path and each optimal node path to be matched.

[0056] In this embodiment, the mapping relationship between the node path relevance and the node path is obtained. Among them, the mapping relationship includes the similarity between the label encodings of the node path and the number of nodes. Since the greater the similarity of the label encoding, the greater the node path relevance, and the more nodes, the weaker the semantic association of the node path, thus making the node path relevance smaller. Therefore, the node path relevance is positively correlated with the similarity between the label encodings of the node path and negatively correlated with the number of nodes, which improves the calculation accuracy of the node path relevance.

[0057] S203: According to the preset graph data, determine the features of the second target node, write the features of the second target node into the first target node, obtain the candidate feature set of the first target node, and traverse all the first target nodes to obtain the candidate feature set of each first target node.

[0058] In step S203, according to the preset graph data, determine the features of the second target node, where the features of the second target node are the features of the node path containing the second target node. Write the features of the second target node into the first target node to obtain the candidate feature set of the first target node.

[0059] In this embodiment, according to the preset graph data, determine the features of the second target node. When determining the features of the second target node, use a language model to process each node path containing the second target node to obtain the features of the second target node. For example, if the second target node is the entity "mobile phone" and there is a path to 2020: mobile phone → release_year → 2020, the language model can be used to transform the path into "release_year: 2020". Traverse each node path containing the second target node to obtain the features of the second target node. Among them, the language model is a neural network model. Or use graph embedding technology to vectorize each node path containing the second target node to obtain the features of the second target node. Other methods can also be used to obtain the corresponding features of the second target node, which is not limited in this embodiment. Write the features of the second target node into the first target node to obtain the candidate feature set of the first target node, and traverse all the first target nodes to obtain the candidate feature set of each first target node.

[0060] In this embodiment, since the second target node and the second target node are matching nodes, the features of the second target node are used as the candidate features of the first target node. When performing data augmentation on the first target node according to the features of the second target node, the accuracy of data augmentation can be improved.

[0061] S204: Obtain a preset large language model, fine-tune the model parameters of the large language model according to the candidate feature set of each first target node to obtain a fine-tuned large language model, and use the fine-tuned large language model to perform data augmentation processing on the graph data to be augmented to obtain the augmented graph data of the graph data to be augmented.

[0062] In step S204, the fine-tuning process is to adjust the model parameters in the preset large language model so that the large language model can learn the feature weights, so as to select the key features required for the downstream task or generate the missing key features to obtain the augmented graph data of the graph data to be augmented.

[0063] In this embodiment, a pre-set large language model is obtained. The large language model can be a lightweight model, such as the Mistral-7B model. When fine-tuning the model parameters of the large language model according to the candidate feature set of each first target node, a graph neural network model is selected as the downstream task model. If the downstream task is a classification task, a graph neural network classification model can be selected as the downstream task model. The candidate feature set is used as the feature of the first target node, and the graph neural network classification model is used for classification processing to output the first classification prediction result. The large language model is used for classification processing to output the second classification prediction result. According to the difference between the first classification prediction result and the second classification prediction result, the model parameters of the large language model are adjusted so that the difference between the first classification prediction result and the second classification prediction result meets the preset requirements. Here, the preset requirements can be that the difference between the first classification prediction result and the second classification prediction result is less than the preset threshold.

[0064] It should be noted that when fine-tuning the large language model, the QLoRA method can be used for fine-tuning. QLoRA allows us to use 4-bit quantization + LoRA adapters, and only a small number of parameters need to be trained to achieve an effect similar to full-parameter fine-tuning. It has been verified that this embodiment can be deployed on a single 24GB GPU and can handle the fine-tuning of graphs with more than 700,000 entities and more than 1.4 million relationship attributes. In addition, as the number of GPUs increases, the fine-tuning time can be linearly reduced. Therefore, this patent has good practicability and scalability.

[0065] The large language model is fine-tuned to obtain a fine-tuned large language model. The fine-tuned large language model learns the feature weights, that is, it can output appropriate features for downstream tasks. The fine-tuned large language model is used to perform data augmentation processing on the graph data to be enhanced, that is, the fine-tuned large language model outputs the corresponding features for each first target node in the graph data to be enhanced and writes them into the corresponding first target nodes, outputs the key features required for downstream tasks, and performs feature enhancement on the first target nodes in the graph data to be enhanced, thereby obtaining enhanced graph data.

[0066] It should be noted that when using the fine-tuned large language model to perform data augmentation processing on the graph data to be enhanced, feature selection and semantic extraction can be performed, that is, the key features of the downstream tasks can be selected from the corresponding candidate feature sets. The formula is as follows: Where is the large language model, are the model parameters of the large language model, is the text corresponding to the feature in the candidate feature set, is the candidate feature set, is the selected key feature, It is the m-th key feature selected from the candidate feature set of node u.

[0067] In this embodiment, when using the fine-tuned large language model to perform data augmentation on the graph data to be augmented, when the corresponding feature does not exist in the candidate feature set of a node, the fine-tuned large language model can perform semantic feature extraction. For example, in the knowledge graph KG, the relevant information of some less well-known people or products is often sparse. At this time, the fine-tuned large language model will try to semantically supplement the attributes that are judged to be important but missing on this entity. For example, if the keywords of paper A are missing, the fine-tuned large language model can extract and generate keywords based on the existing abstract content to supplement the missing key features of the corresponding node.

[0068] After using the fine-tuned large language model to perform data augmentation on the graph data to be augmented and obtaining the augmented graph data of the graph data to be augmented, according to different downstream tasks, the augmented graph data can also be feature-embedded. For example, if the downstream task is a classification task, specific embedding models such as SimCSE can be used to convert the output m key features into feature vectors with a fixed dimension. The dimension of this vector is independent of the number of key features and depends on the embedding model or the GNN classification model, with good scalability. Of course, the Prompt of the large language model can also be modified to directly embed the features as attributes into the graph data and output it as the augmented graph data for use in other scenarios.

[0069] In this embodiment, according to the candidate feature set, data augmentation is performed on the graph data to be augmented, that is, the key features required for the downstream task are output at each first target node, the first target nodes in the graph data to be augmented are feature-augmented, and the fine-tuned large language model can also perform semantic feature extraction to supplement the missing key features of the corresponding nodes, enriching the semantic information of the first target nodes and improving the accuracy of the augmented graph data.

[0070] Optionally, obtaining a preset large language model and fine-tuning the model parameters of the large language model according to the candidate feature set of each first target node to obtain the fine-tuned large language model includes: Obtaining the downstream task model of the large language model; For any candidate feature set, input the features in the candidate feature set into the downstream task model to output a first prediction result, input the features in the candidate feature set into the large language model to output a feature vector, and input the feature vector into the downstream task model to output a second prediction result; Obtaining an objective function, where the objective function is used to represent the conditions that need to be satisfied when fine-tuning the model parameters of the large language model; According to the candidate feature set, the first prediction result, and the second prediction result, combined with the objective function, fine-tune the model parameters of the large language model to obtain the fine-tuned large language model.

[0071] In this embodiment, a corresponding downstream task model is set for the large language model, such as a graph neural network classification model. The graph neural network classification model is used for classification tasks. Input the features in the candidate feature set into the downstream task model to output the first prediction result. Input the features in the candidate feature set into the large language model to output a feature vector. Input the feature vector into the downstream task model to output the second prediction result.

[0072] Obtain the objective function, which is used to represent the conditions that need to be satisfied when fine-tuning the model parameters of the large language model. Among them, the formula of the objective function is as follows: Among them, is the value of the objective function. The condition that needs to be satisfied when fine-tuning the model parameters of the large language model is the value of is the smallest, is the model parameter of the large language model, in the graph data to be enhanced , is the prediction result output after the large language model processes the features in node u, that is, the first prediction result, is the indicator function. When , that is, when the prediction result output after processing the features extracted by the large language model is equal to the prediction result output by the downstream task model, it takes the value of 1, otherwise it takes the value of 0, is the feature vector obtained by using the large language model to extract the features of node u. is the downstream task model, is the large language model, is the prediction result output after using the downstream task model to process the features extracted by the large language model, that is, the second prediction result, is the prediction result output by the downstream task model, that is, the first prediction result. is the feature of node u output by the large language model after setting perturbations to node u, that is, the feature required by the downstream task, , is the noise vector sampled from the uniform distribution U(−1,1). L represents the length of the feature sequence, and d is the dimension of the feature sequence. These two parameters are used to control the noise scale, so as to evaluate the impact of noise on embedding vectors of different lengths and dimensions. is an adjustable parameter that allows the large language model to dynamically adjust the noise intensity according to the training data, so as to more accurately adapt to the downstream classification task. is the set of candidate features for node u. is for determining the features output by the large language model After that, it is the probability that the prediction result output after processing the features extracted by the large language model using the downstream task model is equal to the prediction result output by the downstream task model.

[0073] According to the set of candidate features, the first prediction result and the second prediction result, combined with the objective function, fine-tune the model parameters of the large language model until the value of f is minimized, that is, until the value of the objective function calculated according to the set of candidate features, the first prediction result and the second prediction result is minimized, and the fine-tuned large language model is obtained.

[0074] In this embodiment, the model parameters of the large language model are fine-tuned according to the downstream task. After the model parameters of the large language model are fine-tuned, the large language model can learn the key features required by the downstream task model, so that the large language model can automatically output the key features required by the downstream task, automatically enrich the features of the nodes, and improve the efficiency of data augmentation.

[0075] In this application, align the graph data to be augmented with the preset graph data, determine the matching nodes in the graph data to be augmented and the preset graph data, write the features corresponding to the preset graph data in the matching nodes into the corresponding nodes of the graph data to be augmented, obtain the candidate features of the corresponding nodes of the graph data to be augmented, obtain the preset large language model, fine-tune the large language model according to the set of candidate features, obtain the fine-tuned large language model, and use the fine-tuned large language model to perform data augmentation processing on the graph data to be augmented, and obtain the augmented graph data of the graph data to be augmented. Using cross-domain heterogeneous data for graph data augmentation provides additional semantic information for the nodes, enriches the semantic features of the nodes in the graph data to be augmented, makes the node representation more comprehensive and in-depth, improves the accuracy of the features of the corresponding nodes, and thus improves the accuracy of the augmented graph data.

[0076] Please refer to Figure 3 , Figure 3 is a graph data augmentation device provided by an embodiment of the present invention, and this graph data augmentation device corresponds one-to-one to the sample generation method in the above embodiment. Specifically, please refer to Figure 2 and Figure 2 For the relevant descriptions in the corresponding embodiments. For the sake of illustration, only the parts related to this embodiment are shown. See Figure 3 , the graph data augmentation device 30 includes: an acquisition module 31, a matching module 32, a writing module 33, and an augmentation module 34.

[0077] An acquisition module 31 for acquiring the graph data to be enhanced and preset graph data, where the graph data to be enhanced includes a first node set, and the preset graph data includes a second node set.

[0078] A matching module 32 for, for any first target node in the first node set, determining a second target node in the second node set that matches the first target node.

[0079] A writing module 33 for, according to the preset graph data, determining the features of the second target node, writing the features of the second target node into the first target node to obtain a candidate feature set of the first target node, and traversing all the first target nodes to obtain a candidate feature set for each first target node.

[0080] An enhancement module 34 for obtaining a preset large language model, fine-tuning the model parameters of the large language model according to the candidate feature set of each first target node to obtain a fine-tuned large language model, and using the fine-tuned large language model to perform data enhancement processing on the graph data to be enhanced to obtain enhanced graph data of the graph data to be enhanced.

[0081] Optionally, the above-mentioned matching module 32 includes: A first determination unit for determining a set of nodes to be matched in the second node set that are similar to the first target node.

[0082] A second determination unit for, according to the graph data to be enhanced, determining an optimal first node path connected with the first target node as the start node, and according to the preset graph data, determining an optimal node path to be matched connected with each node to be matched in the set of nodes to be matched as the start node.

[0083] A calculation unit for calculating the relevance between the optimal first node path and each optimal node path to be matched, and according to the relevance, determining the optimal node path to be matched corresponding to the maximum value of the relevance as the target node path to be matched that matches the optimal first node path.

[0084] A third determination unit for determining the node to be matched in the target node path to be matched as the second target node.

[0085] Optionally, the above-mentioned second determination unit includes: A first determination subunit for, according to the graph data to be enhanced, determining multiple initial first node paths connected with the first target node as the start node.

[0086] A first calculation subunit for, for any initial first node path, calculating the first semantic relevance in the initial first node path.

[0087] The first selection subunit is used to select the initial first node path corresponding to the maximum value of the first semantic relevance as the optimal first node path.

[0088] Optionally, the above-mentioned second determination unit includes: The second determination subunit is used to determine multiple initial to-be-matched node paths connected with any to-be-matched node in the to-be-matched node set as the start node according to the preset graph data.

[0089] The second calculation subunit is used to calculate the second semantic relevance in the initial to-be-matched nodes for any initial to-be-matched node path.

[0090] The second selection subunit is used to select the initial to-be-matched node path corresponding to the maximum value of the second semantic relevance as the optimal to-be-matched node path for the corresponding to-be-matched node.

[0091] The traversal subunit is used to traverse each to-be-matched node to obtain the optimal to-be-matched node paths connected with each to-be-matched node in the to-be-matched node set as the start node.

[0092] Optionally, the above-mentioned calculation unit includes: The acquisition subunit is used to acquire the mapping relationship between the node path relevance and the node paths.

[0093] The third calculation subunit is used to calculate the relevance between the optimal first node path and each optimal to-be-matched node path according to the mapping relationship.

[0094] Optionally, the above-mentioned enhancement module 34 includes: The first acquisition unit is used to acquire the downstream task model of the large language model; The output unit is used to input the features in any candidate feature set into the downstream task model for any candidate feature set, output the first prediction result, input the features in the candidate feature set into the large language model, output the feature vector, and input the feature vector into the downstream task model to output the second prediction result.

[0095] The second acquisition unit is used to acquire the objective function, and the objective function is used to characterize the conditions that need to be met when fine-tuning the model parameters of the large language model.

[0096] The fine-tuning unit is used to fine-tune the model parameters of the large language model according to the candidate feature set, the first prediction result and the second prediction result, in combination with the objective function, to obtain the fine-tuned large language model.

[0097] It should be noted that, for the information interaction and execution process between the above-mentioned units, etc., since they are based on the same concept as the method embodiments of the present invention, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be elaborated here.

[0098] Figure 4 is a schematic structural diagram of a computer device provided by an embodiment of the present invention. As Figure 4 shown, the computer device of this embodiment includes: at least one processor ( Figure 4 only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, it implements the steps in any of the above-mentioned method embodiments for graph data enhancement.

[0099] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 4 merely examples of computer devices are not intended to limit the computer device. The computer device may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include a network interface, a display screen, and an input device, etc.

[0100] The so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0101] The memory includes a readable storage medium, an internal memory, etc. Among them, the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be the hard disk of a computer device, and in some other embodiments, it can also be an external storage device of the computer device. For example, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device. Further, the memory can also include both the internal storage unit of the computer device and an external storage device. The memory is used to store the operating system, application programs, a BootLoader, data, and other programs, such as the program code of a computer program. The memory can also be used to temporarily store the data that has been output or will be output.

[0102] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above-mentioned device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0103] All or part of the processes in the above method embodiments of this application can also be completed by a computer program product. When the computer program product runs on a computer device, it enables the computer device to execute and implement the steps in the above method embodiments.

[0104] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0105] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0106] In the embodiments provided in this application, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the device or unit can be in electrical, mechanical or other forms.

[0107] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0108] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.

Claims

1. A graph data augmentation method, characterized in that, The described graph data augmentation method includes: Obtain the graph data to be augmented and preset graph data, where the graph data to be augmented includes a first node set, and the preset graph data includes a second node set; For any first target node in the first node set, determine a second target node in the second node set that matches the first target node; According to the preset graph data, determine the features of the second target node, write the features of the second target node into the first target node to obtain a candidate feature set of the first target node, and traverse all the first target nodes to obtain the candidate feature set of each first target node; Obtain a preset large language model, fine-tune the model parameters of the large language model according to the candidate feature set of each first target node to obtain a fine-tuned large language model, and use the fine-tuned large language model to perform data augmentation processing on the graph data to be augmented to obtain the augmented graph data of the graph data to be augmented.

2. The graph data augmentation method according to claim 1, wherein The determining of the second target node that matches the first target node in the second node set includes: Determine a set of nodes to be matched in the second node set that are similar to the first target node; According to the graph data to be augmented, determine the optimal first node path starting from the first target node, and according to the preset graph data, determine the optimal node path to be matched starting from each node to be matched in the set of nodes to be matched; Calculate the relevance between the optimal first node path and each optimal node path to be matched, and according to the relevance, determine that the optimal node path to be matched corresponding to the maximum value of the relevance is the target node path to be matched that matches the optimal first node path; Determine the node to be matched in the target node path to be matched as the second target node.

3. The graph data enhancement method according to claim 2, wherein The determining of the optimal first node path starting from the first target node according to the graph data to be augmented includes: According to the graph data to be augmented, determine multiple initial first node paths starting from the first target node; For any initial first node path, calculate the first semantic relevance in the initial first node path; Take the initial first node path corresponding to the maximum value of the first semantic relevance as the optimal first node path.

4. The graph data augmentation method according to claim 2, wherein The determining of the optimal node path to be matched starting from each node to be matched in the set of nodes to be matched according to the preset graph data includes: For any node to be matched in the set of nodes to be matched, according to the preset graph data, determine multiple initial node paths to be matched starting from the node to be matched; For any initial node path to be matched, calculate the second semantic relevance in the initial node to be matched; Take the initial node path to be matched corresponding to the maximum value of the second semantic relevance as the optimal node path to be matched for the corresponding node to be matched; Traverse each node to be matched to obtain the optimal node path to be matched starting from each node to be matched in the set of nodes to be matched.

5. The graph data enhancement method according to claim 2, wherein Calculating the relevance between the optimal first node path and each optimal node path to be matched includes: Obtaining the mapping relationship between the node path relevance and the node paths; Calculating the relevance between the optimal first node path and each optimal node path to be matched according to the mapping relationship.

6. The graph data augmentation method according to claim 1, wherein Obtaining a preset large language model and fine-tuning the model parameters of the large language model according to the candidate feature sets of each first target node to obtain a fine-tuned large language model includes: Obtaining the downstream task model of the large language model; For any candidate feature set, inputting the features in the candidate feature set into the downstream task model to output a first prediction result, inputting the features in the candidate feature set into the large language model to output a feature vector, and inputting the feature vector into the downstream task model to output a second prediction result; Obtaining an objective function, where the objective function is used to characterize the conditions to be satisfied when fine-tuning the model parameters of the large language model; Fine-tuning the model parameters of the large language model according to the candidate feature set, the first prediction result, and the second prediction result, in combination with the objective function, to obtain a fine-tuned large language model.

7. A graph data augmentation device, characterized in that, The graph data augmentation device includes: An acquisition module for acquiring the graph data to be augmented and preset graph data, where the graph data to be augmented includes a first node set, and the preset graph data includes a second node set; A matching module for determining, for any first target node in the first node set, a second target node in the second node set that matches the first target node; A writing module for determining the features of the second target node according to the preset graph data, writing the features of the second target node into the first target node to obtain a candidate feature set of the first target node, and traversing all the first target nodes to obtain the candidate feature sets of each first target node; An augmentation module for obtaining a preset large language model, fine-tuning the model parameters of the large language model according to the candidate feature sets of each first target node to obtain a fine-tuned large language model, and using the fine-tuned large language model to perform data augmentation processing on the graph data to be augmented to obtain the augmented graph data of the graph data to be augmented.

8. The figure data enhancement device according to claim 7, characterized in that, The matching module includes: A first determination unit for determining a set of nodes to be matched in the second node set that are similar to the first target node; A second determination unit for determining, according to the graph data to be augmented, the optimal first node path starting from the first target node, and determining, according to the preset graph data, the optimal node paths to be matched starting from each node to be matched in the set of nodes to be matched; A calculation unit for calculating the relevance between the optimal first node path and each optimal node path to be matched, and determining, according to the relevance, the optimal node path to be matched corresponding to the maximum value of the relevance as the target node path to be matched that matches the optimal first node path; A third determination unit, configured to determine the node to be matched in the target node path to be matched as a second target node.

9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the graph data enhancement method according to any one of claims 1 to 6 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the graph data enhancement method according to any one of claims 1 to 6 is implemented.