Graph alignment method and device, and method and device for predicting matching points

By iteratively calculating the distance matrix within and between graphs and updating the alignment matrix, the problems of unstable and low efficiency of graph alignment in the prior art are solved, and more efficient and more accurate graph alignment is achieved.

CN120407862APending Publication Date: 2025-08-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410133506.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-30
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

When existing graph alignment technologies deal with large-scale, high-noise, structural changes or mismatched graph data, the accuracy is unstable, low efficiency, and the alignment effect is affected.

Method used

By determining the adjacency and feature matrix of the source domain graph and the target domain graph, the alignment matrix is initialized, and the calculation of the distance matrix in the graph and the distance matrix between graphs is iteratively performed until the target alignment matrix converges, the alignment matrix is updated to improve accuracy and efficiency.

Benefits of technology

It improves the accuracy and processing efficiency of the graph alignment method, enhances robustness, adapts to various graph data, and improves the accuracy and accuracy of graph alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407862A_ABST
    Figure CN120407862A_ABST
Patent Text Reader

Abstract

The invention provides a graph alignment method and device. The graph alignment method comprises the steps that an adjacent matrix and a feature matrix of a corresponding graph in a source domain graph and a target domain graph are determined, the source domain graph comprises a plurality of source domain nodes, and the target domain graph comprises a plurality of target domain nodes. Initializing the alignment matrix as a target alignment matrix, and iteratively executing the following steps until the maximum number of iterations is reached or the target alignment matrix is converged: determining an intra-graph distance matrix of a corresponding graph based on an adjacent matrix and a feature matrix of the corresponding graph in the source domain graph and the target domain graph; determining an inter-graph distance matrix of the iteration based on the intra-graph distance matrix of the source domain graph, the intra-graph distance matrix of the target domain graph and the target alignment matrix; and based on the inter-graph distance matrix of the current iteration, determining an alignment matrix of the current iteration, and updating the target alignment matrix to the alignment matrix of the current iteration. And aligning the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the target alignment matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and more particularly to a graph alignment method and apparatus, a method and apparatus for predicting matching points, a computing device, a computer-readable storage medium, and a computer program product. Background Art

[0002] In the era of the rapid development of Internet technologies, with the development of deep learning, graph alignment technologies have been widely applied and received extensive attention in many fields (such as social networks, commodity networks, and the expression of protein structures, etc.).

[0003] Common graph alignment technologies mainly display learning node embeddings and minimize the embedding distances between matching nodes. However, since these graph alignment technologies directly align two embedding spaces, the accuracy of the alignment results is unstable, and a poor accuracy will be generated, and the alignment efficiency is too low. In addition, when dealing with graph data with large scale, a lot of noise, structural changes or scale mismatches, the alignment effect may be significantly affected. These problems and drawbacks limit the further development of graph alignment technologies. Summary of the Invention

[0004] In view of this, the present disclosure provides a graph alignment method and apparatus, a method and apparatus for predicting matching points, a computing device, a computer-readable storage medium, and a computer program product, so as to alleviate, mitigate or even eliminate some or all of the above problems and other possible problems.

[0005] According to an aspect of the present disclosure, a graph alignment method is provided, which includes: determining the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph, where the source domain graph includes a plurality of source domain nodes, the target domain graph includes a plurality of target domain nodes, each element in the adjacency matrix represents the connection relationship between nodes in the corresponding graph, and each element in the feature matrix represents the feature of the nodes in the corresponding graph. Initializing an alignment matrix as the target alignment matrix, and iteratively performing the following steps until the maximum number of iterations is reached or the target alignment matrix converges: based on the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph, determining the intra-graph distance matrix of the corresponding graph, where each element in the intra-graph distance matrix represents the association degree between nodes in the corresponding graph; based on the intra-graph distance matrix of the source domain graph, the intra-graph distance matrix of the target domain graph, and the target alignment matrix, determining the inter-graph distance matrix of the current iteration, where each element in the inter-graph distance matrix represents the association degree between the source domain nodes of the corresponding source domain graph and the target domain nodes of the target domain graph; based on the inter-graph distance matrix of the current iteration, determining the alignment matrix of the current iteration, and updating the target alignment matrix to the alignment matrix of the current iteration, where each element in the alignment matrix represents the matching probability between the source domain nodes of the corresponding source domain graph and the target domain nodes of the target domain graph. Aligning the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the target alignment matrix.

[0006] According to some embodiments of the present disclosure, the determining the intra-graph distance matrix of the corresponding graph based on the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph includes: based on the feature matrix of the corresponding graphs in the source domain graph and the target domain graph, determining the node similarity matrix of the corresponding graph, where each element in the node similarity matrix represents the similarity between the features of the nodes in the corresponding graph; based on the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph, determining the splicing matrix of the corresponding graph, where each element in the splicing matrix represents the relationship between a node in the corresponding graph and its neighbor nodes; determining the intra-graph distance matrix of the corresponding graph based on the adjacency matrix, the node similarity matrix, and the splicing matrix of the corresponding graph.

[0007] According to some embodiments of the present disclosure, the determining the intra-graph distance matrix of the corresponding graph based on the adjacency matrix, the node similarity matrix, and the splicing matrix of the corresponding graph includes: determining the weighted sum of the adjacency matrix, the node similarity matrix, and the splicing matrix of the corresponding graph as the intra-graph distance matrix of the corresponding graph.

[0008] According to some embodiments of the present disclosure, determining the splicing matrix of the corresponding graph based on the adjacency matrix and the feature matrix of the corresponding graph in the source domain graph and the target domain graph includes: using a deep learning network to sequentially determine the 1st to Kth order neighbor feature matrices of the corresponding graph based on the adjacency matrix and the feature matrix of the corresponding graph in the source domain graph and the target domain graph, where each element in any Mth order neighbor feature matrix represents the feature of the Mth order neighbor node of the node in the corresponding graph, K is a positive integer not less than 1, and M is a positive integer greater than or equal to 1 and less than or equal to K. The Mth order neighbor node of a node refers to the node separated from this node by M - 1 nodes; splicing the feature matrix of the corresponding graph and the 1st to Kth order neighbor feature matrices of the corresponding graph according to weights; using an activation function to activate the spliced matrix to obtain the splicing matrix of the corresponding graph.

[0009] According to some embodiments of the present disclosure, determining the node similarity matrix of the corresponding graph based on the feature matrix of the corresponding graph in the source domain graph and the target domain graph includes: performing a transpose operation on the feature matrix of the corresponding graph in the source domain graph and the target domain graph to obtain the transposed feature matrix of the corresponding graph; performing an inner product operation on the feature matrix of the corresponding graph and the transposed feature matrix of the corresponding graph to obtain the node similarity matrix of the corresponding graph.

[0010] According to some embodiments of the present disclosure, determining the alignment matrix of the current iteration based on the inter - graph distance matrix of the current iteration includes: determining the optimal transport distance between the inter - graph distance matrix of the current iteration and the target alignment matrix, where the optimal transport distance represents the minimum cost of mapping the inter - graph distance matrix of the current iteration to the target alignment matrix; determining the alignment matrix of the current iteration based on the optimal transport distance.

[0011] According to some embodiments of the present disclosure, determining the optimal transport distance between the inter - graph distance matrix of the current iteration and the target alignment matrix includes: performing data normalization processing on the inter - graph distance matrix of the current iteration to obtain the normalized inter - graph distance matrix of the current iteration; determining the optimal transport distance between the normalized inter - graph distance matrix of the current iteration and the target alignment matrix based on the normalized inter - graph distance matrix of the current iteration.

[0012] According to some embodiments of the present disclosure, performing data normalization processing on the inter - graph distance matrix of the current iteration to obtain the normalized inter - graph distance matrix of the current iteration includes: selecting the maximum value from all elements of the inter - graph distance matrix of the current iteration; for each element in the inter - graph distance matrix of the current iteration, based on the maximum value, linearly normalizing each element in the inter - graph distance matrix of the current iteration into a normalized element to obtain the normalized inter - graph distance matrix of the current iteration.

[0013] According to some embodiments of the present disclosure, aligning the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the target alignment matrix includes: for each source domain node in the source domain graph, selecting the highest alignment weight corresponding to each source domain node from all elements of the target alignment matrix, where the highest alignment weight represents the alignment probability with the highest probability of aligning the source domain nodes in the source domain graph with the target domain nodes in the target domain graph; based on the highest alignment weight corresponding to each source domain node, selecting one target domain node that matches each source domain node respectively to obtain a matching node pair; and based on the matching node pair, aligning the source domain nodes in the source domain graph with the target domain nodes in the target domain graph.

[0014] According to some embodiments of the present disclosure, aligning the source domain nodes in the source domain graph with the target domain nodes in the target domain graph based on the matching node pair includes: sorting the highest alignment weights corresponding to each source domain node from largest to smallest, and selecting a preset number of matching node pairs as a set of labeled node pairs according to the sorted highest alignment weights, where the set of labeled node pairs includes a set of labeled source domain nodes and a set of labeled target domain nodes; based on the set of labeled source domain nodes and the set of labeled target domain nodes, determining a cross-graph cost matrix, where each element in the cross-graph cost matrix represents the similarity between the labeled source domain nodes in the set of labeled source domain nodes and the labeled target domain nodes in the set of labeled target domain nodes; iteratively updating the target alignment matrix based on the cross-graph cost matrix until the target alignment matrix converges to obtain an updated target alignment matrix; and aligning the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the updated target alignment matrix.

[0015] On the other hand, according to the present disclosure, a method for predicting matching points is proposed, which includes: obtaining a source domain graph and a target domain graph, where the source domain graph includes a plurality of source domain nodes and the target domain graph includes a plurality of target domain nodes; aligning the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the graph alignment method according to any one of the foregoing claims to obtain the aligned source domain graph and target domain graph; and predicting the number of matching points between the source domain graph and the target domain graph based on the aligned source domain graph and target domain graph.

[0016] According to another aspect of the present disclosure, a graph alignment device is provided, which includes: a determination module configured to determine the adjacency matrix and the feature matrix of corresponding graphs in a source domain graph and a target domain graph, wherein the source domain graph includes a plurality of source domain nodes, the target domain graph includes a plurality of target domain nodes, each element in the adjacency matrix represents the connection relationship between nodes in the corresponding graph, and each element in the feature matrix represents the feature of a node in the corresponding graph; an iteration module configured to initialize an alignment matrix as a target alignment matrix and iteratively execute the following steps until the maximum number of iterations is reached or the target alignment matrix converges: based on the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph, determine the intra-graph distance matrix of the corresponding graph, and each element in the intra-graph distance matrix represents the association degree between nodes in the corresponding graph; based on the intra-graph distance matrix of the source domain graph, the intra-graph distance matrix of the target domain graph, and the target alignment matrix, determine the inter-graph distance matrix of the current iteration, and each element in the inter-graph distance matrix represents the association degree between the source domain node of the corresponding source domain graph and the target domain node of the target domain graph; based on the inter-graph distance matrix of the current iteration, determine the alignment matrix of the current iteration, and update the target alignment matrix to the alignment matrix of the current iteration, and each element in the alignment matrix represents the matching probability between the source domain node of the corresponding source domain graph and the target domain node of the target domain graph; an alignment module configured to align the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the target alignment matrix.

[0017] According to another aspect of the present disclosure, a device for predicting matching points is provided, which includes: an acquisition module configured to acquire a source domain graph and a target domain graph, wherein the source domain graph includes a plurality of source domain nodes, and the target domain graph includes a plurality of target domain nodes; a graph alignment module configured to use the graph alignment device according to another aspect of the present disclosure to align the source domain nodes in the source domain graph with the target domain nodes in the target domain graph to obtain the aligned source domain graph and target domain graph; a prediction module configured to predict the number of matching points between the source domain graph and the target domain graph based on the aligned source domain graph and target domain graph.

[0018] According to another aspect of the present disclosure, a computing device is provided, including: a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, it causes the processor to execute the graph alignment method according to some embodiments of the present disclosure.

[0019] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which computer-readable instructions are stored, and when the computer-readable instructions are executed, they implement the graph alignment method according to some embodiments of the present disclosure.

[0020] According to another aspect of the present disclosure, a computer program product is provided, including a computer program which, when executed by a processor, implements the graph alignment method according to some embodiments of the present disclosure.

[0021] In the graph alignment method and apparatus according to some embodiments of the present disclosure, after determining the adjacency matrix and the feature matrix of the source domain graph and the target domain graph respectively, by iteratively executing determining the intra-graph distance matrix based on the adjacency matrix and the feature matrix, determining the inter-graph distance matrix based on the intra-graph distance matrix and the target alignment matrix, and determining the updated iteration matrix based on the inter-graph distance matrix, until the maximum number of iterations is reached or the target alignment matrix converges, the various parameters for calculating the target alignment matrix can be continuously updated, so as to improve the accuracy of the finally obtained target alignment matrix. And because the source domain graph and the target domain graph are directly used for calculation (without alignment in the embedding space), the processing efficiency of the graph alignment method is improved, the adaptability of the graph alignment method under various graph data is improved, the robustness of the graph alignment method is enhanced, which is beneficial to improving the efficiency of graph alignment using the target alignment matrix subsequently and increasing the accuracy and precision of graph alignment.

[0022] According to the embodiments described hereinafter, these and other aspects of the present application will be apparent and will be elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] According to the following detailed description and the drawings, the various different aspects, features and advantages of the present disclosure will be readily understood. In the drawings:

[0024] Figure 1 An exemplary implementation environment of the graph alignment method according to some embodiments of the present disclosure is schematically shown;

[0025] Figure 2 A flowchart of the graph alignment method according to some embodiments of the present disclosure is schematically shown;

[0026] Figure 3 A schematic diagram of the principle of the graph alignment method according to some embodiments of the present disclosure is schematically shown;

[0027] Figure 4 A flowchart of the graph alignment method according to some embodiments of the present disclosure is schematically shown;

[0028] Figure 5 A flowchart of the graph alignment method according to some embodiments of the present disclosure is schematically shown;

[0029] Figure 6 A schematic diagram of the principle of the graph alignment method according to some embodiments of the present disclosure is schematically shown;

[0030] Figure 7 A flowchart schematically showing a method for predicting matching points according to some embodiments of the present disclosure;

[0031] Figures 8A - 8C A comparison diagram schematically showing the effects of a graph alignment method according to some embodiments of the present disclosure and a method of the related art;

[0032] Figure 9 An example block diagram schematically showing a graph alignment device according to some embodiments of the present disclosure;

[0033] Figure 10 An example block diagram schematically showing a device for predicting matching points according to some embodiments of the present disclosure; and

[0034] Figure 11 An example block diagram schematically showing a computing device according to some embodiments of the present disclosure.

[0035] It should be noted that the above-mentioned drawings are merely schematic and illustrative, and are not necessarily drawn to scale. Detailed Description of Specific Embodiments

[0036] Several embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings so that those skilled in the art can implement the present disclosure. The present disclosure can be embodied in many different forms and for many different purposes and should not be limited to the embodiments set forth herein. These embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art. The embodiments do not limit the present disclosure.

[0037] It will be understood that although the terms first, second, third, etc. may be used herein to describe various elements, components, and / or parts, these elements, components, and / or parts should not be limited by these terms. These terms are only used to distinguish one element, component, or part from another element, component, or part. Thus, the first element, component, or part discussed below may be referred to as the second element, component, or part without departing from the teachings of the present disclosure.

[0038] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the present disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising", when used in this specification, specify the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0039] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the relevant art and / or the context of this specification, and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0040] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0041] The flowcharts shown in the drawings are merely illustrative and not necessarily inclusive of all promotional information and operations / steps, nor are they necessarily to be executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0042] Those skilled in the art can understand that the drawings are only example interfaces of example embodiments, and the modules or processes in the drawings are not necessarily essential for implementing the present application, and thus cannot be used to limit the protection scope of the present application.

[0043] Artificial Intelligence (AI) utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence. It is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, intelligent transportation, and automatic control.

[0044] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, teaching learning, and active learning.

[0045] Before introducing the embodiments of the present disclosure in detail, for clarity, some related concepts are first explained.

[0046] 1. Optimal Transport (OT): It is a classical optimization problem that aims to measure the distance between two different probability distributions and find the best way to map one distribution to another to minimize the transport cost.

[0047] 2. Gromov-Wasserstein distance (GWD): It is a measure of the distance between two metric spaces that takes into account the correspondence between elements in the spaces and is often used to measure the differences between structured data. The Gromov-Wasserstein distance can consider the similarity between two metric spaces by introducing an auxiliary metric space and taking into account the structural information in both metric spaces.

[0048] 3. Graph Neural Network (GNN): A deep learning model specifically designed to process graph-structured data. The core idea of GNN is to learn node representations by passing and aggregating information between nodes. GNN takes the feature vector of each node as input and updates the node representations through multiple rounds of neighbor information passing. It captures the local and global structures of the graph by propagating information on nodes and learns the relationships and features between nodes. Then it integrates the feature information of nodes with the structural information of the graph to obtain richer representations for tasks such as graph classification, graph generation, and node classification.

[0049] 4. Self-training: A semi-supervised learning strategy. Usually, the model is first trained using data with given labels, then uses unlabeled data, and generates pseudo-labels according to certain rules, and uses these pseudo-labels for supervised learning to further improve the model performance. This method can also be applied to unsupervised situations.

[0050] 5. Graph alignment: Given two attributed graphs G s and G t , graph alignment means finding a one-to-one correspondence between the node sets of these two graphs. An attributed graph is a graph structure where each node in the attributed graph is associated with a set of attributes or features, and these attributes or features can include information such as node labels, numerical values, vectors, etc., which are used to describe the properties or characteristics of the nodes. In addition, multiple graph structures can also be compared and aligned to find their similarities or correspondences.

[0051] 6. Random Walk: A mathematical model that describes the process of transitioning from one state to another in random steps. It has a wide range of applications in many fields. In a random walk, an object starts from the current state and reaches the next state through a series of random moves, and the selection of each state is random and has a certain probability. Specifically, random walks can be divided into discrete and continuous forms.

[0052] In related technologies, graph alignment tasks are mainly divided into two categories: unsupervised tasks and supervised tasks:

[0053] (1) Unsupervised task-based alignment schemes: mainly include unsupervised graph alignment algorithms Walign, GWL, and SLOTAlign. Representative unsupervised methods can be roughly divided into two categories. One category follows the intuitive idea of "Embedd-Then-Cross-Compare", that is, first embed all nodes in two graphs, and then find the solution with the smallest embedding distance corresponding to the matching nodes. Walign is one of the representative algorithms. This algorithm first uses, for example, lightweight graph neural networks (GNNs) to calculate the embedding Z of each node u u , and then takes minimizing the sum of distances between different node embeddings as the objective of the unsupervised graph alignment task. Considering that this objective can be interpreted as a kind of Wasserstein distance (WD), Walign is inspired by Wasserstein GAN and uses a neural network to implement a 1-Lipschitz function f W , and generates new pseudo-corresponding node pairs based on this. Thus, Walign proposes a GAN-based framework, and its discriminator takes minimizing the difference between pseudo-node pair embeddings as the optimization objective. The other category uses optimal transport for graph alignment, and finds that there is a probability match between two distributions on the graph, with better accuracy and robustness. Among them, the representative ones are the GWL algorithm and the SLOTAlign algorithm. GWL combines graph-based optimal transport and node embedding learning, and proposes a graph matching framework. It predicts an alignment matrix T based on the optimal transport solution. For the given learnable node representations Z s ∈ R n1×d (a matrix of n1×d) and Z t ∈ R n2×d (a matrix of n2×d), GWL minimizes the combined objective function of Wasserstein discrepancy and Gromov-Wasserstein discrepancy (GWD). The GWD part in the algorithm manually sets the cross-graph loss matrices C s and C t by combining the adjacency information and cosine distance of node embeddings. Finally, GWL uses the proximal point method to transform the GWD objective into an optimal transport problem, and alternately learns the alignment matrix T and node embeddings Z s and Z t . The SLOTAlign algorithm improves the existing GWD-based methods and designs the internal costs C s and C t more carefully. It introduces a multi-view structure modeling module, which integrates graph structure (i.e., edges), node features, and multi-hop neighborhood information through the learnable parameter β; and adopts a process similar to GWL to learn the parameter β, and alternately uses the alignment matrix T to minimize GWD.

[0054] (2) Alignment scheme based on supervised tasks: mainly includes the supervised graph alignment algorithm PARROT. The supervised graph alignment method mainly focuses on how to make full use of the given anchor information and propagate it to neighboring nodes. The solution based on restart random walk (RWR) is widely adopted, and recent work has also tried to combine the idea of label propagation with optimal transport. As one of the representative supervised algorithms, PARROT combines the above two works. For a given set of anchor points, PARROT first compares the position embeddings and node features of the two graphs based on RWR within the graph, and calculates the benchmark matrix C base , and then performs RWR-based propagation on the product graph of G s and G t to calculate the cross-graph cost matrix C rwr . After obtaining the cross-graph cost C rwr , PARROT can convert the graph alignment problem into an optimal transport problem.

[0055] This application provides a graph alignment method, which enhances graph alignment from the perspective of graph structure. Figure 1 Schematically shows an example implementation environment 100 of the graph alignment method according to some embodiments of the present disclosure. As Figure 1 shown, the implementation environment 100 may include a terminal device 110, a server 120, and a network 130 for connecting the terminal device 110 and the server 120. In some embodiments, the terminal device 110 may be used to implement the graph alignment method according to the present disclosure. For example, the terminal device 110 may be deployed with corresponding programs or instructions for executing various methods provided by the present disclosure. Optionally, the server 120 may also be used to implement various methods according to the present disclosure.

[0056] The terminal device 110 and the third-party terminal device 140 may be any type of mobile computing device, including mobile computers (e.g., personal digital assistants (PDAs), laptop computers, notebook computers, tablet computers, netbooks, etc.), as Figure 1 shown, mobile phones (e.g., cellular phones, smartphones, etc.), wearable computing devices (e.g., smart watches, head-mounted devices, including smart glasses), or other types of mobile devices. In some embodiments, the terminal device 110 may also be a fixed computing device, such as a desktop computer, a game console, a smart TV, etc.

[0057] Server 120 may be a single server or a server cluster, or may be a cloud server or a cloud server cluster capable of providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. It should be understood that the servers mentioned herein are typically server computers with a large amount of memory and processor resources, but other embodiments are also possible. Optionally, server 120 may also be an ordinary desktop computer, which includes a host, a display, etc.

[0058] Examples of network 130 include a local area network (LAN), a wide area network (WAN), a personal area network (PAN), and / or a combination of communication networks such as the Internet. Server 120 and terminal device 110 may include at least one communication interface (not shown) capable of communicating via network 130. Such a communication interface may be one or more of the following: any type of network interface (e.g., network interface card (NIC)), wired or wireless (such as IEEE 802.11 wireless LAN (WLAN)) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth TM interface, near field communication (NFC) interface, etc.

[0059] As Figure 1 shown, terminal device 110 may include a display screen 111 and the end user can interact with terminal application 112 via display screen 111. Terminal device 110 may interact with server 120 via network 130, for example, by sending data to it or receiving data from it. Terminal application 112 may be a native application, a web application, or a mini program (LiteApp, such as a mobile mini program, a WeChat mini program) as a lightweight application. In the case where the terminal application is a native application that needs to be installed, terminal application 112 may be installed in terminal device 110. In the case where terminal application 112 is a web application, terminal application 112 may be accessed through a browser. In the case where terminal application 112 is a mini program, terminal application 112 may be directly opened on terminal device 110 by searching for relevant information of the terminal application (such as the name of the terminal application), scanning the graphic code of the terminal application (such as a barcode, a QR code, etc.), etc., without installing terminal application 112.

[0060] Figure 1The exemplary implementation environment is merely illustrative, and the graph alignment method according to the present disclosure is not limited to the illustrated exemplary implementation environment and server. It should be understood that although the server 120 and the terminal device 110 are shown and described as separate structures in this document, they may be different components of the same computing device. Optionally, all steps of the graph alignment method according to some embodiments of the present disclosure may also be implemented on the server 120 side, or may also be jointly implemented on the terminal device 110 side and the server 120 side.

[0061] Figure 2 Schematically shows a flowchart of a graph alignment method according to some embodiments of the present disclosure. In some embodiments, as Figure 1 shown, the graph alignment method according to the present disclosure may be executed on the terminal device 110 side. In other embodiments, the graph alignment method according to the present disclosure may also be executed in combination by the server 120 and the terminal device 110.

[0062] As Figure 2 shown, the graph alignment method according to some embodiments of the present disclosure may include steps S210 - S230, where step S220 may include sub - steps S220a, S220b, and S220c:

[0063] S210, determine the adjacency matrix and the feature matrix of the corresponding graphs in the source - domain graph and the target - domain graph, where the source - domain graph includes a plurality of source - domain nodes, the target - domain graph includes a plurality of target - domain nodes, each element in the adjacency matrix represents the connection relationship between nodes in the corresponding graph, and each element in the feature matrix represents the feature of the nodes in the corresponding graph;

[0064] S220, initialize the alignment matrix as the target alignment matrix, and iteratively execute the following steps until the maximum number of iterations is reached or the target alignment matrix converges:

[0065] S220a, based on the adjacency matrix and the feature matrix of the corresponding graphs in the source - domain graph and the target - domain graph, determine the intra - graph distance matrix of the corresponding graph, and each element in the intra - graph distance matrix represents the degree of association between nodes in the corresponding graph,

[0066] S220b, based on the intra - graph distance matrix of the source - domain graph, the intra - graph distance matrix of the target - domain graph, and the target alignment matrix, determine the inter - graph distance matrix of this iteration, and each element in the inter - graph distance matrix represents the degree of association between the source - domain nodes of the corresponding source - domain graph and the target - domain nodes of the target - domain graph,

[0067] S220c, based on the inter - graph distance matrix of this iteration, determine the alignment matrix of this iteration, and update the target alignment matrix to the alignment matrix of this iteration, and each element in the alignment matrix represents the probability of matching between the source - domain nodes of the corresponding source - domain graph and the target - domain nodes of the target - domain graph;

[0068] S230. Align the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the target alignment matrix.

[0069] The following will Figure 3 describe steps S210 - S230 in detail, where Figure 3 FIG. schematically shows a schematic diagram of the principle of the graph alignment method according to some embodiments of the present disclosure.

[0070] First of all, it should be noted that the graph alignment method in the embodiments of the present application can be applied to graphs that need to be aligned, such as social network graphs, commodity network graphs, and protein structure graphs. For the sake of easy understanding and explanation, the graph alignment method in the present application will be mainly described below by taking a social network graph as an example. However, it can be understood that the graph alignment method of the present application is not limited to social network graphs, that is, the social network graph does not constitute a limitation on the technical solutions provided by the embodiments of the present invention.

[0071] Currently, with the prosperity of social networks, people are not limited to using only one social network, but will use multiple social networks simultaneously to enjoy more applications. For example, the same user can follow some stars, writers, or people and things of interest on Twitter, and follow different video hosts on YouTube. These social networks are often independent of each other, but can jointly reflect user information, which provides opportunities and challenges for combining heterogeneous social networks and jointly mining the information they hide. Finding the same user in different social networks is the first step in many data mining tasks. For example, by identifying the same user in different social networks, more resources can be combined to mine user information. Such problems are called social network alignment. And a social network can actually be expressed as a graph, so the social network alignment problem can also be expressed as a graph alignment problem. Among them, the actions involving multiple graph alignments are based on the active authorization of the nodes (i.e., users) in the graph after knowing the relevant intentions of the platform or system.

[0072] In step S210, the source domain graph and the target domain graph can be any two graphs that need to be aligned. The source domain graph and the target domain graph are relative concepts used to distinguish two graphs and do not refer to a specific graph. Therefore, graph alignment is essentially to find the corresponding nodes in the two graphs. Therefore, in some embodiments, the source domain graph and the target domain graph can be swapped with each other, that is, when aligning two graphs, either one of the graphs can be arbitrarily determined as the source domain graph, and the remaining graph is the target domain graph. Generally speaking, both the source domain graph and the target domain graph describe the association relationship between entities by defining nodes and edges. Therefore, there are nodes and edges connecting the nodes in both the source domain graph and the target domain graph. Thus, the source domain nodes can be understood as the nodes in the source domain graph, and the target domain nodes can be understood as the nodes in the target domain graph.

[0073] In some embodiments, after obtaining the source domain graph and the target domain graph, the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph can be determined. The adjacency matrix is a square matrix used to represent a finite graph. Each element in the adjacency matrix represents the connection relationship (i.e., whether there is an edge connecting each pair of points) or topological features between the nodes in the corresponding graph; that is to say, the adjacency matrix is a data structure for depicting the relationship between nodes and edges, and its essence is a two-dimensional array, which is suitable for processing the association relationship between the smallest data units. Each element in the feature matrix represents the features of the nodes in the corresponding graph, which includes the attribute features of the nodes themselves (in the application of social network graphs, they can be features corresponding to attributes such as age, gender, and personality signature); for each node in the graph data, its features can be represented as a vector, and these vectors are formed into a matrix row by row, and the size of the feature matrix is the number of nodes multiplied by the feature dimension.

[0074] Step S220 is an iterative step, which can iteratively execute sub-steps S220a, S220b, and S220c until the maximum number of iterations is reached or the target alignment matrix converges. The maximum number of iterations is the preset maximum number of allowed iteration rounds, which can generally be set according to the amount of computation and computational efficiency. The target alignment matrix can be regarded as the output of the iterative process, where the alignment matrix is a matrix used to represent the corresponding relationship or matching relationship between two graphs or two sets of elements, and each element in it represents the probability of matching between the source domain nodes of the corresponding source domain graph and the target domain nodes of the target domain graph. That is to say, an alignment matrix can be returned in each round of iteration until the alignment matrix converges or approximately converges (depending on the required accuracy). Generally, matrix convergence refers to a special matrix transformation, whose characteristic is that as the number of transformations gradually increases, the values in the matrix will gradually approach a stable value, that is, the limit value after the matrix transformation; for example, in large-scale graph processing, it is often necessary to perform multiple transformations such as rotating and translating a node, and the final result will tend to be stable.

[0075] At sub-step S220a, based on the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph, the intra-graph distance matrix of the corresponding graph can be determined. The intra-graph distance matrix (which can also be referred to as the intra-graph cost matrix in this article) is a matrix used to describe the similarity or correlation degree between nodes in a single graph, and each element in it represents the correlation degree between the nodes in the corresponding graph. In other words, the intra-graph cost matrix means that in a graph, for any edge between two nodes, there is a cost value to represent the distance or weight between them, and it can be used to calculate the cost of the optimal matching problem.

[0076] Continue to describe sub-step S220a in detail below, where Figure 4 whereFigure 4 A flowchart of a graph alignment method according to some embodiments of the present disclosure is schematically shown. In some embodiments, sub-step S220a may include:

[0077] S410, based on the feature matrices of the corresponding graphs in the source domain graph and the target domain graph, determine the node similarity matrix of the corresponding graph, where each element in the node similarity matrix represents the similarity between the features of the nodes in the corresponding graph;

[0078] S420, based on the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph, determine the splicing matrix of the corresponding graph, where each element in the splicing matrix represents the relationship between the nodes in the corresponding graph and their neighbor nodes;

[0079] S430, based on the adjacency matrix, the node similarity matrix, and the splicing matrix of the corresponding graph, determine the intra-graph distance matrix of the corresponding graph.

[0080] At step S410, the node similarity matrix may represent the similarity of the node features within the corresponding graph. The node similarity matrix refers to that in a graph, for any two nodes, there is a similarity value to represent the degree of similarity between them. This matrix is usually used to calculate the similarity between nodes for tasks such as graph matching, community detection, and node classification.

[0081] In some embodiments, the node similarity matrix of the corresponding graph may be determined in the following manner: perform a transpose operation on the feature matrices of the corresponding graphs in the source domain graph and the target domain graph to obtain the transposed feature matrices of the corresponding graphs; perform an inner product operation on the feature matrix of the corresponding graph and the transposed feature matrix of the corresponding graph to obtain the node similarity matrix of the corresponding graph. Among them, the transposed feature matrix refers to a new matrix obtained by transposing the original feature matrix for better data analysis and modeling. Then, the inner product of this feature matrix and this transposed feature matrix is used as the node similarity matrix, which can also be referred to as the covariance matrix of the feature matrix, and it reflects the correlation between different features. In this way, the efficiency of determining the node similarity matrix of the corresponding graph can be improved.

[0082] In some embodiments, it can also be obtained by other methods. One method can be to calculate the similarity based on the attributes or structural information between nodes; specifically, if the nodes are represented in the form of vectors, then the similarity between vectors can be used to calculate the similarity between nodes (for example, cosine similarity, Euclidean distance, Manhattan distance, etc. can be used to measure the similarity between nodes). Another method can be based on the network structure information between nodes; for example, the number of common neighbors, etc. can be used to measure the similarity between nodes, which utilizes the relationship between nodes to calculate the similarity. When calculating the node similarity matrix, it is necessary to consider that the number of nodes in the graph is large, so the calculation time and space complexity are important issues. To improve efficiency, approximate calculation methods can be used to calculate the node similarity matrix. For example, methods such as random walk algorithm, locality-sensitive hashing, etc. can be used for approximate calculation, so as to reduce the calculation time and space complexity on the premise of ensuring a certain degree of accuracy.

[0083] At step S420, each element in the splicing matrix can represent the relationship between the nodes in the corresponding graph and their neighbor nodes, which can be used in undirected graphs, directed graphs, weighted graphs, multi-graphs, etc. The splicing matrix can be spliced along the row direction or the column direction of the matrix.

[0084] In some embodiments, the splicing matrix of the corresponding graph can be determined in the following manner: Using a deep learning network, based on the adjacency matrix and feature matrix of the corresponding graph in the source domain graph and the target domain graph, sequentially determine the 1st to Kth order neighbor feature matrices of the corresponding graph, where each element in any Mth order neighbor feature matrix represents the features of the Mth order neighbor nodes of the nodes in the corresponding graph, K is a positive integer not less than 1, and M is a positive integer greater than or equal to 1 and less than or equal to K. The Mth order neighbor nodes of a node refer to the nodes separated from this node by M - 1 nodes; splice the feature matrix of the corresponding graph and the 1st to Kth order neighbor feature matrices of the corresponding graph according to weights; use an activation function to activate the spliced matrix to obtain the splicing matrix of the corresponding graph. The deep learning network can be the GNN network described above or other deep learning models that can be designed to process graph-structured data. In a graph, there may be multiple nodes, and for each node, there may be one or more neighbor nodes, and at the same time, each node may also have Kth order neighbor nodes, where K is a positive integer not less than 1. For example, for a simple graph structure "A - B - C - D", node B is the 1st order neighbor node of node A, node C is the 2nd order neighbor node of node A, and node D is the 3rd order neighbor node of node A. That is to say, the Mth order neighbor nodes of a node refer to the nodes separated from this node by M - 1 nodes, where M is a positive integer greater than or equal to 1 and less than or equal to K. Therefore, each element in the Mth order neighbor feature matrix can represent the features of the Mth order neighbor nodes of the nodes in the corresponding graph, and here M generally takes the maximum value it can take in a graph. Then, the feature matrix of the corresponding graph, the 1st order neighbor feature matrix, the 2nd order neighbor feature matrix,..., the Kth order neighbor feature matrix can be spliced according to weights, where the weights are learnable parameters used to represent the importance and role of each matrix. Then, use an activation function to activate the spliced matrix. The activation function is a non-linear transformation in a neural network that converts the output of a neuron into a non-linear response, and it can be, for example, the RELU function, the Leaky ReLU function, the Tanh function, or the Softmax function, etc. In this way, the node degree information and graph structure information can be better combined, thus ensuring that as many node features as possible are considered.

[0085] At step S430, the intra-graph distance matrix can be determined based on the adjacency matrix, node similarity matrix, and splicing matrix of the corresponding graph. That is to say, the intra-graph distance matrix takes into account the characteristics of the nodes themselves, the relationship between nodes and edges, the characteristics of neighboring nodes, etc. In some embodiments, the weighted sum of the adjacency matrix, node similarity matrix, and splicing matrix of the corresponding graph can be determined as the intra-graph distance matrix of the corresponding graph. That is to say, the intra-graph distance matrix can be determined in the following way: assign corresponding weights to each matrix, and then calculate the weighted sum of all matrices, where the weights represent the importance of the corresponding matrices. In this way, in actual graph alignment, the weights of each item can be balanced according to different objectives and different application scenarios, so that the intra-graph distance matrix can be determined more flexibly and conveniently.

[0086] In this way, by determining the intra-graph distance matrix of the corresponding graph through steps S410 - S430, it is possible to consider both the similarity between the characteristics of the nodes in the corresponding graph and the relationship between the nodes in the corresponding graph and their neighboring nodes, so that the intra-graph distance matrix can make full use of the characteristics and relationships of the nodes and edges in the corresponding graph, thereby better representing the degree of association between nodes.

[0087] At sub-step S220b, the inter-graph distance matrix for this iteration can be determined based on the intra-graph distance matrix of the source domain graph, the intra-graph distance matrix of the target domain graph, and the target alignment matrix. The inter-graph distance matrix (which can also be referred to as the inter-graph cost matrix in this article) is a matrix used to describe the degree of association or similarity between nodes in two graphs, and each element in it represents the degree of association between the source domain node of the corresponding source domain graph and the target domain node of the target domain graph. For example, suppose there are two graphs G s and G t , which are respectively composed of nodes and edges. The inter-graph cost matrix C is a matrix of size (n1 + n2) × (n1 + n2), where n1 and n2 respectively represent the number of nodes in G s and G t . For any i and j, C[i][j] represents the cost or distance between the i-th node in G s and the j-th node in G t . If the i-th node in G s is similar to the j-th node in G t , then C[i][j] should be smaller; if they are not similar, then C[i][j] should be larger. Therefore, first, the intra-graph distance matrix of the source domain graph and the intra-graph distance matrix of the target domain graph can be subtracted and squared, and then combined with the target alignment matrix to determine the inter-graph distance matrix for this iteration. Note that when performing the above operations, the shape and size of the matrices need to be kept consistent first.

[0088] At sub-step S220c, based on the inter-graph distance matrix of the current iteration, the alignment matrix of the current iteration can be determined, and the target alignment matrix can be updated to the alignment matrix of the current iteration. This can be accomplished through the optimal transport algorithm. That is to say, the alignment matrix can be calculated by methods such as iterative optimization algorithms (such as the Sinkhorn algorithm), linear programming methods, or entropy regularization methods. After sub-step S220c ends, check whether the maximum number of iterations is reached or whether the target alignment matrix converges, so as to determine whether to perform a new round of iteration or return the result after the iteration is completed.

[0089] In some embodiments, determining the alignment matrix of the current iteration based on the inter-graph distance matrix of the current iteration can be performed in the following manner: Determine the optimal transport distance between the inter-graph distance matrix of the current iteration and the target alignment matrix, where the optimal transport distance represents the minimum cost of mapping the inter-graph distance matrix of the current iteration to the target alignment matrix; Based on the optimal transport distance, determine the alignment matrix of the current iteration. In determining the optimal transport distance between the inter-graph distance matrix of the current iteration and the target alignment matrix, the target alignment matrix here can be understood as the alignment matrix of the previous round (if it is the first round of iteration, it is the initialized alignment matrix). Then, based on the optimal transport distance, map the elements of the inter-graph distance matrix of the current iteration to the target alignment matrix as the alignment matrix of the current iteration. In this way, the cost of determining the alignment matrix can be reduced.

[0090] In some embodiments, determining the optimal transport distance between the inter-graph distance matrix of the current iteration and the target alignment matrix can be performed in the following manner: Perform data normalization processing on the inter-graph distance matrix of the current iteration to obtain the normalized inter-graph distance matrix of the current iteration; Based on the normalized inter-graph distance matrix of the current iteration, determine the optimal transport distance between the normalized inter-graph distance matrix of the current iteration and the target alignment matrix. Normalization processing refers to scaling the data proportionally so that it falls within a specific range, usually [0,1] or [-1,1]. Normalization can make data with different dimensions comparable and avoid the influence of some feature magnitudes being too large on model training. The way of this normalization processing is not unique as long as it can satisfy mapping the numerical range to a unified comparable interval. In this way, the dispersion degree of the data can be reduced, the stability and reliability of the data can be improved, and thus the accuracy and robustness of the model can be improved, and it can also accelerate model convergence and improve generalization ability.

[0091] In some embodiments, the inter-graph distance matrix of the current iteration can be data-normalized in the following manner: Select the maximum value from all elements of the inter-graph distance matrix of the current iteration; for each element in the inter-graph distance matrix of the current iteration, based on the maximum value, map each element in the inter-graph distance matrix of the current iteration to a normalized element through linear normalization to obtain the normalized inter-graph distance matrix of the current iteration. That is to say, each element in the inter-graph distance matrix of the current iteration can be divided by the maximum value in this matrix, thereby achieving the normalization process. For example, since the graph alignment task does not require a dense matrix, the elements of the original inter-graph distance matrix can be very small values (e.g., 10 -5 ), so after such normalization, all elements can be linearly mapped into the interval [0,1] at the same time, that is, the values of these elements are increased. In this way, the training cost can be reduced and the training speed can be accelerated.

[0092] In some embodiments, in addition to the above normalization method, other methods can also be used, such as min-max normalization, Z-score standardization, Log transformation, Softmax normalization, etc., which will not be elaborated herein. Therefore, a suitable normalization method can be selected according to actual needs and application scenarios.

[0093] The entire iteration process can be executed as Figure 3 shown, which is an unsupervised process. In Figure 3 , after the iteration process starts, the input of this iteration process includes graph G s , graph G t 's adjacency matrix A s , A t and feature matrix X s , X t , as well as the maximum number of iteration rounds maxIter. In the first round of iteration: The alignment matrix T is initialized to a uniform distribution, that is, T i,k = 1 / n1n2, which means that the element in the i-th row and k-th column of the alignment matrix T is set to 1 / n1n2, where n1 is the number of nodes in graph G s and n2 is the number of nodes in graph G t ; the weights β s , β t for calculating the intra-graph distance matrix are initialized to (1,1,1) T , where the superscript T represents the transpose of a matrix or vector; and the iteration round number Iter is set to 1. Next, if Iter < maxIter, the iteration step is executed. Equations 1 and 2 represent calculating the intra-graph distance matrices C s and C tThe method, where the inner product of the feature matrix X and the transpose of the feature matrix X represents the node similarity matrix of the corresponding graph. And GNN(A,X) represents the concatenation matrix of the corresponding graph, which can be calculated by Equation where the first term X in the square brackets represents the feature matrix of the nodes, the second term represents the features of the 1st-order neighbor nodes of a node aggregated, …, the Kth term represents the features of the Kth-order neighbor nodes of a node aggregated, where K is a positive integer not less than 1; then the initial feature matrix, the feature matrix of the 1st-order neighbors, …, the feature matrix of the Kth-order neighbors are concatenated together by the Concat function according to the weight Θ GNN And then it is operated on by the activation function σ. After that, Equation 3 represents the formula for calculating the inter-graph distance matrix of this round of iteration, where u is an element belonging to graph G s and v is an element belonging to graph G t . Equation 4 represents a special normalization process for the inter-graph distance matrix obtained from Equation 3, that is, each element in the matrix is divided by the maximum value in this matrix. Proximal Point in Equation 5 represents the proximal point algorithm, which is an iterative algorithm, that is, the alignment matrix T of the current round and the normalized inter-graph distance matrix C gwd are used to calculate the alignment matrix T of the next round. Equation 6 is the method for calculating the gradient. Each update requires taking the partial derivative of the sum of the Hadamard products of C gwd and T with respect to the parameter β. Then it is judged whether the alignment matrix T converges. If it does not converge, then based on Equation 7 and Equation 8, preparations are made for the next round of iteration, and the parameter β is updated by the gradient descent method, where lr represents the learning rate. Each round of iteration returns an alignment matrix T to represent the matching degree of the nodes in the two graphs. The larger T(i,k) is, the higher the matching degree between node i in graph G s and node k in graph G t . When reaching maxIter or the alignment matrix T converges, the matching degree of the cross-graph node pair (u,v) is determined according to the value of T maxIter (u,v), and finally the target alignment matrix T is returned. After obtaining the target alignment matrix T, for each row in T, find the maximum value in each row, and obtain the index corresponding to this maximum value (that is, the argmax function in Equation 9, where V s represents the node set of G s , and u i represents any one of the nodes), that is, for this node in graph G s , a node that is considered likely to match is found in graph G t . After that, the iterative process ends.

[0094] At step S230, the source domain nodes in the source domain graph can be aligned with the target domain nodes in the target domain graph according to the obtained target alignment matrix. It can be generally understood as: mapping the source domain nodes in the source domain graph into the target domain graph to find the corresponding target domain nodes that match them.

[0095] Continue to describe step S230 in detail below in combination with Figure 5 where Figure 5 FIG. schematically shows a flowchart of a graph alignment method according to some embodiments of the present disclosure. In some embodiments, as Figure 5 shown, step S230 may include:

[0096] S510, for each source domain node in the source domain graph, select the highest alignment weight corresponding to each source domain node from all elements of the target alignment matrix, where the highest alignment weight represents the alignment probability with the highest probability of aligning the source domain node in the source domain graph with the target domain node in the target domain graph;

[0097] S520, based on the highest alignment weight corresponding to each source domain node, select one target domain node that matches each source domain node respectively to obtain a pair of matching nodes;

[0098] S530, based on the pair of matching nodes, align the source domain nodes in the source domain graph with the target domain nodes in the target domain graph.

[0099] At step S510, the purpose of selecting the highest alignment weight is to find the maximum probability of aligning the source domain node with the target domain node, that is, the accuracy of aligning the two nodes. For a target alignment matrix, each source domain node can find its corresponding highest alignment weight from it. The highest alignment weight is a weight value used to measure the alignment degree or correlation between the source domain node and the target domain node in the graph alignment task. In the graph alignment task, a similarity metric (such as cosine similarity, Euclidean distance, etc.) is usually used to calculate the similarity or correlation between the source domain node and the target domain node, and the results of these similarity metrics can be expressed as a weight value, reflecting the alignment degree between the source domain node and the target domain node. The highest alignment weight is the maximum weight value found in the target alignment matrix.

[0100] At step S520, after obtaining the highest alignment weight corresponding to each source domain node, one target domain node that matches each source domain node can be selected in the target domain graph, thereby forming a pair of matching nodes, that is, finding the one-to-one corresponding nodes and mutual relationships in the two graphs. That is to say, the highest alignment weight can be used to determine which target domain node a source domain node should be aligned with.

[0101] At step S530, an alignment operation between the source domain graph and the target domain graph can be implemented based on the matching node pairs to align the source domain nodes with the target domain nodes.

[0102] In this way, by using steps S510 - S530, the matching node pairs can be accurately obtained, so that the source domain graph and the target domain graph can be aligned more quickly by using these matching node pairs.

[0103] In some embodiments, step S530 may include: sorting the highest alignment weights corresponding to each source domain node from largest to smallest, and selecting a preset number of matching node pairs as the labeled node pair set according to the sorted highest alignment weights, where the labeled node pair set includes a labeled source domain node set and a labeled target domain node set; determining a cross - graph cost matrix based on the labeled source domain node set and the labeled target domain node set, where each element in the cross - graph cost matrix represents the similarity between the labeled source domain nodes in the labeled source domain node set and the labeled target domain nodes in the labeled target domain node set; iteratively updating the target alignment matrix based on the cross - graph cost matrix until the target alignment matrix converges to obtain an updated target alignment matrix; and aligning the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the updated target alignment matrix. That is to say, the first q matching node pairs with the largest highest alignment weights can be selected, which can reduce the computational amount while ensuring accuracy. Then, the selected q matching node pairs can be constructed into a labeled node pair set. Then, for example, the labeled source domain node set and the labeled target domain node set can be traversed respectively by the random walk algorithm to determine the cross - graph cost matrix, where each element in the cross - graph cost matrix represents the similarity between each source domain node in the labeled source domain node set and each target domain node in the labeled target domain node set. The random walk algorithm can be the restart random walk (RWR). The restart random walk is an algorithm whose basic idea is to start from a node, and at each step, there are two possible choices: either move to a randomly selected neighbor or jump back to the starting node. This restart random walk algorithm only contains a fixed parameter δ, called the "restart probability", and 1 - δ represents the probability of moving to a neighbor. After the iteration reaches stability, the stable probability vector contains the affinity scores of all nodes in the network with the starting node, and this stable probability vector can be regarded as the "influential impact" exerted by the starting node on the network. After obtaining the cross - graph cost matrix, the target alignment matrix is iteratively updated based on this cross - graph cost matrix until a preset condition is met (for example, the target alignment matrix converges) to obtain an updated target alignment matrix. After that, the alignment operation can be performed according to the updated target alignment matrix. In this way, the alignment matrix can be continuously optimized and updated based on self - training, making the alignment between the source domain nodes and the target domain nodes more accurate and improving the accuracy of the graph alignment method.

[0104] Next, continue to combine with Figure 6 describe the sub-steps of step S530 in detail, where Figure 6 schematically shows a schematic diagram of the principle of the graph alignment method according to some embodiments of the present disclosure. In some embodiments, Figure 6 it may include steps S601 - S608, which can be regarded as a self-training process. This process starts at S601. At S602, the input of the algorithm includes the alignment matrix T obtained from the unsupervised part, and the number q of pseudo-anchors (the range of q can be between 10% and 30% of the number n1 of nodes in the source domain graph G s ), where the pseudo-anchors are matching node pairs used as references, which can be used as labels for the entire self-training process. At S603, pseudo-anchor selection is performed. First, sort { (u i , max k T(i, k)) | u i ∈ G s} in descending order according to max k T(i, k), and then select the top q matching node pairs (u i , M(u i )) with the largest values as pseudo-anchors. At S604, define node similarity instead of distance as the basic case to obtain the benchmark S base , where: in the first term, R represents RWR, and this term represents the importance relative to some nodes in this graph; in the second term, X still represents the feature matrix, and this term still represents similarity; α is a weight, which is an adjustable hyperparameter; τ is a temperature parameter. Considering the regularization of the parameters, there is S base (i, k) ∈ [e -2τ , 1]. At S605, through RWR propagation, similarity-based cost propagation (SCP) can be performed to further obtain the supervised cross-graph cost vec(C int ). According to the definition of RWR, the following formula can be further obtained:

[0105]

[0106] It is a variant of the formula in S605, where δ is the restart probability of RWR, represents the parameter matrix of RWR, vec(C) represents concatenating all columns of C into an n1n2 - dim vector (dim represents the dimension of the matrix), and 1 n1n2 represents a matrix of all 1s. The second term represents the RWR propagation of the personalized vector vec(S base ), that is, the node similarity propagation on . Then use the all-1 matrix 1 n1n2Subtract it from the first term to convert the similarity into a distance representation. And if S base (i, k) ≤ 1 for all i and k, then In S606, PARROT is applied to the self-training phase, but the calculation between the positional embedding of PARROT and the cost COST is suboptimal. The reference matrix C obtained by PARROT through the inter-node positions and features base can be understood as the per-node distance between graphs, where matrix components with small values carry more information; however, by taking the product graph of G s and G t and performing node similarity propagation based on RWR on it, the final propagation cost C is obtained. C rwr , C base can be regarded as the personalized vector of RWR. Therefore, these small values have less actual impact during propagation. In S607, its meaning is similar to that of Equation 9 in Figure 3 . The self-training method returns new aligned node pairs (based on the new target alignment matrix T'); for pseudo-anchors, their matching results can still be modified according to the new alignment matrix. This process ends in S608. In principle, the above self-training process can include any supervised solution based on the optimal transport framework, such as GWL and SLOTAlign, in addition to being used in this application, and can be used to improve existing unsupervised methods based on optimal transport.

[0107] Many real-world problems can be naturally represented and solved as graph structures. Studying graph data from the perspective of subgraphs is an important method for graph data analysis. The core of subgraph analysis is the subgraph counting problem, that is, calculating the number of subgraphs that match between the data graph and the query graph. Its typical application scenarios include analyzing subgraph patterns in protein interaction networks, understanding the biological relationships between proteins, calculating specific subgraph patterns in social networks, calculating the map positioning and matching process, studying the propagation path and scope of information in the network, calculating the similarity between graphs, supporting the retrieval and matching of similar graphs, etc. The representative methods for the current subgraph counting task are: predicting whether a given query graph is a subgraph of the data graph. These algorithms regard the existence or non-existence problem as a binary classification problem and cannot be extended to the subgraph counting problem that is usually regarded as a regression problem. At the same time, the method of regarding subgraph counting as a regression problem and using graph neural networks to learn features only applies the graph neural network to the sub-structures obtained from the query graph and cannot fully exploit the data graph information. Unsupervised graph alignment methods can be directly used for subgraph counting tasks, such as cardinality estimation of the results of graph database queries (which is equivalent to subgraph matching) and query cost estimation of complex queries in similar multimodal databases.

[0108] Therefore,Figure 7 A flowchart of a method for predicting matching points according to some embodiments of the present disclosure is schematically shown. As Figure 7 shown, the method for predicting matching points according to some embodiments of the present disclosure may include steps S710 - S730:

[0109] S710, obtaining a source domain graph and a target domain graph, wherein the source domain graph includes a plurality of source domain nodes, and the target domain graph includes a plurality of target domain nodes;

[0110] S720, aligning the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the foregoing graph alignment method to obtain the aligned source domain graph and target domain graph;

[0111] S730, predicting the number of matching points between the source domain graph and the target domain graph based on the aligned source domain graph and target domain graph.

[0112] Among them, steps S710 and S720 are basically the same as the related descriptions above, so they will not be elaborated here. At step S730, after obtaining the aligned source domain graph and target domain graph, the number of sub - graphs in the target domain graph that match the source domain graph can be counted to predict the number of matching points between the source domain graph and the target domain graph. The methods for sub - graph counting may include, for example, the enumeration method, the graph - search - based method, the sub - graph isomorphism - based method, etc.

[0113] The effects obtained by using the graph alignment method according to some embodiments of the present disclosure will be further described below in conjunction with Figures 8A - 8C further illustration. Figures 8A - 8C A comparison diagram of the effects of using the graph alignment method according to some embodiments of the present disclosure and the methods of related technologies is schematically shown. It should be noted that Figures 8A - 8C it is only schematic and exemplary, and does not imply any limitation to the method of the present disclosure. The test data set is shown in Table 1 below

[0114] Table 1 Test Data Set

[0115]

[0116] Among them, the number of nodes, the number of edges, the feature dimension, and the number of ground truths respectively describe the corresponding features of each data set.

[0117] In Figure 8AAmong them, on the above dataset, a comparison is made with related technologies 1, 2, and 3 of the unsupervised graph alignment representative methods. "Hits@1" represents the node considered to have the most accurate prediction, that is, given 1 node in the source domain graph, and then predicting 1 node in the target domain graph, the probability of this prediction being correct; "Hits@5" represents given 1 node in the source domain graph, and then predicting 5 nodes in the target domain graph, the probability that one of these 5 nodes is correct. Therefore, Hits@5 is larger than Hits@1 because the probability that one of the 5 predicted nodes is correct is higher than the probability that 1 predicted node is correct; "Hits@10" represents given 1 node in the source domain graph, and then predicting 10 nodes in the target domain graph, the probability that one of these 10 nodes is correct; for Hits@1, Hits@5, and Hits@10, the larger the value, the better. "Time" represents the time consumed by the algorithm process, and its unit is seconds, so the smaller the time, the better. It can be seen that the Hits@1, Hits@5, and Hits@10 of the graph alignment method of this application are basically better than those of related technologies 1 - 3, perform better in the key attention index Hits@1, and also take less time. Further, a comparison is also made between the graph alignment method of this application without GNN and the graph alignment method of this application without normalization processing. It can be seen that the performance slightly decreases without GNN or normalization processing, which reflects the importance of GNN and normalization processing.

[0118] In Figure 8B it, the above self-training method is mainly introduced. Compared with Figure 8A it can be seen that the performance of the graph alignment method of this application is better after self-training. And the graph alignment method of this application is also better than the performance of related technologies 2 - 3 after self-training. However, the performance slightly decreases without GNN. "This application + self-training (degree)" means that instead of selecting pseudo anchor points according to the maximum value for self-training, the node with the largest degree is selected for self-training, and its performance is also very good. And the performance of training using the supervised model PARROT decreases compared with the performance of the graph alignment method of this application after self-training. Therefore, it can be seen that the above self-training method can further increase the accuracy of the model and improve the efficiency of the model.

[0119] In Figure 8CAmong them, these several models are still compared. It can be seen that the graph alignment method of the present application represented by the red line is the best, with high precision, less time consumption, and stable training. In other cases, there is a slight decline. For example, although the related art 1 represented by the green line has the least training time, its precision is not high, it oscillates continuously, fluctuates greatly, and has poor stability. Therefore, the graph alignment method of the present application improves the training efficiency while ensuring the precision, enables the alignment matrix to converge through fewer training rounds, and has a shorter training time at the same time.

[0120] In this way, by using the graph alignment method according to some embodiments of the present disclosure, after determining the adjacency matrix and feature matrix of the source domain graph and the target domain graph respectively, by iteratively executing determining the intra-graph distance matrix based on the adjacency matrix and the feature matrix, determining the inter-graph distance matrix based on the intra-graph distance matrix and the target alignment matrix, and determining the updated iterative matrix based on the inter-graph distance matrix, until the maximum number of iterations is reached or the target alignment matrix converges, the various parameters used to calculate the target alignment matrix can be continuously updated, thereby being able to improve the accuracy of the finally obtained target alignment matrix. And since the source domain graph and the target domain graph are directly used for calculation (without alignment in the embedding space), the processing efficiency of the graph alignment method is improved, the adaptability of the graph alignment method under various graph data is improved, the robustness of the graph alignment method is enhanced, which is beneficial to subsequent improving the efficiency of graph alignment using the target alignment matrix and increasing the accuracy and precision of graph alignment.

[0121] Figure 9 An exemplary block diagram of a graph alignment device 900 according to some embodiments of the present disclosure is schematically shown. Figure 9 The graph alignment device 900 shown in Figure 1 can correspond to the terminal device 110 shown in

[0122] As Figure 9As shown in [the figure], the graph alignment device 900 may include a determination module 910, an iteration module 920, and an alignment module 930. The determination module 910 may be configured to determine the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph, where the source domain graph includes a plurality of source domain nodes, the target domain graph includes a plurality of target domain nodes, each element in the adjacency matrix represents the connection relationship between nodes in the corresponding graph, and each element in the feature matrix represents the feature of the nodes in the corresponding graph. The iteration module 920 may be configured to initialize the alignment matrix as the target alignment matrix and iteratively execute the following steps until the maximum number of iterations is reached or the target alignment matrix converges: Based on the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph, determine the intra-graph distance matrix of the corresponding graph, where each element in the intra-graph distance matrix represents the degree of association between nodes in the corresponding graph; Based on the intra-graph distance matrix of the source domain graph, the intra-graph distance matrix of the target domain graph, and the target alignment matrix, determine the inter-graph distance matrix of this iteration, where each element in the inter-graph distance matrix represents the degree of association between the source domain nodes of the corresponding source domain graph and the target domain nodes of the target domain graph; Based on the inter-graph distance matrix of this iteration, determine the alignment matrix of this iteration, and update the target alignment matrix to the alignment matrix of this iteration, where each element in the alignment matrix represents the probability of matching between the source domain nodes of the corresponding source domain graph and the target domain nodes of the target domain graph. The alignment module 930 may be configured to align the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the target alignment matrix.

[0123] Figure 10 FIG. [X] schematically shows an example block diagram of a device 1000 for predicting matching points according to some embodiments of the present disclosure. Figure 10 The device 1000 for predicting matching points shown in [the figure] may correspond to Figure 1 the terminal device 110 shown in [the figure].

[0124] As Figure 10 shown in [the figure], the device 1000 for predicting matching points may include an acquisition module 1010, a graph alignment module 1020, and a prediction module 1030. The acquisition module 1010 may be configured to acquire a source domain graph and a target domain graph, where the source domain graph includes a plurality of source domain nodes and the target domain graph includes a plurality of target domain nodes. The graph alignment module 1020 may be configured to use the graph alignment device 900 to align the source domain nodes in the source domain graph with the target domain nodes in the target domain graph to obtain the aligned source domain graph and target domain graph. The prediction module 1030 may be configured to predict the number of matching points between the source domain graph and the target domain graph based on the aligned source domain graph and target domain graph.

[0125] It should be noted that the above various modules can be implemented in software or hardware or a combination of both. Multiple different modules can be implemented in the same software or hardware structure, or one module can be implemented by multiple different software or hardware structures.

[0126] In the graph alignment device according to some embodiments of the present disclosure, after determining the adjacency matrix and feature matrix of the source domain graph and the target domain graph respectively, by iteratively performing determining the intra-graph distance matrix based on the adjacency matrix and the feature matrix, determining the inter-graph distance matrix based on the intra-graph distance matrix and the target alignment matrix, and determining the updated iteration matrix based on the inter-graph distance matrix, until the maximum number of iterations is reached or the target alignment matrix converges, the various parameters for calculating the target alignment matrix can be continuously updated, thereby improving the accuracy of the finally obtained target alignment matrix. And since the source domain graph and the target domain graph are directly used for calculation (without alignment in the embedding space), the processing efficiency of the graph alignment method is improved, the adaptability of the graph alignment method under various graph data is improved, the robustness of the graph alignment method is enhanced, which is beneficial to subsequent improving the efficiency of graph alignment using the target alignment matrix and increasing the accuracy and precision of graph alignment.

[0127] Figure 11 An example block diagram of a computing device 1100 according to some embodiments of the present disclosure is schematically shown. The computing device 1100 can represent a device for implementing various devices or modules described herein and / or executing various methods described herein. The computing device 1100 can be, for example, a server, a desktop computer, a laptop computer, a tablet, a smart phone, a smart watch, a wearable device, or any other suitable computing device or computing system, which can include various levels of devices from full-resource devices with a large amount of storage and processing resources to low-resource devices with limited storage and / or processing resources. In some embodiments, regarding the Figure 9 described graph alignment device 900 can be implemented in one or more computing devices 1100 respectively.

[0128] As Figure 11 shown, the example computing device 1100 includes a processing system 1101, one or more computer-readable media 1102, and one or more I / O interfaces 1103 that are communicatively coupled to each other. Although not shown, the computing device 1100 may also include a system bus or other data and command transfer systems that couple the various components to each other. The system bus can include any one or a combination of different bus structures, and the bus structures can be such as a memory bus or a memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus using any one of various bus architectures. Alternatively, it can also include such as control and data lines.

[0129] The processing system 1101 represents the functionality of performing one or more operations using hardware. Thus, the processing system 1101 is illustrated as including hardware elements 1104 that can be configured as processors, functional blocks, etc. This can include being implemented in hardware as an application specific integrated circuit or other logic devices formed using one or more semiconductors. The hardware elements 1104 are not limited by the materials from which they are formed or the processing mechanisms employed therein. For example, a processor can be composed of (multiple) semiconductors and / or transistors (e.g., an electronic integrated circuit (IC)). In such a context, the executable instructions of the processor can be electronically executable instructions.

[0130] The computer-readable medium 1102 is illustrated as including a memory / storage device 1105. The memory / storage device 1105 represents the memory / storage device associated with one or more computer-readable media. The memory / storage device 1105 can include volatile media (such as random access memory (RAM)) and / or non-volatile media (such as read-only memory (ROM), flash memory, optical discs, magnetic disks, etc.). The memory / storage device 1105 can include fixed media (e.g., RAM, ROM, fixed hard disk drive, etc.) and removable media (e.g., flash memory, removable hard disk drive, optical disc, etc.). Exemplarily, the memory / storage device 1105 can be used to store the first audio of the first category of users, the queued list of requests, etc. mentioned in the above embodiments. The computer-readable medium 1102 can be configured in various other ways as described further below.

[0131] One or more I / O (input / output) interfaces 1103 represent the functionality that allows a user to type commands and information into the computing device 1100 and also allows information to be displayed to the user and / or sent to other components or devices using various input / output devices. Examples of input devices include a keyboard, a cursor control device (e.g., a mouse), a microphone (e.g., for voice input), a scanner, a touch function (e.g., a capacitive or other sensor configured to detect physical touch), a camera (e.g., which can detect motion not involving touch as a gesture using visible or non-visible wavelengths such as infrared frequencies), a network card, a receiver, etc. Examples of output devices include a display device (e.g., a monitor or a projector), a speaker, a printer, a tactile response device, a network card, a transmitter, etc. Exemplarily, in the embodiments described above, the first category of users and the second category of users can input to initiate requests and enter audio and / or video, etc. through the input interfaces on their respective terminal devices, and can view various notifications and watch videos or listen to audio, etc. through the output interfaces.

[0132] The computing device 1100 also includes a graph alignment application 1106. The graph alignment application 1106 can be stored as computing program instructions in the memory / storage 1105, or it can be hardware or firmware. The graph alignment application 1106 can implement all the functions of the various modules of the graph alignment device 900 described with respect to Figure 9 together with the processing system 1101 and so on.

[0133] Various technologies can be described herein in the general context of software, hardware, components, or program modules. Generally, these modules include routines, programs, objects, elements, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The terms "module", "function", etc. used herein generally refer to software, firmware, hardware, or a combination thereof. The features of the technologies described herein are platform-independent, meaning that these technologies can be implemented on various computing platforms with various processors.

[0134] The implementation of the described modules and technologies can be stored on or transmitted across some form of computer-readable medium. The computer-readable medium can include various media accessible by the computing device 1100. By way of example and not limitation, the computer-readable medium can include "computer-readable storage media" and "computer-readable signal media".

[0135] Contrary to mere signal transmission, carrier, or signal itself, "computer-readable storage media" refers to media and / or devices that can persistently store information, and / or tangible storage devices. Thus, computer-readable storage media refers to non-signal-bearing media. Computer-readable storage media includes hardware such as volatile and non-volatile, removable and non-removable media and / or storage devices implemented in a method or technology suitable for storing information such as computer-readable instructions, data structures, program modules, logic elements / circuits, or other data. Examples of computer-readable storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical storage devices, hard disks, cassette tapes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or other storage devices, tangible media, or articles of manufacture suitable for storing the desired information and accessible by a computer.

[0136] "Computer-readable signal medium" refers to a signal-bearing medium configured to send instructions to a computing device 1100, such as via a network. A signal medium typically can embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, a data signal, or other transmission mechanism. The signal medium also includes any information delivery medium. By way of example, and not limitation, the signal medium includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media.

[0137] As described above, hardware elements 1104 and computer-readable media 1102 represent instructions, modules, programmable device logic, and / or fixed device logic implemented in hardware, which in some embodiments can be used to implement at least some aspects of the techniques described herein. Hardware elements can include integrated circuits or system-on-a-chip, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and components of other hardware devices implemented in silicon or other hardware. In this context, the hardware elements can serve as processing devices that execute program tasks defined by the instructions, modules, and / or logic embodied by the hardware elements, as well as hardware devices that store instructions for execution, e.g., the computer-readable storage media described previously.

[0138] The foregoing combinations can also be used to implement the various techniques and modules described herein. Thus, software, hardware, or program modules and other program modules can be implemented as one or more instructions and / or logic on some form of computer-readable storage medium and / or embodied by one or more hardware elements 1104. Computing device 1100 can be configured to implement specific instructions and / or functions corresponding to the software and / or hardware modules. Thus, for example, by using the computer-readable storage medium of the processing system and / or hardware elements 1104, a module can be implemented at least in part in hardware as a module executable by computing device 1100 as software. The instructions and / or functions can be executed / operable by, for example, one or more computing devices 1100 and / or processing system 1101 to implement the techniques, modules, and examples described herein.

[0139] The techniques described herein can be supported by these various configurations of computing device 1100 and are not limited to the specific examples of the techniques described herein.

[0140] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer programs. For example, embodiments of the present disclosure provide a computer program product that includes a computer program carried on a computer-readable medium, the computer program including program code for performing at least one step in the method embodiments of the present disclosure.

[0141] In some embodiments of the present disclosure, one or more computer-readable storage media are provided, on which computer-readable instructions are stored, and when the computer-readable instructions are executed, a graph alignment method according to some embodiments of the present disclosure is implemented. Each step of the graph alignment method according to some embodiments of the present disclosure can be transformed into computer-readable instructions through programming and thus stored in a computer-readable storage medium. When such a computer-readable storage medium is read or accessed by a computing device or a computer, the computer-readable instructions therein are executed by a processor on the computing device or the computer to implement the method according to some embodiments of the present disclosure.

[0142] In the description of this specification, the descriptions of the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0143] Any process or method description represented in a flowchart or described in other ways herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations, where the functions can be executed in an order not shown or discussed (including in a substantially simultaneous manner according to the functions involved or in a reverse order), which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.

[0144] The logic and / or steps represented in a flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0145] It should be understood that various parts of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, any one of the following techniques well known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays, field programmable gate arrays, and the like.

[0146] Those of ordinary skill in the art of this technology can understand that all or part of the steps of the above method embodiments can be completed by hardware related to program instructions, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0147] In addition, in each embodiment of the present disclosure, each functional unit can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0148] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0149] Although the present disclosure has been described in connection with some embodiments, it is not intended to be limited to the specific forms set forth herein. Rather, the scope of the present disclosure is limited only by the appended claims. Additionally, although individual features may be included in different claims, these may possibly be advantageously combined, and inclusion in different claims does not imply that a combination of features is not feasible and / or advantageous. The order of features in the claims does not imply that the features must work in any specific order. Further, in the claims, the word "comprising" does not exclude other elements, and the terms "a" or "an" do not exclude a plurality. The reference signs in the claims are provided only as illustrative examples and should not be construed as limiting the scope of the claims in any way.

[0150] It will be understood that, in the specific embodiments of the present disclosure, various data for graph alignment are involved. When the embodiments described in the present disclosure involving such data are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.

Claims

1. A graph alignment method, characterized in that, The method includes: Determine the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph. The source domain graph includes a plurality of source domain nodes, and the target domain graph includes a plurality of target domain nodes. Each element in the adjacency matrix represents the connection relationship between nodes in the corresponding graph, and each element in the feature matrix represents the feature of the nodes in the corresponding graph; Initialize the alignment matrix as the target alignment matrix, and iteratively execute the following steps until the maximum number of iterations is reached or the target alignment matrix converges: Based on the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph, determine the intra-graph distance matrix of the corresponding graph. Each element in the intra-graph distance matrix represents the degree of association between nodes in the corresponding graph; Based on the intra-graph distance matrix of the source domain graph, the intra-graph distance matrix of the target domain graph, and the target alignment matrix, determine the inter-graph distance matrix of the current iteration. Each element in the inter-graph distance matrix represents the degree of association between the source domain nodes of the corresponding source domain graph and the target domain nodes of the target domain graph; Based on the inter-graph distance matrix of the current iteration, determine the alignment matrix of the current iteration, and update the target alignment matrix to the alignment matrix of the current iteration. Each element in the alignment matrix represents the matching probability between the source domain nodes of the corresponding source domain graph and the target domain nodes of the target domain graph; Align the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the target alignment matrix.

2. The method according to claim 1, wherein The step of determining the intra-graph distance matrix of the corresponding graph based on the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph includes: Based on the feature matrix of the corresponding graphs in the source domain graph and the target domain graph, determine the node similarity matrix of the corresponding graph. Each element in the node similarity matrix represents the similarity between the features of the nodes in the corresponding graph; Based on the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph, determine the splicing matrix of the corresponding graph. Each element in the splicing matrix represents the relationship between the nodes in the corresponding graph and their neighbor nodes; Determine the intra-graph distance matrix of the corresponding graph based on the adjacency matrix, the node similarity matrix, and the splicing matrix of the corresponding graph.

3. The method according to claim 2, wherein The step of determining the intra-graph distance matrix of the corresponding graph based on the adjacency matrix, the node similarity matrix, and the splicing matrix of the corresponding graph includes: Determine the weighted sum of the adjacency matrix, the node similarity matrix, and the splicing matrix of the corresponding graph as the intra-graph distance matrix of the corresponding graph.

4. The method according to claim 2, characterized in that, The step of determining the splicing matrix of the corresponding graph based on the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph includes: Use a deep learning network to sequentially determine the 1st to Kth order neighbor feature matrices of the corresponding graph based on the adjacency matrix and the feature matrix of the corresponding graphs in the source domain graph and the target domain graph. Each element in any Mth order neighbor feature matrix represents the feature of the Mth order neighbor nodes of the nodes in the corresponding graph. K is a positive integer not less than 1, and M is a positive integer greater than or equal to 1 and less than or equal to K. The Mth order neighbor nodes of a node refer to the nodes separated from this node by M - 1 nodes; Concatenate the feature matrix of the corresponding graph and the 1st to Kth order neighbor feature matrices of the corresponding graph according to weights; Use an activation function to activate the concatenated matrix to obtain the concatenated matrix of the corresponding graph.

5. The method according to claim 2, wherein Determining the node similarity matrix of the corresponding graph based on the feature matrices of the corresponding graphs in the source domain graph and the target domain graph includes: Perform a transpose operation on the feature matrices of the corresponding graphs in the source domain graph and the target domain graph to obtain the transposed feature matrices of the corresponding graphs; Perform an inner product operation on the feature matrix of the corresponding graph and the transposed feature matrix of the corresponding graph to obtain the node similarity matrix of the corresponding graph.

6. The method according to claim 1, characterized in that Determining the alignment matrix for the current iteration based on the inter-graph distance matrix for the current iteration includes: Determine the optimal transport distance between the inter-graph distance matrix for the current iteration and the target alignment matrix, where the optimal transport distance represents the minimum cost of mapping the inter-graph distance matrix for the current iteration to the target alignment matrix; Determine the alignment matrix for the current iteration based on the optimal transport distance.

7. The method according to claim 6, characterized in that, The determining the optimal transport distance between the inter-graph distance matrix for the current iteration and the target alignment matrix includes: Perform data normalization on the inter-graph distance matrix for the current iteration to obtain the normalized inter-graph distance matrix for the current iteration; Based on the normalized inter-graph distance matrix for the current iteration, determine the optimal transport distance between the normalized inter-graph distance matrix for the current iteration and the target alignment matrix.

8. The method according to claim 7, characterized in that, The performing data normalization on the inter-graph distance matrix for the current iteration to obtain the normalized inter-graph distance matrix for the current iteration includes: Select the maximum value from all elements of the inter-graph distance matrix for the current iteration; For each element in the inter-graph distance matrix for the current iteration, map each element in the inter-graph distance matrix for the current iteration to a normalized element by linear normalization based on the maximum value to obtain the normalized inter-graph distance matrix for the current iteration.

9. The method according to claim 1, characterized in that Aligning the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the target alignment matrix includes: For each source domain node in the source domain graph, select the highest alignment weight corresponding to each source domain node from all elements of the target alignment matrix, where the highest alignment weight represents the alignment probability with the highest probability of aligning the source domain node in the source domain graph with the target domain node in the target domain graph; Based on the highest alignment weight corresponding to each source domain node, select a target domain node that matches each source domain node respectively to obtain a pair of matching nodes; Based on the pair of matching nodes, align the source domain nodes in the source domain graph with the target domain nodes in the target domain graph.

10. The method according to claim 9, characterized in that, The aligning the source domain nodes in the source domain graph with the target domain nodes in the target domain graph based on the pair of matching nodes includes: Sort the highest alignment weights corresponding to each source domain node in descending order, and select a preset number of matching node pairs as the set of labeled node pairs according to the sorted highest alignment weights. The set of labeled node pairs includes a set of labeled source domain nodes and a set of labeled target domain nodes; Based on the set of labeled source domain nodes and the set of labeled target domain nodes, determine a cross-graph cost matrix, where each element in the cross-graph cost matrix represents the similarity between a labeled source domain node in the set of labeled source domain nodes and a labeled target domain node in the set of labeled target domain nodes; Iteratively update the target alignment matrix based on the cross-graph cost matrix until the target alignment matrix converges to obtain an updated target alignment matrix; Align the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the updated target alignment matrix.

11. A method for predicting matching points, characterized in that, The method includes: Obtain a source domain graph and a target domain graph, where the source domain graph includes a plurality of source domain nodes, and the target domain graph includes a plurality of target domain nodes; Align the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the graph alignment method described in any one of the preceding claims to obtain an aligned source domain graph and target domain graph; Based on the aligned source domain graph and target domain graph, predict the number of matching points between the source domain graph and the target domain graph.

12. A graph alignment device, comprising: A determination module configured to determine the adjacency matrix and the feature matrix of the corresponding graph in the source domain graph and the target domain graph, where the source domain graph includes a plurality of source domain nodes, the target domain graph includes a plurality of target domain nodes, each element in the adjacency matrix represents the connection relationship between nodes in the corresponding graph, and each element in the feature matrix represents the feature of the nodes in the corresponding graph; An iteration module configured to initialize the alignment matrix as the target alignment matrix and iteratively execute the following steps until the maximum number of iterations is reached or the target alignment matrix converges: Based on the adjacency matrix and the feature matrix of the corresponding graph in the source domain graph and the target domain graph, determine the intra-graph distance matrix of the corresponding graph, where each element in the intra-graph distance matrix represents the correlation degree between nodes in the corresponding graph; Based on the intra-graph distance matrix of the source domain graph, the intra-graph distance matrix of the target domain graph, and the target alignment matrix, determine the inter-graph distance matrix of this iteration, where each element in the inter-graph distance matrix represents the correlation degree between the corresponding source domain node of the source domain graph and the target domain node of the target domain graph; Based on the inter-graph distance matrix of this iteration, determine the alignment matrix of this iteration and update the target alignment matrix to the alignment matrix of this iteration, where each element in the alignment matrix represents the matching probability between the corresponding source domain node of the source domain graph and the target domain node of the target domain graph; An alignment module configured to align the source domain nodes in the source domain graph with the target domain nodes in the target domain graph according to the target alignment matrix.

13. A device for predicting matching points, comprising: An acquisition module, configured to acquire a source domain graph and a target domain graph, wherein the source domain graph includes a plurality of source domain nodes, and the target domain graph includes a plurality of target domain nodes; A graph alignment module, configured to align the source domain nodes in the source domain graph with the target domain nodes in the target domain graph by using the graph alignment device according to claim 12, so as to obtain the aligned source domain graph and target domain graph; A prediction module, configured to predict the number of matching points between the source domain graph and the target domain graph based on the aligned source domain graph and target domain graph.

14. A computing device, characterized in that, The computing device includes: A memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, it causes the processor to execute the method according to any one of claims 1-11.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed, the method according to any one of claims 1-11 is implemented.

16. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-11 is implemented.