Sampling-based model training method, data classification method and related equipment

By building a lightweight target graph in graph neural network training, the neighbor explosion problem is solved through sampling operations of nodes and edges, and the training speed is improved and the classification accuracy is guaranteed.

CN120279350APending Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410028321.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

There is a neighbor explosion problem in the existing graph neural network (GNN) training, resulting in large amounts of training data and long time, and the existing scheme has high computing time and storage space costs, and weak expression capabilities.

Method used

By constructing a sampling-based model training method, including sampling nodes and edges in the initial graph, a lightweight target map is generated, which is used to train the target classification model, avoid neighbor explosion problems, and classify data through the pre-trained target classification model.

Benefits of technology

While ensuring a certain accuracy, the training speed and data classification efficiency are improved, the training data volume is reduced, and the model training speed and timeliness of classification results are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279350A_ABST
    Figure CN120279350A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a sampling-based model training method, a data classification method and related equipment. The sampling-based model training method, the data classification method and the related equipment can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic and auxiliary driving. The model training method comprises the steps of obtaining a target map of training data, and performing model training based on the target map to obtain a model for executing a data classification task; the target graph is obtained through construction operation: obtaining an initial graph of training data, wherein the graph comprises sample nodes corresponding to all samples in the training data and a first edge connected with the sample nodes; node sampling is carried out based on the connection relation of sample nodes in the initial graph, and a sample node set comprising at least two layers in the target graph is obtained; and for nodes in each adjacent layer in the sample node set, carrying out edge sampling based on a connection object of the first edge in the initial graph to obtain a second edge connected with the sample nodes of the adjacent layer in the target graph. According to the method, the model training time can be shortened by constructing the lightweight graph training model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology. Specifically, this application relates to a sampling-based model training method, a data classification method, and related devices. Background Art

[0002] In recent years, remarkable achievements have been made in the field of artificial intelligence. Among them, graph neural networks (GNNs) are a type of deep learning model used to process graph-structured data. Entities in graph-structured data are represented as nodes, and the relationships between entities are represented as edges. The goal of GNNs is to learn useful representations from graph-structured data and use these representations for various tasks, such as data classification.

[0003] However, there is a problem of neighbor explosion in existing GNN training, resulting in a large amount of data to be trained and a long training time. Summary of the Invention

[0004] Embodiments of this application provide a sampling-based model training method, a data classification method, and related devices to solve the above at least one technical problem. The technical solutions are as follows:

[0005] In a first aspect, embodiments of this application provide a sampling-based model training method, including:

[0006] Obtain a target graph corresponding to training data to perform model training based on the target graph and obtain a target classification model for performing a data classification task;

[0007] Wherein, the target graph is obtained by performing the following construction operations:

[0008] Obtain an initial graph corresponding to training data; the initial graph includes sample nodes corresponding to each sample in the training data and first edges connecting the sample nodes;

[0009] Perform a node sampling operation based on the connection relationship of sample nodes in the initial graph to obtain a sample node set including at least two layers in the target graph;

[0010] For each adjacent layer of sample nodes in the sample node set, perform an edge sampling operation based on the connection objects of the first edges in the initial graph to obtain second edges connecting the sample nodes in the adjacent layer in the target graph.

[0011] In a feasible embodiment, the performing a node sampling operation based on the connection relationship of sample nodes in the initial graph to obtain a sample node set including at least two layers in the target graph includes at least one of the following:

[0012] For the construction of each layer in the target graph, based on the sample nodes in the previous layer, at least one sample node adjacent to the sample node is obtained from the initial graph to form the sample nodes of the current layer.

[0013] Based on the initial nodes sampled from the initial graph and the connection relationships of the included sample nodes, node sampling is performed from the initial graph to obtain the sample nodes of each layer in the target graph.

[0014] In a feasible embodiment, the construction of each layer in the target graph, based on the sample nodes in the previous layer, obtaining at least one sample node adjacent to the sample node from the initial graph to form the sample nodes of the current layer, includes:

[0015] Repeatedly execute the following operation for determining the nodes of the current layer until the number of constructed layers reaches the preset number of layers:

[0016] For each first sample node in the previous layer, based on the connection relationships of the sample nodes in the initial graph, obtain a preset first sampling number of second sample nodes adjacent to the first sample node;

[0017] Take the union of the first sample node and the second sample nodes as the sample nodes of the current layer;

[0018] Wherein, when the previous layer is the first layer for constructing the target graph, the sample nodes of this layer include the sample nodes sampled from the initial graph based on the mini-batch sampling method.

[0019] In a feasible embodiment, the node sampling is performed from the initial graph based on the initial nodes sampled from the initial graph and the connection relationships of the included sample nodes to obtain the sample nodes of each layer in the target graph, including:

[0020] Sample third sample nodes from the initial graph based on the mini-batch sampling method, and construct the output layer of the target graph based on the third sample nodes;

[0021] For each third sample node, randomly obtain a preset second sampling number of fourth sample nodes corresponding to the third sample node based on the connection relationships of the sample nodes in the initial graph;

[0022] Based on the union of the third sample nodes and the fourth sample nodes, construct the other layers in the target graph except the output layer.

[0023] In a feasible embodiment, for each sample node in each adjacent layer of the sample node set, based on the connection objects of the first edges in the initial graph, an edge sampling operation is performed to obtain the second edges connecting the sample nodes in this adjacent layer in the target graph, including:

[0024] For any two sample nodes located between adjacent layers in the sample node set, if there is a first edge connecting the two sample nodes in the initial graph, then connect the two sample nodes to form the third edges connecting the sample nodes between different layers in the target graph;

[0025] Starting from the output layer of the target graph, the following first sampling operation is repeatedly executed for the third edges until all the second edges connecting the sample nodes between all layers are obtained:

[0026] For the starting node corresponding to the third edge in the current adjacent layer, obtain a preset third sampling number of edges in the third edges of the next adjacent layer whose arrival node is the starting node as the second edges connecting the sample nodes between different layers in the next adjacent layer.

[0027] In a feasible embodiment, for each sample node in each adjacent layer of the sample node set, based on the connection objects of the first edges in the initial graph, an edge sampling operation is performed to obtain the second edges connecting the sample nodes in this adjacent layer in the target graph, including:

[0028] Starting from the output layer of the target graph, the following second sampling operation is repeatedly executed until all the second edges connecting the sample nodes between all layers are obtained:

[0029] For the sample nodes between the current adjacent layers, obtain a preset fourth sampling number of the first edges connecting the sample nodes in the initial graph, and construct the second edges connecting the sample nodes in this adjacent layer.

[0030] In a feasible embodiment, training the model based on the target graph to obtain a target classification model for performing data classification tasks includes:

[0031] Sampling the sample nodes in the target graph to obtain the first seed nodes;

[0032] When training the initial classification model based on the training data, in each training round, sample the sample nodes in the target graph to obtain the second seed nodes for the current training round, and determine the target training samples for the current training round based on the first seed nodes and the second seed nodes, and train the initial classification model through the target training samples to train and obtain a target classification model for performing data classification tasks.

[0033] In a feasible embodiment, sampling for sample nodes in the target graph includes at least one of the following:

[0034] Based on a first sampling rate, uniformly and randomly sample sample nodes in the target graph;

[0035] For each sample node in the target graph, determine the node degree of the sample node based on a second edge connected to the sample node; sample the sample nodes in the target graph based on the second sampling rate and the node degree; the second sampling rate is proportional to the node degree.

[0036] In a feasible embodiment, sampling for sample nodes in the target graph to obtain second seed nodes for the current training round includes:

[0037] Set the first seed nodes in the target graph to a non-sampling state, and sample sample nodes in the target graph to obtain second seed nodes for the current training round; the second seed nodes have no intersection with the first seed nodes.

[0038] In a second aspect, an embodiment of the present application provides a data classification method, including:

[0039] In response to a data classification operation, classify data corresponding to the operation through a pre-trained target classification model;

[0040] Wherein, the target classification model is trained by the method provided in the first aspect.

[0041] In a third aspect, an embodiment of the present application provides a sampling-based model training device, including:

[0042] An acquisition module, configured to acquire a target graph corresponding to training data, so as to perform model training based on the target graph to obtain a target classification model for performing a data classification task;

[0043] Wherein, the device further includes a construction module, configured to perform the following construction operation to obtain the target graph:

[0044] Acquire an initial graph corresponding to training data; the initial graph includes sample nodes corresponding to each sample in the training data and first edges connecting the sample nodes;

[0045] Based on the connection relationship of sample nodes in the initial graph, perform a node sampling operation to obtain a sample node set including at least two layers in the target graph;

[0046] For each sample node in each adjacent layer of the sample node set, an edge sampling operation is performed based on the connection objects of the first edges in the initial graph, and second edges connected between the sample nodes in this adjacent layer in the target graph are obtained.

[0047] In a fourth aspect, an embodiment of the present application provides a data classification device, including:

[0048] A classification module, configured to, in response to a data classification operation, classify data corresponding to the operation through a pre-trained target classification model;

[0049] Wherein, the target classification model is trained by the method provided in the first aspect.

[0050] In a fifth aspect, an embodiment of the present application provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the methods provided in the first aspect and the second aspect above.

[0051] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the methods provided in the first aspect and the second aspect above are implemented.

[0052] In a seventh aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the methods provided in the first aspect and the second aspect above are implemented.

[0053] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are:

[0054] In a first aspect, an embodiment of the present application provides a sampling-based model training method, which can construct a lightweight graph for model training to improve the training speed while ensuring a certain accuracy. Specifically, the target graph corresponding to the training data used in training can be obtained by performing a sampling operation on the basis of an initial graph, where the initial graph can include sample nodes corresponding to each sample in the training data and first edges connected between the sample nodes. On this basis, the present application can perform a node sampling operation based on the connection relationship of the sample nodes in the initial graph to obtain a sample node set including at least two layers in the target graph, completing the construction of the nodes. That is, the present application does not need to save the edges connected between the sample nodes during the construction of the nodes, and can also effectively avoid the neighbor explosion problem. Subsequently, for each adjacent layer of sample nodes in the sample nodes, an edge sampling operation can be performed based on the connection objects of the first edges in the initial graph to obtain second edges connected between the sample nodes in the adjacent layer in the target graph. That is, in order to avoid the problem of a large amount of training data caused by too many edges in a dense graph during the construction of the edges, the present application can reduce the required training data volume through the edge sampling operation to effectively accelerate the training.

[0055] In a second aspect, an embodiment of the present application provides a data classification method. Specifically, when responding to a data classification operation, the pre-trained target classification model can be used to classify the data corresponding to the operation. Among them, the target classification model can be trained by using the sampling-based model training method provided in the first aspect, and its training process can improve the training speed while ensuring a certain accuracy, which is beneficial to ensuring the accuracy of the classification result and the timeliness of the response when classifying data through the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description in the embodiments of the present application.

[0057] Figure 1 It is a flowchart of a sampling-based model training method provided by an embodiment of the present application;

[0058] Figure 2 It is a flowchart of a data classification method provided by an embodiment of the present application;

[0059] Figure 3 It is a schematic diagram of an initial graph provided by an embodiment of the present application;

[0060] Figure 4 It is a schematic diagram of a hierarchy and a block provided by an embodiment of the present application;

[0061] Figure 5 It is a schematic diagram of node construction provided by an embodiment of the present application;

[0062] Figure 6 Another schematic diagram of node construction provided by an embodiment of the present application;

[0063] Figure 7 A schematic diagram of constructing a target graph provided by an embodiment of the present application;

[0064] Figure 8 A schematic diagram of the first prediction accuracy data provided by an embodiment of the present application;

[0065] Figure 9 A schematic diagram of the second prediction accuracy data provided by an embodiment of the present application;

[0066] Figure 10 A comparison diagram of the effect of convergence speed provided by an embodiment of the present application;

[0067] Figure 11 A comparison diagram of the effect of training time provided by an embodiment of the present application;

[0068] Figure 12 A schematic diagram of a sampling-based model training device provided by an embodiment of the present application;

[0069] Figure 13 A schematic diagram of a data classification device provided by an embodiment of the present application;

[0070] Figure 14 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0071] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.

[0072] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude being implemented as other features, information, data, steps, operations, elements, components and / or their combinations supported by the technical field of the present invention. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".

[0073] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0074] Artificial Intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making.

[0075] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include, for example, sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the base model, can be widely applied to downstream tasks in various major directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0076] In the embodiments of this application, machine learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pre-trained model is the latest development result of deep learning, integrating the above technologies.

[0077] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence-generated content (AIGC), conversational interactions, smart healthcare, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role. The solution provided in this application can adapt to data in different fields for processing, such as finance, transportation, education, instant messaging, etc., express the corresponding data through graph data, and perform classification tasks according to the requirements.

[0078] In the existing GNN training methods, various solutions have been provided to solve the neighbor explosion problem, but the existing solutions still have the following problems:

[0079] (1) The computational time and storage space costs are relatively high. Node sampling still cannot avoid the neighbor explosion problem. Although layer sampling limits the number of nodes in each layer, this upper bound is often set relatively high. To bridge the gap between minibatch training and fullbatch (using all training data at once for model training in machine learning) training, subgraph sampling methods often choose to construct larger subgraphs. At the same time, subgraph sampling will retain the local structure of the original graph, resulting in a large number of edges needing to be stored when constructing the MPG (Message Passing Graph).

[0080] (2) The expressive ability is relatively weak. The MPG constructed by node sampling and layer sampling contains too few edges and the structure is extremely sparse, resulting in nodes being unable to obtain enough information from the sampled neighbors. At the same time, when a node has a high degree in the original graph, sampling only a few neighbor nodes for message propagation may lead to a large bias, affecting the final accuracy.

[0081] In view of at least one technical problem existing in the above-mentioned prior art, an embodiment of the present application provides a sampling-based model training method, specifically a method for training a model by constructing a lighter and more expressive message propagation graph.

[0082] The technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below by describing several exemplary embodiments. It should be noted that the following embodiments can be referred to, borrowed from, or combined with each other. The same terms, similar features, and similar implementation steps in different embodiments will not be described repeatedly.

[0083] The sampling-based model training method in the embodiments of the present application will be described below.

[0084] Specifically, the execution subject of the method provided by the embodiment of the present application can be a terminal or a server; the terminal (which can also be referred to as a device) can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart voice interaction device (such as a smart speaker), a wearable electronic device (such as a smart watch), a vehicle-mounted terminal, a smart home appliance (such as a smart TV), an AR / VR device, etc., but is not limited thereto. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers (such as a distributed cloud storage system), or a cloud server providing cloud computing and cloud storage services.

[0085] Specifically, as Figure 1 shown, the sampling-based model training method includes step S101:

[0086] Step S101: Obtain a target graph corresponding to the training data, and train a model based on the target graph to obtain a target classification model for performing a data classification task.

[0087] Among them, the training data obtained is different according to different application scenarios. For example, in the instant messaging scenario, when learning the object data of an instant messaging program, the features related to the object can constitute various information to be learned in the training data. When classifying using the model trained based on the training data, it can be divided according to a certain preference or operation habit of the object.

[0088] Among them, the target graph can be an MPG, including multiple layers of nodes with no edges between the nodes within a layer, while there are edges between the nodes of different layers to identify the path of node message propagation through the edges. The model can be a deep learning model for processing graph-structured data, such as a GNN, or it can also be a GCN (Graph Convolutional Networks), GraphSAGE (Graph Sample and aggregate), GAT (Graph Attention Network), etc.

[0089] Among them, the target graph is obtained by performing the construction operations of the following steps A1 - A3:

[0090] Step A1: Obtain an initial graph corresponding to the training data; the initial graph includes sample nodes corresponding to each sample in the training data and first edges connecting the sample nodes.

[0091] Optionally, as Figure 3 shown, when reflecting the training data through the initial graph, it can be the obtained graph G(V, E), where each node v has an initial feature. The GNN can map the initial feature to information in other feature spaces through operations such as aggregating neighbors and linear transformation, and obtain the initial graph G including the sample nodes V and the first edges E connecting the sample nodes. Obtaining the graph-structured data of a certain data using the GNN (such as obtaining the initial graph G) can refer to the related technology.

[0092] In one example, assuming the training data is data corresponding to a family tree, then each sample node can correspond to a family member, and the first edge can correspond to the relationship between family members; for example, the sample node V1 can be the father, and the sample node V4 can be the grandfather, then the first edge connecting the sample nodes V1 and V4 can indicate the existence of a father-son relationship between the two sample nodes.

[0093] Step A2: Perform a node sampling operation based on the connection relationship of the sample nodes in the initial graph to obtain a sample node set including at least two layers in the target graph.

[0094] Among them, through the connection relationship of the sample nodes in the initial graph, the association relationship between any node and other nodes can be known. For example, Figure 3 shown, it can be known that the neighbor nodes of the sample node V1 include the sample nodes V2, V3, and V4, while there is no association relationship between the sample node V1 and the sample node V5.

[0095] Among them, the sampling operation of nodes can be performed on the sampled nodes included in the initial graph. For example, it can be performed on some given, already sampled, or arbitrary sampled nodes V. Based on the connection relationships of the sample nodes in the initial graph, other nodes having connection relationships with them are sampled. Exemplarily, as Figure 3 shown, assuming that the sampling operation of nodes is performed based on the sample node V2, then based on the Figure 3 connection relationships of the sample nodes in this, when sampling based on the sample node V2, its neighbor nodes V1, V3, and V4 can be obtained first, and then based on the neighbor node V1, its neighbor nodes V2, V3, and V4 can be obtained, and so on. Optionally, at this time, the sample node V2 can be used as a sample node in one layer of the sample node set, and the sample nodes V2, V1, V3, and V4 can be used as sample nodes in another layer. Among them, when sampling neighbor nodes based on the sample node V2, the number of sampled sample nodes can be adjusted according to actual needs, such as based on the sparsity of the graph, the data volume size of the graph structure data, etc.

[0096] Step A3: For the sample nodes in each adjacent layer of the sample node set, perform an edge sampling operation based on the connection objects of the first edges in the initial graph to obtain the second edges connecting the sample nodes in this adjacent layer in the target graph.

[0097] Optionally, as Figure 4 shown, the adjacent layers can refer to layer 0 and layer 1, or layer 1 and layer 2. Optionally, the adjacent layers can also be understood using the concept of a block. That is, Figure 4 there can be two blocks, where block 2 includes layer 1 and layer 2. Correspondingly, for the Figure 4 example sample node set, it is necessary to process two adjacent layers separately.

[0098] Optionally, the connection objects of the first edges in the initial graph can know the two sample nodes connected by any first edge. For example, Figure 3 shown, there are 7 edges. For the leftmost first edge, it can be known that its connection objects are the sample node V1 and the sample node V3; for the rightmost first edge, it can be known that its connection objects are the sample node V4 and the sample node V5.

[0099] Optionally, the edge sampling operation can be performed on the first edges included in the initial graph. Exemplarily, when targeting the Figure 4 sample nodes in the adjacent layers (layer 1 and layer 2) shown, by Figure 3The connection objects on the first side shown can obtain the first sides connected between sample node V1 and sample node V2, between sample node V1 and sample node V3, between sample node V1 and sample node V4, between sample node V2 and sample node V1, between sample node V2 and sample node V3, and between sample node V2 and sample node V4. Subsequently, sampling can be performed on the corresponding connected edges to obtain the second edges connected between the sample nodes in the adjacent layers in the target graph. For example, any two first edges are respectively obtained for sample node V1 and sample node V2 as the second edges.

[0100] In a feasible embodiment, in step A2, the sampling operation of nodes is performed based on the connection relationship of sample nodes in the initial graph to obtain a sample node set including at least two layers in the target graph, including at least one of the following steps A21 - step A22:

[0101] Step A21: For the construction of each layer in the target graph, based on the sample nodes in its upper layer, at least one sample node adjacent to the sample node is obtained from the initial graph to form the sample nodes in the current layer.

[0102] Optionally, as Figure 5 shown, when constructing layer 1 (the middle layer), based on the sample nodes V1 and V2 in layer 2, at least one sample node adjacent to sample nodes V1 and V2 is respectively obtained from the initial graph to form the sample nodes in the current layer (layer 1).

[0103] Exemplarily, the sample nodes (neighbor nodes) adjacent to sample node V1 include V2, V3, and V4; the sample nodes (neighbor nodes) adjacent to sample node V2 include V1, V3, and V4; at this time, the sample node set of layer 1 can be constructed based on sample nodes V1, V2, V3, and V4. Figure 5 Illustratively, one adjacent sample node is respectively obtained for each sample node in layer 2 to construct the sample nodes in layer 1. For example, sample node V1 obtains the adjacent sample node V2, and sample node V2 obtains the adjacent sample node V3. At this time, layer 1 as Figure 5 shown can be constructed based on sample nodes V1, V2, and V3.

[0104] Optionally, in step A21, for the construction of each layer in the target graph, based on the sample nodes in its upper layer, at least one sample node adjacent to the sample node is obtained from the initial graph to form the sample nodes in the current layer, including:

[0105] Repeatedly execute the determination operation of the current layer nodes in the following steps A211 - step A212 until the number of constructed layers reaches the preset number of layers:

[0106] Step A211: For each first sample node in the upper layer, based on the connection relationship of the sample nodes in the initial graph, obtain a preset first sampling quantity of second sample nodes adjacent to the first sample node; wherein, when the upper layer is the first layer (such as the output layer - layer 2) for constructing the target graph, the sample nodes of this layer include the sample nodes sampled from the initial graph based on the minibatch sampling method.

[0107] Step A212: Use the union of the first sample nodes and the second sample nodes as the sample nodes of the current layer.

[0108] Optionally, the minibatch sampling can be performed by sampling the sample nodes included in the initial graph through sampling methods such as uniform sampling and sampling by node degree to obtain sample nodes with a smaller data volume; wherein, the size of the minibatch can be selected based on information such as training, the size of the graph structure data, and the sparsity of the graph. Exemplarily, as Figure 5 shown, layer 2 is obtained by collecting minibatch: {1, 2} from the initial graph, and at this time, layer 2 includes sample nodes V1 and V2.

[0109] Exemplarily, assume the preset number of layers is 3, as Figure 5 shown, then the sample nodes of layer 1 can be determined based on layer 2, and then the sample nodes of layer 0 can be determined based on layer 1. Among them, assume the first sampling quantity is 1. When determining the sample nodes of layer 1 based on the sample nodes in layer 2, one sample node V2 adjacent to the sample node V1 in layer 2 can be obtained (see the connection relationship of the sample nodes in the initial graph, and one of the sample nodes V2, V3, and V4 can be determined as the sampled sample node), and one sample node V3 adjacent to the sample node V2 in layer 2 can be obtained (see the connection relationship of the sample nodes in the initial graph, and one of the sample nodes V1, V3, and V4 can be determined as the sampled sample node). At this time, the first sample nodes include V1 and V2, and the second sample nodes include V2 and V3. Then the union of the first sample nodes and the second sample nodes is {V1, V2, V3}. Therefore, the sample nodes of layer 1 include V1, V2, and V3. Based on the same processing logic, the sample nodes of layer 0 can be determined based on the sample nodes of layer 1, and a sample node set of the target graph (including 3 layers) as shown on the right side in Figure 5 is formed.

[0110] Step A22: Based on the initial nodes sampled from the initial graph and the connection relationship of the included sample nodes, perform node sampling from the initial graph to obtain the sample nodes of each layer in the target graph.

[0111] Optionally, the construction of the sample nodes in the target graph can start from the initial nodes sampled from the initial graph, asFigure 6 As shown, initial nodes can be sampled from the initial graph first, and then based on the sampled sample nodes V1 and V2 (the nodes shown by the dashed lines in the figure) and the connection relationships of the sample nodes included in the initial graph, the sample nodes of each layer in the target graph are determined.

[0112] Optionally, in step A22, based on the initial nodes sampled from the initial graph and the connection relationships of the included sample nodes, node sampling is performed from the initial graph to obtain the sample nodes of each layer in the target graph, including steps A221 - A223:

[0113] Step A221: Sample third sample nodes from the initial graph based on a mini - batch sampling method, and construct the output layer of the target graph based on the third sample nodes.

[0114] Optionally, the sampling of mini - batch can be to sample the sample nodes included in the initial graph through sampling methods such as uniform sampling and sampling according to node degrees to obtain third sample nodes with a smaller data volume; among them, the size of the mini - batch can be selected based on information such as training, the size of the graph structure data, and the sparsity of the graph. Exemplarily, as Figure 6 shown, its layer 2 is obtained by collecting mini - batch: {1, 2} from the initial graph. At this time, the third sample nodes V1 and V2 are included in layer 2 (output layer).

[0115] Step A222: For each third sample node, based on the connection relationships of the sample nodes in the initial graph, randomly obtain a preset number of second - sampled fourth sample nodes corresponding to the third sample node.

[0116] Optionally, when obtaining the fourth sample nodes based on the third sample nodes, a random walk method can be used, combined with the currently preset number of second samplings (the length of the walk) for sampling sample nodes. Exemplarily, as Figure 6 shown, assuming that the current number of second samplings is 3, when sampling based on the third sample node V1, the paths of the walk can be obtained: V1 - V2 - V3, V1 - V3 - V2, V1 - V3 - V4, V1 - V4 - V2, V1 - V4 - V3, etc. At this time, one of the paths obtained from the walk can be randomly selected, and the sample nodes corresponding to the path are confirmed as the fourth sample nodes. For example, when randomly obtaining the path V1 - V3 - V2, the fourth sample nodes corresponding to the third sample node V1 include V2 and V3; correspondingly, Figure 6 in the example, the fourth sample nodes corresponding to the third sample node V2 include V3 and V4 (the walk path is V2 - V4 - V3).

[0117] Step A223: Based on the union of the third sample nodes and the fourth sample nodes, construct the other layers in the target graph except the output layer.

[0118] Optionally, taking the example in step A222 as an illustration, at this time the third sample nodes include V1 and V2, and the fourth sample nodes include V2, V3, and V4. Therefore, the union of the two is {V1, V2, V3, V4}. At this time, except for the output layer ( Figure 6 the layer 2 shown), the sample nodes in the other layers of the target graph include V1, V2, V3, and V4. Figure 6 The target graph in the example in includes 3 layers, that is, except for layer 2, the sample nodes included in layer 0 and layer 1 are the same.

[0119] In a feasible embodiment, in step A3, for the sample nodes in each adjacent layer in the sample node set, based on the connection objects of the first edges in the initial graph, perform edge sampling operations to obtain the second edges connected between the sample nodes in this adjacent layer in the target graph, including steps A301 - A302:

[0120] Step A301: For any two sample nodes located between adjacent layers in the sample node set, if there is a first edge connected between the two sample nodes in the initial graph, then connect the two sample nodes to form the third edge connecting the sample nodes between different layers in the target graph.

[0121] Optionally, as Figure 7 shown, during the node construction process, after obtaining the sample node set of the target graph, when constructing edges, for any two sample nodes located between adjacent layers in the sample node set, such as the sample nodes V1 and V2 between layer 2 and layer 1, the sample nodes V3 and V4 between layer 1 and layer 0, etc., determine whether there is a first edge connected between the two sample nodes from the initial graph. As Figure 7 shown in the leftmost schematic diagram, it can be determined that there are first edges connecting V1 and V2, and V3 and V4 in the current example. Then, connect the sample nodes V1 and V2, and connect the sample nodes V3 and V4 in the target graph to obtain the third edges in the target graph (it should be noted that in this step, all two sample nodes located between adjacent layers can be processed, and the processing of V1 and V2, V3 and V4 above is only an example).

[0122] Optionally, the third edge is a directed edge. For the adjacent layers of layer 1 and layer 2, the sample nodes included in layer 1 are the starting nodes, and the sample nodes included in layer 2 are the arrival nodes. Then, when constructing the third edge, start from the sample nodes in layer 1 and point to the sample nodes in layer 2.

[0123] Step A302: Starting from the output layer of the target graph, repeatedly perform the following first sampling operation for the third edge until the second edges connecting the sample nodes between all layers are obtained:

[0124] For the starting nodes corresponding to the third edge in the current adjacent layers, obtain, from the third edges in the next adjacent layer, a preset number of three sampled edges whose arrival nodes are the starting nodes as the second edges connecting the sample nodes in different layers in the next adjacent layer.

[0125] Optionally, in order to consider that for a graph that is itself dense, the number of edges included in the target graph InducedMPG (the included edges are called third edges) constructed through the above Step A301 is too large, which is not conducive to accelerating training. Therefore, on this basis, the resampling technique can be adopted to reduce the third edges included in the target graph to obtain the final target graph Final MPG (the included edges are called second edges).

[0126] During the resampling process, as Figure 7 shown in the first branch of, assuming that the number of three sampled edges is 2, for the starting nodes (sample nodes V1, V2, and V3 in layer 1) corresponding to the third edge in the current adjacent layers (layer 2 and layer 1), 3 edges whose arrival nodes are the starting nodes can be obtained from the third edges in the next adjacent layer (layer 1 and layer 0) (for example, for the sample node V1 in layer 1, 3 edges can be obtained from the 4 third edges connecting the sample nodes V1, V2, V3, and V4 in layer 0 to the sample node V1 in layer 1 between layer 1 and layer 0) as the second edges connecting the sample nodes in different layers in the next adjacent layer (layer 1 and layer 0) (for example, retain the three edges V1-V1, V2-V1, and V3-V1).

[0127] In a feasible embodiment, in Step A3, for the sample nodes in each adjacent layer in the sample node set, based on the connection objects of the first edges in the initial graph, an edge sampling operation is performed to obtain the second edges connecting the sample nodes in this adjacent layer in the target graph, including Step A311:

[0128] Step A311: Starting from the output layer of the target graph, repeatedly perform the following second sampling operation until the second edges connecting the sample nodes between all layers are obtained: For the sample nodes between the current adjacent layers, obtain a preset number of four sampled first edges connecting these sample nodes from the initial graph, and construct the second edges connecting the sample nodes in this adjacent layer.

[0129] Optionally, compared with the edge construction scheme provided in step A301 and step A302, in the process of edge construction in the embodiments of the present application, it is also possible to directly obtain the final target graph FinalMPG after resampling without constructing the induced graph Induced MPG, as shown in Figure 7 the schematic illustration of the dashed arrow in

[0130] Optionally, as shown in Figure 7 the second branch of Figure 7 assuming that the fourth sampling quantity is 3, starting from the output layer (layer 2), for the sample nodes between the current adjacent layers (layer 2 and layer 1), obtain 3 first edges connecting between these sample nodes from the initial graph, and directly construct the second edges connecting between the sample nodes of the adjacent layers in the sample node set of the target graph, such as the 3 second edges connecting the sample nodes V1, V2, and V3 in layer 1 to the sample node V1 in layer 2. Based on this edge construction logic, the second edges between layer 1 and layer 0 can be obtained. It can be understood that in the second branch shown in

[0131] since the sample nodes included in other layers except the output layer (layer 2) are the same, after performing the edge construction operation on two blocks, the second edges of the entire target graph can be obtained. In order to enable the model to learn richer information, the edge construction logic can be separately executed for different blocks (different adjacent layers) to obtain possibly the same or possibly different second edges for adjacent layers with the same sample node structure.

[0131] In a feasible embodiment, based on the target graph, model training is performed to obtain a target classification model for performing data classification tasks, including step B1 - step B2:

[0132] Step B1: Sample the sample nodes in the target graph to obtain the first seed nodes.

[0133] Step B2: When training the initial classification model based on the training data, in each training round, sample the sample nodes in the target graph to obtain the second seed nodes for the current training round, and determine the target training samples for the current training round based on the first seed nodes and the second seed nodes, and train the initial classification model through the target training samples to train a target classification model for performing data classification tasks.

[0134] In the embodiments of the present application, the first seed node (seed1) obtained in step B1 can be a fixed node during model training, that is, each training epoch of model training involves the first seed node. On this basis, in order to shorten the model training time while ensuring sufficient training, in each training epoch, sampling can be performed on the sample nodes in the target graph to obtain a second seed node (seed2), so as to determine the target training samples (including sample nodes and second edges) of the current training epoch based on the static first seed node and the dynamic second seed node. Optionally, the number of iterations of the model = epoch * the number of iterations in each epoch. The seed node can indicate the sample node used for loss calculation.

[0135] In the implementation of this application, a part of the nodes seed1 are fixed before training, and a part of the nodes seed2 are resampled in each epoch. Then, during the training process, seed1 and seed2 are combined as the seeds for model training in the current epoch, which can shorten the model training time while ensuring sufficient training.

[0136] Optionally, the training of the model can be based on the target graph or the initial graph, that is, the sampling of the first seed node and / or the second seed node can be performed in the graph structure data of the initial graph.

[0137] In a feasible embodiment, sampling the sample nodes in the target graph in step B1 and / or step B2 includes at least one of the following steps C1-step C2:

[0138] Step C1: Based on the first sampling rate, perform uniform random sampling on the sample nodes in the target graph.

[0139] Step C2: For each sample node in the target graph, determine the node degree of the sample node based on the second edge connected to the sample node; based on the second sampling rate and the node degree, perform sampling on the sample nodes in the target graph; the second sampling rate is proportional to the node degree.

[0140] Optionally, in the sampling of the first seed node and the second seed node, any sampling operation provided by the above steps C1 and C2 can be used. Among them, the first sampling rate and the second sampling rate can be preset values, empirical values or hyperparameters, such as parameters that can be adjusted according to the training program of the model, the time consumed for model training, etc.; the first sampling rate and the second sampling rate can be the same or different.

[0141] Among them, the node degree can indicate the number of edges connected to a certain sample node. For example Figure 7For the target graph (the rightmost message propagation graph) shown in the first branch, the node degree of the sample node V2 in layer 1 can be 5. Optionally, in the method of node sampling based on node degree (Degree-based Sampling, DBS), the node sampling probability is proportional to the node degree.

[0142] In a feasible embodiment, in step B2, sampling for the sample node in the target graph to obtain the second seed nodes for the current training round includes:

[0143] Setting the first seed nodes in the target graph to a non-sampling state, and sampling for the sample nodes in the target graph to obtain the second seed nodes for the current training round; the second seed nodes have no intersection with the first seed nodes.

[0144] Optionally, in order to improve the training efficiency and avoid sampling the same sample nodes as the first seed nodes during the sampling of sample nodes in the current training round, after sampling a fixed number of first seed nodes before training, the corresponding sample nodes can be set to a non-sampling state in the target graph, so that the same sample nodes as the first seed nodes cannot be obtained during the sampling of the second seed nodes.

[0145] In the embodiments of the present application, seeds are resampled in each epoch, and this method can be called ER (Epoch Resampling). Considering that if only some nodes are selected to calculate the loss, the information of other nodes may be lost. Therefore, seeds can be resampled at the beginning of each epoch. Compared with the method of sampling a batch of seeds before the entire training and keeping them fixed in all epochs, Epoch Resampling can utilize more node information, enabling the model to learn more sufficient information in each epoch and improving the training efficiency of the model.

[0146] The embodiments of the present application also propose a node sampling method based on DBS, namely PreFER (Prioritize Fixed EpochResampling): At the beginning of the entire training, a part of the nodes, seed1 (the first stage), is preferentially selected, and then seed2 (the second stage) is alternately selected within each epoch. Seed1 and seed2 are combined as the seeds for this epoch to calculate the loss. Optionally, sampling by node degree is used in the first stage, and uniform sampling is used in the second stage.

[0147] The effects achievable by the embodiments of the present application will be described below in combination with experimental data.

[0148] Specifically, the time efficiency and prediction accuracy of PreFER and other node classification algorithms were compared on four datasets.

[0149] (1) Training time

[0150] As shown in Table 1 below (s represents single label, m represents multi label, and FS (Feature Size) represents the initial feature dimension of nodes):

[0151] Table 1

[0152] Dataset Number of Nodes Number of Edges Number of Labels FS Ogbn - arxiv 166,243 1,166,243 47(s) 128 Reddit 232,965 11,606,919 41(s) 602 Yelp 716,847 6,977,410 100(m) 300 Amazon 1,598,960 132,169,734 107(m) 200

[0153] The baselines of node classification algorithms include GraphSAGE (Graph Sample and aggregate), GNS (Graph Neural Samlpe), ClusterGCN (Cluster-GraphConvolutional Network), and GraphSAINT (Graph Sampling Based InductiveLearning Method). In the embodiments of this application, the same GNN architecture is adopted, and the parameters of the model when it reaches the optimal accuracy on the validation set are saved. When the model is lower than the current optimal accuracy on the validation set for multiple times, it can be determined that the model converges, and the total training time at this time is obtained.

[0154] (2) Prediction accuracy

[0155] As shown in Table 2 below (1 represents the optimal result, and 2 represents the sub-optimal result):

[0156] Table 2

[0157]

[0158]

[0159] Combined with Table 2 above, from the overall data, the algorithm PreFER provided in the embodiments of this application can maintain a certain accuracy or only have partial losses. Relatively speaking, the algorithms ClusterGCN and GraphSAINT have more accuracy losses.

[0160] (3) Training time

[0161] As Figure 11As shown, the total training time when different algorithms converge on four datasets is presented (GSA, CLU, and GSAI represent GraphSAGE, ClusterGCN, and GraphSAINT respectively). Compared with GraphSAGE and GNS, PreFER provided by the embodiments of the present application can achieve a speed improvement of 1 - 1.5x. Although ClusterGCN and GraphSAINT have shorter training times, they have relatively more accuracy loss, and this accuracy is very difficult to make up for in actual operations.

[0162] Next, the data classification method in the embodiments of the present application will be described.

[0163] Specifically, the execution subject of the method provided by the embodiments of the present application can be a terminal or a server; the terminal (which can also be referred to as a device) can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart voice interaction device (such as a smart speaker), a wearable electronic device (such as a smart watch), a vehicle-mounted terminal, a smart home appliance (such as a smart TV), an AR / VR device, etc., but is not limited thereto. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers (such as a distributed cloud storage system), or a cloud server providing cloud computing and cloud storage services.

[0164] Specifically, as Figure 2 shown, the data classification method includes step S201: In response to a data classification operation, classify the data corresponding to the operation through a pre-trained target classification model.

[0165] Among them, the target classification model is trained by the model training method based on sampling provided in the above embodiments.

[0166] Exemplarily, considering that a new interaction function is launched on a certain instant messaging platform, the update of this function has a great impact on the object interaction data on the instant messaging platform. At this time, the model can be retrained based on the latest data stored on the platform to make the division of the characteristics, preferences, etc. of certain objects by the model more accurate. Among them, the data of the instant messaging platform can first use the related technology GNN to construct graph-structured data to obtain an initial graph. On this basis, considering that the large amount of data may lead to a long training time of the model, the above method can be used to construct a lightweight and powerful target graph (message propagation graph) for model training based on the initial graph. On this basis, when performing model training, the training operations provided in step B1 - step B2 can be executed based on the initial graph or the target graph to train the model and obtain the final target classification model.

[0167] To better illustrate the method provided by the embodiments of the present application, a feasible application example is given for illustrative purposes.

[0168] Application Example 1 (Node Construction Scheme 1 in the Target Graph)

[0169] Adopt Algorithm 1: Layer-wise Construction

[0170] Obtain parameters: the number of GNN layers λ, the minibatch node L λ-1 , and adopt the number of fanouts s1;

[0171] Limit condition (output): the node set of each layer is {L1|l = 0,..., λ - 1};

[0172] At the setting (l = 0,..., λ - 2), determine the sample nodes corresponding to each layer i in the target graph.

[0173] Optionally, assume that the node set Li of the i-th layer has been sampled. For each sample node in Li, sample s1 neighbor nodes. Concatenate all the neighbor nodes and the sample nodes in Li to form the node set Li-1 of the (i - 1)-th layer. Exemplarily, as in Figure 5 , since fanout = 1, each of the sample nodes V1 and V2 in the second layer samples 1 neighbor node (i.e., nodes V2 and V3 are obtained). Then the nodes V2 and V3 and the nodes V1 and V2 are concatenated to form {1, 2, 3}, which constitutes the node set of the first layer. Compared with the prior art, in one aspect, setting a relatively small value for fanout, which can be 1 or 2, can effectively achieve the lightweight of the graph; on the other hand, during the node construction stage, the edges between layers do not need to be saved, while in the prior art, due to the need to know the propagation relationship between nodes, the edge information needs to be saved.

[0174] Application Example 2 (Node Construction Scheme 2 in the Target Graph)

[0175] Adopt Algorithm 2: Global Construction

[0176] Obtain parameters: the number of GNN layers λ, the minibatch node L λ-1 ;

[0177] Limit condition (output): the node set of each layer is {L1|l = 0,..., λ - 1};

[0178] At the setting (l = 0,..., λ - 2) and L all = V λ-1Under this condition, determine the sample nodes corresponding to each layer i in the target graph.

[0179] Optionally, taking each sample node in the minibatch as an initial node, perform random walks for several times. Concatenate all the encountered nodes and the nodes in the minibatch to form a node set Lall, and set the node sets of all layers (except the output layer) in the MPG to Lall. For example, in Figure 6 , first perform a random walk on each of the V1 and V2 nodes, with a length of 3. Then concatenate all the encountered nodes, i.e., {1, 2, 3, 4}. Finally, set the node sets of each layer (except the output layer) of the MPG to {1, 2, 3, 4}. The node construction method provided in the embodiments of the present application can completely avoid the neighbor explosion problem.

[0180] Application Example 3 (Edge Construction Scheme in the Target Graph)

[0181] Adopt Algorithm 3: Resampling technique

[0182] Obtain parameters: the number of GNN layers λ, the node set of each layer is {L1|l = 0,..., λ - 1}, and the adopted quantity is fanouts2;

[0183] Limit condition (output): the final target graph Final MPG S;

[0184] Under the setting (l = 1,..., λ - 1), and (l = 0,..., λ - 2), determine the second edges between the sample nodes connected in different layers in the target graph.

[0185] Optionally, add second edges between the node sets of adjacent layers. First, according to the structure of the initial graph (original graph), supplement all the edges between adjacent layers, that is, if there are a node u in the i-th layer and a node v in the i + 1-th layer, and there is an edge (u, v) in the original graph, then add it between the i-th and i + 1-th layers. The MPG constructed in this way is called Induced MPG, as in Figure 7As shown. However, considering that for a graph that is dense in itself, the Induced MPG contains too many edges, which is not conducive to accelerating training. Therefore, this application can use the Resample technique to process it to obtain the final Final MPG. During the resampling process, Fi and Ei are used to represent the edges in Induced MPG Blocki and Final MPG Blocki respectively. Assume that Ei+1 has been sampled, that is, the edges connecting the i-th layer and the (i + 1)-th layer. Then, the departure nodes Li corresponding to these edges Ei+1 are obtained. For each node v in Li, s2 edges with v as the arrival node are selected from Fi, and these edges are concatenated to form Ei. Among them, the value of s2 can be set to be much larger than the value of s1 during the node set construction process, such as it can be set to about 5 - 10.

[0186] Optionally, during a sampling process, it is not necessary to explicitly construct the Induced MPG. Instead, the same result as the above example can be achieved by sampling the intersection of the neighbor set of the original graph and the node set of the next layer.

[0187] In the embodiments of this application, the above algorithm that combines layer-based node construction and edge construction can be called RALC (Layer-wise Construction + Resample), and the algorithm that combines global graph-based node construction and edge construction can be called RAGC (Global Construction + Resample).

[0188] The following shows a comparison of the time efficiency and prediction accuracy of RALC, RAGC and other node classification algorithms on six datasets. The detailed information of the datasets is shown in Table 3 below:

[0189] Table 3

[0190] Dataset Number of Nodes Number of Edges Number of Labels FS Flicker 89,250 989,006 7(s) 500 Ogbn - arxiv 169,343 2,484,941 40(s) 128 Reddit 232,965 23,446,803 41(s) 602 Yelp 716,847 13,954,819 100(m) 300 Amazon 1,569,960 264,339,468 107(m) 200 Ogbn - products 2,449,029 126,167,053 47(s) 100

[0191] The node classification algorithm baselines include GraphSAGE, GNS, LABOR, GraphSAINT. For fair comparison, the same GNN architecture and hyperparameter settings are applied to all models: for the datasets Flickr, Ogbn-arxiv, Reddit, 4-layer SAGEConv is used with a hidden layer dimension of 256; for other datasets, 4-layer SAGEConv is used with a hidden layer dimension of 256; for GraphSAGE, LABOR and GNS, 2 neighbors are sampled for each node; for RALC and RAGC, s1 and s2 are set to 1 and 5 respectively. To reduce the impact of batch size on training time, the batch size is fixed at 2048.

[0192] Regarding the prediction accuracy, such asFigure 8 Table 4 shown in Figure 9 Table 5 shown in, which shows the F1-Score, training time of different algorithms on six datasets, and the improvement ratio compared with GraphSAGE, where (1) represents the optimal result and (2) represents the sub-optimal result. RALC and RAGC have a speed improvement of 1.5x - 3.5x compared with GraphSAGE, and the accuracy is equal to or exceeds the optimal result.

[0193] Regarding the convergence speed, as Figure 10 shown, it shows the graph of the validation set accuracy of various algorithms changing with time on six datasets. The closer the curve is to the upper left corner, the faster the convergence speed and the higher the accuracy. Obviously, RALC and RAGC are superior to other algorithms in terms of convergence speed.

[0194] Application Example 4 (Model Training Scheme Based on Seed Node Sampling)

[0195] Adopt Algorithm 4: PreFER

[0196] Input: Training graph G(V, E) - can be the initial graph or the target graph; GNN model g, loss function f, prediction function h

[0197] Parameters: The first-stage sampling rate α and the second-stage sampling rate β

[0198] Output: The model accuracy of the node classification task

[0199] Under the condition of determining the node degree d(v) of each node v based on G, and , in the first stage, the first seed node Train1 = Sample(G, p, n1) can be obtained, where the parameter p can refer to the normalization of d(v) under traversing all sample nodes. On this basis, the basis for subsequent second seed node sampling is G~ = G - Train1 (that is, other sample nodes except the first seed node). Then, in the second stage, the second seed node Train2 = Sample(G~, n2) can be obtained in each epoch. Correspondingly, the target training sample is SampleTrain = Concat(Train1, Train2), its mini-batch sampling minibatchs is to construct minibatch(G, SampleTrain), and the blocks Blocks can be to construct Block(G, minibatchs). The loss function Loss = ∑ b∈Blocks f(g(b)) / len(Blocks), and use the gradient descent method to update the parameters and save the model parameter W.

[0200] Optionally, in the testing phase, the updated parameter W is loaded into the model g, such that the accuracy is h(g, test nodes).

[0201] The model training method provided by the embodiments of this application can facilitate the rapid convergence and accuracy improvement of training.

[0202] It should be noted that, in the optional embodiments of this application, for the data involved (such as training data, initial graph, target graph, and other related data), when the above embodiments of this application are applied to specific products or technologies, permission or consent from the user is required, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions. That is to say, if the embodiments of this application involve data related to the user, these data need to be obtained with the authorization and consent of the user and comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0203] The embodiments of this application provide a sampling-based model training device, as Figure 12 shown. The sampling-based model training device 100 may include: an acquisition module 101.

[0204] Among them, the acquisition module 101 is configured to acquire a target graph corresponding to training data, and perform model training based on the target graph to obtain a target classification model for performing a data classification task;

[0205] Among them, the device 100 further includes a construction module 102, which is configured to perform the following construction operations to obtain the target graph:

[0206] Acquire an initial graph corresponding to training data; the initial graph includes sample nodes corresponding to each sample in the training data and first edges connected between the sample nodes;

[0207] Based on the connection relationship of the sample nodes in the initial graph, perform a node sampling operation to obtain a sample node set including at least two layers in the target graph;

[0208] For each adjacent layer of sample nodes in the sample node set, based on the connection objects of the first edges in the initial graph, perform an edge sampling operation to obtain second edges connected between the sample nodes in the adjacent layer in the target graph.

[0209] The embodiments of this application provide a data classification device, as Figure 13 shown. The data classification device 200 may include: a classification module 201.

[0210] Among them, the classification module 201 is configured to, in response to a data classification operation, classify the data corresponding to the operation through a pre-trained target classification model;

[0211] Among them, the target classification model is trained by the model training method based on sampling provided in the above embodiments.

[0212] The device according to the embodiment of the present application can execute the method provided in the embodiment of the present application, and the implementation principle is similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed function description of each module of the device, reference can be specifically made to the description in the corresponding method shown above, and details are not described herein again.

[0213] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0214] The modules involved in the embodiments of the present application can be implemented by software. Among them, the name of the module does not constitute a limitation on the module itself in some cases. For example, the acquisition module can also be described as "the module for acquiring the target graph corresponding to the training data", "the first module", etc.

[0215] An electronic device is provided in the embodiments of the present application, including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the model training method and the data classification method based on sampling. Compared with the related art, it can achieve:

[0216] First aspect, an embodiment of the present application provides a sampling-based model training method, which can construct a lightweight graph for model training to improve the training speed while ensuring a certain accuracy. Specifically, the target graph corresponding to the training data used in the training can be obtained by performing a sampling operation on the basis of the initial graph, where the initial graph can include sample nodes corresponding to each sample in the training data and first edges connecting the sample nodes. On this basis, the present application can perform a node sampling operation based on the connection relationship of the sample nodes in the initial graph to obtain a sample node set including at least two layers in the target graph, completing the construction of the nodes. That is, the present application does not need to save the edges connecting the sample nodes during the construction of the nodes, and can also effectively avoid the neighbor explosion problem. Subsequently, for each adjacent layer of sample nodes in the sample nodes, an edge sampling operation can be performed based on the connection objects of the first edges in the initial graph to obtain second edges connecting the sample nodes in the adjacent layer in the target graph. That is, in order to avoid the problem of a large amount of training data caused by too many edges in a dense graph during the construction of the edges, the present application can reduce the required training data volume through the edge sampling operation to effectively accelerate the training.

[0217] Second aspect, an embodiment of the present application provides a data classification method. Specifically, when responding to a data classification operation, a pre-trained target classification model can be used to classify the data corresponding to the operation. Among them, the target classification model can be trained by using the sampling-based model training method provided in the first aspect, and its training process can improve the training speed while ensuring a certain accuracy, which is beneficial to ensuring the accuracy of the classification result and the timeliness of the response of data classification by the model.

[0218] In an alternative embodiment, an electronic device is provided, as Figure 14 shown. Figure 14 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.

[0219] The processor 4001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 4001 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0220] The bus 4002 can include a path for transmitting information between the above components. The bus 4002 can be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 14 only a thick line is used to represent it herein, but it does not mean that there is only one bus or one type of bus.

[0221] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.

[0222] The memory 4003 is used to store the computer program for implementing the embodiments of the present application, and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.

[0223] Among them, the electronic device includes but is not limited to: a terminal, a server.

[0224] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0225] The embodiments of the present application further provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0226] In the description and claims of the present application and the above drawings, terms such as "first", "second", "third", "fourth", "1", "2", etc. (if any) are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than the illustrated or textually described order.

[0227] It should be understood that although the flowchart of the embodiments of the present application indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless otherwise clearly stated in this application, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage of these sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0228] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, other similar implementation means based on the technical idea of the present application also belong to the protection scope of the embodiments of the present application.

Claims

1. A sampling-based model training method, characterized in that, Including: Obtaining a target graph corresponding to training data, and based on the target graph, performing model training to obtain a target classification model for performing a data classification task; Wherein, the target graph is obtained by performing the following construction operations: Obtaining an initial graph corresponding to training data; the initial graph includes sample nodes corresponding to each sample in the training data and first edges connected between the sample nodes; Performing a node sampling operation based on the connection relationship of the sample nodes in the initial graph to obtain a sample node set including at least two layers in the target graph; For the sample nodes in each adjacent layer in the sample node set, performing an edge sampling operation based on the connection objects of the first edges in the initial graph to obtain second edges connected between the sample nodes in the adjacent layer in the target graph.

2. The method according to claim 1, wherein The performing a node sampling operation based on the connection relationship of the sample nodes in the initial graph to obtain a sample node set including at least two layers in the target graph includes at least one of the following: For the construction of each layer in the target graph, based on the sample nodes in the previous layer, obtaining at least one sample node adjacent to the sample node from the initial graph to form the sample nodes in the current layer; Based on the initial nodes sampled from the initial graph and the connection relationship of the included sample nodes, performing node sampling from the initial graph to obtain the sample nodes in each layer of the target graph.

3. The method according to claim 2, characterized in that, The for the construction of each layer in the target graph, based on the sample nodes in the previous layer, obtaining at least one sample node adjacent to the sample node from the initial graph to form the sample nodes in the current layer includes: Repeatedly performing the following determination operation of the current layer nodes until the constructed number of layers reaches a preset number of layers: For each first sample node in the previous layer, based on the connection relationship of the sample nodes in the initial graph, obtaining a preset first sampling number of second sample nodes adjacent to the first sample node; Taking the union of the first sample nodes and the second sample nodes as the sample nodes in the current layer; Wherein, when the previous layer is the first layer for constructing the target graph, the sample nodes in this layer include the sample nodes sampled from the initial graph based on a mini-batch sampling method.

4. The method according to claim 2, wherein The based on the initial nodes sampled from the initial graph and the connection relationship of the included sample nodes, performing node sampling from the initial graph to obtain the sample nodes in each layer of the target graph includes: Sampling third sample nodes from the initial graph based on a mini-batch sampling method, and constructing the output layer of the target graph based on the third sample nodes; For each third sample node, based on the connection relationship of the sample nodes in the initial graph, randomly obtaining a preset second sampling number of fourth sample nodes corresponding to the third sample node; Based on the union of the third sample nodes and the fourth sample nodes, constructing other layers in the target graph except the output layer.

5. The method according to claim 1, characterized in that, The for the sample nodes in each adjacent layer in the sample node set, performing an edge sampling operation based on the connection objects of the first edges in the initial graph to obtain second edges connected between the sample nodes in the adjacent layer in the target graph includes: For any two sample nodes in the sample node set that are located between adjacent layers, if there is a first edge connecting the two sample nodes in the initial graph, then connect the two sample nodes to form a third edge connecting the sample nodes between different layers in the target graph; Starting from the output layer of the target graph, repeatedly perform the following first sampling operation on the third edge until the second edges connecting the sample nodes between all layers are obtained: For the starting node corresponding to the third edge in the current adjacent layers, obtain a preset number of third sampled edges in the next adjacent layer whose arrival node is the starting node as the second edges connecting the sample nodes between different layers in the next adjacent layer.

6. The method according to claim 1, wherein The operation of sampling edges for each adjacent layer of sample nodes in the sample node set based on the connection objects of the first edges in the initial graph to obtain the second edges connecting the sample nodes in this adjacent layer in the target graph includes: Starting from the output layer of the target graph, repeatedly perform the following second sampling operation until the second edges connecting the sample nodes between all layers are obtained: For the sample nodes between the current adjacent layers, obtain a preset number of fourth sampled first edges connecting the sample nodes in the initial graph and construct the second edges connecting the sample nodes in this adjacent layer.

7. The method according to claim 1, characterized in that The training the target classification model for performing the data classification task based on the target graph includes: Sampling the sample nodes in the target graph to obtain the first seed nodes; When training the initial classification model based on the training data, in each training round, sampling the sample nodes in the target graph to obtain the second seed nodes for the current training round, and determining the target training samples for the current training round based on the first seed nodes and the second seed nodes, and training the initial classification model through the target training samples to train the target classification model for performing the data classification task.

8. The method according to claim 7, characterized in that The sampling the sample nodes in the target graph includes at least one of the following: Based on the first sampling rate, uniformly and randomly sampling the sample nodes in the target graph; For each sample node in the target graph, determining the node degree of the sample node based on the second edges connected to the sample node; Based on the second sampling rate and the node degree, sampling the sample nodes in the target graph; The second sampling rate is proportional to the node degree.

9. The method according to claim 7, wherein The sampling the sample nodes in the target graph to obtain the second seed nodes for the current training round includes: Setting the first seed nodes in the target graph to a non-sampling state and sampling the sample nodes in the target graph to obtain the second seed nodes for the current training round; the second seed nodes have no intersection with the first seed nodes.

10. A data classification method, characterized in that, including: In response to the data classification operation, classifying the data corresponding to the operation through the pre-trained target classification model; Wherein, the target classification model is trained by the method described in any one of claims 1-9.

11. A sampling-based model training device, characterized in that including: An acquisition module, configured to acquire a target graph corresponding to training data, so as to perform model training based on the target graph to obtain a target classification model for performing a data classification task; Wherein, the apparatus further includes a construction module, configured to perform the following construction operations to obtain the target graph: Acquire an initial graph corresponding to the training data; the initial graph includes sample nodes corresponding to each sample in the training data and first edges connected between the sample nodes; Perform a node sampling operation based on the connection relationship of the sample nodes in the initial graph to obtain a sample node set including at least two layers in the target graph; For the sample nodes in each adjacent layer in the sample node set, perform an edge sampling operation based on the connection objects of the first edges in the initial graph to obtain second edges connected between the sample nodes in the adjacent layer in the target graph.

12. A data classification device, characterized in that, Includes: A classification module, configured to, in response to a data classification operation, classify data corresponding to the operation through a pre-trained target classification model; Wherein, the target classification model is trained by the method according to any one of claims 1-9.

13. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the method according to any one of claims 1-10.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-10.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-10.