A method for identifying a target group and a computer-readable storage medium
By constructing a dynamic graph neural network model, combining time and spatial attributes, the problem of insufficient accuracy and flexibility in target group identification in the existing technology is solved, and efficient target group dynamic tracking and automatic clustering are achieved.
Patent Information
- Application Number
- CN202210273388.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-03-18
AI Technical Summary
The prior art lacks a method of automatically clustering users while applying graph neural networks on dynamic graphs, resulting in low accuracy and insufficient flexibility in target group identification.
The dynamic graph neural network model is used to model social accounts, and by constructing a dynamic graph neural network model to classify the target group, and combining time and spatial attributes, the community discovery problems on the dynamic graph are realized.
It improves the efficiency and accuracy of dynamic tracking of target groups, and can automatically cluster without pre-acquisition of user tags, improving the accuracy and efficiency of identification.
Smart Images

Figure CN114626473B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target group recognition, and particularly to a method for recognizing a target group and a computer-readable storage medium. Background Art
[0002] There has been a great deal of interest in the potential of graph neural networks (GNNs) in graph learning; with the increasing development of the network economy and network media, the research on target groups on social networks has become increasingly important. The dynamic tracking method of target groups in two dimensions of time and space can be studied. The time dimension is reflected in the order of the information published by users, while the space dimension is reflected in the change of user positions. In the dynamic tracking of target groups, traditional methods often do not consider spatio-temporal types. Only the expression of users is modeled through the graph convolutional neural network GCN, and then the distance function is used to obtain the target group.
[0003] In view of the multi-source and spatio-temporal characteristics of social media data, a method for dynamically tracking target groups based on streaming data is proposed to capture the temporal changes of target groups on social media or in the geospatial. At the individual level, based on historical data (including GPS data, timestamps, text messages, behaviors such as likes / forwards, etc.), the social behaviors of groups and the text information attached to the social behaviors are mainly analyzed, the changes in group positions are captured and displayed, and the tracking of target groups is realized. Specifically, how to quickly obtain an effective expression of the target group has become an important research topic.
[0004] Existing methods can be roughly divided into two categories: statistical models and classification models. A major disadvantage of statistical models is poor flexibility. Generally speaking, statistical models need to model the generation process of dynamic graphs, but the actual situation is often complex. Therefore, statistical models based on the generation process modeling often show low accuracy; relatively speaking, the flexibility of classification models is higher. They can embed dynamic graph neural network methods. Since they do not model the generation process of graphs, their accuracy is higher than that of statistical models. However, the biggest problem with classification models is that there is no community constraint on the classification of points, that is, it does not force the points within the discovered community to be closely connected while the connections outside the community are sparse.
[0005] In summary, there is a lack of a method in the prior art to automatically cluster users while applying graph neural networks on dynamic graphs.
[0006] The disclosure of the above background art content is only used to assist in understanding the concept and technical solution of the present invention, and it does not necessarily belong to the prior art of this patent application. Without clear evidence that the above content was publicly available on the filing date of this patent application, the above background art should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention
[0007] The present invention provides a target group identification method and a computer-readable storage medium to solve existing problems.
[0008] To solve the above problems, the technical solution adopted by the present invention is as follows:
[0009] A target group identification method includes the following steps: S1: Model the social accounts of users and obtain whether there is information flow in the social accounts within a time window; S2: Construct a dynamic graph neural network model to classify the social accounts to obtain different clusters; S3: Identify the target group from the clusters.
[0010] Preferably, the social account is modeled as model V = {v1, v2,..., v n}, where v i is the social account; within a time window, within a time window, the graph structure A t of a single time slice is used to represent whether there is information flow in the social account within the time window: indicates that there is information flowing from v i to v j ; indicates that there is no information flowing from v i to v j ; Whether there is information flow in the social account within multiple time windows is represented as a set {A 1 , A 2 ,..., A m}.
[0011] Preferably, according to the set {A 1 , A 2 ,..., A m}, the model V = {v1, v2,..., v n} is divided into K non-overlapping subsets where, i≠j, and K is a positive integer.
[0012] Preferably, the non-overlapping subsets are modeled as a community discovery problem on a dynamic graph; a machine learning model is used to learn the dynamic graph pooling scheme where, C ij represents the probability that v i belongs to the jth community, is an n*K matrix.
[0013] Preferably, the dynamic graph neural network model is constructed to process the dynamic graph; the dynamic graph is input into the dynamic graph neural network model to obtain dynamic graph pooling.
[0014] Preferably, the dynamic graph neural network model includes a graph neural network model; inputting one piece of the time slice data into the graph neural network model outputs a high-dimensional representation, specifically as follows:
[0015] H t = GNN(A t )
[0016] where H is a set of individual user representations in the current time graph.
[0017] Use an additive network to accumulate all the representations of the dynamic graph, specifically as follows:
[0018] C = ReGNN(A 1 , A 2 ,..., A m )
[0019] where C is the output of the dynamic neural network.
[0020] Preferably, the graph neural network model is a graph convolutional neural network or a graph isomorphism network model.
[0021] Preferably, the optimization function of the dynamic graph neural network model is:
[0022]
[0023] where D is the degree matrix. Given the adjacency matrix A, D is a diagonal matrix, L is the Laplacian matrix and L = D - A;
[0024] The meanings of K, k, and T are respectively.
[0025] Preferably, there are multiple time slices in the dynamic graph, that is, {A 1 , A 2 ,..., A m}, that is, there are also multiple corresponding degree matrices {D 1 , D 2 ,..., D m} and Laplacian matrices {L 1 , L 2 ,..., L m}. Take the means of the degree matrix and the Laplacian matrix in multiple time windows as the optimized proxy.
[0026] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of any of the preceding methods are implemented.
[0027] The beneficial effects of the present invention are as follows: A target group recognition method and a computer-readable storage medium are provided. By modeling the problem of target group recognition with both spatial and temporal attributes as a community discovery problem of a dynamic graph, a dynamic graph neural network model is constructed. The dynamic graph neural network is used to learn better node representations for the dynamically changing graph, combining temporal data and spatial data organically to improve the efficiency and accuracy of dynamic tracking of target groups. Description of the Drawings
[0028] Figure 1 It is a schematic diagram of a target group recognition method in an embodiment of the present invention.
[0029] Figure 2 It is a schematic diagram of the problem of dynamic graph pooling in an embodiment of the present invention.
[0030] Figure 3 It is a schematic diagram of a dynamic graph in an embodiment of the present invention. Detailed Embodiments
[0031] In order to make the technical problems, technical solutions and beneficial effects to be solved by the embodiments of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0032] It should be noted that when an element is referred to as being "fixed to" or "disposed on" another element, it can be directly on the other element or indirectly on the other element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element. In addition, the connection can be for a fixing function or for a circuit connection function.
[0033] It should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the embodiments of the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation of the present invention.
[0034] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0035] The present invention relates to multiple network models, specifically as follows:
[0036] Dynamic graph neural network, Recurrent GNN network;
[0037] Recurrent neural network model (RNN, Recurrent Neural Network);
[0038] Graph neural network model (GNN, Graph Neural Network);
[0039] Graph convolutional networks (GCN);
[0040] Graph isomorphism network model (GIN, Graph Isomorphism Network);
[0041] Diffusion Convolutional Recurrent Neural Network (DCRNN).
[0042] As Figure 1 shown, the present invention provides a target group identification method, including the following steps:
[0043] S1: Model the social accounts of users and simultaneously obtain whether there is information flow in the social accounts within the time window;
[0044] S2: Construct a dynamic graph neural network model to classify the social accounts to obtain different clusters;
[0045] S3: Identify the target group from the clusters.
[0046] In the prior art, when modeling social accounts, the time factor is not considered, and only static graphs are established. The present invention comprehensively considers the spatial and temporal attributes of group identification, uses a dynamic graph neural network to learn better point expressions for dynamic graphs with structural changes, and organically combines time data and spatial data to improve the efficiency and accuracy of dynamic tracking of target groups.
[0047] The present invention does not require pre-acquiring user labels and uses a dynamic graph neural network to automatically cluster users. The problem of group identification with both spatial and temporal attributes in the present invention can be modeled as a community discovery problem on a dynamic graph, fusing the process of obtaining points and the process of clustering, and realizing group identification by obtaining clustering from the point expressions. The following will elaborate specifically.
[0048] In a specific embodiment, the user's social account is modeled as model V = {v1, v2,..., v n}, where v i is the social account;
[0049] Within a time window, within a time window, the graph structure A t of a single time slice is used to represent whether there is information flow in the social account within the time window:
[0050] Then it means that within the time window, there is information flowing from v i to v j ;
[0051] Then it means that within the time window, there is no information flowing from v i to v j ;
[0052] Then whether there is information flow in the social account within multiple time windows is represented as a set {A 1 , A 2 ,..., A m}.
[0053] In a specific embodiment, the information flow can be information forwarding on Weibo or WeChat information forwarding, etc.
[0054] The community discovery problem of the dynamic graph can be modeled as, given the set {A 1 , A 2 ,..., A m} and model V, it is necessary to find a partition of model y into K non - overlapping subsets And given i≠j; that is, according to the set {A 1 , A 2 ,..., A m} partition the model V = {v1, v2,..., v n} into K non - overlapping subsets where, i≠j, and K is a positive integer.
[0055] Modeling the partition into non - overlapping subsets as a community discovery problem on a dynamic graph; learning the dynamic graph pooling scheme with a machine learning model
[0056] where, C ij represents the probability that v i belongs to the j - th community, is an n*K matrix.
[0057] As shown Figure 2 in the figure, the left side of the figure is the dynamic graph input, where the nodes are {a, b, c, d, e, f}, and the edges can be information forwarding within a period of time, so the edges are dynamically changing. The right side of the figure is the learning objective, which is to learn a dynamic graph pooling scheme. Since the nodes {a, e} are assigned to a community, they are pooled into an abstract node A. Similarly, the nodes {b, c, d} are pooled into an abstract node B, and the node {f} is pooled into an abstract node C.
[0058] A research topic highly relevant to the present invention is the node classification problem on dynamic graphs. However, the community discovery problem on dynamic graphs is significantly different from node classification. The difference is that the community discovery problem on dynamic graphs requires that the connections within each community be tight, while the connections between different communities be as sparse as possible. Based on this characteristic, the community discovery problem on dynamic graphs is generally considered more difficult than the classification problem on dynamic graphs.
[0059] The following introduces the dynamic graph neural network Recurrent GNN (ReGNN). ReGNN utilizes the characteristics of both the RNN and GNN network models to process dynamic graphs. Specifically, a dynamic graph neural network model is constructed to process the dynamic graph; the dynamic graph is input into the dynamic graph neural network model to obtain dynamic graph pooling.
[0060] The dynamic graph neural network model includes a graph neural network model; when a time slice of data is input into the graph neural network model, a high-dimensional representation is output, specifically as follows:
[0061] H t = GNN(A t )
[0062] where H is the set of individual user representations in the current time graph, and A t is the graph structure of a single time slice, that is, the adjacency matrix.
[0063] An additive network is used to accumulate all the representations of the dynamic graph, specifically as follows:
[0064] C = ReGNN(A 1 , A 2 ,..., A m )
[0065] where C is the output of the dynamic neural network.
[0066] In a specific embodiment, the graph neural network model is a graph convolutional neural network or a graph isomorphism network model.
[0067] The present invention models the problem of discovering abnormal groups in online public opinion as a community discovery problem on a dynamic graph, and proposes a dynamic graph pooling method to handle this problem, improving the accuracy and efficiency of identifying target groups.
[0068] Another core component of the dynamic graph neural network is the optimization objective. Since the problem of discovering communities on a dynamic graph is generally considered an unsupervised problem, the present invention needs to construct a reasonable optimization objective. In the present invention, the classical graph balanced cut normalized cut is relaxed so that it can be used as the optimization function of the dynamic graph neural network. Specifically, the formula for the classical graph balanced cut is as follows:
[0069]
[0070] where vol(V k ) represents the sum of all edges belonging to community V k , represents the complement of the vertex set V k , represents the sum of the edges connecting community V k and other vertices. Figuratively speaking, it is hoped that fewer edges are "cut off" so that the dynamic graph vertex set is divided into several independent parts. By introducing the output C of the dynamic graph neural network and the relaxation condition, that is, C can be any real number, the above graph balanced cut formula is transformed into the optimization function of the dynamic graph neural network model as:
[0071]
[0072] where D is the degree matrix. Given the adjacency matrix A, D is a diagonal matrix, L is the Laplacian matrix and L = D - A; k is a variable, and T is the transpose algorithm.
[0073] In the dynamic graph, there are multiple time slices, that is, {A 1 , A 2 ,..., A m}, that is, there are also multiple corresponding degree matrices {D 1 , D 2 ,..., D m} and Laplacian matrices {L 1 , L 2 ,..., L m}. The mean values of the degree matrix and the Laplacian matrix in multiple time windows are taken as the optimization proxy.
[0074] The present invention proposes an optimization objective for dynamic graph pooling, which relaxes the classical graph balanced cut; through the above optimization objective, the present invention unifies the model and the optimization objective, and usually this model is the dynamic graph pooling method.
[0075] In another embodiment of the present invention, other dynamic graph neural networks such as DCRNN are used to model user expressions, and then a distance function is used to characterize the target group, such as the K-nearest neighbor method (KNN); moreover, the pooling scheme of the present invention can be incorporated into the entire neural network, resulting in better effects.
[0076] As Figure 3 shown, taking the Weibo data in the first quarter of 2015 in Yangpu District, Shanghai as an example, a dynamic graph is constructed. 40,935 users are obtained through data preprocessing. Let K = 32, that is, the 40,935 users are divided into 32 communities, and the true labels of the communities are extracted from text keywords. The method of the present invention is compared with the GCN+KNN method, that is, first, user expressions are obtained through GCN, and then the gang clustering results are obtained through KNN, as shown in Table 1.
[0077] Table 1 Recognition target accuracy
[0078] Method Name Accuracy Rate GCN+KNN 19.1% The Method of the Present Invention 34.7%
[0079] As can be seen from the above table, the accuracy of the present invention is significantly higher than that of the existing GCN+KNN method. The method of the present invention can well combine the advantages of the current dynamic graph neural network to obtain a more accurate characterization of communities or events; in addition, the present invention integrates the process of obtaining points and the clustering process, and does not require pre-acquiring user labels, so it is also faster in terms of efficiency.
[0080] The embodiment of the present application further provides a control device, including a processor and a storage medium for storing a computer program; wherein, the processor is configured to execute at least the method as described above when executing the computer program.
[0081] The embodiment of the present application further provides a storage medium for storing a computer program, and when the computer program is executed, it executes at least the method as described above.
[0082] The embodiment of the present application further provides a processor, and the processor executes a computer program and at least executes the method as described above.
[0083] The storage medium can be implemented by any type of volatile or non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM), a synchronous static random access memory (SSRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a sync link dynamic random access memory (SLDRAM), a direct rambus random access memory (DRRAM). The storage medium described in the embodiments of the present invention is intended to include but not limited to these and any other suitable types of memory.
[0084] In several embodiments provided by the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.
[0085] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0086] In addition, in each embodiment of the present invention, the various functional units can all be integrated in one processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit; the above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0087] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media that can store program codes such as removable storage devices, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs.
[0088] Alternatively, if the above integrated units of the present invention are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as removable storage devices, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0089] The methods disclosed in several method embodiments provided by this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0090] The features disclosed in several product embodiments provided by this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0091] The features disclosed in several method or device embodiments provided by this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0092] The above content is a further detailed description of the present invention in combination with specific preferred implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those skilled in the technical field to which the present invention pertains, without departing from the concept of the present invention, several equivalent substitutions or obvious variations can be made, and as long as the performance or use is the same, they should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for identifying a target group, characterized in that, It includes the following steps: S1: Model the user's social account and obtain whether there is information flow in the social account within the time window; S2: Construct a dynamic graph neural network model to classify the social account to obtain different clusters; S3: Identify the target group from the clusters; where: Model the social account as model V = {v1, v2, …, v n}, where v i is the social account; Within a said time window, use the graph structure A of a single time slice t to represent whether there is information flow of the social account within the said time window: It means that there is information flowing from v i to v j ; It means that no information flows from v i to v j ; Whether there is information flow in the social account within multiple time windows is represented as a set {A 1 , A 2 ,..., A m}; According to the set {A 1 , A 2 ,..., A m}, the model V = {v1, v2,..., v n} is divided into K non - overlapping subsets wherein, i≠j, and K is a positive integer; Modeling the segmentation into disjoint subsets as a community discovery problem on a dynamic graph; learning a dynamic graph pooling scheme with a machine learning model Among them, C ij represents the probability that v i belongs to the j-th community, is an n*K matrix; Construct the dynamic graph neural network model to process the dynamic graph; input the dynamic graph into the dynamic graph neural network model to obtain dynamic graph pooling; The dynamic graph neural network model includes a graph neural network model; Input a piece of the time slice data into the graph neural network model to output a high-dimensional representation, specifically as follows: H t = GNN(A t ) Wherein, H is the set of expressions of a single user in the current time graph; Use an additive network to accumulate all the expressions of the dynamic graph, specifically as follows: c = ReGNN(A 1 , A 2 ,..., A m ) Wherein, C is the output of the dynamic graph neural network; The optimization function of the dynamic graph neural network model is: Wherein, D is the degree matrix, given the adjacency matrix A, D is a diagonal matrix, L is the Laplacian matrix and L = D - A, k is a variable, and T is the transpose algorithm.
2. The target group identification method according to claim 1, wherein The graph neural network model is a graph convolutional neural network or a graph isomorphism network model.
3. The target group identification method according to claim 1, wherein There are multiple time slices in the dynamic graph, namely {A 1 , A 2 ,..., A m}, that is, there are multiple corresponding degree matrices {D 1 , D 2 ,..., D m} and Laplacian matrices {L 1 , L 2 ,..., L m}. Take the mean values of the degree matrix and the Laplacian matrix in multiple time windows as the optimized proxy.
4. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-3.
Citation Information
Patent Citations
Microblog specific event attention group identification method
CN111026976A
Network user group division method and system
CN112712115A