Entity dependency relationship mining method and device based on time sequence diagram, medium and product
By dynamically segmenting the time period to be mined and computing node encoding vectors of the timing chart, the problem of poor timing chart dependency mining in the existing technology is solved, and more flexible and accurate entity dependency mining is achieved.
Patent Information
- Application Number
- CN202510279447.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The existing time-series graph dependency mining methods have problems such as time partition inadaptation, spatial feature limitations and lack of modeling of timing information, resulting in poor mining of entity dependencies.
By obtaining the relationship diagram of each moment in the period to be mined, the node characteristic change rate matrix is calculated, the period to be mined dynamically divides the period to be mined, the node's initial and target encoding vectors are determined, and the entity dependency relationship is determined using the Pearson correlation coefficient.
A flexible time partitioning strategy is implemented, which can better capture the dynamic changes in the timing chart, accurately identify the micro-dynamic relationship between nodes and edges, and model the timing dependence from a long span, improving the mining effect of entity dependencies.
Smart Images

Figure CN120216781A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of entity relationship mining, and in particular to an entity dependency relationship mining method, device, medium and product based on a time series graph. Background Art
[0002] Graph models have been widely used in complex information modeling due to their powerful capabilities in representing entities, attributes, and relationships between entities. A time series graph is an important extension of the graph model and is widely used in large knowledge bases in the fields of social networks, transportation networks, and financial markets. The entities in these knowledge bases are not static but evolve over time, which requires the ability to accurately and completely capture and update information to support downstream decision-making and fact-checking, thereby improving data consistency and integrity.
[0003] In a dynamic environment, changes in entities and relationships need to be captured and updated in real time to maintain the accuracy and reliability of data. The process of mining the dependency relationships of a time series graph involves identifying and extracting those patterns that repeatedly appear in the time series, and these patterns can reveal the internal laws of entity behavior. For example, by analyzing the time series graph of user interactions in a social network, we can identify the patterns of the formation and dissolution of user groups, which is crucial for social network analysis and community detection. In bioinformatics, by analyzing the time series graph of a biological network, the changing rules of gene expression over time can be revealed, which is of great value for disease research and drug development.
[0004] To effectively mine the dependency relationships of entities in these time series graphs, a variety of algorithms and technologies have been developed. However, existing dependency relationship mining methods often face the following challenges: (1) Inadaptability of time partitioning: Most time series graph dependency relationship mining methods use a fixed time window to process time series data, and this approach is difficult to adaptively adjust the time partitioning according to the dynamic characteristics of different nodes or graphs. (2) Limitations of spatial features: Traditional graph convolution methods usually rely on static graph structures and are difficult to fully capture the time series dependency relationships between nodes. (3) Lack of modeling of time series information: Many existing methods fail to effectively integrate long-term time series dependency information, thereby limiting the mining effect of time series graph dependency relationships.
[0005] In summary, the current mining effect of entity dependency relationships is poor, and it is impossible to well realize the mining of entity dependency relationships. Summary of the Invention
[0006] The purpose of the present application is to provide an entity dependency relationship mining method, device, medium and product based on a time series graph to solve the problem of poor mining effect of entity dependency relationships.
[0007] To achieve the above purpose, the present application provides the following solutions:
[0008] In a first aspect, the present application provides a method for mining entity dependency relationships based on a timing diagram, including:
[0009] Obtain the relationship diagrams at each moment in the period to be mined; the relationship diagram includes multiple nodes and multiple edges; one node corresponds to one entity, and one edge corresponds to a direct relationship between two entities; the number of nodes and the corresponding entities in the relationship diagrams at each moment are the same;
[0010] Based on the relationship diagrams at each moment respectively, determine the node feature matrix corresponding to the corresponding moment; the elements in the node feature matrix are the interaction times between the entities corresponding to two nodes;
[0011] Starting from the second moment, determine any moment as the current moment, and determine the node feature change rate matrix at the current moment according to the node feature matrix at the current moment and the node feature matrix at the previous moment;
[0012] Based on the node feature change rate matrices at each moment, segment the period to be mined to obtain multiple sub-periods, and determine any sub-period in the period to be mined as the target sub-period, and determine any node as the current node;
[0013] Based on the relationship diagrams at each moment in the target sub-period, determine the initial node coding vector of the current node in the target sub-period and the target time decay function values between the current node and its corresponding neighbor nodes in the target sub-period;
[0014] Based on the target time decay function values between the current node and its corresponding neighbor nodes in the target sub-period, update the initial node coding vector of the current node in the target sub-period to obtain the target node coding vector of the current node in the target sub-period;
[0015] Based on the target node coding vector of the current node in the target sub-period and the target node coding vector in the corresponding adjacent sub-period, determine the fused node coding vector of the current node in the target sub-period;
[0016] Based on the fused node coding vectors of any two nodes in the target sub-period, determine the Pearson correlation coefficient between the two nodes in the target sub-period;
[0017] Based on the Pearson correlation coefficient between the two nodes in the target sub-period, determine the dependency relationship between the entities corresponding to the two nodes in the target sub-period; the dependency relationship is dependent or independent.
[0018] Optionally, segmenting the period to be mined based on the node feature change rate matrices at each moment to obtain multiple sub-periods includes:
[0019] Determine all moments when the node feature change rate matrix exceeds the change rate threshold as time segmentation points;
[0020] Divide the to-be-mined time period according to all time segmentation points to obtain a plurality of sub-time periods.
[0021] Optionally, based on the relationship graph of each moment in the target sub-time period, determine the initial node encoding vector of the current node in the target sub-time period and the target time decay function values between the current node and its corresponding neighbor nodes in the target sub-time period, including:
[0022] Determine the initial node encoding vector of the current node in the target sub-time period according to the interaction situation of the entity corresponding to the current node and the entities corresponding to other nodes in the relationship graph at all moments in the target sub-time period;
[0023] Determine the nodes with edges to the current node as the neighbor nodes of the current node, and determine any moment as the first moment;
[0024] Calculate the initial time decay function values between the current node and its corresponding neighbor nodes at the first moment according to the central moment, the first moment and the time decay hyperparameter of the target sub-time period;
[0025] Take the average of the initial time decay function values between the current node and its corresponding neighbor nodes at all moments in the target sub-time period to obtain the target time decay function values between the current node and its corresponding neighbor nodes in the target sub-time period.
[0026] Optionally, based on the target time decay function values between the current node and its corresponding neighbor nodes in the target sub-time period, update the initial node encoding vector of the current node in the target sub-time period to obtain the target node encoding vector of the current node in the target sub-time period, including:
[0027] Use multi-layer graph convolution to update the initial node encoding vector of the current node in the target sub-time period for a preset maximum number of convolution layers according to the target time decay function values between the current node and its corresponding neighbor nodes in the target sub-time period, to obtain the target node encoding vector of the current node in the target sub-time period.
[0028] Optionally, based on the target node encoding vector of the current node in the target sub-time period and the target node encoding vector in the corresponding adjacent sub-time period, determine the fused node encoding vector of the current node in the target sub-time period, including:
[0029] Determine any sub-time period from the 2nd sub-time period to the second-to-last sub-time period in the to-be-mined time period as the current sub-time period;
[0030] Determine the fused node encoding vector of the current node in the current sub-period based on the target node encoding vector of the current node in the current sub-period, the target node encoding vector in the previous sub-period, and the target node encoding vector in the next sub-period;
[0031] Determine the fused node encoding vector of the current node in the first sub-period based on the target node encoding vector of the current node in the first sub-period and the target node encoding vector in the second sub-period;
[0032] Determine the fused node encoding vector of the current node in the last sub-period based on the target node encoding vector of the current node in the last sub-period and the target node encoding vector in the penultimate sub-period.
[0033] Optionally, determining the fused node encoding vector of the current node in the current sub-period based on the target node encoding vector of the current node in the current sub-period, the target node encoding vector in the previous sub-period, and the target node encoding vector in the next sub-period includes:
[0034] Determine the forward node encoding vector of the current node in the current sub-period according to the target node encoding vector of the current node in the current sub-period and the target node encoding vector in the previous sub-period;
[0035] Determine the backward node encoding vector of the current node in the current sub-period according to the target node encoding vector of the current node in the current sub-period and the target node encoding vector in the next sub-period;
[0036] Determine the fused node encoding vector of the current node in the current sub-period according to the forward node encoding vector and the backward node encoding vector of the current node in the current sub-period.
[0037] Optionally, determining the dependency relationship between the entities corresponding to two nodes in the target sub-period based on the Pearson correlation coefficient between the two nodes in the target sub-period includes:
[0038] Judge whether the Pearson correlation coefficient between the two nodes in the target sub-period exceeds a preset threshold;
[0039] If so, determine the dependency relationship between the entities corresponding to the two nodes in the target sub-period as dependent;
[0040] If not, determine the dependency relationship between the entities corresponding to the two nodes in the target sub-period as independent.
[0041] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the method for mining entity dependency relationships based on a time series graph described in any one of the above.
[0042] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for mining entity dependency relationships based on a timing diagram described in any one of the above is implemented.
[0043] In a fourth aspect, the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for mining entity dependency relationships based on a timing diagram described in any one of the above is implemented.
[0044] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:
[0045] The present application discloses a method, apparatus, medium, and product for mining entity dependency relationships based on a timing diagram. By dynamically segmenting the period to be mined according to the node feature change rate matrix, a flexible time partitioning strategy is implemented to cope with the changes in timing data at different time scales, overcoming the limitations of the fixed time window method and being able to better capture the dynamic changes in the timing diagram. Based on the target time decay function values between the current node and its corresponding neighbor nodes in the target sub-period, the initial node encoding vector of the current node in the target sub-period is updated, and the local timing dependency relationships are mined to obtain the target node encoding vector of the current node in the target sub-period, which helps to accurately identify the microscopic dynamic relationships between the nodes and edges in the graph. Based on the target node encoding vector of the current node in the target sub-period and the target node encoding vector in the corresponding adjacent sub-period, the fusion node encoding vector of the current node in the target sub-period is determined, modeling the timing dependency from a long time span and capturing the macroscopic laws of the changes in node and edge attributes over time. The mining effect of entity dependency relationships is improved. Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0047] Figure 1 It is a schematic flowchart of the method for mining entity dependency relationships based on a timing diagram provided by an embodiment of the present application;
[0048] Figure 2 It is a schematic diagram of the relational graph structure;
[0049] Figure 3 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed Embodiments
[0050] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0051] The purpose of the present application is to provide a method, device, medium and product for mining entity dependence relationships based on sequence diagrams, aiming to improve the mining effect of entity dependence relationships.
[0052] To make the above objects, features and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0053] In an exemplary embodiment, as Figure 1 shown, a method for mining entity dependence relationships based on sequence diagrams is provided, including:
[0054] Step 1: Obtain the relationship diagrams at each moment in the period to be mined; the relationship diagram includes multiple nodes and multiple edges; one node corresponds to one entity, and one edge corresponds to the direct relationship between two entities; the number of nodes and the corresponding entities in the relationship diagrams at each moment are the same.
[0055] Specifically, the relationship diagram at a certain moment is as Figure 2 shown, Figure 2 in which A, B, C, and D are the nodes corresponding to 4 entities respectively. The edge between A and B represents the direct relationship between the entities corresponding to A and B, the edge between A and D represents the direct relationship between the entities corresponding to A and D, and the edge between B and C represents the direct relationship between the entities corresponding to B and C.
[0056] In practical applications, the relationship diagrams at each moment in the period to be mined are determined based on the publicly available dataset ICEWS data containing N events.
[0057] First, the event set in the publicly available dataset ICEWS data is defined as D = {(t n , s n , r n , d n , o n )|n = 1, 2,..., N}, where t n is the moment of the nth event, s n is the initiating entity of the nth event, r n is the direct relationship (CAMEO code) between entities of the nth event, d n is the target entity of the nth event, on The text description for the nth event, where N is the total number of events in the dataset.
[0058] Then, by processing the publicly available ICEWS dataset, the temporal event quadruples ε are extracted as shown in formula (1), and redundant fields are removed to ensure the structured representation of the data.
[0059] ε = {(s n , r n , t n , d n ) | n = 1, 2, …, N} (1)
[0060] Secondly, for the extracted temporal event quadruples, they need to be de-duplicated, missing values repaired, and entities normalized to improve the data quality for better data processing.
[0061] Finally, the cleaned temporal event quadruples are expressed as a timestamped graph structure, i.e., the relationship graph at each moment in the period to be mined. The relationship graphs at each moment in the period to be mined form a temporal graph set G, which is defined as follows: G = (V, E, R, T). Here, V is the set of nodes in the temporal graph set; E is the set of edges in the temporal graph set; R is the set of direct relationships of entities; T is the set of moments.
[0062] Step 2: Based on the relationship graph at each moment respectively, determine the node feature matrix corresponding to that moment; the elements in the node feature matrix are the interaction times between the entities corresponding to two nodes.
[0063] Specifically, when the relationship graph is Figure 2 , the node feature matrix Z is expressed as:
[0064]
[0065] Among them, each element in Z represents the interaction times between the entities corresponding to the two letters in the subscript.
[0066] Step 3: Starting from the second moment, take any moment as the current moment, and determine the node feature change rate matrix at the current moment according to the node feature matrix at the current moment and the node feature matrix at the previous moment.
[0067] Specifically, the calculation formula for the node feature change rate matrix is:
[0068] ΔX(T i ) = X(T i ) - X(T i-1 ) (3)
[0069] Among them, ΔX(T i ) is the node feature change rate matrix at the i-th moment T iThe node feature change rate matrix; X(T i ) is the node feature matrix at the i-th moment T i ; X(T i-1 ) is the node feature matrix at the (i - 1)-th moment T i-1 .
[0070] Step 4: Based on the node feature change rate matrices at each moment, segment the period to be mined to obtain multiple sub-periods, and determine any sub-period in the period to be mined as the target sub-period, and determine any node as the current node.
[0071] As an optional implementation manner, Step 4 includes:
[0072] Step 41: Determine all moments when the node feature change rate matrix exceeds the change rate threshold as time segmentation points.
[0073] Step 42: According to all time segmentation points, segment the period to be mined to obtain multiple sub-periods.
[0074] Specifically, according to all time segmentation points M1, M2, …, M k , segment the period to be mined to obtain multiple sub-periods. Each sub-period corresponds to a sub-graph set, so k - 1 relationship graphs g1, g2, …, g k-1 of sub-periods are obtained. Among them, M1 is the first time segmentation point, that is, the initial moment of the period to be mined; M2 is the second time segmentation point; M k is the k-th time segmentation point, that is, the end moment of the period to be mined; g1 is the relationship graph of the first sub-period; g2 is the relationship graph of the second sub-period; g k-1 is the relationship graph of the (k - 1)-th sub-period. The b-th sub-period is represented as [M b , M b+1 ), M b is the b-th time segmentation point; M b+1 is the (b + 1)-th time segmentation point; the relationship graph set of the b-th sub-period is represented as g b .
[0075] Step 5: Based on the relationship graphs at each moment in the target sub-period, determine the initial node coding vector of the current node in the target sub-period and the target time decay function values between the current node and its corresponding neighbor nodes in the target sub-period.
[0076] As an optional implementation manner, Step 5 includes:
[0077] Step 51: Determine the initial node coding vector of the current node in the target sub-period according to the interaction situation between the entity corresponding to the current node and the entities corresponding to other nodes in the relationship graphs at all moments in the target sub-period.
[0078] Specifically, when the set of relationship diagrams in the b-th sub-period contains three relationship diagrams as shown in Figure 2 , and the edges between each node remain unchanged, only the type of the direct relationship represented by the edges changes. Among them, in the relationship diagram at the first moment in the b-th sub-period: the direct relationship between the entities corresponding to A and B is 01000, the direct relationship between the entities corresponding to A and D is 00100, and the direct relationship between the entities corresponding to B and C is 10000. In the relationship diagram at the second moment in the b-th sub-period: the direct relationship between the entities corresponding to A and B is 01000, the direct relationship between the entities corresponding to A and D is 00100, and the direct relationship between the entities corresponding to B and C is 10000. In the relationship diagram at the third moment in the b-th sub-period: the direct relationship between the entities corresponding to A and B is 00001, the direct relationship between the entities corresponding to A and D is 00100, and the direct relationship between the entities corresponding to B and C is 00010. At this time, the initial node encoding vector of node A is (00100, [01000 + 00001], 0) = (00100, 01001, 0). The initial node encoding vector of node B is (0, [01000 + 00001], [10000 + 00010]) = (0, 01001, 10010). The initial node encoding vector of node C is (0, 0, [10000 + 00010]) = (0, 0, 10010). The initial node encoding vector of node D is (00100, 0, 0).
[0079] Step 52: Determine the nodes with edges to the current node as the neighbor nodes of the current node, and determine any moment as the first moment.
[0080] Step 53: Calculate the initial time decay function values between the current node and its corresponding neighbor nodes at the first moment according to the central moment, the first moment, and the time decay hyperparameter of the target sub-period.
[0081] Specifically, the calculation formula for the initial time decay function value is:
[0082]
[0083] where is the initial time decay function value of node v c and neighbor node v d at the m-th moment in the b-th sub-period; is the central moment of the b-th sub-period; is the m-th moment in the b-th sub-period; τ is the time decay hyperparameter, which controls the influence degree of time.
[0084] Step 54: Calculate the average of the initial time decay function values at all times in the target sub-period between the current node and its corresponding neighbor nodes to obtain the target time decay function value between the current node and its corresponding neighbor nodes in the target sub-period.
[0085] Step 6: Based on the target time decay function values between the current node and its corresponding neighbor nodes in the target sub-period, update the initial node encoding vector of the current node in the target sub-period to obtain the target node encoding vector of the current node in the target sub-period.
[0086] As an optional implementation manner, Step 6 includes:
[0087] Step 61: Use multi-layer graph convolution to update the initial node encoding vector of the current node in the target sub-period for a preset maximum number of convolution layers according to the target time decay function values between the current node and its corresponding neighbor nodes in the target sub-period, to obtain the target node encoding vector of the current node in the target sub-period.
[0088] Specifically, the update formula for multi-layer graph convolution is:
[0089]
[0090] Where, is the node encoding vector of node v c after the (l + 1)-th layer of graph convolution in the b-th sub-period; σ(·) is the activation function; W (l) is the learnable transformation weight matrix of the l-th layer of graph convolution; is the node encoding vector of node v c after the l-th layer of graph convolution in the b-th sub-period; N(v c ) is the set of all neighbor nodes of node v c ; is the target time decay function value between node v c and neighbor node v d in the b-th sub-period; is the node encoding vector of neighbor node v c of node v d after the l-th layer of graph convolution in the b-th sub-period.
[0091] Step 7: Based on the target node encoding vector of the current node in the target sub-period and the target node encoding vector in the corresponding adjacent sub-period, determine the fused node encoding vector of the current node in the target sub-period.
[0092] As an optional implementation manner, Step 7 includes:
[0093] Step 71: Determine any sub - period from the second sub - period to the penultimate sub - period in the period to be mined as the current sub - period.
[0094] Step 72: Determine the fused node encoding vector of the current node in the current sub - period based on the target node encoding vector of the current node in the current sub - period, the target node encoding vector in the previous sub - period, and the target node encoding vector in the next sub - period.
[0095] As an optional implementation manner, step 72 includes:
[0096] Step 721: Determine the forward node encoding vector of the current node in the current sub - period according to the target node encoding vector of the current node in the current sub - period and the target node encoding vector in the previous sub - period.
[0097] Specifically, for node v y The calculation formula for the forward node encoding vector in the b - th sub - period is:
[0098]
[0099] Where, is the forward node encoding vector of node v y in the b - th sub - period; LSTM(·) is the long short - term memory network; is the target node encoding vector of node v y in the (b - 1) - th sub - period, and L is the preset maximum number of convolutional layers; is the target node encoding vector of node v y in the b - th sub - period.
[0100] Step 722: Determine the backward node encoding vector of the current node in the current sub - period according to the target node encoding vector of the current node in the current sub - period and the target node encoding vector in the next sub - period.
[0101] Specifically, for node v y The calculation formula for the backward node encoding vector in the b - th sub - period is:
[0102]
[0103] Where, is the backward node encoding vector of node v y in the b - th sub - period; is the target node encoding vector of node v y in the (b + 1) - th sub - period.
[0104] Step 723: Determine the fused node encoding vector of the current node in the current sub-period based on the forward node encoding vector and backward node encoding vector of the current node in the current sub-period.
[0105] Specifically, for node v y The calculation formula for the fused node encoding vector in the b-th sub-period is:
[0106]
[0107] where, o y,b is the fused node encoding vector of node v y in the b-th sub-period; concat(·) is the concatenation operation.
[0108] Step 73: Determine the fused node encoding vector of the current node in the first sub-period based on the target node encoding vector of the current node in the first sub-period and the target node encoding vector of the current node in the second sub-period.
[0109] Specifically, for node v y The calculation formula for the fused node encoding vector in the first sub-period is:
[0110]
[0111] where, o y,1 is the fused node encoding vector of node v y in the first sub-period; is the target node encoding vector of node v y in the first sub-period; is the target node encoding vector of node v y in the second sub-period.
[0112] Step 74: Determine the fused node encoding vector of the current node in the last sub-period based on the target node encoding vector of the current node in the last sub-period and the target node encoding vector of the current node in the penultimate sub-period.
[0113] Specifically, for node v y The calculation formula for the fused node encoding vector in the last sub-period is:
[0114]
[0115] where, o y,k-1 is the fused node encoding vector of node v y in the last sub-period (i.e., the (k - 1)-th sub-period); is the target node encoding vector of node v y in the penultimate sub-period (i.e., the (k - 2)-th sub-period); For node v y The target node encoding vector in the last sub-period (i.e., the (k - 1)-th sub-period).
[0116] Step 8: Based on the fused node encoding vectors of any two nodes in the target sub-period, determine the Pearson correlation coefficient between the two nodes in the target sub-period.
[0117] Specifically, the calculation formula of the Pearson correlation coefficient is:
[0118]
[0119] where is the Pearson correlation coefficient between node v x and v y in the b-th sub-period; con(·) is the covariance; o x,b is the fused node encoding vector of node v x in the b-th sub-period; ζ(·) is the standard deviation.
[0120] Step 9: Based on the Pearson correlation coefficient between two nodes in the target sub-period, determine the dependency relationship between the entities corresponding to the two nodes in the target sub-period; the dependency relationship is dependent or independent.
[0121] As an alternative implementation, Step 9 includes:
[0122] Step 91: Determine whether the Pearson correlation coefficient between two nodes in the target sub-period exceeds a preset threshold.
[0123] Step 92: If so, determine the dependency relationship between the entities corresponding to the two nodes in the target sub-period as dependent.
[0124] Step 93: If not, determine the dependency relationship between the entities corresponding to the two nodes in the target sub-period as independent.
[0125] In an exemplary embodiment, a computer device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the method for mining entity dependency relationships based on a time series graph.
[0126] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the method for mining entity dependency relationships based on a time series graph is implemented.
[0127] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the method for mining entity dependency relationships based on a time series graph is implemented.
[0128] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal, and its internal structure diagram may be as follows Figure 3 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for mining entity dependency relationships based on a timing diagram.
[0129] Those skilled in the art can understand that Figure 3 the structure shown in
[0130] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0131] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0132] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0133] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0134] In this text, specific examples are used to illustrate the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for mining entity dependency relationships based on a time series graph, characterized in that: The entity dependency mining method based on the time series graph includes: Obtain a relationship graph at each moment in the period to be mined; the relationship graph includes multiple nodes and multiple edges; one node corresponds to one entity, and one edge corresponds to a direct relationship between two entities; the number of nodes and corresponding entities in the relationship graph at each moment are the same; Based on the relationship graphs at each time, a node feature matrix at the corresponding time is determined; the elements in the node feature matrix are the number of interactions between entities corresponding to the two nodes; Starting from the second moment, any moment is determined as the current moment, and the node feature change rate matrix of the current moment is determined according to the node feature matrix of the current moment and the node feature matrix of the previous moment; The time period to be mined is divided based on the node feature change rate matrix at each moment to obtain multiple sub-periods, and any sub-period in the time period to be mined is determined as a target sub-period, and any node is determined as a current node; Based on the relationship diagram at each time in the target sub-period, determine the initial node encoding vector of the current node in the target sub-period and the target time attenuation function value between the current node and each corresponding neighbor node in the target sub-period; Based on the target time decay function value between the current node and the corresponding neighbor nodes in the target sub-period, the initial node encoding vector of the current node in the target sub-period is updated to obtain the target node encoding vector of the current node in the target sub-period; Determine a fused node coding vector of the current node in the target sub-period based on the target node coding vector of the current node in the target sub-period and the target node coding vector in the corresponding adjacent sub-period; Based on the fusion node encoding vectors of any two nodes in the target sub-period, determine the Pearson correlation coefficient between the two nodes in the target sub-period; Based on the Pearson correlation coefficient between the two nodes in the target sub-period, the dependency relationship between the entities corresponding to the two nodes in the target sub-period is determined; the dependency relationship is dependency or non-dependence.
2. The entity dependency mining method based on time series graph according to claim 1 is characterized in that: The time period to be mined is divided based on the node feature change rate matrix at each moment to obtain multiple sub-periods, including: The moments when the node feature change rate matrix exceeds the change rate threshold are determined as time segmentation points; The time period to be mined is divided according to all time division points to obtain a plurality of sub-time periods.
3. The entity dependency mining method based on time series graph according to claim 1 is characterized in that: Based on the relationship graph at each time in the target sub-period, the initial node encoding vector of the current node in the target sub-period and the target time attenuation function value between the current node and each corresponding neighbor node in the target sub-period are determined, including: Determine the initial node encoding vector of the current node in the target sub-period according to the interaction between the entity corresponding to the current node and the entities corresponding to other nodes in the relationship graph at all times in the target sub-period; Determine the nodes that have edges with the current node as neighbor nodes of the current node, and determine any moment as the first moment; According to the central moment, the first moment and the time decay hyperparameter of the target sub-period, the initial time decay function value between the current node and each corresponding neighbor node at the first moment is calculated; The initial time decay function values between the current node and the corresponding neighbor nodes at all times in the target sub-period are averaged to obtain the target time decay function value between the current node and the corresponding neighbor nodes in the target sub-period.
4. The entity dependency mining method based on time series graph according to claim 1 is characterized in that: Based on the target time decay function value between the current node and the corresponding neighbor nodes in the target sub-period, the initial node encoding vector of the current node in the target sub-period is updated to obtain the target node encoding vector of the current node in the target sub-period, including: By using multi-layer graph convolution, according to the target time attenuation function value between the current node and the corresponding neighbor nodes in the target sub-period, the initial node coding vector of the current node in the target sub-period is updated with a preset maximum number of convolution layers to obtain the target node coding vector of the current node in the target sub-period.
5. The entity dependency mining method based on time series graph according to claim 1 is characterized in that: Determining a fusion node coding vector of the current node in the target sub-period based on the target node coding vector of the current node in the target sub-period and the target node coding vector in the corresponding adjacent sub-period includes: Determine any sub-period from the second sub-period to the second-to-last sub-period in the period to be mined as the current sub-period; Determine a fusion node coding vector of the current node in the current sub-period based on the target node coding vector of the current node in the current sub-period, the target node coding vector in the previous sub-period, and the target node coding vector in the next sub-period; Determine a fusion node coding vector of the current node in the first sub-period based on the target node coding vector of the current node in the first sub-period and the target node coding vector in the second sub-period; Based on the target node coding vector of the current node in the last sub-period and the target node coding vector in the second to last sub-period, the fusion node coding vector of the current node in the last sub-period is determined.
6. The entity dependency mining method based on time series graph according to claim 5 is characterized in that: Determining a fusion node coding vector of the current node in the current sub-period based on a target node coding vector of the current node in the current sub-period, a target node coding vector in the previous sub-period, and a target node coding vector in the next sub-period includes: Determine a forward node encoding vector of the current node in the current sub-period according to the target node encoding vector of the current node in the current sub-period and the target node encoding vector in the previous sub-period; Determine a backward node encoding vector of the current node in the current sub-period according to the target node encoding vector of the current node in the current sub-period and the target node encoding vector in the next sub-period; According to the forward node coding vector and the backward node coding vector of the current node in the current sub-period, a fused node coding vector of the current node in the current sub-period is determined.
7. The entity dependency mining method based on time series graph according to claim 1 is characterized in that: Based on the Pearson correlation coefficient between the two nodes in the target sub-period, the dependency relationship between the entities corresponding to the two nodes in the target sub-period is determined, including: Determine whether the Pearson correlation coefficient between two nodes in the target sub-period exceeds a preset threshold; If yes, the dependency relationship between the entities corresponding to the two nodes in the target sub-period is determined as dependency; If not, the dependency relationship between the entities corresponding to the two nodes in the target sub-period is determined to be non-dependent.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the entity dependency mining method based on a timing graph as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the entity dependency mining method based on the timing graph described in any one of claims 1 to 7 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the entity dependency mining method based on the timing graph described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method for collaborative filtering recommendation based on interest changes and trust relations
CN106570090A
Entity vector determining method and apparatus, and information retrieval method and apparatus
CN108717407A
Data processing method and device, equipment and storage medium
CN115033791A
Personalized service recommendation method based on Web
CN117216403A
Mapping Entities to Accounts for De-Anonymization of Online Activity
US20230180214A1