A visualization analysis system and method for time-series knowledge graphs

By combining a storyline with a node linking graph visualization analysis system, the problem of insufficient utilization of time information in existing technologies is solved, enabling the display of entity and relationship change trends and user-friendly knowledge graph understanding.

CN115168601BActive Publication Date: 2025-12-02ZHEJIANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210724550.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-23
Publication Date
2025-12-02
Estimated Expiration
2042-06-23

AI Technical Summary

Technical Problem

Existing knowledge graph visualization technologies do not make full use of time information, making it difficult to observe the changing trends of entities and relationships. They also have a high barrier to entry for users and are not convenient for exploring unfamiliar knowledge graphs.

Method used

Design a visualization analysis system and method for time-series knowledge graphs. Combining storylines and point-line graphs, users can iteratively select entities, relationships, and time points. The system automatically generates visualization charts and provides descriptive text to highlight graph differences and trends.

Benefits of technology

By combining storylines with node link diagrams in a visual format, the differences in the graph are highlighted, helping users observe the changing trends of entities and relationships, lowering the barrier to understanding, and enabling them to quickly comprehend unfamiliar knowledge graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115168601B_ABST
    Figure CN115168601B_ABST
Patent Text Reader

Abstract

This invention discloses a visualization analysis system and method for time-series knowledge graphs. Users iteratively select entities, relationships, and time points of interest within the time-series knowledge graph. Based on the user's selections, the system automatically generates a visualization chart combining storylines and point-line graphs, displaying the topological structure and temporal changes of the corresponding entities and relationships in the graph. Descriptive text is also generated to supplement the visualization chart. This invention meets the visualization needs of time-series knowledge graphs, effectively reduces the difficulty of exploring time-series knowledge graphs, enhances users' perception of temporal changes in the graph, and promotes the research and application of time-series knowledge graphs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of time-series knowledge graph visualization, and more particularly to a visualization analysis system and method for time-series knowledge graphs. Background Technology

[0002] Knowledge graphs are an important branch of artificial intelligence and the cornerstone of machine cognitive intelligence. In 2012, Google released a large-scale knowledge graph for internet search, marking the birth of the knowledge graph. In the following few years, a wealth of theoretical and technical research results emerged. Upper-level applications such as data analysis, intelligent recommendation, and decision support all placed demands on knowledge graphs. Knowledge graphs have been applied in numerous fields, including e-commerce (such as Alibaba's e-commerce knowledge graph), healthcare (such as the Linked Life Data project), and finance (such as the Kensho financial knowledge engine), demonstrating the vibrant development of knowledge graph research.

[0003] With the rapid development of knowledge graph technology, the demand for knowledge graph visualization has emerged. Knowledge graph visualization helps people understand the relationships between entities more intuitively by displaying the internal topological structure of the knowledge graph. Existing knowledge graph visualization technologies generally involve first modeling the triple information in the knowledge graph, then using graph layout algorithms such as force-directed layout to lay out the graph, and finally rendering the layout result as a visualization of the knowledge graph.

[0004] In recent years, knowledge graph researchers have discovered the crucial importance of temporal information within knowledge graphs. On the one hand, some structured knowledge is only valid within specific timeframes; on the other hand, many facts in knowledge graphs dynamically change over time. Fully utilizing this temporal information allows for better modeling of the dynamic topological structure between entities, understanding the temporal trends of entity and relation changes, and facilitating the construction, completion, and reasoning of knowledge graphs. To this end, knowledge graph researchers have proposed temporal knowledge graphs composed of a quadruple of (subject, relation, object, and temporal information), and have conducted research from various perspectives, including temporal information encoding, temporal relation dependencies, and temporal logical reasoning.

[0005] However, most existing knowledge graph visualization technologies are designed for static knowledge graphs and do not make full use of the temporal information in the knowledge graph or improve them for time-series knowledge graphs.

[0006] The baseline approach for visualizing time-series knowledge graphs is to use static knowledge graph visualization techniques to model and lay out the time-series knowledge graph, and finally add the time information directly to the corresponding layout result.

[0007] Furthermore, application publication number CN114036311A discloses a time-series visualization development method based on knowledge graphs, the steps of which include: obtaining data requests; generating query statements from the data requests using query templates, and performing queries based on the query statements; sorting the query results according to time nodes, and rendering a timeline based on the time series; obtaining data requests for time nodes on the timeline, and querying data that matches the time nodes; indexing and marking the data corresponding to the queried time nodes, and performing visualization rendering on the data; and outputting the rendered data.

[0008] The above methods all have the following limitations:

[0009] First, the differences in the graphs at different time points are not obvious, making it difficult to observe the changing trends of entities and relationships, and users find it difficult to discover topological structures with strong correlations in the changing trends.

[0010] Secondly, it requires users to have some understanding of the graph structure being analyzed, which makes it difficult to get started and hinders users from exploring unfamiliar knowledge graphs. Summary of the Invention

[0011] To address the shortcomings of existing technologies, this invention proposes a visualization analysis system and method for time-series knowledge graphs. Users can iteratively select entities, relationships, and time points of interest in the time-series knowledge graph. Based on the user's selection, the system automatically generates a visualization chart that combines storylines and point-line graphs to show the user the topological structure and temporal changes of the corresponding entities and relationships in the graph. At the same time, descriptive text is generated as a supplement to the visualization chart.

[0012] The objective of this invention is achieved through the following technical solution:

[0013] A visualization and analysis system for time-series knowledge graphs, the system comprising:

[0014] The overview generation module generates a dataset overview based on the overview configuration data.

[0015] The storyline generation module generates storylines based on storyline configuration data.

[0016] The text generation module generates descriptive text based on storyline configuration data;

[0017] The canvas module displays the system-generated overview, storyline, and text, and responds to user interactions by updating the overview and storyline configuration data; it consists of a configuration panel, an overview panel, and a storyline panel.

[0018] The configuration panel is used to receive user modifications to the overview configuration data and storyline configuration data.

[0019] The overview panel is used to display an overview view, receive entities selected by the user, and initialize storyline configuration data;

[0020] The storyline panel is used to display the storyline view and receive user interactions with entities, relationships, and time points.

[0021] Furthermore, the storyline panel is divided into a timeline, a static section, and a chronological section. The static section is used to display static relationships, while the chronological section is used to display chronological relationships and event relationships.

[0022] A visualization analysis method for time-series knowledge graphs, implemented based on a visualization analysis system, comprising:

[0023] The system generates an overview view based on the overview configuration data input by the user and displays it to the user;

[0024] The system initializes storyline configuration data based on the entity selected by the user in the overview view; then it generates a storyline view and descriptive text based on the storyline configuration data and displays them to the user.

[0025] The system updates the storyline configuration data based on the user's interactions with entities, relationships, and time points in the storyline view, and then generates a storyline view and descriptive text based on the storyline configuration data, which are then displayed to the user.

[0026] Furthermore, the overview configuration data includes time span segmentation method, entity encoding method, and area map encoding method;

[0027] The storyline configuration data includes monitoring status flags, selected entity sets, monitored entity sets, visible entity sets, selected relationship sets, visible relationship sets, selected time point sets, visible time point sets, and manipulated time points.

[0028] Furthermore, the system generates an overview view based on the overview configuration data input by the user, specifically including:

[0029] First, the overall time span of the dataset is segmented, and the segmented time periods are mapped onto the y-axis. Then, the information within each time period is encoded by area, and the encoded values ​​are mapped onto the x-axis to create an area map. Finally, the entities existing within each time period are encoded, with the encoded values ​​mapped to text size and the entity category mapped to text color. A word cloud is then drawn within the corresponding time period of the area map.

[0030] Furthermore, the specific sub-steps for generating the storyline view based on the storyline configuration data are as follows:

[0031] (1) Calculate the visible set;

[0032] ① Initialize the visible entity set with the selected entity set, the visible relation set with the selected relation set, and the visible time point set with the selected time point set;

[0033] ② If currently under monitoring, add the start time, start time minus unit time, end time, and end time plus unit time of all entities in the monitored entity set and their associated non-static relationships to the visible time point set; add all entities in the monitored entity set and entities reachable within the expansion step size to the visible entity set; add the pairwise relationships between all entities in the visible entity set to the visible relationship set.

[0034] ③ Add the subjects and objects of all relations in the selected relation set to the visible entity set;

[0035] (2) Calculate the storyline;

[0036] Calculate the line order of the entity storyline; calculate the line order of all storylines; calculate the storyline layout; expand the storyline layout.

[0037] (3) Calculate the diagram layout along the storyline;

[0038] Traverse the set of visible time points in chronological order. At any given time point, the subgraph to be laid out includes all visible relationships that are newly appearing at that time point or disappearing at the next time point, as well as the associated entities of these relationships. After moving the positions of entities or relationships several times, the graph layout on the storyline at that time point is obtained while minimizing the objective function under the constraints.

[0039] Each relationship corresponds to a line segment from the subject position to the relationship position and a line segment from the relationship position to the object position. The constraint is that the entity or relationship to be laid out must fall on the corresponding position of its storyline at that time point on the y-axis and fall within a limited width on the x-axis. The limited width is the width of each subgraph. The objective function is the sum of the number of intersections between the two line segments corresponding to the relationship to be laid out and the two line segments corresponding to other relationships, as well as the number of intersections between the bounding boxes corresponding to other entities or relationships to be laid out.

[0040] (4) Calculate the static diagram layout

[0041] The subgraph to be laid out in the static graph includes all static relations in the visible relation set and the associated entities of these relations. There are no constraints on the position of these relations on the y-axis. If an entity is a static entity, there are no constraints on the position of the entity on the y-axis. Otherwise, the entity falls on the y-axis position corresponding to its storyline at the manipulation time point. If the corresponding storyline does not exist at the manipulation time point, the entity is placed above or below the inner canvas depending on whether the corresponding storyline has not appeared or has disappeared. The other constraints and optimization objective function are the same as the graph layout calculated on the storyline.

[0042] Furthermore, the sub-steps for generating descriptive text based on the storyline configuration data are as follows:

[0043] (1) Preprocessing: Organize and supplement each selected set to obtain the text generation start time point, text generation end time point, text generation entity set, and text generation relation set. If the data is insufficient to generate text, the text generation will end.

[0044] (2) Serialization: Based on time information, graph topology, and user operation sequence, the entities and relations in the text generation entity set and text generation relation set are sorted to obtain an ordered list of entities and entity relationships, so that the final generated text is orderly, organized, and consistent with the user's intent.

[0045] (3) Template filling: Use the given template and combination rules to convert the serialization result into descriptive text.

[0046] Furthermore, the specific sub-steps of the serialization are as follows:

[0047] (a) Calculate the priority of entities, relations, and time sequence;

[0048] For each entity in the text-generated entity set, its weight is (centrality - selection order / size of the text-generated entity set). The entity with higher weight has higher priority. If the weights are the same, the entity selected earlier has higher priority. For each type of relation in the text-generated relation set, the entity with fewer relations of the same type in the text-generated relation set has higher priority. For each relation in the text-generated relation set, the entity selected earlier has higher priority.

[0049] (b) Divide the entities in the text-generated entity set into several clusters;

[0050] Each non-static entity is an independent cluster; two static entities that are associated by a static relationship are grouped into the same cluster; the entity with the highest priority in each cluster is the root entity of that cluster;

[0051] (c) Calculate the time point set and divide the non-static entities and non-static relationships into several time point buckets;

[0052] List all the time points in the text generation entity set and text generation relation set that are associated with each entity and relation within the time span formed by the start time and end time of text generation; use these time points as buckets to classify the non-static entities and non-static relations in the text generation entity set and text generation relation set into the buckets associated with them, and if multiple time points are associated, classify them into the bucket with the earlier time sequence.

[0053] (d) Process each bucket in chronological order;

[0054] The entities and relationships within each bucket are further divided into several entity buckets;

[0055] Each entity bucket is processed sequentially according to its corresponding entity priority.

[0056] Furthermore, the specific sub-steps for further dividing the entities and relationships within a bucket into several entity buckets are as follows:

[0057] Entities are assigned to their corresponding buckets; relations are attached to the root entity of the cluster in which their subject and object belong, and are assigned to the bucket corresponding to the root entity of the cluster. If the root entity does not exist, a new bucket is created for the corresponding entity.

[0058] Furthermore, the sub-steps for processing each entity bucket according to its priority are as follows:

[0059] If the entity corresponding to the current entity bucket is not dependent, then jump to (d.2.4) to process the entity bucket that is not dependent; otherwise, if the entity corresponding to the current entity bucket is in the bucket at the current time point, then jump to (d.2.2) to process the static entity bucket; otherwise, jump to (d.2.1) to process the non-static entity bucket.

[0060] The entity that can be attached to the current entity bucket refers to the entity being in the bucket at the current time point or the entity being an unvisited static entity.

[0061] (d.2.0) The process for handling relationships is as follows:

[0062] Given an entity and several relations, the relations are first grouped according to their relation categories. Relations of different categories with the same related entity are merged into another group. The groups are sorted by category priority between groups and by relation priority within groups. Finally, the tuple consisting of the given entity and relation sequence is added to the serialization result, and the given relation is marked as visited.

[0063] (d.2.1) The process for handling non-static entity buckets is as follows:

[0064] The relationships to be processed are: several relationships belonging to the current entity bucket, several static relationships with the entity corresponding to the current entity bucket as the subject, and several static relationships with the entity corresponding to the current entity bucket as the object and the subject as a static entity; use the relationship processing method in (d.2.0) to process the relationships to be processed; the extended entities are another entity associated with the relationships to be processed, and these entities are unvisited static entities, and the root entity of their cluster cannot be the entity bucket to be processed; jump to (d.2.3) to process the extended entities;

[0065] (d.2.2) The process for handling static entity buckets is as follows:

[0066] If a relationship exists within a bucket and all relationships are attached to the same entity, or if no relationship exists within a bucket and a unique entity exists, then that entity is designated as the entry entity; otherwise, the entity corresponding to the current entity bucket is designated as the entry entity. The cluster to which the entity corresponding to the current entity bucket belongs is designated as the current cluster. The entry entity is added to the candidate set, and the remaining entities in the current cluster are added to the residual set. The entity with the highest priority is selected from the candidate set, and it is marked as visited. The relationships to be processed are those that are associated with the entity but have not been visited and are in the current entity bucket, as well as the static relationships that are associated with the entity and another entity in the current cluster but have not been visited. The above relationships to be processed are processed using the relationship processing method in (d.2.0). Entities associated with the above relationships to be processed and in the residual set are removed from the residual set and added to the candidate set. The above operations are repeated until the candidate set is empty. The extended entities are other entities associated with relationships within the entity bucket, and these entities are unvisited static entities, and the root entity of their cluster cannot be the entity bucket to be processed. Jump to (d.2.3) to process the extended entities.

[0067] (d.2.3) The process for handling extended static entities is as follows:

[0068] For several static entities to be processed, entities belonging to the same cluster are grouped into the same bucket, and then further divided into several static entity buckets; each static entity bucket is processed sequentially according to the priority of the corresponding entity in step (d.2.2);

[0069] (d.2.4) The process for handling unattached physical buckets is as follows:

[0070] For several relations within an entity bucket, if the root entity of the cluster to which the relation is associated can be attached, the relation is assigned to the corresponding entity bucket. If the corresponding entity bucket does not exist, a new entity bucket is created and added to the sequence of entity buckets to be processed. For the remaining relations that have not been assigned to other entity buckets, they are first grouped according to the entities to which the relation is attached, and then sorted between groups according to the priority of the entities. Then, the method for processing relations in (d.2.0) is used to process each group of relations in turn. Finally, the current entity bucket processing flow ends.

[0071] The beneficial effects of this invention are as follows:

[0072] This invention uses a visualization format combining storylines and node linking graphs to present temporal knowledge graphs to users, highlighting the differences in the graph at different time points. This helps users observe the changing trends of entities and relationships and discover topological structures with strong correlations in these trends. Based on user interaction, this invention generates descriptive text with reasonable word order and clear logic, lowering the threshold for understanding the graph structure and helping users quickly understand knowledge graphs in unfamiliar fields. Attached Figure Description

[0073] Figure 1This is a schematic diagram of the system panel for a visualization analysis system for time-series knowledge graphs.

[0074] Figure 2 A flowchart illustrating a visualization analysis method for time-series knowledge graphs;

[0075] Figure 3 This is a flowchart of the serialization algorithm. Detailed Implementation

[0076] The present invention will be described in detail below with reference to the accompanying drawings and preferred embodiments. The purpose and effects of the present invention will become clearer. It should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0077] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. Terms such as "association" as used in this invention mean that the two entities appearing before or after the term are respectively the subject and object of a certain relationship, or that the entities appearing before or after the term are the subject or object of the relationship appearing before or after the term, or that the start or end time of the entities or relationship appearing before or after the term is the point in time before or after the term.

[0078] It should be noted that this embodiment requires the construction of a specific data format.

[0079] The data required in this embodiment can be divided into two categories: entity data and relational data. The basic attributes of entity data are start time and end time; the basic attributes of relational data are start time, end time, subject, and object. Both entity data and relational data are further divided into three subcategories: static, time-series, and event.

[0080] in:

[0081] Static entities It is an entity that exists continuously but does not carry time information; its start and end times are undefined; such as the moon.

[0082] Time-series entities An entity is an entity that exists within a certain time period and carries time information. Its basic attributes define a start time and / or an end time; for example, Neil Alden Armstrong existed from August 5, 1930 to August 25, 2012.

[0083] Event Entity It is an entity that exists only at a certain point in time and carries time information. Its start time and end time are at the same point in time. It is generally abstracted from a certain event, such as the Armstrong moon landing event, which occurred on July 20, 1969.

[0084] static relationshipIt is a relationship that has always existed but does not contain time information; that is, its start and end times are undefined.

[0085] For example (Neil Alden Armstrong, nationality: American);

[0086] Temporal Relationship It is a relationship that exists within a certain time period and carries time information. Its basic attributes define a start time and / or an end time, such as (Neil Alden Armstrong, Professor of Aeronautical Engineering, University of Cincinnati), which exists from 1971 to 1979.

[0087] Event Relationship It is a relationship that exists only at a certain point in time and contains time information, that is, its start time and end time are at the same point in time, such as (Neil Alden Armstrong, landing, moon), which exists on July 20, 1969.

[0088] The basic attributes required for each type of data are as follows:

[0089]

[0090] Existing knowledge graphs have many data structures. When using the visualization analysis system and method for time-series knowledge graphs of this invention, the existing knowledge graphs need to be converted into the data structure required by this invention. Here, a feasible data structure conversion method is provided for two types of knowledge graphs: time-series knowledge graphs constructed in the format of (subject, relation, object, time information) quadruples and general knowledge graphs with time information constructed in the format of (subject, relation, object) triples.

[0091] For a time-series knowledge graph constructed in the format of (subject, relation, object, time information) quadruple, it can be converted into the required data format through the following steps: extract all original entities as time-series entities that start at negative infinity and end at positive infinity, and extract all original relations as time-series relations.

[0092] For a general knowledge graph with time information constructed in the (subject, relation, object) triple format, it can be converted into the required data format through the following steps: Determine the classes of the original entities extracted as entities, the classes of the original entities extracted as relations, the original relations relating to time information, the original relations relating to subjects and relations, and the original relations relating to objects and relations. Based on the above information, extract some original entities as entities and supplement their time information; extract some original entities as relations, specify their subjects and objects, and supplement their time information.

[0093] The meanings of some technical terms involved in the system and method of this invention are explained below.

[0094] Monitoring status flag: Indicates whether the user is in monitoring status;

[0095] Selected entity set: The collection of entities that are selected;

[0096] Monitoring entity set: The collection of entities being monitored;

[0097] Visible entity set: The collection of entities currently displayed to the user;

[0098] Selected set of relations: The set of relations that are selected;

[0099] Visible relationship set: The set of relationships currently displayed to the user;

[0100] Selected time point set: The set of time points selected by the user;

[0101] Visible time point set: The set of time points currently displayed to the user;

[0102] Manipulate Time Point: The earliest time point visible to the user in the timeline section of the storyline panel. It is used to manipulate the layout of the static parts of the graph, and changes as the user scrolls the timeline.

[0103] The visualization analysis system for time-series knowledge graphs of the present invention includes the following modules:

[0104] (1) Overview generation module, which generates a dataset overview based on overview configuration data;

[0105] (2) Storyline generation module, which generates storylines based on storyline configuration data;

[0106] (3) Text generation module, which generates descriptive text based on storyline configuration data;

[0107] (4) The canvas module displays the system-generated overview, storyline, and text, and responds to user interactions by updating the overview configuration data and storyline configuration data; such as Figure 1 As shown, it is divided into a configuration panel, an overview panel, and a storyline panel. The configuration panel receives user modifications to the overview and storyline configuration data. The overview panel displays an overview view, receives user-selected entities, and initializes the storyline configuration data. The storyline panel displays the storyline view and receives user interactions with entities, relationships, and time points. The storyline panel is further divided into a timeline, a static section, and a chronological section. The static section displays static relationships, and the chronological section displays chronological and event relationships. Users can click on entities, relationships, and time points in this view to modify the selected entity set, monitored entity set, selected relationship set, and selected time point set, and scroll the timeline to modify the manipulated time points.

[0108] The visualization analysis method for time-series knowledge graphs of the present invention is as follows: Figure 2 As shown, it includes the following steps:

[0109] Step 1: The system generates an overview view based on the overview configuration data input by the user and displays it to the user;

[0110] Step 2: The system initializes the storyline configuration data based on the entity selected by the user in the overview view; then it generates a storyline view and descriptive text based on the storyline configuration data and displays them to the user.

[0111] Step 3: The system updates the storyline configuration data based on the user's interactions with entities, relationships, and time points in the storyline view, and then generates a storyline view and descriptive text based on the storyline configuration data, which are then displayed to the user.

[0112] For example, in this embodiment, the overview configuration data includes time span segmentation method, entity encoding method, area map encoding method, etc.

[0113] For example, in this embodiment, the time span segmentation method can be to divide the time span of the dataset into a specified number of segments, or to segment it according to a specified step size; the entity encoding method is the number of non-static relationships that an entity has within a time period; and the area map encoding method is the number of non-static relationships that exist within a time period.

[0114] The sub-steps in step one where the system generates the overview view based on the overview configuration data input by the user are as follows:

[0115] First, the overall time span of the dataset is segmented, and the segmented time periods are mapped onto the y-axis. Then, the information within each time period is encoded by area, and the encoded values ​​are mapped onto the x-axis to create an area map. Finally, the entities existing within each time period are encoded, with the encoded values ​​mapped to text size and the entity category mapped to text color. Word clouds are then drawn using wordcloud2.js within the corresponding time period of the area map.

[0116] The storyline configuration data includes monitoring status flags, selected entity sets, monitored entity sets, visible entity sets, selected relationship sets, visible relationship sets, selected time point sets, visible time point sets, and manipulated time points.

[0117] Step Two: The system initializes the storyline configuration data based on the entity selected by the user in the overview view as follows:

[0118] The system clears the selected entity set, monitored entity set, selected relationship set, and selected time point set, adds the entities selected by the user in the overview view to the selected entity set and monitored entity set, and sets the monitoring status flag to true.

[0119] The methods for generating storyline views from system storyline configuration data in steps two and three are as follows:

[0120] (3-1.1) Calculate the visible set:

[0121] (3-1.1.1) Initialize the visible entity set with the selected entity set, initialize the visible relation set with the selected relation set, and initialize the visible time point set with the selected time point set.

[0122] (3-1.1.2) If in monitoring state, add the start time, start time-unit time, end time, and end time+unit time of all entities and their associated non-static relationships in the monitoring entity set to the visible time point set; add all entities in the monitoring entity set and entities reachable within the expansion step size to the visible entity set; add the pairwise relationships between all entities in the visible entity set to the visible relationship set.

[0123] Preferably, the expansion step size is set to 1. The expansion step size is the minimum number of relationships required for one entity to traverse to another, as given by the user, and can be defined by the user in the configuration panel.

[0124] (3-1.1.3) Adds the subjects and objects of all relations in the selected relation set to the visible entity set.

[0125] (3-1.2) Calculate the storyline:

[0126] (3-1.2.1) Calculate the line order of the entity storyline:

[0127] Let each entity in the visible entity set belong to a group. Iterate through the set of visible time points in chronological order; at any given time point:

[0128] (a) If there is a newly emerging (not existing at the previous time point) or soon disappearing (not existing at the next time point) relationship between two entities in the visible entity set, then merge the two groups corresponding to these two entities into one group;

[0129] (b) If two entities in the visible entity set had a relationship at the previous time point, and that relationship disappeared at the current time point, then the two groups that were merged into the same group at the previous time point will split back into two groups.

[0130] (c) Record the current grouping status as the grouping information at that point in time.

[0131] Using the line movement interaction steps of a storyline visualization layout generation method disclosed in application publication number CN109068152A, the method involves moving the storyline several times. Under the constraint of the continuity of the storyline order corresponding to the same group of entities at each time point, the method minimizes the number of intersections of all storylines, thereby obtaining the line order of each entity in the visible entity set as the storyline in the vertical direction.

[0132] (3-1.2.2) Calculate the line order of all storylines:

[0133] Traverse the set of visible time points in chronological order. At any given time point, based on the line order of the entity storyline, for any entity storyline, if there are several non-static relationships with that entity as the subject in the set of visible relationships, then the storylines corresponding to these relationships are inserted before and after the entity storyline in order according to the storyline order of their objects and their position relative to the entity storyline.

[0134] For example, if the line order of the entity storyline at a certain point in time is [E1,E2,E3,E4], and it is evident that there are non-static relations R1:(E3,r1,E1), R2:(E3,r2,E2), and R3:(E3,r3,E4) in the relation set, then the line order after inserting the relation storyline is [E1,E2,R1,R2,E3,R3,E4]. Here, R represents the relation, the first E in parentheses represents the subject, r represents the relation category of relation R, and the second E represents the object.

[0135] (3-1.2.3) Calculate the storyline layout:

[0136] After several moves of the storyline, the layout height is less than h, and the distance between any two storylines is greater than or equal to d. l The distance between any two storylines in different groups is greater than or equal to d. g Under the constraint that the order of the storylines remains unchanged, the storyline layout can be obtained by minimizing the number of times the storylines bend at adjacent time points.

[0137] Preferred, d g =2d l h = max(number of storylines * d) l +Number of groups * d g *2).

[0138] Where, d l d represents the interval within a group. g The spacing between groups; these two parameters can be defined by the user in the configuration panel. h is the system-defined internal canvas height.

[0139] (3-1.2.4) Expanding the storyline layout:

[0140] The canvas currently used for drawing the timeline is designated as the inner canvas, with blank areas extending a certain height at the top and bottom for drawing extended storylines.

[0141] Traverse the set of visible time points in chronological order. At any given time point, if several new storylines appear at the next time point, extend the corresponding storylines to the current time point, and sort them according to their order at the next time point. l The storylines are spaced out above the inner canvas; if several storylines existed at the previous time point but disappeared at the current time point, then the corresponding storylines are extended to the current time point, according to the order of the corresponding storylines at the previous time point, d l The storyline spacing is located below the inner canvas.

[0142] (3-1.3) Calculate the diagram layout along the storyline:

[0143] Traverse the set of visible time points in chronological order. At any given time point, the subgraph to be laid out includes all visible relationships that are newly appearing at that time point or disappearing at the next time point, as well as the associated entities of these relationships. After moving the positions of entities or relationships several times, the graph layout on the storyline at that time point is obtained while minimizing the objective function under the constraints.

[0144] in,

[0145] The constraints are: the entities or relationships to be laid out must fall on the y-axis at the corresponding position of their storyline at that point in time, and fall within a defined width on the x-axis. The defined width is the width of each subplot, which can be set by the user in the configuration panel.

[0146] Because each relationship corresponds to a line segment from the subject's position to the relationship's position and a line segment from the relationship's position to the object's position.

[0147] The objective function is:

[0148] The sum of the number of intersections between the two line segments corresponding to the relationship to be laid out and the two line segments corresponding to other relationships, as well as the number of intersections between the bounding boxes corresponding to other entities or relationships to be laid out.

[0149] (3-1.4) Calculate the static diagram layout

[0150] The subgraphs to be laid out in the static graph include all static relationships in the visible relationship set and the associated entities of these relationships. The positions of these relationships on the y-axis are not constrained. If an entity is static, its position on the y-axis is not constrained; otherwise, the entity falls on the y-axis position corresponding to its storyline at the manipulation time point. If the corresponding storyline does not exist at the manipulation time point, the entity is placed above or below the internal canvas, depending on whether the corresponding storyline has not appeared or has disappeared. Other constraints and optimization objectives are the same as in step (3-1.3), and the static graph layout is calculated using the method in step (3-1.3).

[0151] The descriptive text generation method described in step three is as follows:

[0152] (3-2.1) Preprocessing: Organize and supplement each selected set to obtain data such as the start time point of text generation, the end time point of text generation, the set of entities generated by text generation, and the set of relations generated by text generation. If the data is insufficient to generate text, then the text generation will end.

[0153] For example, the preprocessing flow in this embodiment is as follows: Record the earliest and latest time points in the selected time point set as the start time point and end time point, respectively. If the selected time point set is empty, record the earliest and latest time points in the dataset as the start time point and end time point, respectively. Initialize the text generation entity set using the selected entity set and the text generation relation set using the selected relation set. If the text generation relation set is empty, add the association relationships of all entities in the text generation entity set to the text generation relation set. Add the subjects and objects of all relations in the text generation relation set to the text generation entity set. If the text generation entity set is empty, do not generate text and end step three.

[0154] (3-2.2) Serialization: Based on time information, graph topology, and user operation sequence, sort the entities and relations in the text generation entity set and text generation relation set to obtain an ordered list of entities and entity associations, so that the final generated text is orderly, organized and consistent with the user's intent.

[0155] The serialization algorithm in this embodiment is as follows: Figure 3 As shown:

[0156] (a) Calculate the priority of entities and relations regardless of time sequence:

[0157] For each entity in the text-generated entity set, its weight is: centrality - selection order / size of the text-generated entity set. The entity with higher weight has higher priority. If the weights are the same, the entity selected earlier has higher priority.

[0158] For each type of relation in the text generation relation set, those with fewer relations of the same type in the text generation relation set have higher priority.

[0159] For each relation in the text-generated relation set, the earlier relations are selected, the higher their priority.

[0160] For example, (Faxi Temple, located in Hangzhou) and (West Lake, located in Hangzhou) belong to the same type of relationship.

[0161] (b) Divide the entities in the text-generated entity set into several clusters:

[0162] Each non-static entity is an independent cluster; two static entities associated by a static relationship are grouped into the same cluster. The entity with the highest priority in each cluster is the root entity of that cluster.

[0163] (c) Calculate the time point set and divide the non-static entities and non-static relationships into several time point buckets:

[0164] List the time points of all entities and relationships associated with the text-generated entity set and text-generated relation set, retaining those time points within the time span formed by the start and end times of text generation. Using these time points as buckets, group the non-static entities and non-static relationships in the text-generated entity set and text-generated relation set into the buckets associated with them. If multiple time points are associated, group them into the bucket with the earlier time sequence.

[0165] (d) Process each bucket in chronological order:

[0166] (d.1) Further divide the entities and relations within the bucket into several entity buckets:

[0167] Entities are assigned to their corresponding buckets; relations are attached to the root entity of the cluster in which their subject and object belong, and are assigned to the bucket corresponding to the root entity of the cluster. If the root entity does not exist, a new bucket is created for the corresponding entity.

[0168] (d.2) Process each entity bucket in order of priority according to the corresponding entity:

[0169] If the entity corresponding to the current entity bucket is not available, then jump to (d.2.4); if the entity corresponding to the current entity bucket is available, determine whether the entity corresponding to the current entity bucket is in the bucket at the current time point. If it is, then jump to (d.2.2); if it is not, then jump to (d.2.1).

[0170] The entity that can be attached to the current entity bucket refers to the entity being in the bucket at the current time point or the entity being an unvisited static entity.

[0171] (d.2.0) Handling Relationships:

[0172] Given an entity and several relations, the relations are first grouped according to their relation categories. Relations of different categories with the same related entity are merged into another group. The groups are sorted by category priority between groups and by relation priority within groups. Finally, the tuple consisting of the given entity and relation sequence is added to the serialization result, and the given relation is marked as visited.

[0173] (d.2.1) Handling non-static entity buckets:

[0174] Non-static entity buckets include the corresponding entities and the relationships to be processed within the bucket.

[0175] The relationships to be processed include several relationships belonging to the current entity bucket, several static relationships with the corresponding entity of the current entity bucket as the subject, and several static relationships with the corresponding entity of the current entity bucket as the object and the subject as a static entity. The relationships to be processed are handled using method (d.2.0).

[0176] Extended entities are other entities associated with the relationship to be processed, and these entities are unaccessed static entities, and the root entity of their cluster cannot be the bucket of the entity to be processed; jump to (d.2.3) to process the extended entities.

[0177] (d.2.2) Handling static entity buckets:

[0178] If a bucket contains relations and all relations depend on the same entity, or if a bucket contains no relations and a unique entity exists, then that entity is designated as the entry entity. Otherwise, the entity corresponding to the current entity's bucket is designated as the entry entity. The cluster to which the entity corresponding to the current entity's bucket belongs is designated as the current cluster.

[0179] ① Add the entry entity to the candidate set, and add the remaining entities of the current cluster to the residual set.

[0180] ② Select the entity with the highest priority from the candidate set, and record that the entity has been visited. The relationships to be processed are the relationships that are related to the entity but have not been visited and are in the current entity bucket, and the static relationships that are related to the entity and another entity in the current cluster but have not been visited. Use the method in (d.2.0) to process the above relationships to be processed.

[0181] ③ Remove entities that are associated with the above-mentioned relationships to be processed and are in the residual set from the residual set and add them to the candidate set.

[0182] ④ Repeat steps ② and ③ until the candidate set is empty.

[0183] At this point, the extended entity is another entity associated with a relationship within the entity bucket, and these entities are unaccessed static entities, whose root entity cannot be the entity bucket to be processed. Jump to (d.2.3) to process the extended entity.

[0184] (d.2.3) Handling extended entities:

[0185] Expand the entity into several static entities to be processed. Entities belonging to the same cluster are grouped into the same bucket, and then further divided into several static entity buckets. Process each static entity bucket in sequence according to the priority of the corresponding entity, using the same processing method as (d.2.2). End the processing flow of the current entity bucket.

[0186] (d.2.4) Handling unattached physical buckets:

[0187] For several relations within an entity bucket, if the root entity of the cluster corresponding to another entity associated with the relation can be attached, then the relation is assigned to the corresponding entity bucket. Even if there is no corresponding entity bucket, a new corresponding entity bucket is added and added to the sequence of entity buckets to be processed, so that all relations can be assigned to the corresponding entity bucket.

[0188] For the remaining relations not yet assigned to other entity buckets, first group them by the other entity they are associated with, then sort them by priority. Finally, process each group of relations sequentially using method (d.2.0). End the current entity bucket processing flow.

[0189] (3-2.3) Template filling: Use a given template and combination rules to convert the serialization result into descriptive text.

[0190] Templates and combination rules need to be customized based on the dataset.

[0191] For example, for the relation (A, Professor of Aeronautical Engineering, B, ST-ED), the corresponding template is "A is Professor of Aeronautical Engineering of B from ST to ED"; for the serialization result

[0192]

[0193]

[0194] The template and combination rules can be used to generate descriptive text such as "Neil Alden Armstrong, an American, was a professor of aeronautical engineering at the University of Cincinnati from 1971 to 1979."

[0195] It will be understood by those skilled in the art that the above descriptions are merely preferred examples of the invention and are not intended to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art can still modify the technical solutions described in the foregoing examples or make equivalent substitutions for some of the technical features. All modifications and equivalent substitutions made within the spirit and principles of the invention should be included within the scope of protection of the invention.

Claims

1. A visualization analysis system for time-series knowledge graphs, characterized in that, The system includes: The overview generation module generates a dataset overview based on the overview configuration data. The storyline generation module generates storylines based on storyline configuration data. The text generation module generates descriptive text based on storyline configuration data; The canvas module displays the system-generated overview, storyline, and text, and responds to user interactions by updating the overview and storyline configuration data; it consists of a configuration panel, an overview panel, and a storyline panel. The configuration panel is used to receive user modifications to the overview configuration data and storyline configuration data. The overview panel is used to display an overview view, receive entities selected by the user, and initialize storyline configuration data; The storyline panel is used to display the storyline view and receive user interactions with entities, relationships, and time points.

2. The visualization analysis system for time-series knowledge graphs according to claim 1, characterized in that, The storyline panel is divided into a timeline, a static section, and a chronological section. The static section is used to display static relationships, while the chronological section is used to display chronological relationships and event relationships.

3. A visualization analysis method for time-series knowledge graphs, characterized in that, This method is implemented based on the visualization analysis system of any one of claims 1 to 2, and the method includes: The system generates an overview view based on the overview configuration data input by the user and displays it to the user; The system initializes storyline configuration data based on the entity selected by the user in the overview view; then it generates a storyline view and descriptive text based on the storyline configuration data and displays them to the user. The system updates the storyline configuration data based on the user's interactions with entities, relationships, and time points in the storyline view, and then generates a storyline view and descriptive text based on the storyline configuration data, which are then displayed to the user.

4. The visualization analysis method for time-series knowledge graphs according to claim 3, characterized in that, The overview configuration data includes time span segmentation method, entity encoding method, and area map encoding method; The storyline configuration data includes monitoring status flags, selected entity sets, monitored entity sets, visible entity sets, selected relationship sets, visible relationship sets, selected time point sets, visible time point sets, and manipulated time points.

5. The visualization analysis method for time-series knowledge graphs according to claim 3, characterized in that, The system generates an overview view based on the overview configuration data input by the user, specifically including: First, the overall time span of the dataset is segmented, and the segmented time periods are mapped onto the y-axis. Then, the information within each time period is encoded by area, and the encoded values ​​are mapped onto the x-axis to create an area map. Finally, the entities existing within each time period are encoded, with the encoded values ​​mapped to text size and the entity category mapped to text color. A word cloud is then drawn within the corresponding time period of the area map.

6. The visualization analysis method for time-series knowledge graphs according to claim 3, characterized in that, The specific sub-steps for generating a storyline view based on storyline configuration data are as follows: (1) Calculate the visible set; ① Initialize the visible entity set with the selected entity set, the visible relation set with the selected relation set, and the visible time point set with the selected time point set; ② If currently under monitoring, add the start time, start time minus unit time, end time, and end time plus unit time of all entities in the monitored entity set and their associated non-static relationships to the visible time point set; add all entities in the monitored entity set and entities reachable within the expansion step size to the visible entity set; add the pairwise relationships between all entities in the visible entity set to the visible relationship set. ③ Add the subjects and objects of all relations in the selected relation set to the visible entity set; (2) Calculate the storyline; Calculate the line order of the entity storylines; calculate the line order of all storylines; calculate the storyline layout; expand the storyline layout; (3) Calculate the diagram layout along the storyline; Traverse the set of visible time points in chronological order. At any given time point, the subgraph to be laid out includes all visible relationships that are newly appearing at that time point or disappearing at the next time point, as well as the associated entities of these relationships. After moving the positions of entities or relationships several times, the graph layout on the storyline at that time point is obtained while minimizing the objective function under the constraints. Each relationship corresponds to a line segment from the subject position to the relationship position and a line segment from the relationship position to the object position. The constraint is that the entity or relationship to be laid out must fall on the corresponding position of its storyline at that time point on the y-axis and fall within a limited width on the x-axis. The limited width is the width of each subgraph. The objective function is the sum of the number of intersections between the two line segments corresponding to the relationship to be laid out and the two line segments corresponding to other relationships, as well as the number of intersections between the bounding boxes corresponding to other entities or relationships to be laid out. (4) Calculate the static diagram layout The subgraph to be laid out in the static graph includes all static relations in the visible relation set and the associated entities of these relations. There are no constraints on the position of these relations on the y-axis. If an entity is a static entity, there are no constraints on the position of the entity on the y-axis. Otherwise, the entity falls on the y-axis position corresponding to its storyline at the manipulation time point. If the corresponding storyline does not exist at the manipulation time point, the entity is placed above or below the inner canvas depending on whether the corresponding storyline has not appeared or has disappeared. The other constraints and optimization objective function are the same as the graph layout calculated on the storyline.

7. The visualization analysis method for time-series knowledge graphs according to claim 3, characterized in that, The sub-steps for generating descriptive text based on storyline configuration data are as follows: (1) Preprocessing: Organize and supplement each selected set to obtain the text generation start time point, text generation end time point, text generation entity set, and text generation relation set. If the data is insufficient to generate text, the text generation will end. (2) Serialization: Based on time information, graph topology, and user operation sequence, the entities and relations in the text generation entity set and text generation relation set are sorted to obtain an ordered list of entities and entity relationships, so that the final generated text is orderly, organized, and consistent with the user's intent. (3) Template filling: Use the given template and combination rules to convert the serialization result into descriptive text.

8. The visualization analysis method for time-series knowledge graphs according to claim 7, characterized in that, The specific sub-steps of the serialization process are as follows: (a) Calculate the priority of entities, relations, and time sequence; For each entity in the text-generated entity set, its weight is (centrality - selection order / size of the text-generated entity set). The entity with higher weight has higher priority. If the weights are the same, the entity selected earlier has higher priority. For each type of relation in the text generation relation set, those with fewer relations of the same type in the text generation relation set have higher priority. For each relation in the text generation relation set, the earlier relations are selected with higher priority. (b) Divide the entities in the text-generated entity set into several clusters; Each non-static entity is an independent cluster; two static entities that are associated by a static relationship are grouped into the same cluster; the entity with the highest priority in each cluster is the root entity of that cluster; (c) Calculate the time point set and divide the non-static entities and non-static relationships into several time point buckets; List all the time points in the text generation entity set and text generation relation set that are associated with each entity and relation within the time span formed by the start time and end time of text generation; use these time points as buckets to classify the non-static entities and non-static relations in the text generation entity set and text generation relation set into the buckets associated with them, and if multiple time points are associated, classify them into the bucket with the earlier time sequence. (d) Process each bucket in chronological order; The entities and relationships within each bucket are further divided into several entity buckets; Each entity bucket is processed sequentially according to its corresponding entity priority.

9. The visualization analysis method for time-series knowledge graphs according to claim 8, characterized in that, The specific sub-steps for further dividing the entities and relationships within a bucket into several entity buckets are as follows: Entities are assigned to their corresponding buckets; relations are attached to the root entity of the cluster in which their subject and object belong, and are assigned to the bucket corresponding to the root entity of the cluster. If the root entity does not exist, a new bucket is created for the corresponding entity.

10. The visualization analysis method for time-series knowledge graphs according to claim 8, characterized in that, The sub-steps for processing each entity bucket according to its priority are as follows: If the entity corresponding to the current entity bucket is not dependent, then jump to (d.2.4) to process the entity bucket that is not dependent; otherwise, if the entity corresponding to the current entity bucket is in the bucket at the current time point, then jump to (d.2.2) to process the static entity bucket; otherwise, jump to (d.2.1) to process the non-static entity bucket. The entity that can be attached to the current entity bucket refers to the entity being in the bucket at the current time point or the entity being an unvisited static entity. (d.2.0) The process for handling relationships is as follows: Given an entity and several relations, first group the relations according to their categories. Then, merge different categories of relations with the same related entity into another group. Sort between groups according to category priority and within groups according to the priority of the relations themselves. Finally, add the tuple consisting of the given entity and relation sequence to the serialization result and mark the given relation as visited. (d.2.1) The process for handling non-static entity buckets is as follows: The relationships to be processed are: several relationships belonging to the current entity bucket, several static relationships with the entity corresponding to the current entity bucket as the subject, and several static relationships with the entity corresponding to the current entity bucket as the object and the subject as a static entity. Use the method for handling relationships in (d.2.0) to process the relationship to be processed; the extended entity is another entity associated with the relationship to be processed, and these entities are unvisited static entities, and the root entity of their cluster cannot be the bucket of the entity to be processed; jump to (d.2.3) to process the extended entity; (d.2.2) The process for handling static entity buckets is as follows: If a relationship exists within a bucket and all relationships are attached to the same entity, or if no relationship exists within a bucket and a unique entity exists, then that entity is designated as the entry entity; otherwise, the entity corresponding to the current entity bucket is designated as the entry entity. The cluster to which the entity corresponding to the current entity bucket belongs is designated as the current cluster. The entry entity is added to the candidate set, and the remaining entities in the current cluster are added to the residual set. The entity with the highest priority is selected from the candidate set, and it is marked as visited. The relationships to be processed are those that are associated with the entity but have not been visited and are in the current entity bucket, as well as the static relationships that are associated with the entity and another entity in the current cluster but have not been visited. The above relationships to be processed are processed using the relationship processing method in (d.2.0). Entities associated with the above relationships to be processed and in the residual set are removed from the residual set and added to the candidate set. The above operations are repeated until the candidate set is empty. The extended entities are other entities associated with relationships within the entity bucket, and these entities are unvisited static entities, and the root entity of their cluster cannot be the entity bucket to be processed. Jump to (d.2.3) to process the extended entities. (d.2.3) The process for handling extended static entities is as follows: For several static entities to be processed, entities belonging to the same cluster are grouped into the same bucket, and then further divided into several static entity buckets; each static entity bucket is processed sequentially according to the priority of the corresponding entity in step (d.2.2); (d.2.4) The process for handling unattached physical buckets is as follows: For several relations within an entity bucket, if the root entity of the cluster to which the relation is associated can be attached, the relation is assigned to the corresponding entity bucket. If the corresponding entity bucket does not exist, a new entity bucket is created and added to the sequence of entity buckets to be processed. For the remaining relations that have not been assigned to other entity buckets, they are first grouped according to the entities to which the relation is attached, and then sorted between groups according to the priority of the entities. Then, the method for processing relations in (d.2.0) is used to process each group of relations in turn. Finally, the current entity bucket processing flow ends.

Citation Information

Patent Citations

  • A method for generating storyline visual layout

    CN109068152A

  • Time sequence visualization development method and system based on knowledge graph

    CN114036311A

  • Visualizing method for knowledge genealogy

    CN102779143A

  • Systems, methods, and computer readable media for visualization of semantic information and inference of temporal signals indicating salient associations between life science entities

    US20180082197A1