Space-time correlation degree measurement method and system, storage medium and device

By employing a rule-based and dependency-syntax-based spatiotemporal correlation metric calculation method, the problem of underutilization of spatiotemporal information in existing technologies is solved, enabling more accurate spatiotemporal knowledge representation and deep integration of entity information, thereby enhancing the application capabilities of geographic knowledge.

CN117708269BActive Publication Date: 2026-04-14CHINA THREE GORGES UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA THREE GORGES UNIV
Filing Date
2023-11-03
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing knowledge graphs do not fully consider the accurate retrieval and adaptive representation of spatiotemporal information in entity objects, lack a unified organizational expression framework and a reliable structured acquisition scheme, resulting in independent spatiotemporal distribution and limited relevance and interaction, which hinders the application of geographical knowledge and the progress of social services.

Method used

A spatiotemporal correlation metric calculation method based on rules and dependency syntax is adopted. This method involves preprocessing open-domain text sequences, performing syntactic analysis by formulating a rule matching table, extracting candidate triples and spatiotemporal information, calculating the spatiotemporal correlation value, and generating a spatiotemporal knowledge tuple expression model.

Benefits of technology

By comprehensively considering the context and spatiotemporal information of the text, accurately calculating spatiotemporal correlation, obtaining a more accurate spatiotemporal knowledge representation model, deeply integrating textual semantic information, and improving the integration and association capabilities of entity information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117708269B_ABST
    Figure CN117708269B_ABST
Patent Text Reader

Abstract

The application discloses a rule and dependency grammar-based spatiotemporal correlation degree calculation method, which comprises the following steps: obtaining an open field text sequence, and preprocessing the text sequence to obtain a preprocessed text sequence; performing syntactic analysis on the preprocessed text sequence, formulating a rule matching table, and matching the syntactic analysis result according to the rule matching table to obtain candidate triplets and spatiotemporal information; calculating a spatiotemporal correlation degree value according to a dependency path of a relation predicate in the candidate triplets and the spatiotemporal information, a dependency distance of a head entity and the spatiotemporal information, a spatiotemporal feature of head and tail entity types, and a spatiotemporal semantic feature of the relation predicate; defining two threshold values, comparing the spatiotemporal correlation degree value with the two threshold values, determining a classification result of the spatiotemporal correlation degree, and generating a knowledge expression model. The application accurately calculates spatiotemporal correlation by comprehensively considering the context and spatiotemporal information of the text, thereby obtaining a spatiotemporal knowledge expression model and more accurately conveying the semantics of the text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to open-domain text processing technology, and more particularly to spatiotemporal correlation measurement methods, systems, storage media, and devices. Background Technology

[0002] With the rapid development of Earth observation, navigation, information, and communication technologies, spatiotemporal data is growing exponentially, leading to a growing phenomenon known as "data overload." Simultaneously, related challenges, such as "information overload" and difficulties in knowledge retrieval, are becoming increasingly apparent, as the sheer volume of data makes the extraction of valuable knowledge challenging. This situation poses a significant challenge to traditional geographic information analysis and service frameworks.

[0003] Existing knowledge graphs, or geographical knowledge graphs, typically treat spatiotemporal information as entity attributes, failing to fully consider its crucial role in accurate retrieval and adaptive representation of entity objects. Simultaneously, current text-based information mining methods are underutilized, lacking a unified organizational framework and reliable structured acquisition schemes. Furthermore, spatiotemporal distribution remains relatively independent and dispersed, with limited direct relevance and interaction, and a lack of effective means to integrate and associate entity information. These challenges directly hinder the application of geographical knowledge and limit the progress of social services. Summary of the Invention

[0004] To address the aforementioned problems, this invention provides a method for calculating spatiotemporal correlation metrics based on rules and dependency syntax, comprising the following steps:

[0005] S1. Obtain the open-domain text sequence and preprocess the text sequence to obtain the preprocessed text sequence;

[0006] S2. Formulate a rule matching table, perform syntactic analysis on the preprocessed text sequence, and match the syntactic analysis results according to the rule matching table to obtain candidate triples and spatiotemporal information;

[0007] S3. Based on the candidate triples and spatiotemporal information, obtain the dependency path value between the relation predicate and the spatiotemporal information, the dependency distance between the head entity and the spatiotemporal information, the spatiotemporal feature value of the head and tail entity types, and the spatiotemporal semantic feature value of the relation predicate in the candidate triples.

[0008] S4. Calculate the spatiotemporal correlation value based on the dependency path value between the relational predicate and spatiotemporal information, the dependency distance between the head entity and spatiotemporal information, the spatiotemporal feature value of the head and tail entity types, and the spatiotemporal semantic feature value of the relational predicate in the candidate triples.

[0009] S5. Define two thresholds, compare the spatiotemporal correlation value with the two thresholds, determine the classification result of spatiotemporal correlation, and generate a spatiotemporal knowledge tuple expression model.

[0010] Furthermore, step S2 specifically includes:

[0011] S21. Based on Chinese syntax rules, define syntax rules and form a rule matching table;

[0012] The preprocessed text sequence was subjected to syntactic analysis using the LTP natural language processing tool to obtain syntactic analysis results, which included: syntactic structure and syntactic parse tree.

[0013] S22. Match the syntactic analysis results according to the rule matching table to obtain candidate triples and spatiotemporal information.

[0014] Further, in step S3, the dependency path values ​​between the relational predicates and spatiotemporal information in the candidate triples are obtained as follows:

[0015] Obtain the node positions of the relational predicate and spatiotemporal information in the parse tree of the candidate triplet, calculate the dependency path length between the two node positions (relational predicate node and spatiotemporal information node) and the number of path branches in the parse tree, and obtain the dependency path value between the relational predicate and spatiotemporal information. The expression is:

[0016]

[0017] Where PredSTDep represents the dependency path value between the relational predicate and spatiotemporal information, and Ways Pred-TMP|LOC Ways represents the length of the dependency path between two nodes. Sent This represents the number of path branches in the syntactic parsing tree.

[0018] Furthermore, in step S3, the dependency distance between the head entity and the spatiotemporal information is specifically obtained as follows:

[0019] To obtain the word segmentation distance between the head entity and spatiotemporal information in the text sequence, as well as the total number of words in the text sequence, the dependency distance between the head entity and spatiotemporal information is calculated using the following expression:

[0020]

[0021] Wherein, SubSTDis represents the dependency distance between the head entity and spatiotemporal information, and Words SubEntity-TMP|LOC Words represent the word segmentation spacing between the head entity and spatiotemporal elements in the text sequence. Sent This indicates the total number of words in the text sequence.

[0022] Furthermore, in step S3, the spatiotemporal characteristic values ​​of the head and tail entity types are obtained as follows:

[0023] Construct a spatiotemporal correlation scoring table for entity types, map the types of the head and tail entities in the candidate triples to the spatiotemporal correlation scoring table for entity types, and determine the spatiotemporal characteristic values ​​of the head and tail entity types.

[0024] The spatiotemporal correlation of different entity types is quantitatively evaluated using an expert scoring method, resulting in a spatiotemporal correlation value for each entity type, expressed as follows:

[0025]

[0026] Wherein, SubtypeScore represents the spatiotemporal correlation value of entity type, Score(i) represents the score of expert i on the spatiotemporal correlation of entity type, and n represents the total number of experts;

[0027] Download the knowledge graph knowledge base of Peking University Chinese Encyclopedia, obtain the entity-type triple table, and perform concept matching between the head and tail entities of the candidate triples and the entity types in the entity-type triple table to determine the head and tail entity types.

[0028] By mapping the head and tail entity types to the spatiotemporal correlation values ​​of entity types as quantitatively assessed by experts, the spatiotemporal characteristic values ​​of the head and tail entity types are obtained.

[0029] Furthermore, in step S3, the spatiotemporal semantic feature values ​​of the relational predicate are obtained as follows:

[0030] Relational predicates and spatiotemporal text were obtained from the Baidu Encyclopedia corpus using natural language processing tools;

[0031] The word vectors of each relational predicate obtained from the Baidu Encyclopedia corpus are obtained using the Word2Vec model, and the word vectors are clustered into different classes using the Kmeans algorithm.

[0032] Obtain the total number of spatiotemporal statements and the total number of text statements for each category of relational predicates, calculate the spatiotemporal proportion of each category, and obtain the spatiotemporal semantic knowledge base of relational predicates. The expression is as follows:

[0033]

[0034] Where PredScore represents the spatiotemporal semantic value of a certain type of relational predicate, and Sum(a i Sum(b) represents the total number of single category relation predicates in the spatiotemporal text. j ) represents the total number of single category relation predicates in all texts, n represents the total number of spatiotemporal text statements in the text corpus, and m represents the total number of text statements in the text corpus;

[0035] The relation predicates of candidate triples are obtained, and the similarity calculation method based on Word2Vec is used to compare them with the relation predicates in the spatiotemporal semantic knowledge base of relation predicates. The class of the relation predicate with the highest matching degree is found, and the spatiotemporal semantic score of the relation predicate corresponding to the class is the spatiotemporal semantic feature value of the relation predicate of the candidate triple.

[0036] Furthermore, step S4 specifically includes:

[0037] Based on the dependency path values ​​between the relational predicate and spatiotemporal information, the dependency distance between the head entity and spatiotemporal information, the spatiotemporal feature values ​​of the head and tail entity types, and the spatiotemporal semantic feature values ​​of the relational predicate in the candidate triples, the spatiotemporal correlation value is obtained, expressed as:

[0038] S overall =W a ×PredSTDep+W b ×SubSTDis+W c ×SubtypeScore+W d

[0039] ×PredScore

[0040] Where S overall W represents the spatiotemporal correlation value. a The weight of PredSTDep, where PredSTDep represents the dependency path value between the relational predicate and spatiotemporal information, W b The weight of SubSTDis, where SubSTDis represents the dependency distance between the header entity and spatiotemporal information, W c The weight of SubtypeScore, where SubtypeScore represents the spatiotemporal characteristic value of the head and tail entity types, is W. d The weight of PredScore is given by PredScore, which represents the spatiotemporal semantic feature value of the candidate triple relation predicate.

[0041] This invention also proposes a spatiotemporal correlation measurement calculation system based on rules and dependency syntax, including a data acquisition unit, a rule matching unit, a spatiotemporal correlation degree value calculation unit, and a spatiotemporal knowledge tuple expression model generation unit;

[0042] The data acquisition unit is used to acquire open-domain text sequences and preprocess the text sequences to obtain preprocessed text sequences.

[0043] The rule matching unit is used to formulate a rule matching table, perform syntactic analysis on the preprocessed text sequence, and match the syntactic analysis results according to the rule matching table to obtain candidate triples and spatiotemporal information.

[0044] The spatiotemporal correlation value calculation unit is used to obtain the dependency path value between the relation predicate and the spatiotemporal information, the dependency distance between the head entity and the spatiotemporal information, the spatiotemporal feature value of the head and tail entity types, and the spatiotemporal semantic feature value of the relation predicate in the candidate triples based on the candidate triples and spatiotemporal information.

[0045] The spatiotemporal correlation value is calculated based on the dependency path value between the relational predicate and spatiotemporal information, the dependency distance between the head entity and spatiotemporal information, the spatiotemporal feature value of the head and tail entity types, and the spatiotemporal semantic feature value of the relational predicate in the candidate triples.

[0046] The spatiotemporal knowledge tuple expression model generation unit is used to compare the spatiotemporal correlation value with two defined thresholds to determine the classification result of the spatiotemporal correlation and generate the spatiotemporal knowledge tuple expression model.

[0047] The present invention also proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described spatiotemporal correlation metric calculation method.

[0048] The present invention also proposes an electronic device, including a processor and a memory, wherein the processor and the memory are interconnected, wherein the memory is used to store a computer program, the computer program including computer-readable instructions, and the processor is configured to invoke the computer-readable instructions to execute the above-described spatiotemporal correlation metric calculation method.

[0049] The beneficial effects of the technical solution provided by this invention are:

[0050] This invention, based on the inherent characteristics of spatiotemporal knowledge, constructs an open-domain text corpus containing spatiotemporal information through statistical analysis of sentence structure using rules and dependency grammar. It employs dependency syntax analysis to determine the semantic relationships between entity relations and spatiotemporal information, and establishes a set of rules combining knowledge types and syntactic analysis to extract entity relations, deeply integrating textual semantic information. The model comprehensively evaluates the text content and its features such as entities and relational predicates, and determines the degree of spatiotemporal relevance by setting appropriate thresholds. This method, by comprehensively considering the text's context and spatiotemporal information, accurately calculates spatiotemporal relevance, thereby obtaining a spatiotemporal knowledge representation model that more accurately conveys the text's semantics. Attached Figure Description

[0051] Figure 1 This is a flowchart of a spatiotemporal correlation metric calculation method based on rules and dependency syntax, according to an embodiment of the present invention.

[0052] Figure 2 This is a block diagram of an electronic device according to an exemplary embodiment of the present invention. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0054] A flowchart of a spatiotemporal correlation metric calculation method based on rules and dependency syntax according to an embodiment of the present invention is shown below. Figure 1 Specifically, it includes the following steps:

[0055] S1. Obtain the open-domain text sequence and preprocess the text sequence to obtain the preprocessed text sequence.

[0056] The open-domain text sequence in this embodiment of the invention is a general open-domain text sequence, not limited by a vertical domain. Conventional preprocessing is performed on the text sequence, such as removing non-text portions, performing word segmentation, and removing stop words from the segmentation results.

[0057] S2. Formulate a rule matching table, perform syntactic analysis on the preprocessed text sequence, and match the syntactic analysis results according to the rule matching table to obtain candidate triples and spatiotemporal information.

[0058] Syntactic analysis delves into each component of the preprocessed text sequence. Each component is assigned a specific syntactic label, clarifying its grammatical function and its role within the sentence framework. To ensure accurate capture of head and tail entities, their relationships, and related spatiotemporal information within the current context, precise rules must be designed based on different sentence patterns throughout the entity relation extraction process.

[0059] Specifically:

[0060] S21. Based on Chinese syntax rules, define syntax rules and form a rule matching table;

[0061] The preprocessed text sequence was subjected to syntactic analysis using the LTP natural language processing tool to obtain syntactic analysis results, which included: syntactic structure and syntactic parse tree.

[0062] S22. Match the syntactic analysis results according to the rule matching table to obtain candidate triples and spatiotemporal information.

[0063] S3. Based on the candidate triples and spatiotemporal information, obtain the dependency path value between the relation predicate and the spatiotemporal information, the dependency distance between the head entity and the spatiotemporal information, the spatiotemporal feature value of the head and tail entity types, and the spatiotemporal semantic feature value of the relation predicate in the candidate triples.

[0064] Specifically:

[0065] (1) Obtain the node positions of the relational predicate and spatiotemporal information in the parse tree of the candidate triplet, calculate the dependency path length between the two node positions of the relational predicate node and the spatiotemporal information node, as well as the number of path branches in the parse tree, and obtain the dependency path value between the relational predicate and the spatiotemporal information. The expression is:

[0066]

[0067] Where PredSTDep represents the dependency path value between the relational predicate and spatiotemporal information, and Ways Pred-TMP|LOC Ways represents the length of the dependency path between two nodes. Sent Represents the number of path branches in the syntactic parsing tree;

[0068] (2) Obtain the word segmentation distance between the head entity and spatiotemporal information in the text sequence, as well as the total number of words in the text sequence, to obtain the dependency distance between the head entity and spatiotemporal information. The expression is:

[0069]

[0070] Wherein, SubSTDis represents the dependency distance between the head entity and spatiotemporal information, and Words SubEntity-TMP|LOC Words represent the word segmentation spacing between the head entity and spatiotemporal elements in the text sequence. Sent Indicates the total number of words in the text sequence;

[0071] (3) Construct a spatiotemporal correlation scoring table for entity types, map the types of the head and tail entities in the candidate triples to the spatiotemporal correlation scoring table of entity types, and determine the spatiotemporal characteristic values ​​of the head and tail entity types.

[0072] (3.1) The spatiotemporal correlation of different entity types is quantitatively evaluated using an expert scoring method to form the spatiotemporal correlation value of the entity type, as expressed below:

[0073]

[0074] Wherein, SubtypeScore represents the spatiotemporal correlation value of the entity type, Score(i) represents the score of expert i on the spatiotemporal correlation of the entity type, and n represents the total number of experts.

[0075] (3.2) Download the knowledge graph knowledge base of Peking University Chinese Encyclopedia, obtain the entity-type triple table, perform concept matching between the head and tail entities of the candidate triples and the entity types in the entity-type triple table, and determine the head and tail entity types.

[0076] (3.3) Map the head and tail entity types to the spatiotemporal correlation values ​​of entity types quantitatively evaluated by experts to obtain the spatiotemporal characteristic values ​​of the head and tail entity types.

[0077] (4) Obtain the spatiotemporal semantic knowledge base of relational predicates, calculate the similarity between the relational predicates of candidate triples and the relational predicates in the spatiotemporal semantic knowledge base of relational predicates, and obtain the spatiotemporal semantic feature value of relational predicates;

[0078] (4.1) Use natural language processing tools to obtain relational predicates and spatiotemporal text from the Baidu Encyclopedia corpus;

[0079] (4.2) Use the Word2Vec model to obtain the word vector of each relation predicate obtained from the Baidu Encyclopedia corpus, and use the Kmeans algorithm to cluster the word vectors and divide them into different classes;

[0080] (4.3) Obtain the total number of spatiotemporal statements and the total number of text statements for each category of relational predicates, calculate the spatiotemporal proportion of each category, and obtain the spatiotemporal semantic knowledge base of relational predicates. The expression is as follows:

[0081]

[0082] Where PredScore represents the spatiotemporal semantic value of a certain type of relational predicate, and Sum(a i Sum(b) represents the total number of single category relation predicates in the spatiotemporal text. j ) represents the total number of single category relation predicates in all texts, n represents the total number of spatiotemporal text statements in the text corpus, and m represents the total number of text statements in the text corpus.

[0083] The relation predicates of candidate triples are obtained, and the similarity calculation method based on Word2Vec is used to compare them with the relation predicates in the spatiotemporal semantic knowledge base of relation predicates. The class of the relation predicate with the highest matching degree is found, and the spatiotemporal semantic score of the relation predicate corresponding to the class is the spatiotemporal semantic feature value of the relation predicate of the candidate triple.

[0084] Step 3 involves a comprehensive analysis of the text content and the extraction of relevant features, such as head and tail entity types, relational predicate semantics, dependency distances between head entities and spatiotemporal information, and dependency paths connecting relational predicates and spatiotemporal information. By using mathematical expressions to measure the influence of spatiotemporal information on text semantics, spatiotemporal relevance can be effectively captured, and its contribution to text semantics can be assessed.

[0085] S4. Based on the dependency path values ​​between the relational predicate and spatiotemporal information, the dependency distance between the head entity and spatiotemporal information, the spatiotemporal feature values ​​of the head and tail entity types, and the spatiotemporal semantic feature values ​​of the relational predicate in the candidate triples, calculate the spatiotemporal correlation value, expressed as:

[0086] S overall =W a ×PredSTDep+Wb ×SubSTDis+W c ×SubtypeScore+W d

[0087] ×PredScore

[0088] Where S overall W represents the spatiotemporal correlation value. a The weight of PredSTDep, where PredSTDep represents the dependency path value between the relational predicate and spatiotemporal information, W b The weight of SubSTDis, where SubSTDis represents the dependency distance between the header entity and spatiotemporal information, W c The weight of SubtypeScore, where SubtypeScore represents the spatiotemporal characteristic value of the head and tail entity types, is W. d The weight of PredScore is given by PredScore, which represents the spatiotemporal semantic feature value of the candidate triple relation predicate.

[0089] S5. Define two thresholds, compare the spatiotemporal correlation value with the two thresholds, determine the classification result of spatiotemporal correlation, and generate a spatiotemporal knowledge tuple expression model.

[0090] In a further embodiment, two thresholds, K1 and K2 (0 < K1 < K2), are defined. Comparing the spatiotemporal correlation value with the values ​​of K1 and K2 determines the degree of correlation with spatiotemporal information. K1 identifies weak and moderate correlations, while K2 identifies moderate and strong correlations. Finally, based on the classification results of the spatiotemporal correlation, a knowledge representation model is generated as the final output of the model.

[0091] Depending on the degree of spatiotemporal correlation, the structured description of knowledge tuples has different expression models, mainly divided into the following five representation forms:

[0092] (1) The spatiotemporal knowledge tuple representation model that directly expresses time and space entities is as follows:

[0093] STtuple= <S,P T ,T o >

[0094] STtuple= <S,P L ,L o >

[0095] Where STtuple = <S,P T ,T o > represents a spatiotemporal knowledge tuple of a time entity, S represents the head entity, P T The verb indicating time relationship, To Indicates time, STtuple = <S,P L ,L o > Represents a spatiotemporal knowledge tuple, P L L is a spatial relational predicate. o Representation space.

[0096] (2) The spatiotemporal knowledge tuple representation model that is strongly correlated with both time and space is as follows:

[0097] STtuple= <S,P,O,(P T ,T)|T o ,(P L ,L)|L o >

[0098] Where S, P, and O represent the head entity, relational predicate, and tail entity, respectively. T ,T)|T o Indicates the relevant time, (P) L ,L)|L o Indicates the relevant space.

[0099] (3) The spatiotemporal knowledge tuple representation model that is only strongly related to time or space is as follows:

[0100] STtuple= <S,P,O,(P T ,T)|T o >

[0101] STtuple= <S,P,O,(P L ,L)|L o >

[0102] Where STtuple = <S,P,O,(P T ,T)|T o > This represents a spatiotemporal knowledge tuple representation model with strong temporal correlation, where S, P, and O represent the head entity, relational predicate, and tail entity, respectively. (P T ,T)|T o Indicates the relevant time, STtuple = <S,P,O,(P L ,L)|L o > Represents a spatiotemporal knowledge tuple representation model with strong spatial correlation, (P L ,L)|L o Indicates the relevant space.

[0103] (4) The spatiotemporal knowledge tuple representation model with spatiotemporal degree correlation is as follows:

[0104] STtuple = {<S,P,O> ,P={A T A L,...,A n}|T o ,L o}

[0105] Where, P = {A} L A L ,...,A n} is the core knowledge relation predicate, A T As a time attribute, A L An represents spatial attributes, while other attributes are represented by An. In other words, time and space are recorded as general attribute information in different tuples.

[0106] (5) Geological knowledge with weak spatiotemporal correlation is basically unrelated to spatiotemporal information, and its knowledge graph representation model is shown in the following formula:

[0107] STtuple=<S,P,O>

[0108] This invention also proposes a spatiotemporal correlation measurement calculation system based on rules and dependency syntax, including a data acquisition unit, a rule matching unit, a spatiotemporal correlation degree value calculation unit, and a spatiotemporal knowledge tuple expression model generation unit;

[0109] The data acquisition unit is used to acquire open-domain text sequences and preprocess the text sequences to obtain preprocessed text sequences.

[0110] The rule matching unit is used to formulate a rule matching table, perform syntactic analysis on the preprocessed text sequence, and match the syntactic analysis results according to the rule matching table to obtain candidate triples and spatiotemporal information.

[0111] The spatiotemporal correlation value calculation unit is used to obtain the dependency path value between the relation predicate and the spatiotemporal information, the dependency distance between the head entity and the spatiotemporal information, the spatiotemporal feature value of the head and tail entity types, and the spatiotemporal semantic feature value of the relation predicate in the candidate triples based on the candidate triples and spatiotemporal information.

[0112] The spatiotemporal correlation value is calculated based on the dependency path value between the relational predicate and spatiotemporal information, the dependency distance between the head entity and spatiotemporal information, the spatiotemporal feature value of the head and tail entity types, and the spatiotemporal semantic feature value of the relational predicate in the candidate triples.

[0113] The spatiotemporal knowledge tuple expression model generation unit is used to compare the spatiotemporal correlation value with two defined thresholds to determine the classification result of the spatiotemporal correlation and generate the spatiotemporal knowledge tuple expression model.

[0114] In one exemplary embodiment, a computer-readable storage medium is included, which stores a computer program that, when executed by a processor, implements the above-described spatiotemporal correlation metric calculation method.

[0115] Please see Figure 2 In one exemplary embodiment, the device further includes an electronic device including at least one processor, at least one memory, and at least one communication bus.

[0116] The memory stores a computer program, which includes computer-readable instructions. The processor calls the computer-readable instructions stored in the memory through the communication bus to execute the above-mentioned spatiotemporal correlation metric calculation method.

[0117] Following the specific implementation plan above, taking "As of 2021, Harbin City has jurisdiction over 9 districts" as an example, the text sequence first needs to be preprocessed, including word segmentation, to obtain the following result: "As of / v2021 / nt, / wpHarbin City / ns jurisdiction / v 9 / m units / q districts / n". Based on syntactic rules, candidate triples (Harbin City, jurisdiction, 9 districts) are generated, and the spatiotemporal information "As of 2021" is extracted. Next, feature calculations are performed on the extracted results, including calculating the dependency path value between the relational predicate "jurisdiction" and the spatiotemporal information "As of 2021", the dependency distance value between the head entity "Harbin City" and the spatiotemporal information "As of 2021", the spatiotemporal feature value of the head entity "Harbin City" and the tail entity "9 districts", and the spatiotemporal semantic feature value of the relational predicate "jurisdiction". These calculation results will be used to generate spatiotemporal correlation values. Finally, this spatiotemporal correlation value is compared with pre-set thresholds K1 and K2 to determine the degree of correlation between the candidate triples (Harbin City, jurisdiction, 9 municipal districts) in the statement "As of 2021, Harbin City has jurisdiction over 9 municipal districts" and the spatiotemporal information "As of 2021". Based on the final spatiotemporal correlation value, a spatiotemporal knowledge tuple representation model can be generated to describe the relationship between the candidate triples (Harbin City, jurisdiction, 9 municipal districts) and the spatiotemporal information "As of 2021".

[0118] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0119] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In the unit claims listing several devices, several of these devices may be embodied by the same hardware item. The use of the terms first, second, and third, etc., does not indicate any order and can be interpreted as identifiers.

[0120] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for calculating spatiotemporal correlation metrics based on rules and dependency syntax, characterized in that, Includes the following steps: S1. Obtain the open-domain text sequence and preprocess the text sequence to obtain the preprocessed text sequence; S2. Formulate a rule matching table, perform syntactic analysis on the preprocessed text sequence, and match the syntactic analysis results according to the rule matching table to obtain candidate triples and spatiotemporal information; S3. Based on the candidate triples and spatiotemporal information, obtain the dependency path value between the relation predicate and the spatiotemporal information, the dependency distance between the head entity and the spatiotemporal information, the spatiotemporal feature value of the head and tail entity types, and the spatiotemporal semantic feature value of the relation predicate in the candidate triples. S4. Calculate the spatiotemporal correlation value based on the dependency path value between the relational predicate and spatiotemporal information, the dependency distance between the head entity and spatiotemporal information, the spatiotemporal feature value of the head and tail entity types, and the spatiotemporal semantic feature value of the relational predicate in the candidate triples. S5. Define two thresholds, compare the spatiotemporal correlation value with the two thresholds, determine the classification result of spatiotemporal correlation, and generate a spatiotemporal knowledge tuple expression model.

2. The spatiotemporal correlation metric calculation method based on rules and dependency syntax according to claim 1, characterized in that, Step S2 is as follows: S21. Based on Chinese syntax rules, define syntax rules and form a rule matching table; The preprocessed text sequence was subjected to syntactic analysis using the LTP natural language processing tool to obtain syntactic analysis results, which included: syntactic structure and syntactic parse tree. S22. Match the syntactic analysis results according to the rule matching table to obtain candidate triples and spatiotemporal information.

3. The spatiotemporal correlation metric calculation method based on rules and dependency syntax according to claim 2, characterized in that, In step S3, the dependency path values ​​between the relational predicates and spatiotemporal information in the candidate triples are obtained as follows: Obtain the node positions of the relational predicate and spatiotemporal information in the parse tree of the candidate triplet, calculate the dependency path length between the two node positions (relational predicate node and spatiotemporal information node) and the number of path branches in the parse tree, and obtain the dependency path value between the relational predicate and spatiotemporal information. The expression is: in, This represents the dependency path value between the relational predicate and spatiotemporal information. This indicates the length of the dependency path between two nodes. This represents the number of path branches in the syntactic parsing tree.

4. The spatiotemporal correlation metric calculation method based on rules and dependency syntax according to claim 1, characterized in that, In step S3, the dependency distance between the head entity and the spatiotemporal information is specifically obtained as follows: To obtain the word segmentation distance between the head entity and spatiotemporal information in the text sequence, as well as the total number of words in the text sequence, the dependency distance between the head entity and spatiotemporal information is calculated using the following expression: in, This indicates the dependency distance between the head entity and spatiotemporal information. This indicates the word segmentation spacing between the head entity and the spatiotemporal context in the text sequence. This indicates the total number of words in the text sequence.

5. The spatiotemporal correlation metric calculation method based on rules and dependency syntax according to claim 1, characterized in that, In step S3, the spatiotemporal characteristic values ​​of the head and tail entity types are obtained as follows: Construct a spatiotemporal correlation scoring table for entity types, map the types of the head and tail entities in the candidate triples to the spatiotemporal correlation scoring table for entity types, and determine the spatiotemporal characteristic values ​​of the head and tail entity types. The spatiotemporal correlation of different entity types is quantitatively evaluated using an expert scoring method, resulting in a spatiotemporal correlation value for each entity type, expressed as follows: in, The spatiotemporal association value representing the entity type, This represents the score given by expert i for the spatiotemporal relevance of entity types. Indicates the total number of experts; Download the knowledge graph knowledge base of Peking University Chinese Encyclopedia, obtain the entity-type triple table, and perform concept matching between the head and tail entities of the candidate triples and the entity types in the entity-type triple table to determine the head and tail entity types. By mapping the head and tail entity types to the spatiotemporal correlation values ​​of entity types as quantitatively assessed by experts, the spatiotemporal characteristic values ​​of the head and tail entity types are obtained.

6. The spatiotemporal correlation metric calculation method based on rules and dependency syntax according to claim 1, characterized in that, In step S3, the spatiotemporal semantic feature values ​​of the relational predicate are obtained as follows: Relational predicates and spatiotemporal text were obtained from the Baidu Encyclopedia corpus using natural language processing tools; The word vectors of each relational predicate obtained from the Baidu Encyclopedia corpus are obtained using the Word2Vec model, and the word vectors are clustered into different classes using the Kmeans algorithm. Obtain the total number of spatiotemporal statements and the total number of text statements for each category of relational predicates, calculate the spatiotemporal proportion of each category, and obtain the spatiotemporal semantic knowledge base of relational predicates. The spatiotemporal semantic value expression of relational predicates is as follows: in, Represents the spatiotemporal semantic value of a certain type of relational predicate. This represents the total number of single category relation predicates in the spatiotemporal text. This represents the total number of single category relation predicates in all texts. This indicates the total number of spatiotemporal text statements in the text corpus. This indicates the total number of text sentences in the text corpus; The relation predicates of candidate triples are obtained, and the similarity calculation method based on Word2Vec is used to compare them with the relation predicates in the spatiotemporal semantic knowledge base of relation predicates. The spatiotemporal semantic value of the relation predicate of the class with the highest matching degree is the spatiotemporal semantic feature value of the relation predicate of the candidate triple.

7. The spatiotemporal correlation metric calculation method based on rules and dependency syntax according to claim 1, characterized in that, Step S4 is as follows: Based on the dependency path values ​​between the relational predicate and spatiotemporal information, the dependency distance between the head entity and spatiotemporal information, the spatiotemporal feature values ​​of the head and tail entity types, and the spatiotemporal semantic feature values ​​of the relational predicate in the candidate triples, the spatiotemporal correlation value is obtained, expressed as: in Indicates the spatiotemporal correlation value. for The weight it accounts for express , for The weight it accounts for The dependency distance value between the head entity and spatiotemporal information. for The weight it accounts for The spatiotemporal characteristic value representing the type of the head and tail entity. for The weight it accounts for The spatiotemporal semantic feature value represents the candidate triple relation predicate.

8. A spatiotemporal correlation metric calculation system based on rules and dependency syntax, characterized in that, It includes a data acquisition unit, a rule matching unit, a spatiotemporal correlation value calculation unit, and a spatiotemporal knowledge tuple expression model generation unit; The data acquisition unit is used to acquire open-domain text sequences and preprocess the text sequences to obtain preprocessed text sequences. The rule matching unit is used to formulate a rule matching table, perform syntactic analysis on the preprocessed text sequence, and match the syntactic analysis results according to the rule matching table to obtain candidate triples and spatiotemporal information. The spatiotemporal correlation value calculation unit is used to obtain the dependency path value between the relation predicate and the spatiotemporal information, the dependency distance between the head entity and the spatiotemporal information, the spatiotemporal feature value of the head and tail entity types, and the spatiotemporal semantic feature value of the relation predicate in the candidate triples based on the candidate triples and spatiotemporal information. The spatiotemporal correlation value is calculated based on the dependency path value between the relational predicate and spatiotemporal information, the dependency distance between the head entity and spatiotemporal information, the spatiotemporal feature value of the head and tail entity types, and the spatiotemporal semantic feature value of the relational predicate in the candidate triples. The spatiotemporal knowledge tuple expression model generation unit is used to compare the spatiotemporal correlation value with two defined thresholds to determine the classification result of the spatiotemporal correlation and generate the spatiotemporal knowledge tuple expression model.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-7.

10. An electronic device, characterized in that, The device includes a processor and a memory interconnected thereto, wherein the memory is used to store a computer program, the computer program including computer-readable instructions, and the processor is configured to invoke the computer-readable instructions to perform the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Text-oriented method for extracting cognitive relationships among knowledge topics

    CN110188347A

  • Multi-stage recognition of named entities in natural language text based on morphological and semantic features

    US20170364503A1