Knowledge graph analysis method based on grammar and semantic analysis

Through quasi-real-time word frequency update and dependency syntax analysis based on memory curves, the knowledge graph is constructed, and the existing models cannot be customized training and cross-domain application is solved, and efficient semantic role annotation and personalized analysis are achieved.

CN120278250AInactive Publication Date: 2025-07-08XIHUA UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510425403.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing semantic role labeling model cannot provide customized training corpus and continuously improve performance, and cannot be applied across fields and cannot meet the needs of different users.

Method used

A quasi-real-time word frequency update mechanism based on memory curves is adopted to generate phrases and build an entity database through dependent syntax analysis, and a knowledge graph is generated by combining grammar and semantic analysis, and a statement standardization process and semantic database construction are carried out.

Benefits of technology

Cross-domain semantic analysis is realized, word frequency is updated dynamically, and the effectiveness and adaptability of semantic role annotations are improved, and the personalized needs of different users are met.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses a knowledge graph analysis method based on grammar and semantic analysis, and the method is characterized in that the method comprises the steps: S1, carrying out the word segmentation of a text in the field, intention, text and appearing time defined by the text, and carrying out the statistics of the words to form a quasi-real-time lexicon; s2, in different fields and different intentions, words in the quasi-real-time lexicon form phrases one by one through dependency syntactic analysis, and a phrase intention search library is obtained; s3, defining the real-time word library and the phrase awareness search library as grammar and semantic analysis map entity data, and constructing an entity database; s4, the user inputs an original statement into the entity database for analysis; s5, obtaining part-of-speech and syntactic relationships from the original statement source input by the user by using a dependency syntactic analysis method, and packaging the syntactic relationships into a knowledge graph; s6, performing statement standardization processing; and S7, obtaining the processed knowledge graph and the analysis model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of syntax and semantic analysis, and particularly relates to a knowledge graph analysis method based on syntax and semantic analysis. Background Art

[0002] Descriptions of various events in human life (ranging from a small action to a historical event) exist in large quantities in natural language, and also include descriptions of the time, place, participating roles, status, and relationships between events of the event. With the rise of Internet-related technologies, people increasingly rely on the network to obtain information, and the information on the Internet shows characteristics such as massive, explosive growth, and redundancy. In order to better monitor and utilize the information, enabling machines to analyze events in text, research on sentence analysis oriented to events has become increasingly important. Sentence analysis refers to analyzing the functions and semantics of each component in a sentence, transforming the linear word order between words in the input sentence into a non-linear data structure.

[0003] A knowledge graph is a modern theory that combines the theories and methods of disciplines such as applied mathematics, graphics, information visualization technology, and information science with methods such as bibliometric citation analysis and co-occurrence analysis, and uses a visualized graph to vividly display the core structure, development history, frontier fields, and overall knowledge architecture of a discipline to achieve the purpose of multi-disciplinary integration. It displays complex knowledge fields through data mining, information processing, knowledge measurement, and graph drawing, reveals the dynamic development laws of knowledge fields, and provides practical and valuable references for disciplinary research.

[0004] Currently, the main theories in the field of natural language processing regarding sentence analysis include: dependency syntax, the formal grammar theory developed by Chomsky, namely phrase structure grammar and its extensions, such as: lexical functional grammar, functional unification grammar, generalized phrase structure grammar, head-driven phrase structure grammar. The ideas of these methods are all based on English grammar knowledge, and do not divide the components in a sentence into events and event roles from the perspective of understanding events and analyze the relationships between them. Currently, most research on events focuses on identifying and extracting events and event role extraction from text, event-based automatic summarization, and text automatic generation, etc. These researches all urgently require the support of the sentence analysis method based on event structure of the present invention.

[0005] Semantic role labeling is a core technology in natural language processing. Traditionally, semantic role labeling uses trained part-of-speech tagging models, dependency parsing models, etc. to analyze the semantic roles in sentences. However, these models are scattered and not in the same system. In addition, existing semantic role labeling can only provide a trained system, which cannot meet the different needs of users to provide different types of training corpora, nor can it allow users to continuously improve the efficiency by themselves.

[0006] Therefore, it is necessary to provide a semantic analysis method based on a knowledge graph to solve the above technical problems. Summary of the Invention

[0007] To solve the above problems, the purpose of this application is to provide a knowledge graph analysis method based on syntax and semantic analysis.

[0008] This application adopts the following technical solutions to solve the above technical problems:

[0009] 1. A knowledge graph analysis method based on syntax and semantic analysis, including:

[0010] Step S1: For the domain, intention, text, and occurrence time defined by the text, perform word segmentation on the text and respectively count the total number of times the words obtained after word segmentation appear in different domains and different intentions. This total number changes according to the law of the memory curve, and the decayed total number plus the number of times of reappearance is used as the current word frequency of the word, forming a quasi-real-time thesaurus.

[0011] Step S2: In different domains and different intentions, respectively use dependency parsing to form the words in the quasi-real-time thesaurus into phrases, and change the number of times the phrases appear according to the law of the memory curve to form the frequency of the phrases. Count the frequency of the same phrase in different domains to obtain the correlation degree of the phrase in different domains, and obtain a phrase intention search library.

[0012] Step S3: Define the real-time thesaurus and the phrase awareness search library as syntax and semantic analysis graph entity data, and construct an entity database.

[0013] Step S4: The user inputs the original statement into the entity database for analysis.

[0014] Step S5: Use the dependency parsing method to obtain the part of speech and syntactic relationship of the original statement source input by the user, and encapsulate the syntactic relationship into a knowledge graph.

[0015] Step S6: Perform statement standardization processing.

[0016] Step S7: Obtain the processed knowledge graph and analysis model.

[0017] 2. The knowledge graph analysis method based on syntax and semantic analysis according to claim 1, characterized in that, in step S4, the original statement further includes the specific steps of generating the corpus:

[0018] S4.1: Word segmentation;

[0019] S4.2: Part-of-speech tagging;

[0020] S4.3: Syntactic analysis;

[0021] S4.4: Semantic and semantic role analysis.

[0022] 3. The knowledge graph analysis method based on syntax and semantic analysis according to claim 1, characterized in that, in step S5, the steps of encapsulating the syntactic relationship are:

[0023] S5.1: Parse the original statement source and generate multiple semantic vectors by parsing it;

[0024] S5.2: Perform data fusion on the generated semantic vectors, mine the correlation relationships between the fused semantic vectors, and generate the relationships between the semantic vectors;

[0025] S5.3: Construct a semantic database from the fused semantic vectors and the relationships between the semantic vectors, store it in the server, and perform quality evaluation on the semantic database to complete the construction of the knowledge graph.

[0026] On the basis of conforming to the common knowledge in the art, the above preferred conditions can be combined arbitrarily to obtain the preferred embodiments of the present application.

[0027] The positive and progressive effects of the present application are as follows: The present application has dynamic word frequency adjustment: a quasi-real-time word frequency update mechanism based on the memory curve, rather than the traditional static statistics. Cross-domain phrase association: Generate phrases through dependency syntactic analysis, and combine cross-domain frequency statistics to quantify the intention relevance of the phrases. Syntax and semantics fusion: Directly encapsulate the results of dependency syntactic analysis into the knowledge graph structure, rather than relying only on semantic embedding or rule matching. Detailed implementation manners

[0028] The present application will be further described below by way of examples, but the present application is not limited to the scope of the described examples. The experimental methods without specific conditions noted in the following examples are carried out according to the conventional methods and conditions, or selected according to the product specifications.

[0029] The experimental methods used in the following examples are all conventional methods unless otherwise specified.

[0030] Example 1

[0031] 1. A knowledge graph analysis method based on syntax and semantic analysis, characterized by comprising:

[0032] Step S1: For the domain, intention, text, and occurrence time defined by the text, perform word segmentation on the text and respectively count the total number of times the words obtained after word segmentation appear in different domains and different intentions. The total number of times changes according to the law of the memory curve, and the decayed total number of times plus the number of times of reappearance is used as the current word frequency of the word, forming a quasi-real-time thesaurus;

[0033] Step S2: In different domains and different intentions, respectively use dependency syntax analysis to form the words in the quasi-real-time thesaurus into phrases, and change the number of times the phrases appear according to the law of the memory curve to form the frequency of the phrases. Count the frequency of the same phrase in different domains to obtain the degree of association of the phrase in different domains, and obtain a phrase intention search library;

[0034] Step S3: Define the real-time thesaurus and the phrase awareness search library as syntax and semantic analysis graph entity data, and construct an entity database;

[0035] Step S4: The user inputs the original statement into the entity database for analysis;

[0036] In step S4, the original statement further includes the specific steps of generating corpus:

[0037] S4.1: Word segmentation;

[0038] S4.2: Part-of-speech tagging;

[0039] S4.3: Syntax analysis;

[0040] S4.4: Semantic and semantic role analysis.

[0041] In step S5, the steps of encapsulating the syntactic relationship are:

[0042] S5.1: Parse the original statement source and generate multiple semantic vectors from it;

[0043] S5.2: Perform data fusion on the generated semantic vectors, mine the association relationships between the fused semantic vectors, and generate the relationships between the semantic vectors;

[0044] S5.3: Package the generated relationship between semantic vectors into a knowledge graph.

[0045] S5.4: Standardize the generated knowledge graph to obtain a processed knowledge graph and an analysis model.

[0046] S5.5: Use the processed knowledge graph and analysis model to perform analysis and generate results.

[0047] S5.3: Construct a semantic database based on the relationships between the fused semantic vectors and the semantic vectors, store it in the server, and perform quality assessment on the semantic database to complete the construction of the knowledge graph.

[0048] Step S6: The operation of performing statement standardization is as follows:

[0049] In order to more efficiently select multiple reference standard statements from a preset database, a text retrieval method can be used to select multiple reference standard statements from the preset database. Specifically, the statement to be parsed can be matched with multiple standard statements in the preset database according to the results of word segmentation, entity annotation, and part-of-speech annotation in the first processing result to obtain multiple matching values. Then, determine the reference standard statements corresponding to the first matching values higher than the first preset threshold among the multiple matching values to obtain multiple reference standard statements.

[0050] Step S7: Obtain the processed knowledge graph and analysis model.

[0051] This application performs word segmentation on the text information through the knowledge graph semantic model to obtain at least one word; respectively obtain the characteristics of the at least one word; the characteristics of words can be divided into major categories such as nouns, verbs, adjectives, numerals, quantifiers, pronouns, adverbs, prepositions, conjunctions, auxiliary words, interjections, and onomatopoeia, and continue to be subdivided into multiple small categories within the major categories. For example, nouns can be divided into person nouns, thing nouns, time nouns, location nouns, relationship nouns, etc. By obtaining the characteristics of words, the meaning of words can be better understood; specifically as follows: your (personal pronoun) right (thing noun) requirement (thing noun) what (interrogative pronoun). The specific example of the classical semantic analysis method of the knowledge graph is also applicable to the traditional Chinese medicine classical semantic analysis system combining the knowledge graph in this embodiment. Through the foregoing detailed description of the semantic analysis method combining the knowledge graph, those skilled in the art can clearly know the combination of the knowledge graph in this embodiment. Therefore, for the sake of simplicity of the specification, it will not be elaborated here. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0052] Finally, it should also be noted that in this application, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device.

[0053] Although the present application has been disclosed above through the description of specific embodiments of the present application, it should be understood that those skilled in the art can design various modifications, improvements or equivalents to the present application within the spirit and scope of the appended claims. These modifications, improvements or equivalents should also be considered to be included within the scope claimed by the present application.

Claims

1. A knowledge graph analysis method based on syntax and semantic analysis, characterized in that Including: Step S1: Segment the text according to the subject field, meaning, length, and occurrence time of the text, and respectively count the total number of times the words obtained after segmentation appear in different fields and with different intents. The total number of times changes according to the law of the memory curve. The decayed total number of times plus the number of times of reappearance is used as the current word frequency of the word, forming a quasi-real-time thesaurus; Step S2: Use dependency syntactic analysis to extract grammatical structures such as subject, predicate, and object in different fields and with different intents, generate phrases, and change the number of times the phrases appear according to the law of the memory curve to form the frequency of the phrases. Count the frequencies of the same phrase in different fields to obtain the correlation degree of the phrase in different fields, and obtain a phrase intent search library; Step S3: Define the real-time thesaurus and the phrase awareness search library as grammatical and semantic analysis graph entity data, and construct an entity database; Step S4: The user inputs the original statement into the entity database for analysis; Step S5: Use the dependency syntactic analysis method to obtain the part of speech and syntactic relationship of the original statement source, and encapsulate the syntactic relationship into a knowledge graph; Step S6: Perform statement standardization processing; Step S7: Obtain the processed knowledge graph and analysis model.

2. The knowledge graph analysis method based on syntax and semantic analysis according to claim 1, wherein In Step S4, the original statement further includes the specific steps for generating the corpus: S4.1: Word segmentation; S4.2: Part-of-speech tagging; S4.3: Syntactic analysis; S4.4: Semantic and semantic role analysis.

3. The knowledge graph analysis method based on syntax and semantic analysis according to claim 1, wherein In Step S5, the steps for encapsulating the syntactic relationship are: S5.1: Parse the original statement source and generate multiple semantic vectors by parsing it; S5.2: Perform data fusion on the generated semantic vectors, and mine the correlation relationship between the fused semantic vectors to generate the relationship between the semantic vectors; S5.3: Construct a semantic database from the fused semantic vectors and the relationship between the semantic vectors, store it in the server, and perform quality evaluation on the semantic database to complete the construction of the knowledge graph.

Citation Information

Patent Citations

  • Multi-round semantic analysis method based on dependency syntax analysis and Chinese grammar

    CN111984778A

  • Data integration method based on knowledge graph

    CN118333059A

  • Software defect knowledge-oriented knowledge search method

    WO2021008180A1