Hierarchical semantic analysis method and system based on semantic mode

Through the hierarchical semantic analysis method based on semantic patterns, the problems of fuzzy and insufficient hierarchical processing of sub-modes in the existing technology are solved, efficient and interpretable deep semantic understanding and reasoning are achieved, the accuracy and robustness of semantic analysis are improved, and it is suitable for fields such as text analysis and information retrieval.

CN120579552APending Publication Date: 2025-09-02GUANGZHOU ZHIYAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510681411.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

In the semantic analysis, existing natural language processing technologies have problems such as fuzzy sub-modes, insufficient hierarchical processing, high cost of generative models and poor semantic analysis, making it difficult to achieve efficient, interpretable and low-cost deep semantic understanding and reasoning.

Method used

A hierarchical semantic analysis method based on semantic patterns is adopted, and a low-dependence reasoning mechanism that is adapted across domains is constructed by defining structured semantic patterns. A semantic pattern library is used for layer-by-layer matching, candidate parsing tree nodes are generated, hierarchical semantic structures are constructed, and they are converted into predicate logical expressions to support inference.

Benefits of technology

It improves the accuracy and flexibility of semantic analysis, can handle the hierarchical nested structure of complex sentences, supports chapter understanding and reasoning, provides a solid foundation for natural language processing, and is suitable for text analysis, information retrieval, and intelligent question and answer fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579552A_ABST
    Figure CN120579552A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, in particular to a hierarchical semantic analysis method and system based on a semantic schema, comprising the following steps: automatically acquiring a semantic schema, the semantic schema being composed of a left meaning result, a right grammar component and a semantic component, the right comprising grammar semantic components such as a subject, a predicate, an object and the like, and the semantic component being a semantic component; the left part represents the meaning result of the right part, the left meaning result and the right part jointly form a meaning structure, and semantic nesting is realized based on the meaning structure; performing layer-by-layer matching on the input text based on the semantic pattern library, generating candidate parse tree nodes and constructing a hierarchical semantic structure; and converting the hierarchical semantic structure into a predicate logic expression to form a reasonable logic expression. According to the method, the dependence of a traditional method on domain knowledge is broken through, the robustness, flexibility and interpretability of semantic analysis are remarkably improved, efficient support is provided for tasks such as chapter understanding, ambiguity resolution and causal reasoning, and meanwhile the data and computing power cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a hierarchical semantic analysis method and system based on semantic patterns. Background Art

[0002] With the continuous development of artificial intelligence (AI) technology, the demand for a deeper understanding and effective processing of natural language is growing. Currently, natural language processing (NLP) technology is being applied in a variety of fields, such as intelligent customer service, information retrieval, and machine translation, but it still faces numerous challenges. Within the field of NLP, semantic analysis is a core research direction, aiming to enable learning and reasoning based on natural language understanding. Traditional semantic analysis methods, such as structural pattern recognition, dependency syntax, case grammar, frame semantics, and generative AI, have achieved some success.

[0003] Structural pattern recognition (syntactic pattern recognition), proposed by Professor Fu Jingsun, describes complex patterns through the formation of multi-level structures using sub-patterns. This technology can be used to recognize speech, semantics, and images. For example, advances in computer vision technology have led to new discoveries in image semantic recognition, enabling better identification of complex image structures. In specific scenarios, speech recognition systems utilize multi-level structures to describe speech patterns, aiding in the accurate recognition of voice commands and improving interaction efficiency. In visual inspection of industrial product defects, defects can be quickly located by constructing product structural models.

[0004] Dependency syntax, proposed by L. Tesnière, reveals syntactic structure through word dependencies and is a mainstream syntactic analysis method. It is widely integrated into natural language processing tools such as NLTK and Stanford CoreNLP. It has a wide range of applications. For example, intelligent writing assistance software uses dependency syntax to analyze sentences for grammatical errors, while search engines use it to understand the structure of user queries and improve search relevance.

[0005] Case grammar and frame semantics were proposed by Charles Fillmore. The former focuses on the semantic relationship between nouns and verbs, while the latter emphasizes the representation of cognitive outcomes based on a network of scenario concepts. They are widely used in the field of information extraction. Case grammar can be used to extract key semantic relationships from text, such as extracting information such as the subject and object of an event from news reports. In the early semantic analysis stage of machine translation, it can be used to assist in determining the semantic roles of words, providing support for more accurate translation. Frame semantics also has a wide range of applications. For example, in the field of intelligent customer service, by building a customer service scenario framework, the user's inquiry intent can be understood and answers can be quickly matched. In text classification tasks, the text is divided into categories based on the scenario framework involved, improving classification accuracy. In recent years, generative models, trained on massive amounts of data, have made significant progress in natural language processing, semantic analysis and comprehension, learning, and reasoning. They can handle a wide range of semantic analysis tasks, understand semantic relationships and roles, and recognize different expressions of the same meaning. They possess excellent contextual understanding, allowing them to grasp the overall meaning of complex text within the context of the original text. They also possess the ability to generalize, enabling them to tackle new tasks across diverse domains in few-shot and zero-shot scenarios. Leveraging self-supervised learning, they can automatically extract linguistic knowledge from vast amounts of text, reducing reliance on manual annotation and enabling continuous training and evolution with new data. They have also demonstrated promising results in logical and commonsense reasoning tasks. For example, DeepSeek-R1 leverages reinforcement learning to apply reasoning capabilities across tasks, enabling it to draw reasonable inferences based on existing knowledge.

[0006] However, with the development of technology, the demand for cross-domain understanding, learning and reasoning based on formal and structured representation is increasing, and the limitations of existing technologies are gradually becoming apparent.

[0007] 1. Sub-pattern problem in structural pattern recognition Existing structural pattern recognition relies on domain-specific sub-pattern extraction, and the original feature selection relies on designer experience. The lack of generalized methods has led to limited cross-domain applications and hindered the development and application of structural pattern recognition.

[0008] 2. Insufficient hierarchical processing of dependency syntax It can only analyze direct dependencies between words, but cannot reflect the multi-layered nested structure of sentences. It also lacks analysis of the nested clauses of complex sentences, making it unable to handle the hierarchical nested structures of complex sentences. For example, for the phrase "companies with net assets of more than 20 billion yuan and listed five years ago," it can only identify the attributive relationship between "listed" and "company," but cannot deeply analyze the semantic roles and nested relationships at each level.

[0009] 3. Limitations of Case Syntax and Frame Semantics Case grammar relies too much on verbs, lacks structural analysis of attributive modification within noun phrases, lacks formal features, and has not formed a formal semantic combination and reasoning system; frame semantics is mainly used to describe cognitive outcomes, and does not clarify how natural language expresses complex meanings through hierarchical nested combinations, making it difficult to support deep understanding and reasoning of natural language.

[0010] 4. Drawbacks of Generative Models Generative AI faces problems such as reliance on massive amounts of data for training, high costs, hallucinations in generated content, and a lack of deep thinking and reasoning capabilities. For example, its reliability is poor during semantic analysis, and it is difficult to accurately analyze complex semantic relationships and structures. In the credibility assessment of passage semantic understanding, its reliability is only 40%. In the natural language understanding link, the "hallucination problem" is more prominent, often generating content that contradicts the facts, and lacking the ability to proactively clarify user ambiguous queries. It is easy to answer based on guesswork and give results that do not meet requirements. From a learning perspective, model training is extremely dependent on massive amounts of data and powerful computing power, resulting in high costs, and the learning process is greatly affected by data quality and distribution, resulting in low efficiency in learning specific domain knowledge. In terms of reasoning, when models handle complex reasoning tasks, they are prone to drawing erroneous or unreasonable conclusions, and the reasoning process is like a "black box" that is difficult to explain, adding obstacles to application and optimization.

[0011] Therefore, how to achieve efficient, explainable and low-cost deep semantic understanding and reasoning is a technical problem to be solved. Summary of the Invention

[0012] The problem to be solved by the present invention is to provide a hierarchical semantic analysis method and system based on semantic patterns. By defining structured semantic patterns, cross-domain adaptation and low-dependence reasoning mechanisms, the problems of fuzzy sub-patterns, insufficient hierarchical processing, high cost of generating models and poor robustness of semantic analysis in traditional technologies are solved.

[0013] In view of the shortcomings of the prior art, the present invention solves the technical problems by adopting a technical solution: a hierarchical semantic analysis method based on semantic patterns, comprising the following steps: Automatically obtain a semantic model, wherein the semantic model is composed of a left-hand semantic result and a right-hand grammatical component and a semantic component, wherein the right-hand part includes grammatical and semantic components such as a subject, a predicate, and an object, and the left-hand part represents the semantic result of the right-hand part. The left-hand semantic result and the right-hand component together constitute a semantic structure, and semantic nesting is achieved based on the semantic structure; The input text is matched layer by layer based on the semantic pattern library to generate candidate parse tree nodes and construct a hierarchical semantic structure. The root node of the parse tree is the overall semantics of the sentence, the non-leaf nodes are the left part of the semantic pattern, and the leaf nodes are the instantiation results of the semantic components. The parse tree supports cross-sentence context association and the derivation of multi-layer nested structures. The hierarchical semantic structure is converted into a predicate logic expression, including: identifying the semantic roles and relationships of each node in the parse tree, generating entity, attribute and relationship assertions, and retaining the hierarchical nested structure through logical operators to form a reasonable logical expression.

[0014] Preferably, the semantic patterns include but are not limited to the following types: Event pattern: used to describe the semantic relationship between action subject, predicate and object; Modification pattern: used to analyze the modification relationship between attributives, adverbials and central words; Group pattern: used to identify component combinations in parallel, progressive or selective relationships; Synthetic mode: used to process a combination of several parallel components.

[0015] Preferably, the method for constructing the semantic pattern library includes but is not limited to the following steps: Identify the semantic patterns contained in simple sentences: Using the role recognition model, first identify the roles of simple sentences without nested components to obtain the semantic patterns of simple sentences.

[0016] Pattern comparison method: Based on natural language expressions with known semantic patterns, find similar expressions using different nested components, and reasonably infer that the different nested components used in similar expressions are also legal expressions; Structural deduction method: Based on the heuristic rules of legal phrases, the component structure is divided and combined with the role recognition model to determine the semantic pattern implied by the phrase; Bootstrapping: Exploiting the morphological similarity of isomorphic semantic patterns, new patterns are generated using known patterns as parameterized templates; Element equivalence method: Establish equivalent semantic patterns of different expressions through corpus retrieval and element comparison.

[0017] Preferably, the following modules are included: Tuple module: used for text segmentation, named entity recognition and processing token segmentation ambiguity; Sentence block splitting module: uses punctuation marks, named entity boundaries and special structures to divide logical sentence blocks, and handles non-standard punctuation and format control characters; Pattern matching module: matches the input text layer by layer based on the semantic pattern library, generates candidate parse tree nodes and constructs a hierarchical semantic structure; Context window processing module: processes the semantic associations between short-range cross-sentence blocks in the text at the chapter level and optimizes the parsing boundaries through the context window; Likelihood analysis module: performs candidate pattern splicing and immature branch triggering on unmatched inputs, and extracts the maximum likelihood parse tree; Ambiguity resolution module: optimizes parsing results using strategies such as semantic hierarchy minimum priority and mutual information maximization; Multimodal fusion module: converts speech and image features into structural patterns and performs collaborative reasoning with natural language parsing results; Dynamic update module: expands the ontology vocabulary through fine-tuning of pre-trained models and automatically extracts new semantic patterns based on pattern inference algorithms; Reasoning and generation module: converts the parse tree into predicate logic expressions, supporting semantic query, semantic reasoning, and causal learning.

[0018] Preferably, the specific implementation of the sentence block splitting module includes: Identify punctuation marks in named entities to avoid mis-splitting; Merge multiple physical sentence blocks caused by non-standard punctuation into a single logical sentence block; Process the content in brackets and quotation marks, and determine the sentence block affiliation based on semantic logic.

[0019] Preferably, the nested matching of the pattern matching module includes the following steps: Traverse the token sequence in the order of the text and match it layer by layer with the patterns in the semantic pattern library; Generate a parse tree node for the successfully matched pattern, where its child node is the right semantic component and its parent node is the left semantic result; Matching is performed recursively on the nested structure until the leaf nodes are indecomposable semantic components or domain terms.

[0020] The beneficial effects of the present invention are as follows: Solve sub-mode related problems It clarifies that semantic patterns are subpatterns mentioned in structural pattern recognition. Semantic patterns can be categorized into event patterns, modifier patterns, group patterns, and narrative patterns. This addresses the question of what constitutes a subpattern in structural pattern recognition. Furthermore, it provides a domain-independent subpattern extraction method, namely, an ontology lexicon and semantic pattern acquisition method. This will help advance the development of structural pattern recognition and make its applications in multimodal fields such as natural language processing, speech, graphics, and images more versatile and efficient.

[0021] (2) Improving semantic analysis capabilities Compared with dependency syntax, the semantic analysis based on semantic patterns in the present invention can ensure the identification of all semantic levels, and use the same mechanism to produce unified results for each semantic layer. It can better handle the polysemy and polysemy problems of a word, and can also realize real-time guessing, allow word order changes, eliminate interference, and realize equivalent knowledge transfer through semantic reasoning, thereby improving the accuracy, flexibility and robustness of semantic analysis.

[0022] 3. Supporting passage comprehension and cognitive reasoning Based on hierarchical semantic structures, fully equivalent logical expressions can be generated, making structured learning and formal reasoning based on the results of natural language understanding possible. This helps achieve accurate understanding of entire paragraphs and chapters, providing a more solid foundation for the application of natural language processing in text analysis, information retrieval, intelligent question answering, and other fields. It also lays a solid foundation for various cognitive tasks such as causal learning, logical reasoning, critical thinking, and metacognitive learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a system module architecture diagram of the present invention; Figure 2 This is a flow chart of the ontology vocabulary construction of the present invention; Figure 3 It is a flow chart of semantic pattern acquisition of the present invention. DETAILED DESCRIPTION

[0024] The present invention will be described in further detail below with reference to the accompanying drawings and specific embodiments. The embodiments of the present invention are provided for purposes of illustration and description and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments are chosen and described to better illustrate the principles of the invention and its practical application, and to enable those skilled in the art to understand the invention and design various embodiments with various modifications suitable for specific applications.

[0025] like Figure 1-3 As shown, the present invention provides a hierarchical semantic analysis method and system based on semantic patterns: 1. Core Concepts and Infrastructure 1. Semantic Pattern Used to describe a model that expresses complex semantics based on simple semantics (basic semantics).

[0026] The semantic pattern is divided into a left part and a right part. The right part is composed of grammatical and semantic components. The left part is the semantic result of the right part. The semantic result of the left part and the components of the right part together constitute a semantic structure. Semantic nesting is achieved based on the semantic structure.

[0027] The types of semantic patterns include event pattern (s pattern), modification pattern (m pattern), group pattern (g pattern), and narrative pattern (c pattern).

[0028] The event pattern is used to describe the semantic relationship between the subject, predicate and object of an action; the modification pattern is used to analyze the modification relationship between attributives, adverbials and central words; the group pattern is used to identify the combination of components in parallel, progressive or selective relationships; the narrative pattern is used to process several parallel components, which is a combination of terms, phrases or even sentences.

[0029] For example, the event pattern “@f|subj: @sb + p: invention + obj: @device” represents the declarative sentence “Edison invented the incandescent lamp”, where “@f” is the left part and the right part contains semantic components such as subject, predicate, and object.

[0030] 2. Hierarchical semantic structure According to the relevant theories of cognitive linguistics, in natural language, a sentence represents a complete meaning and can be reduced to a tree. This tree is the semantic parsing tree (hereinafter referred to as "parsing tree"), and the structure contained in it is the hierarchical semantic structure.

[0031] In the parsing tree, the root node represents the entire sentence, and the non-leaf nodes and their child nodes form semantic levels, with each level corresponding to a semantic pattern. The parent node is on the left and the child node is on the right, forming a multi-layer structure through semantic nesting.

[0032] For example, for the sentence "Edison invented the incandescent lamp that lights up the night", its parsing tree structure is as follows: name="@f", semp_type="s", role="", ko="@f", hw="invented", seq=0, text= "Edison invented the incandescent lamp that lights up the night" ├ name="@sb", semp_type="d", role="subj", ko="@sb", hw="Edison", seq=0, text= "Edison" ├ name="#v", semp_type=="m", role="p", ko="#v", hw="invented", seq=0,text= "invented" │ ├ name="#v", semp_type=="d", role="p", ko="#v", hw="invented", seq=0,text= "invented" │ └ name="&了", semp_type=="d", role="", ko="&了", hw="了", seq=0,text= "了" └ name="@sth", semp_type="s", role="obj", ko="@sth", hw="incandescent lamp",seq=0, text= "the incandescent lamp that lights up the night" ├ name="&将", semp_type=="d", role="", ko="&将", hw="将", seq=0,seq=0, text= "将" ├ name="@sth", semp_type=="d", role="obj", ko="@sth", hw="night",seq=0, text= "night" ├ name="#v", semp_type=="d", role="p", ko="#v", hw="illuminate", seq=0,text= "illuminate" ├ name="&的", semp_type=="d", role="", ko="&的", hw="的", seq=0,text= "的" └ name="@sth", semp_type=="d", role="tool", ko="@sth", hw="incandescent lamp",seq=0, text= "incandescent lamp" Note: (1)The text of the root node is "Edison invented the incandescent lamp that illuminates the night", indicating that the root node represents the whole sentence. ko = "@f", indicating that the semantic combination result of the child nodes of this node is "@f", where f is the abbreviation of "fact", indicating that this sentence is a declarative sentence. Note: Usually, the value of the ko field is the same as the value of the name field. The name field is used to distinguish different nodes when displaying the tree structure.

[0033] (2)The values of the name fields of the child nodes of the root node are successively: "@sb", "#v", "@sth", and the values of the role fields are successively: subj, p, obj, indicating that the pattern implied by "Edison invented the incandescent lamp that illuminates the night" is "@f|subj:@sb + p:invent + obj:@sth". The rest of the levels follow the same pattern.

[0034] Note 1: "p" is the abbreviation of "predicate", representing the predicate. "subj" is the abbreviation of "subject", representing the subject. "obj" is the abbreviation of "object", representing the object.

[0035] Note 2: The predicate is a special role, so it is marked differently from other roles in the pattern: In the pattern, the component serving as the predicate is usually replaced by hw (hw is the abbreviation of the headword. If the predicate component is a dictionary word, the hw is the same as the value of the text field. If the predicate component is a nested component, hw is the headword of the nested component. See the subsequent "Supplementary Note on Headwords" for details), rather than being replaced by ko like other components. In this example, the predicate component is denoted as "p:invent", while other components are denoted as "subj:@sb" and "obj:@sth" respectively.

[0036] (3)Explanation of the attributes of the parsing tree nodes: 1)name: The name of the parsing tree node, usually the same as the ko field.

[0037] 2) semp_type: The type of semantic pattern. Here, semp is the abbreviation of semantic pattern. Common types of semantic patterns include: s pattern: Used to represent sentences that construct spatio-temporal events. Usually, verbs, adjectives, index words, attributive words or relational words are used as predicates, or questions or phrases derived from these sentences. s is the abbreviation of sentence.

[0038] m pattern: Refers to the pattern where one component modifies another component in the form of a prefix or suffix. For example, adverbs modify verbs, adjectives modify nouns, and suffix words of verbs (such as "zhe", "le", "guo"), etc. The m pattern is also called the modification pattern, and m is the abbreviation of modification.

[0039] g pattern: Composed of several components combined to represent a specific meaning, where there is no predicate, such as time, quantity, etc. g is the abbreviation of group.

[0040] c pattern: Refers to the pattern contained in the combined narration. c is the abbreviation of combination.

[0041] d pattern: Refers to dictionary words (registered words). That is, dictionary words are a special pattern. d is the abbreviation of dictionary.

[0042] 3) role: The role played by the current component.

[0043] 4) ko: ko is the abbreviation of kind-of.

[0044] If the current component is a dictionary word, the value of this field is the category of the current component. See the subsequent "Supplementary Note on ko".

[0045] If the current component is a nested component (such as a sentence or a phrase), this field is the ko of the central word of the nested component.

[0046] 5) hw: Head word, hw is the abbreviation of head word.

[0047] If the current node is a dictionary word, hw is the same as text.

[0048] If the current component is a nested component (such as a sentence or a phrase), this field is the central word of the nested component. See the "Supplementary Note on Head Word". 6) seq: The serial number of the current node among its sibling nodes, starting from 0.

[0049] 7) text: The substring in the input text corresponding to the current node.

[0050] (4) Supplementary Note on ko: The KO of the node "Edison" (i.e., the node whose text field value is "Edison", the same below) is "@sb", where "sb" is the abbreviation of "somebody", indicating that "Edison" is a person's name, and the "@" prefix indicates that this is an ontology category, usually used as a semantic component; The KO of the node "incandescent lamp" is "@sth", indicating that "incandescent lamp" is a kind of thing, the "@" prefix indicates that this is an ontology category, and "sth" is the abbreviation of "something"; The KO of the node "invent" is "#v", indicating that "invent" is a verb, and the "#" prefix indicates that this is a grammatical word, where v is the abbreviation of "verb"; The KO of the node "了" is "&了", and the "&" prefix indicates that this is a fixed word. Usually, the KOs of common grammatical words are represented in a fixed form. Similarly, in this example, there are also "&的" and "&将".

[0051] (5) Supplementary note on the central word: <00001​​​​​​​​​​​​​​​​​​​​​​​​(3) Pattern Matching: Submit tokens or named entities in the order of text, try to match them with known semantic patterns, generate candidate patterns, rewrite the patterns, and build parse tree nodes.

[0058] (4) Context window processing: When processing chapter-level text, the context window is used to determine the semantic pattern attribution across sentence blocks, correct the boundaries of hierarchical semantic structures, and solve problems such as inaccurate punctuation.

[0059] (5) Likelihood analysis and ambiguity resolution: Perform likelihood analysis on inputs that cannot be directly matched, try to splice candidate patterns or trigger immature branches; use strategies such as fewer semantic levels first, left-side rule first, and mutual information to eliminate ambiguity and ensure accurate parsing results.

[0060] 2. Semantic Pattern Extraction Method Semantic patterns, sub-patterns that are not related to domain knowledge. The method of extracting semantic patterns mainly realizes the extraction of sub-patterns that are not related to domain knowledge by constructing an ontology vocabulary and obtaining semantic patterns. The details are as follows: (1) Ontology vocabulary construction 1) Term Identification: Identifying unregistered terms from a large corpus is the foundation for building an ontology lexicon. First, collect existing annotated corpora and datasets containing text from various fields. For each term, generate annotated samples containing the term, the context in which the term appears, and the category to which the term belongs. Term identification is a binary classification task, marking true if the term is present and false if it is not, thus generating annotated data for training.

[0061] Next, select an appropriate machine learning algorithm to train the model, such as a support vector machine (SVM) or a multi-layer perceptron (MLP). Furthermore, by fine-tuning the BERT model, it can also be applied to term classification tasks, leveraging its pre-trained parameters and language understanding capabilities to improve the accuracy of term recognition.

[0062] 2) Term Category Identification: Term category identification involves classifying terms into different categories, such as entity classes (persons, things, events) and pattern classes (actions, behaviors, states, relationships, changes, attributes, roles, indicators). This helps the system more accurately understand and process textual information. This also requires the collection of annotated corpus and datasets to generate annotated data. The sample includes the term, the context in which the term occurs, and the category to which the term belongs (entity class, pattern class, or counterexample).

[0063] During the modeling phase, machine learning algorithms such as support vector machines, decision trees, and neural networks are selected to accurately determine the scope of a term by learning its context. Pre-trained language models can also be fine-tuned and applied to term category identification tasks through feature extraction. As new data emerges, the trained model can be used to identify more terms from the newly added corpus. Highly confident terms can be selected, and the annotated data can be further expanded to iteratively train the model, thereby continuously improving its accuracy.

[0064] 3) Term relationship identification: The relationship between terms in the ontology vocabulary is very important for semantic understanding and reasoning, mainly including hyponym and hyponym relationships, synonym relationships, etc., and whole-part relationships, direction relationships, affiliation relationships, ownership, use, source, upstream and downstream relationships, etc. can also be introduced as needed.

[0065] Taking the identification of hyponyms and hyponyms as an example, we first collected existing annotated corpora and datasets. We also generated annotated data based on typical sentence patterns such as "x is a type of y, x can be divided into types y and z, x belongs to category y, x falls within the category y, x is a subcategory of y, and x is a specific thing in the field of y." The sample was a BIO sample (BIO stands for Begin, Inside, and Outside, respectively). We used labels such as B_hypernym, B_hyponym, I_hypernym, I_hyponym, and O to annotate the hypernym and hyponym fragments in the sample input.

[0066] When modeling, choose models such as CRF (Conditional Random Field) and biLSTM (Bidirectional Long Short-Term Memory). These models effectively learn from sequence features and contextual information in text, accurately identifying hyponymous and hyponymous relationships between terms. Pre-trained models can also be fine-tuned to adapt them to the task of identifying term relationships. Leverage the trained model to identify more term relationships from new corpus, and select high-confidence samples to expand the annotated dataset to further optimize the model.

[0067] Similar methods can also be used to identify other relationships, and corresponding samples and models can be constructed according to the characteristics of different relationships for identification.

[0068] (2) Semantic pattern acquisition Knowledge about semantic patterns can be acquired by following these steps: 1) Preparation: Role identification modeling, using the Chinese information database from HowNet as seed knowledge. Prepare corpus and identify simple and complex sentences within it. Use compound narrative recognition rules to convert compound narratives into non-compound narratives, thereby increasing the number of simple sentences. This process involves two stages: preparation and modeling.

[0069] a) Preparation: Use the Chinese information database of CNKI as the seed knowledge. CNKI contains rich semantic knowledge and lexical relationships. In addition, the corresponding natural language text can be obtained using known semantic patterns to construct context samples for role recognition. The sample composition is as follows: Input: The text as the context, the words or phrases (or sentences) that play the role Ri in the text, and the words that play the predicate. Output: The role name Ri as the output.

[0070] In this way, role recognition is converted into a multi-classification problem. Using the same method, a separate model can be trained for predicate determination.

[0071] b) Modeling: Use the labeled data to train the selected model, such as CRF-biLSM, RNN, etc. It is also possible to fine-tune the pre-trained model. During inference, when inputting the sentence for which the semantic pattern is to be recognized currently, as well as the target word and predicate for which the role is to be determined, the role played by the target word can be obtained.

[0072] 2) Semantic pattern recognition of simple sentences: Based on role recognition, first identify the roles of simple sentences (without nested components) to obtain the semantic patterns of simple sentences.

[0073] a) Judgment of simple sentences: Based on rules, determine whether a sentence is a simple sentence. For example, simple sentences usually do not contain "de", do not contain ",", and do not have more than two verbs, etc. Through these rules, simple sentences can be screened out to provide a basis for subsequent analysis.

[0074] b) Identify the semantic patterns implied in simple sentences: Based on the role recognition model, identify the predicates and their various roles in simple sentences to obtain the semantic patterns implied in simple sentences.

[0075] 3) Phrase pattern recognition: Based on the semantic patterns obtained above, use the pattern comparison method to obtain the nested components (i.e., phrases) in complex sentences, then use the structure derivation method to obtain the component structure of the phrases, and then combine role recognition to obtain the patterns of phrases.

[0076] a) Pattern comparison method: If the semantic pattern implied by a certain expression is known, once different expressions formed by using nested components based on this expression are found, it can be reasonably inferred that the nested components are also legal expressions.

[0077] For example: If you know "I guess this experiment is very important to the paper", then when you see "I guess the information he looked up in the library yesterday is very important to the paper", you can know that "the information he looked up in the library yesterday" replacing "this experiment" is also a legal expression (here it is a phrase with "information" as the central word) and plays the same role as "this experiment".

[0078] b) Structural deduction method: used to construct the component structure and semantic pattern of a phrase.

[0079] For example, for a given expression, such as "1911 Nobel Prize in Chemistry", its component division can be obtained through the following process: Find the legal terms "Nobel Prize in Chemistry, Nobel Prize, Chemistry" from the term and phrase database Prize, Nobel, Chemistry, Prize".

[0080] When initialized, enclose all known terms or phrases in "{}".

[0081] Then, we deduce the structure based on heuristic rules. Rule 1: If ABC, BC, and C are all legal phrases, and if AC is also a legal phrase, then ABC -> {C|A + {C|B + C}}} holds. Rule 2: For an input sequence S, if s can be partitioned into A1, A2, ...An, and A1, A2, ...An are all legal phrases, and there is no overlap between A1, A2, ...An, then the partitioning of the input sequence into A1, A2, ...An is reasonable, denoted as S -> {A1 + A2 + ...An}. Using these rules, we can deduce the 1911 Nobel Prize in Chemistry and determine its component structure.

[0082] c) Then, using the predicate recognition and role recognition models, the roles of each semantic component are determined to obtain the semantic pattern of the phrase.

[0083] 4) Bootstrapping: Using isomorphism substitution method, more semantic patterns are obtained based on known semantic patterns.

[0084] The principle of isomorphism substitution is that if there is a strong similarity between the context of the known pattern word and the newly added pattern word, then we can consider referring to the role of the known pattern to construct the pattern of the newly added pattern word.

[0085] For example, the "Solution" pattern is similar to the "Process, Clear, Dissolve, Resolve, Improve" pattern. Therefore, the "Solution" pattern can be substituted into the generated "Process, Clear, Dissolve, Resolve, Improve" pattern.

[0086] For example, the "discovery" pattern is similar to the "perception, discovery" pattern. Therefore, the "discovery" pattern can be substituted into the generated pattern related to "perception, discovery".

[0087] 5) Equivalent expression: Use the factor comparison method to establish an equivalent relationship between known patterns.

[0088] The principle of the element comparison method is that based on the elements and context of a known fact (the source expression), through corpus retrieval or online search, expressions with the same meaning but different forms (the target expression) can be obtained, namely equivalent expressions. The semantic pattern of the source expression and the semantic pattern of the target expression then constitute an equivalent pattern.

[0089] For example, based on "A was born in Pennsylvania", we may find "A's birthplace is Pennsylvania", so we can know that the semantic pattern "@f|subj: @sb + p: born + in + obj: @sw" and the semantic pattern "@f|o: @sb + of + a: birthplace + is + v: @sw" constitute an equivalent pattern.

[0090] In summary, based on role recognition modeling, we first identify semantic patterns in simple sentences, then phrase patterns. Using isomorphic substitution, we bootstrap further semantic patterns, and finally, compare elements to obtain equivalent patterns. This allows us to form a hierarchical semantic structure through nesting, enabling us to process a variety of complex nested natural sentences. Furthermore, based on this hierarchical semantic structure and context, we can perform long-range semantic analysis such as default recovery and coreference resolution, enabling accurate understanding of paragraphs and even chapters.

[0091] 3. System Modules and Functions 1. Ontology vocabulary construction and management module Responsible for the construction, updating, and maintenance of the ontology vocabulary. During the construction phase, machine learning models, such as fine-tuning the BERT model, are used to accurately identify unregistered terms from a large corpus. Combined with remote supervision and manual annotation, the module determines the categories and categories of terms, as well as their relationships, such as hyponymy and synonymy. As new data continues to emerge, this module will automatically identify more terms and their relationships from the newly added corpus, select high-confidence samples to expand the annotated dataset, and dynamically update and maintain the ontology vocabulary to ensure its timeliness and accuracy.

[0092] 2. Semantic pattern discovery and management module It is used to automatically discover new semantic patterns from corpus. It first identifies the semantic patterns of simple sentences. Then, based on these existing semantic patterns, it applies pattern comparison to identify nested components in complex sentences. For example, given the semantic pattern of "I like red apples," when encountering "I like the red apples my mother bought," it can identify "my mother bought" as a new nested component. It then uses structural inference to construct the component structure and semantic patterns of these nested components, such as determining the component division structure of "my mother bought apples." Then, through bootstrapping and isomorphic substitution, it generates more new patterns based on known similar patterns, such as generating a "like" pattern from a "like" pattern. It can also use element comparison to establish equivalence between known patterns, such as determining the semantic pattern equivalence between "I ate the apple" and "the apple was eaten by me," thereby continuously enriching the semantic pattern library.

[0093] 3. Semantic parsing engine module (1) Tuple module It implements text segmentation and named entity recognition, handles token segmentation ambiguity, and provides a foundational unit for subsequent analysis. It can accurately break down input text into the smallest analysis units. For example, for the sentence "On October 1, 2024, a grand event was held in Region A," it can accurately identify "October 1, 2024" as the time entity and "Region A" as the location entity, providing a clear data foundation for subsequent sentence segmentation and pattern matching.

[0094] (2) Sentence block splitting module The purpose of the sentence chunking module is to accurately break down the input text into logically complete sentence chunks. Each sentence chunk contains a relatively independent semantic unit, providing clear and explicit basic data for subsequent pattern matching and semantic analysis. These logical sentence chunks can better reflect the semantic structure of the text and facilitate a deeper understanding of the text content.

[0095] Therefore, the sentence chunk splitting module should not only be divided according to punctuation marks and named entity boundaries, but also consider the processing of special structures and format control characters, especially the processing of non-standard punctuation marks.

[0096] Segmentation based on punctuation and named entity boundaries: This approach uses punctuation recognition, such as periods, commas, and semicolons, combined with named entity boundary information, to perform a preliminary segmentation of the text. Using the named entity recognition results, we can avoid splitting entities into different sentence chunks, thus ensuring the rationality of the sentence chunking.

[0097] Special structure processing: In texts, special structures such as brackets and quotation marks are relatively common. If not handled properly, it will affect the accuracy of sentence segmentation. For brackets, the sentence segmentation module will determine the logical relationship between the content in the brackets and the context. If the brackets are a supplement to the previous content, the brackets and the content they modify will be divided into the same sentence block; if the brackets contain independent information, they will be processed as a separate sentence block. For quotation, if the quotation marks contain a complete quoted sentence, it will be treated as an independent sentence block; if the quotation marks contain a phrase or vocabulary, the sentence blocks will be reasonably divided according to its role in the sentence and the context. For example, in the sentence "He said: "Today's meeting is very important", "Today's meeting is very important" will be identified as an independent sentence block.

[0098] Format Control Characters and Non-Standard Punctuation Processing: In actual text, redundant format control characters often appear, especially carriage return and line feed characters. The sentence block splitting module can accurately identify these redundant format control characters and ignore their impact on sentence block division, ensuring the logical coherence of the text. At the same time, for non-standard use of punctuation marks, such as multiple periods in a row, or confusion between commas and periods, the module will make judgments based on the semantics and grammatical structure of the text, merging multiple physical sentence blocks that logically belong to a complete sentence into a single logical sentence block, preventing the incorrect use of punctuation marks from interfering with the understanding of the text content.

[0099] (3) Pattern matching module The module matches the input text layer by layer based on the semantic pattern library, generates candidate parse tree nodes, and constructs a hierarchical semantic structure. This module analyzes the input text based on various patterns in the semantic pattern library, such as event patterns and modification patterns. When processing the sentence "Xiao Ming is reading a book in the library," it quickly matches the event pattern "@f|subj: @sb +loc: @loc + p: reading + obj: @bk," identifies "subj: Xiao Ming," "p: reading," "obj: books," and "loc: library," creates corresponding parse tree nodes, and constructs a hierarchical semantic structure.

[0100] (4) Context window recognition processing This module processes semantic connections across short-range cross-sentence blocks within a text at the chapter level, optimizing parsing boundaries through a context window to enhance chapter comprehension. When processing a chapter, a context window is maintained. For example, for a sentence like "Apple has released a new phone with powerful performance," this module can use the context window to determine that the "it" in the second sentence refers to the "new phone" in the first sentence. This allows it to optimize parsing boundaries, accurately grasp semantic connections across sentence blocks, and enhance overall chapter comprehension.

[0101] (5) Likelihood analysis module Since actual input may contain word order changes, missing content, interference characters, or unregistered words or entries, the system needs to perform likelihood analysis on the candidate semantic patterns in the schedule area corresponding to the context window, thereby extracting a likelihood parse tree from the context window.

[0102] The specific measures are as follows: 1) Try to combine candidate patterns in the schedule area to determine the maximum likelihood matching structure. This strategy is suitable for situations where there are word order changes and interference characters in the actual input.

[0103] 2) Try to trigger the likelihood of the immature branch covering the entire context window. This strategy is suitable for situations where the actual input has missing content compared to the known pattern.

[0104] 3) Based on the context, candidate semantic patterns are used to guess candidate suspected unregistered words or unregistered entries.

[0105] A more complicated situation is that the recovered parsing structure is only a nested layer. In this case, the branch needs to be submitted to the schedule area again, which is a relatively time-consuming process.

[0106] (6) Ambiguity resolution and optimization module We employ various strategies to resolve ambiguity and optimize the various possible parsing results during pattern matching. We prioritize parsing results with simpler semantic hierarchies, using strategies like maximizing mutual information to measure the degree of association between words and select the parsing path with the highest mutual information value.

[0107] 4. Contextual Association Analysis Module Conduct long-range contextual analysis within the text. This refers to spanning paragraphs and chapters. Specifically, this involves clarifying the specific objects referred to by pronouns in the text through reference resolution, as well as restoring missing information to improve the semantic structure. For example, if the preceding paragraph mentions "he bought a book" and the subsequent paragraph mentions "he gave the book to her," the "book" in "he gave the book to her" can be accurately understood to refer to the "book" in "he bought a book," making the semantics clearer and more accurate.

[0108] 5. Multimodal Fusion Module This module can fuse natural language with other modal information, such as speech and images. When processing speech input, the speech is first converted into a structural pattern, and then processed using the system's semantic analysis process. At the same time, information such as the intonation and speaking speed of the speech are combined to assist in semantic understanding. For image information, image features, such as the shape, color, and position of the object, are extracted to form a structural pattern, which is then associated and matched with natural language descriptions to achieve collaborative understanding of multimodal information. For example, when the input is "Describe the location of the red car in the picture", the module can fuse the visual features of the car in the image with the text instructions, accurately understand and answer the question, and provide support for multimodal learning and reasoning.

[0109] 6. Reasoning and Generation Module This module converts hierarchical semantic structures into predicate logic expressions (FOLEs), supporting applications such as semantic querying, semantic reasoning, and causal learning. It achieves the conversion of natural language into structured knowledge, and vice versa. For example, taking the sentence "Xiao Ming ate an apple, and the apple was red," this module converts its hierarchical semantic structure into the first-order predicate logic expression "Eat(subj: Xiao Ming, obj: ?x) ∧ isa(?x, apple) ∧ Red(subj: ?x)." This supports semantic queries such as "Query the color of what Xiao Ming ate," as well as reasoning based on this information, such as inferring that Xiao Ming ate a red object. This provides a foundation for applications such as causal learning.

[0110] 1. Implementation steps of semantic analysis method (1) Input processing: Convert PDF, DOCX and other format documents into plain text through tuples The module segments words and identifies named entities, generating a token array and a named entity list.

[0111] (2) Sentence segmentation: Split the text into logical Sentence blocks, handle special structures such as brackets and quotation marks, and ensure the accuracy of sentence blocks.

[0112] (3) Pattern matching and parse tree construction: Traverse tokens and sentence blocks from left to right, attempting to match them with semantic patterns in the pattern library. For successfully matched patterns, create parse tree nodes, with child nodes representing the semantic and grammatical components on the right side of the pattern and parent nodes representing the semantic results on the left side. For example, when processing "Edison invented the incandescent lamp," match the event pattern and construct a parse tree with "@f" as the parent node and child nodes "subj: Edison," "p: invention," and "obj: incandescent lamp."

[0113] 1) Context window processing: When processing a passage, the context window is maintained and sentence blocks are gradually added. When the current sentence block cannot match the content in the window, the existing parse tree is output, the window is cleared, and the new sentence block is processed to ensure correct semantic association across sentence blocks.

[0114] 2) Ambiguity resolution and result optimization: For multiple possible parse trees, the optimal parse result is selected through strategies such as prioritizing fewer semantic levels and maximizing mutual information, and reference resolution and default recovery are performed to improve the semantic structure.

[0115] 2. System module implementation details (1) Ontology lexicon construction module: This module uses machine learning models for term recognition and category classification, builds training data using a combination of remote supervision and manual annotation, and supports dynamic lexicon updates. For example, by fine-tuning the BERT model, the accuracy of term recognition can be improved, and the ontology lexicon can be automatically expanded.

[0116] (2) Semantic Pattern Library Management Module: This module stores various semantic patterns and their parameters, and supports adding, deleting, modifying, and querying patterns. It automatically discovers new semantic patterns and enriches the pattern library through pattern comparison and structure deduction algorithms.

[0117] (3) Parsing engine module: Based on semantic patterns and ontology lexicon, it implements hierarchical semantic parsing of text and generates a parsing tree list. It supports multiple parsing strategies to adapt to the complexity of different input texts.

[0118] Taking "Kingfa Technology's earnings per share in Q4 2024 is 3.5 yuan" as an example, the hierarchical semantic analysis process is explained: 1. Tuple formation and named entity recognition: The following sentence is used for word segmentation: "Kingfa Technology", "2024Q4", "earnings per share", "3.5 yuan", and "Kingfa Technology" is identified as the company name, "2024Q4" as the time period, "earnings per share" as the indicator, and "3.5 yuan" as the value.

[0119] 2. Pattern matching: Match the oav fact pattern "@f|o: @stock + time: @time + of + a: @attribute + @comparison operator + val: @value" to determine "o: Kingfa Technology", "time: 2024Q4", "a: earnings per share", and "val: 3.5 yuan".

[0120] 3. Parse tree construction: Generate a parse tree with the root node semp_type as "s.ia" (a declarative sentence with an indicator word predicate), hw as "earnings per share", and child nodes containing various semantic components and their roles.

[0121] 4. Result conversion: Convert the parse tree into a first-order predicate logic expression to support subsequent semantic queries (such as "Query Kingfa Technology's earnings per share for Q4 2024") and reasoning (such as comparing earnings changes across quarters and performing causal analysis or trend prediction).

[0122] Through hierarchical semantic parsing and cross-domain semantic model adaptation, the present invention significantly improves the depth and robustness of semantic analysis, reduces dependence on massive data and domain knowledge, supports complex nested structure parsing and multimodal collaborative reasoning, and provides efficient and explainable semantic processing capabilities for scenarios such as intelligent customer service, information retrieval, and text comprehension, while having the advantages of low cost and strong generalization.

Claims

1. A hierarchical semantic analysis method based on semantic patterns, characterized by: The following steps are involved: Automatically obtain a semantic model, wherein the semantic model is composed of a left-hand semantic result and a right-hand grammatical component and a semantic component, wherein the right-hand part includes grammatical and semantic components such as a subject, a predicate, and an object, and the left-hand part represents the semantic result of the right-hand part. The left-hand semantic result and the right-hand component together constitute a semantic structure, and semantic nesting is achieved based on the semantic structure; Based on the semantic pattern library, the input text is matched layer by layer to generate candidate parse tree nodes and construct a hierarchical semantic structure. The root node of the parse tree is the overall semantics of the sentence, the non-leaf nodes are the left part of the semantic pattern, and the leaf nodes are the instantiation results of the semantic components. The parse tree supports cross-sentence context association and multi-layer nested structure deduction; The hierarchical semantic structure is converted into a predicate logic expression, including: identifying the semantic roles and relationships of each node in the parse tree, generating entity, attribute and relationship assertions, and retaining the hierarchical nested structure through logical operators to form a reasonable logical expression.

2. The hierarchical semantic analysis method based on semantic patterns according to claim 1, characterized in that: The semantic patterns include but are not limited to the following types: Event pattern: used to describe the semantic relationship between action subject, predicate and object; Modification pattern: used to analyze the modification relationship between attributives, adverbials and central words; Group pattern: used to identify component combinations in parallel, progressive, or selective relationships; Synthetic mode: used to process a combination of several parallel components.

3. The hierarchical semantic analysis method based on semantic patterns according to claim 1, characterized in that: The method for constructing the semantic pattern library includes but is not limited to the following steps: Identify the semantic patterns contained in simple sentences: Using the role recognition model, first identify the roles of the unnested components of simple sentences to obtain the semantic patterns of simple sentences; Pattern comparison method: Based on natural language expressions with known semantic patterns, find similar expressions using different nested components, and reasonably infer that the different nested components used in similar expressions are also legal expressions; Structural deduction method: Based on the heuristic rules of legal phrases, the component structure is divided and combined with the role recognition model to determine the semantic pattern implied by the phrase; Bootstrapping: Exploiting the morphological similarity of isomorphic semantic patterns, new patterns are generated using known patterns as parameterized templates; Element equivalence method: Establish equivalent semantic patterns of different expressions through corpus retrieval and element comparison.

4. A hierarchical semantic analysis system based on semantic patterns, characterized by: Includes the following modules: Tuple module: used for text segmentation, named entity recognition, and handling token segmentation ambiguity; Sentence block splitting module: uses punctuation marks, named entity boundaries and special structures to divide logical sentence blocks, and handles non-standard punctuation and format control characters; Pattern matching module: matches the input text layer by layer based on the semantic pattern library, generates candidate parse tree nodes and constructs a hierarchical semantic structure; Context window processing module: This module processes the semantic associations between short-range cross-sentence blocks in the text at the chapter level and optimizes the parsing boundaries through the context window. Likelihood analysis module: performs candidate pattern splicing and immature branch triggering on unmatched inputs, and extracts the maximum likelihood parse tree; Ambiguity resolution module: optimizes parsing results using strategies such as semantic hierarchy minimum priority and mutual information maximization; Multimodal fusion module: converts speech and image features into structural patterns and performs collaborative reasoning with natural language parsing results; Dynamic update module: expands the ontology vocabulary through fine-tuning of pre-trained models and automatically extracts new semantic patterns based on pattern inference algorithms; Reasoning and Generation Module: Converts hierarchical semantic structures into predicate logic expressions, supporting semantic query, semantic reasoning, and causal learning.

5. The hierarchical semantic analysis system based on semantic patterns according to claim 4, characterized in that: The specific implementation of the sentence block splitting module includes: Identify punctuation marks in named entities to avoid mis-splitting; Merge multiple physical sentence blocks caused by non-standard punctuation into a single logical sentence block; Process the content in brackets and quotation marks, and determine the sentence block affiliation based on semantic logic.

6. The hierarchical semantic analysis system based on semantic patterns according to claim 4, characterized in that: The nested matching of the pattern matching module includes the following steps: Traverse the token sequence in the order of the text and match it layer by layer with the patterns in the semantic pattern library; Generate a parse tree node for the successfully matched pattern, where its child node is the right semantic component and its parent node is the left semantic result; Matching is performed recursively on the nested structure until the leaf nodes are indecomposable semantic components or domain terms.