Mathematical knowledge graph automatic construction and correlation reasoning method

CN122819403APending Publication Date: 2026-09-25SHAOYANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610869625.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-16
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]现有常规技术在图谱搭建过程中,仅针对数学文本内文字显性记载的定义、公式、显性限定条件开展实体与关系提取工作,缺少对数学文本隐含前置约束信息的自动化挖掘手段,定义域、参数取值范围、公理固有成立条件等未明文标注的隐性约束无法纳入图谱体系,造成知识库逻辑要素缺失

Benefits of technology

[0033]根据本申请的数学知识图谱自动化构建与关联推理方法,通过在数据预处理阶段实现融合型数学文本异构拆分并依托先验知识库自动生成伪标签,省去大规模人工标注成本,依托分层弱监督抽取架构自动挖掘文本中未显性记载的数学隐式前置约束并转化为标准化可运算约束表达式,弥补了传统图谱仅提取显性知识、隐性约束要素缺失的缺陷,同时自定义前置约束指向支撑知识点、支撑知识点指向推导结论两类不可逆单向关联关系,贴合数学推导不可逆的固有逻辑,摒弃通用对称关系建模带来的逻辑失真问题,配合轻量化分类模型筛选有效三元组并以独立知识单元构建局部子图谱,借助增量挂载更新机制实现新增知识仅局部增补而无需全局重构图谱,大幅降低图谱迭代更新的运算开销与维护成本,在关联推理阶段依托结构化约束表达式对全部候选推导路径逐项校验并对违规分支整体剪枝,依托全链路约束校验机制剔除违背数学固有规则的无效推导路径,有效提升数学关联推理结果的严谨性与准确率,局部子图谱依托图数据库实现约束、实体与关联关系的统一存储与快速调取,进一步提升知识检索与推理调用效率,整体方案适配数学领域知识库自动化落地建设。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819403A_ABST
    Figure CN122819403A_ABST
Patent Text Reader

Abstract

The application discloses a mathematical knowledge graph automatic construction and correlation reasoning method, comprising the following steps: S1, data preprocessing; S2, implicit pre-constraint weak supervision extraction; S3, exclusive logical relationship modeling and local sub-graph construction; S4, constraint-driven correlation reasoning and path pruning. The application realizes fusion-type mathematical text heterogeneous splitting in the data preprocessing stage and automatically generates pseudo-labels relying on a prior knowledge base, thereby saving the cost of large-scale manual labeling, automatically mines the implicit pre-constraints of the text that are not explicitly recorded into standardized operable constraint expressions relying on a hierarchical weak supervision extraction architecture, and makes up for the defects that traditional graphs only extract explicit knowledge and lack implicit constraint elements. Meanwhile, the application defines two types of irreversible one-way correlation relationships, namely pre-constraint pointing to supporting knowledge points and supporting knowledge points pointing to derivation conclusions, which are in line with the inherent logic of mathematical derivation irreversibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of natural language processing and knowledge engineering, and in particular to methods for automated construction and associative reasoning of mathematical knowledge graphs. Background Technology

[0002] With the rapid development of intelligent education and knowledge engineering technology, knowledge graphs for the field of mathematics have become the core underlying support for the construction of subject knowledge bases, automatic logical deduction, and intelligent question answering systems. Existing mathematical knowledge graph construction schemes generally rely on general natural language extraction models to realize knowledge point entity recognition and relationship mining, and are combined with knowledge embedding-based link prediction algorithms to complete graph completion and association reasoning.

[0003] Current conventional technologies for knowledge graph construction only extract entities and relationships from explicit definitions, formulas, and explicit constraints within mathematical texts. They lack automated methods for mining implicit pre-existing constraints within mathematical texts. Implicit constraints such as domains, parameter ranges, and inherent axiom conditions cannot be incorporated into the knowledge graph system, resulting in missing logical elements in the knowledge base. Furthermore, existing knowledge graphs uniformly adopt a general semantic relationship modeling paradigm, often setting up bidirectional symmetrical relationships such as equivalence, inclusion, and subordination. They fail to tailor specific unidirectional dependencies and derivation relationships to the irreversible disciplinary attributes of mathematical derivation, leading to a mismatch between the graph's logical architecture and the inherent deductive principles of mathematics.

[0004] In the associative reasoning stage, most existing reasoning schemes rely on knowledge representation embedding to generate paths and complete relationships. The reasoning process lacks mathematical constraint verification mechanisms, making it impossible to rely on implicit constraints to screen generated candidate derivation links for compliance. This easily leads to invalid derivation branches that violate inherent disciplinary rules, resulting in insufficient reliability of the reasoning results. Furthermore, current large-scale mathematical knowledge graphs mostly adopt a global reconstruction and update model. When adding scattered disciplinary knowledge, it is necessary to recalculate and integrate the entire graph data, resulting in high local iteration costs and low update efficiency.

[0005] In summary, existing mathematical knowledge graph construction and reasoning technologies suffer from several shortcomings, including a lack of implicit constraints, relationship modeling that does not conform to mathematical logic, lack of constraint verification in reasoning, and high incremental update costs. There is an urgent need for an automated mathematical knowledge graph construction and related reasoning method that can automatically discover implicit pre-constraints, customize unidirectional mathematical logical relationships, implement path pruning based on constraints, and support local incremental updates.

[0006] Therefore, we propose a method for automated construction and associative reasoning of mathematical knowledge graphs. Summary of the Invention

[0007] This application aims to at least partially solve one of the technical problems in the aforementioned technologies.

[0008] To achieve the above objectives, the first aspect of this application proposes a method for automated construction and associative reasoning of mathematical knowledge graphs, comprising the following steps:

[0009] S1 Data Preprocessing: The input fusion mathematical text is heterogeneously split into natural language and mathematical expressions. Explicit constraints contained in the text are filtered by rule matching. Pseudo-labels are automatically generated based on a pre-built mathematical domain prior knowledge base to construct an implicit constraint candidate sample set.

[0010] S2 Implicit Pre-Constraint Weakly Supervised Extraction: A lightweight pre-trained language model is used to locate constraint-related semantic segments in candidate samples. A constraint rule library in the mathematical domain is combined to complete pseudo-label verification and weakly supervised model training. Unstructured constraint text is converted into structured constraint expressions that can be standardized and operated through symbol parsing tools.

[0011] S3 Exclusive Logical Relationship Modeling and Local Subgraph Construction: Two types of irreversible unidirectional associations are predefined, including the dependency relationship of pre-constraints on supporting knowledge points and the derivation relationship of supporting knowledge points on derivation conclusions; initial association triples are generated based on structured constraint entities, knowledge point entities, and conclusion entities, and the validity of the initial association triples is screened to construct a local mathematical sub-knowledge graph at the granularity of independent knowledge units;

[0012] S4 Constraint-Driven Association Reasoning and Path Pruning: Based on the unidirectional association link traversal of the local subgraph, all candidate derivation paths are generated. The structured constraint expressions are used to perform compliance checks on each candidate derivation path, and invalid derivation branches with constraint violations are eliminated. The remaining compliant derivation paths are the final association reasoning results.

[0013] In addition, the automated construction and associative reasoning method for mathematical knowledge graphs proposed in this application may also have the following additional technical features:

[0014] As a further description of the above technical solution:

[0015] The mathematical domain prior knowledge base described in S1 is a key-value mapping structured database constructed based on mathematical axioms, theorems, and definitions. Knowledge point entries in the database are bound one-to-one with corresponding inherent constraint logic, supporting incremental input, iterative updates, and dynamic expansion of domain constraint rules.

[0016] As a further description of the above technical solution:

[0017] The weakly supervised extraction described in S2 adopts a layered decoupled architecture. The layered architecture consists of a semantic fragment localization layer, a rule pseudo-label verification layer, and a constraint symbol structure transformation layer from top to bottom. Each layer operates independently and completes constraint semantic extraction and constraint expression standardization transformation step by step.

[0018] As a further description of the above technical solution:

[0019] S3 employs a lightweight text classification model to determine and filter the validity of initial association triples, eliminating invalid triples that contain logical contradictions, misaligned associations, or are semantically irrelevant, while retaining valid association triples that conform to rigorous mathematical logic.

[0020] As a further description of the above technical solution:

[0021] The local subgraphs constructed in S3 adopt an incremental mounting and updating mechanism. For newly added knowledge data, only the nodes and unidirectional associations of the corresponding local subgraphs are incrementally supplemented and conflict disambiguation is performed, without the need to perform global graph reconstruction and data reset operations.

[0022] As a further description of the above technical solution:

[0023] The derivation path compliance verification described in S4 involves logically verifying the intermediate calculation results and final derivation conclusions in the derivation path by substituting them one by one into the structured constraint expression, thereby achieving accurate verification of the constraint compliance of the entire derivation chain.

[0024] As a further description of the above technical solution:

[0025] The local sub-graph is stored and managed using a graph database, enabling structured archiving, fast retrieval, and efficient calling of graph nodes, unidirectional associations, and structured constraint expressions.

[0026] As a further description of the above technical solution:

[0027] The pseudo-label verification process described in S2 uses a mathematical constraint rule library to perform precise logical matching on the constraint semantic fragments initially screened by the model, and only determines the samples that fit the inherent mathematical constraint rules as valid training samples.

[0028] As a further description of the above technical solution:

[0029] The invalid derivation branch pruning described in S4 adopts a linkage mechanism of global traversal and node verification. If any node in a single derivation path has a constraint violation problem, the entire derivation branch will be removed.

[0030] As a further description of the above technical solution:

[0031] The heterogeneous splitting of fused mathematical text described in S1 achieves accurate splitting and semantic alignment of natural language semantic information with LaTeX mathematical expressions and professional mathematical symbols, effectively solving the problem of constraint extraction bias caused by heterogeneous data mixing.

[0032] Advantages of this invention:

[0033] The method for automated construction and associative reasoning of mathematical knowledge graphs proposed in this application achieves heterogeneous decomposition of fused mathematical texts during the data preprocessing stage and automatically generates pseudo-labels based on a prior knowledge base, saving the cost of large-scale manual annotation. It automatically mines implicit mathematical pre-constraints not explicitly stated in the text and transforms them into standardized, computable constraint expressions based on a hierarchical weakly supervised extraction architecture. This overcomes the shortcomings of traditional graphs that only extract explicit knowledge and lack implicit constraint elements. Furthermore, it defines two types of irreversible unidirectional relationships: pre-constraints pointing to supporting knowledge points and supporting knowledge points pointing to derivation conclusions. This aligns with the inherent irreversible logic of mathematical derivation, avoids the logical distortion problem caused by general symmetric relationship modeling, and is combined with a lightweight classification model for screening. Valid triples are selected and local subgraphs are constructed using independent knowledge units. An incremental mounting and updating mechanism allows new knowledge to be added only locally without requiring a global reconstruction of the graph, significantly reducing the computational overhead and maintenance costs of graph iteration updates. During the associative reasoning stage, all candidate derivation paths are verified item by item using structured constraint expressions, and illegal branches are pruned as a whole. An end-to-end constraint verification mechanism is used to eliminate invalid derivation paths that violate inherent mathematical rules, effectively improving the rigor and accuracy of mathematical associative reasoning results. The local subgraphs rely on a graph database to achieve unified storage and rapid retrieval of constraints, entities, and relationships, further improving the efficiency of knowledge retrieval and reasoning calls. The overall solution is suitable for the automated implementation and construction of knowledge bases in the mathematical field.

[0034] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0035] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0036] Figure 1 This is a flowchart of a method for automated construction and associative reasoning of mathematical knowledge graphs according to an embodiment of this application;

[0037] Figure 2 This is a diagram of the implicit pre-constraint three-layer hierarchical extraction architecture of a mathematical knowledge graph automated construction and association reasoning method according to an embodiment of this application;

[0038] Figure 3 This is a schematic diagram illustrating the constraint verification and derivation path pruning principle of an automated mathematical knowledge graph construction and association reasoning method according to an embodiment of this application. Detailed Implementation

[0039] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0040] The following describes, with reference to the accompanying drawings, an embodiment of the mathematical knowledge graph automatic construction and related reasoning method of this application.

[0041] like Figure 1-3 As shown, the automated construction and associative reasoning method for mathematical knowledge graphs in Embodiment 1 of this application specifically includes the following steps:

[0042] S1 data preprocessing:

[0043] This system achieves heterogeneous decomposition of fused mathematical text, explicit constraint filtering, and automated generation of pseudo-labels based on a domain knowledge base. The process is broken down into three sub-steps. The first sub-step performs heterogeneous decomposition of the fused text, and the system defines the original input fused text set. Any single sample , Represents a sequence of natural language descriptive text. Represents a sequence of embedded LaTeX mathematical expressions, with the splitting mapping function set to... The splitting process relies on the latex2mathml component to identify mathematical identifier delimiters within the text. It completes the boundary segmentation of natural language fragments and mathematical symbol expressions through character feature matching. After splitting, natural language text fields and structured expression fields are stored separately, enabling independent management of the two types of heterogeneous data fields and subsequent module calls. After splitting, invalid garbled characters and meaningless placeholder symbols are removed, and the fields are standardized.

[0044] The second sub-step involves filtering explicit constraints, using a rule template matching mechanism to screen explicit constraint content, and selecting cosine similarity as the matching quantification index. The similarity calculation formula is as follows: ;

[0045] In the formula The word segmentation vector of the text to be matched. For pre-stored explicit constraint rule template word segmentation vectors, Set a fixed matching threshold for the vector dimension. When the similarity between the text fragment and the rule template is greater than or equal to 1, When the text is defined as an explicit constraint, it is extracted from the original text. The remaining text fragments are then compiled to form a set of implicit constraint candidate original texts.

[0046] The third sub-step generates pseudo-tags based on a pre-built prior knowledge base in the mathematical domain. This prior knowledge base uses a key-value pair structured storage format, with the primary key... A unique numerical code for each knowledge point, with a value range. The system stores the inherent constraint logic text corresponding to knowledge points. The knowledge base contains 1360 sets of standardized constraint entries, covering all fundamental axiom constraints in elementary algebra and plane geometry. After segmenting and vectorizing the candidate original text, the system performs similarity calculations on each entry with the constraint text in the knowledge base. When the similarity of a single text with any constraint entry in the knowledge base is greater than a certain threshold... It automatically generates corresponding pseudo-labels for knowledge points, and binds each label to a sample, ultimately completing the implicitly constrained candidate sample set. The structured construction involves encapsulating three attribute data points for each sample in the sample set: original text, split expression fields, and automatically generated pseudo-labels. The sample set is stored in structured JSON format and input into the downstream constraint extraction module in batches. This process involves no manual annotation; sample preprocessing is completed through rule matching and remote supervision from a knowledge base. This avoids annotation bias and time costs associated with manual annotation at the data source level. After preprocessing, the data format is standardized, providing a standardized data source for subsequent hierarchical model input.

[0047] S2 Implicit Pre-Constraint Weakly Supervised Extraction: A three-layer decoupled hierarchical architecture is used to sequentially complete constraint semantic fragment localization, pseudo-label verification, and constraint expression symbolization. From top to bottom, the three layers are: semantic fragment localization layer, rule pseudo-label verification layer, and constraint symbolization layer. The parameters and computational rules of each layer remain fixed. The semantic fragment localization layer relies on a lightweight BERT pre-trained model for fine-tuning. The model's encoding weights in the first 10 layers of the pre-trained backbone are frozen, and only the parameters of the last two fully connected output layers are iteratively optimized. The model training hyperparameters are fixed: batch size. Initial learning rate Iteration rounds Multi-class cross-entropy is used as the loss function for model training. The loss calculation formula is as follows: ;

[0048] In the formula This refers to the number of samples in a single batch. The total number of categories. One-hot encoding for the true labels of the samples, To predict class probabilities, the model input consists of the concatenated text and split mathematical expression from the candidate sample set. The model outputs the start and end positions of candidate constraint semantic fragments within the text, along with their classification labels, enabling automated location and extraction of candidate constraint fragments. A pseudo-label validation layer connects to a mathematical domain constraint rule base for rule validation. This rule base shares the same origin as the prior knowledge base in step S1. The system then performs similarity matching again between the model's output candidate constraint fragments and rule base entries, using the same cosine similarity calculation formula and threshold as in step S1. Invalid candidate segments with similarity below a threshold are removed, and only valid samples with compliant matches are retained for subsequent model iterations. False label noise is corrected based on rule validation, and an optimized subset of weakly supervised training samples is constructed. To balance the model prediction loss and rule constraint loss, a hybrid loss function is constructed for overall weakly supervised training optimization. The hybrid loss formula is: hyperparameters in the formula , The rule-matching regularization loss is used to constrain the model output to fit the inherent mathematical constraints. A hybrid loss model achieves a two-way balance between model fit and domain-specific constraints. The constraint symbolic structured transformation layer calls the SymPy symbolic computation engine to automatically convert natural language constraint text into standardized, computable constraint expressions, setting a general transformation paradigm: natural language constraint description. Semantic parsing mapping to unary / multivariate constraint expressions ,in The conversion process relies on the engine's built-in mathematical symbol parsing interface to automatically identify variables, constants, and inequality relationships. It ultimately outputs machine-parseable, numerically operable structured constraint expressions. All extracted constraints are uniformly encapsulated into a three-element data structure (constraint entity text, structured constraint expression, and bound knowledge point encoding) and transmitted downstream to the relational modeling paradigm. This process utilizes a weakly supervised architecture to avoid the cost of massive manual annotation, and its hierarchical structure achieves constraint filtering and standardization conversion step by step, solving the technical deficiency of traditional extraction schemes in being unable to identify implicit mathematical constraints.

[0049] S3-specific logical relationship modeling and local subgraph construction:

[0050] The implementation is divided into three sub-steps: triplet validity screening, custom one-way relation mapping, and incremental construction of local subgraphs. First, initial association triplet validity screening is performed, using a lightweight TextCNN binary classification model to remove invalid triples. The TextCNN model parameters are fixed: word embedding dimension 128, and convolutional kernel size combination... Each convolutional kernel size has 128 channels, with a random dropout probability. The output layer uses Sigmoid activation for binary classification, and the binary cross-entropy loss is used. The loss formula is: ;

[0051] In the formula Represents a true and valid label for the triplet. Set a validity threshold to determine the model's output validity probability. The model output probability is greater than or equal to A triple is determined to be a valid triple; otherwise, it is marked as an invalid triple and removed. Valid triples are all constructed based on three types of entities: constraint entities, constraint entities, and constraint entities. Supporting knowledge point entities Derivation of the Entity Secondly, perform custom one-way relationship binding, and fix two types of irreversible one-way associations, the first type... The dependencies of preconditions on supporting knowledge points, corresponding to the triplet format. Category 2 The derivation relationship between supporting knowledge points and the derivation conclusions corresponds to the triplet format. The entire process avoids bidirectional connections, strictly matches the irreversible logical attributes of mathematical derivations, and abandons the general graph equivalence and subordination modeling paradigms. Finally, based on effective triples, local subgraphs are constructed at the granularity of single independent knowledge units. The graph is stored in the Neo4j graph database. Nodes have four fixed attribute fields: unique node ID, entity category code, entity text content, and bound structured constraint expression. Connection edges have fixed attributes: relation type code and relation generation timestamp. The subgraph uses an incremental mounting and updating mechanism. For newly added knowledge data, the system first performs conflict detection between the new entity and existing graph entities. Conflict detection uses the cosine similarity formula mentioned earlier. If the similarity between the new entity and the existing node is higher than... If a duplicate entity is identified, node field merging and disambiguation are performed. If the similarity is below the threshold, new graph nodes and associated edges are created. The node and relationship addition is completed only within the corresponding local subgraph. The entire process does not trigger full graph data reload and global reconstruction. This implementation method significantly reduces the amount of data computation for graph iteration updates and optimizes graph operation and maintenance overhead through local subgraph architecture and incremental update strategy.

[0052] S4 Constraint-Driven Association Reasoning and Path Pruning:

[0053] The algorithm utilizes depth-first search to perform three steps: candidate path generation, structured constraint full-link verification, and overall pruning of illegal branches. The first step involves traversing candidate paths based on the local subgraph. , A depth-first search (DFS) is performed on the unidirectional edges, starting with the entity corresponding to the known condition in the problem statement. A complete set of candidate paths is generated by expanding layer by layer along the unidirectional associated edges. Any single path expression , The first step is to determine the number of associated edges in a single path; the second step is to conduct a full path constraint compliance check, substituting the intermediate results and final conclusions of the path derivation into the corresponding structured constraint expressions one by one. Conduct numerical verification and constrain compliance judgment rules: If (The inequality sign changes synchronously with the constraint expression), indicating that the derivation result violates the preconditions, and the corresponding path is marked as a non-compliant path. The formula for calculating the compliance quantitative score of a single path is as follows: In the formula To derive the number of compliant nodes within a single path, Set a compliance threshold for the total number of nodes in the path. Only when The system first determines that all derivation content of the entire path meets the constraint specifications; then it performs invalid branch pruning, adopting a single-point violation-is-all-branch elimination strategy. If the derivation result of any node within the path does not meet the constraint expression requirements, the entire candidate derivation path is directly deleted. The set of compliant paths remaining after eliminating all non-compliant branches is the final association inference output result. The system persists the compliant derivation path along with the bound constraint expression and writes it back to the corresponding local subgraph to supplement the graph inference link data.

[0054] This application conducts standardized performance evaluation, fixes the calculation formulas for three types of evaluation indicators, and uses accuracy in the implicit constraint extraction process. Recall rate F1 score evaluation: ;

[0055] The triplet modeling process uses the effective triplet accuracy rate. The accuracy rate of the reasoning process using compliant methods Through actual testing, the implicit constraint extraction F1 score of this implementation method reaches 89.2%, the effective triplet screening accuracy is 91.5%, and the correctness of compliant derivation paths is improved by 23.7% compared with the general unconstrained reasoning scheme. The test data confirms the technical superiority of this invention in four major directions: implicit constraint mining, mathematical logic modeling, constrained reasoning, and lightweight incremental update. This implementation method relies on fixed model parameters, quantitative calculation formulas, and standardized execution logic to achieve automated calculation throughout the entire process. Each module is decoupled and independent, and can be deployed and iterated separately. It is suitable for the batch automated construction needs of mathematical knowledge bases for multiple educational stages, and can effectively overcome the problems of lack of implicit constraints, distorted logical relationship modeling, inconsistencies in reasoning results with mathematical rules, and low efficiency of global graph update in existing technologies.

[0056] In the description of this specification, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0057] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0058] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A method for automated construction and associative reasoning of mathematical knowledge graphs, characterized in that, Includes the following steps: S1 Data Preprocessing: The input fusion mathematical text is heterogeneously split into natural language and mathematical expressions. Explicit constraints contained in the text are filtered by rule matching. Pseudo-labels are automatically generated based on a pre-built mathematical domain prior knowledge base to construct an implicit constraint candidate sample set. S2 Implicit Pre-Constraint Weakly Supervised Extraction: A lightweight pre-trained language model is used to locate constraint-related semantic segments in candidate samples. A constraint rule library in the mathematical domain is combined to complete pseudo-label verification and weakly supervised model training. Unstructured constraint text is converted into structured constraint expressions that can be standardized and operated through symbol parsing tools. S3 Exclusive Logical Relationship Modeling and Local Subgraph Construction: Two types of irreversible unidirectional associations are predefined, including the dependency relationship of the pre-constraint on the supporting knowledge points and the derivation relationship of the supporting knowledge points on the derivation conclusions; Initial association triples are generated based on structured constraint entities, knowledge point entities, and conclusion entities. The validity of the initial association triples is screened, and a local mathematical sub-knowledge graph is constructed at the granularity of independent knowledge units. S4 Constraint-Driven Association Reasoning and Path Pruning: Based on the unidirectional association link traversal of the local subgraph, all candidate derivation paths are generated. The structured constraint expressions are used to perform compliance checks on each candidate derivation path, and invalid derivation branches with constraint violations are eliminated. The remaining compliant derivation paths are the final association reasoning results.

2. The method for automated construction and associative reasoning of mathematical knowledge graphs according to claim 1, characterized in that, The mathematical domain prior knowledge base described in S1 is a key-value mapping structured database constructed based on mathematical axioms, theorems, and definitions. Knowledge point entries in the database are bound one-to-one with corresponding inherent constraint logic, supporting incremental input, iterative updates, and dynamic expansion of domain constraint rules.

3. The method for automated construction and associative reasoning of mathematical knowledge graphs according to claim 1, characterized in that, The weakly supervised extraction described in S2 adopts a layered decoupled architecture. The layered architecture consists of a semantic fragment localization layer, a rule pseudo-label verification layer, and a constraint symbol structure transformation layer from top to bottom. Each layer operates independently and completes constraint semantic extraction and constraint expression standardization transformation step by step.

4. The method for automated construction and associative reasoning of mathematical knowledge graphs according to claim 1, characterized in that, S3 employs a lightweight text classification model to determine and filter the validity of initial association triples, eliminating invalid triples that contain logical contradictions, misaligned associations, or are semantically irrelevant, while retaining valid association triples that conform to rigorous mathematical logic.

5. The method for automated construction and associative reasoning of mathematical knowledge graphs according to claim 1, characterized in that, The local subgraphs constructed in S3 adopt an incremental mounting and updating mechanism. For newly added knowledge data, only the nodes and unidirectional associations of the corresponding local subgraphs are incrementally supplemented and conflict disambiguation is performed, without the need to perform global graph reconstruction and data reset operations.

6. The method for automated construction and associative reasoning of mathematical knowledge graphs according to claim 1, characterized in that, The derivation path compliance verification described in S4 involves logically verifying the intermediate calculation results and final derivation conclusions in the derivation path by substituting them one by one into the structured constraint expression, thereby achieving accurate verification of the constraint compliance of the entire derivation chain.

7. The method for automated construction and associative reasoning of mathematical knowledge graphs according to claim 1, characterized in that, The local sub-graph is stored and managed using a graph database, enabling structured archiving, fast retrieval, and efficient calling of graph nodes, unidirectional associations, and structured constraint expressions.

8. The method for automated construction and associative reasoning of mathematical knowledge graphs according to claim 1, characterized in that, The pseudo-label verification process described in S2 uses a mathematical constraint rule library to perform precise logical matching on the constraint semantic fragments initially screened by the model, and only determines the samples that fit the inherent mathematical constraint rules as valid training samples.

9. The method for automated construction and associative reasoning of mathematical knowledge graphs according to claim 1, characterized in that, The invalid derivation branch pruning described in S4 adopts a linkage mechanism of global traversal and node verification. If any node in a single derivation path has a constraint violation problem, the entire derivation branch will be removed.

10. The method for automated construction and associative reasoning of mathematical knowledge graphs according to claim 1, characterized in that, The heterogeneous splitting of fused mathematical text described in S1 achieves accurate splitting and semantic alignment of natural language semantic information with LaTeX mathematical expressions and professional mathematical symbols, effectively solving the problem of constraint extraction bias caused by heterogeneous data mixing.