A knowledge graph construction method and system for subject education resources
By using neural sequence annotation and multi-scale semantic coverage curvature index calculation, the problems of cross-version concept alignment and relation extraction in the construction of knowledge graphs for multi-version subject education resources were solved, forming an integrated knowledge system that adapts to the updates of education resources and enhancing the interpretability and accuracy of the knowledge graph.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YUNNAN NORMAL UNIV
- Filing Date
- 2026-03-20
- Publication Date
- 2026-06-02
Smart Images

Figure CN122132583A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of knowledge graph construction technology, and more specifically, to a method and system for constructing knowledge graphs for subject-specific educational resources. Background Technology
[0002] Multiple version updates of subject-specific educational resources are commonplace in the education field. Textbooks are revised and iterated along with curriculum standards, and supplementary teaching materials and online question banks are continuously optimized and supplemented. Different versions of resources carry subject knowledge in a scattered manner. Constructing cross-version knowledge graphs can integrate scattered knowledge, establish a unified knowledge system, and provide support for applications such as intelligent teaching and personalized learning. This is an important requirement for the current intelligent development of education.
[0003] However, existing knowledge graph construction methods have significant limitations. These methods often directly extract knowledge points from educational resources and simply group them into graph nodes. Cross-version concept alignment relies solely on word-based matching or single-scale semantic similarity calculation, while relation extraction focuses on local text features or single node embedding, without considering the special structure and multi-version variation characteristics of educational resources.
[0004] In actual updates, concepts from multiple versions of resources are prone to granularity drift. The same concept may be split due to adjustments in teaching needs; for example, the properties of functions in the old version may be split into independent concepts such as monotonicity and parity in the new version. It may also be merged or renamed, leading to differences in the semantic coverage and expansion methods of the same concept across different versions. Existing methods lack the ability to capture the multi-scale semantic coverage of concepts and lack effective means to quantify granularity differences. Alignment based solely on literal meaning or single-scale semantic similarity easily leads to mismatches between parent and child concepts, and between merged and split concepts. Furthermore, relation extraction fails to consider granularity differences, applying fine-grained relational relationships to coarse-grained concepts, resulting in relation mismatches and a chaotic graph structure. These problems make nodes in cross-version graphs incomparable and relations inaccurate, hindering the effective integration of knowledge from multiple versions of resources and making it difficult to support reliable educational applications. Summary of the Invention
[0005] This invention provides a method and system for constructing knowledge graphs for subject-based educational resources, thereby solving the technical problems mentioned in the background.
[0006] This invention provides a method for constructing knowledge graphs for subject-based educational resources, comprising the following steps: Step S101: Perform neural sequence annotation on multi-source subject education resources, locate the text span of concept mentions, and simultaneously extract the structural positioning information of each text span in the resource chapter level; Step S102: Extract a multi-scale context view based on the structural positioning information, and encode the multi-scale context view into a multi-scale distributed vector set by combining the scale embedding vector. At the same time, encode the resource fragments to generate resource fragment vectors. Step S103: Calculate the average intrinsic dimension that varies with the neighborhood scale based on the nearest neighbor distance sequence of the vector set, and perform second-order difference summation on the average intrinsic dimension change curve to obtain the multi-scale semantic coverage curvature index. Step S104: Map the multi-scale semantic coverage curvature index to a query vector and input it into the curvature modulation aggregation network. Calculate the attention weights of each point in the vector set. Perform weighted aggregation on the vector set accordingly to generate concept node vectors and output granular scalars simultaneously. Step S105: Calculate the cosine similarity between concept node vectors of different versions, and subtract the absolute value of the difference between the corresponding multi-scale semantic coverage curvature exponents as curvature penalty terms. Establish cross-version mapping edges based on the calculation results. Step S106: The concept node vector, granularity scalar and multi-scale semantic coverage curvature index are concatenated into a comprehensive feature vector to predict and establish concept relationship edges. At the same time, evidence edges are established according to the matching degree between the concept node vector and the resource fragment vector, and merged with cross-version mapping edges to generate a knowledge graph.
[0007] This invention provides a knowledge graph construction system for subject-based educational resources, comprising: The structural positioning information extraction module performs neural sequence annotation on multi-source subject education resources, locates the text span of concept mentions, and simultaneously extracts the structural positioning information of each text span in the resource chapter level; The resource fragment vector generation module extracts a multi-scale context view based on structural positioning information, and encodes the multi-scale context view into a multi-scale distributed vector set by combining the scale embedding vector. At the same time, it encodes the resource fragments to generate resource fragment vectors. The semantic coverage curvature index calculation module calculates the average intrinsic dimension that varies with the neighborhood scale based on the nearest neighbor distance sequence of the vector set, and performs second-order difference summation on the average intrinsic dimension change curve to obtain the multi-scale semantic coverage curvature index. The concept node vector generation module maps the multi-scale semantic coverage curvature index to query vectors and inputs them into the curvature modulation aggregation network. It calculates the attention weights of each point in the vector set and performs weighted aggregation on the vector set to generate concept node vectors, and outputs granular scalars simultaneously. The cross-version mapping edge construction module calculates the cosine similarity between concept node vectors of different versions and subtracts the absolute value of the difference between the corresponding multi-scale semantic coverage curvature exponents as a curvature penalty term, and establishes cross-version mapping edges based on the calculation results. The knowledge graph generation module concatenates concept node vectors, granular scalars, and multi-scale semantic coverage curvature indices into a comprehensive feature vector to predict and establish concept relationship edges. At the same time, it establishes evidence edges based on the matching degree between concept node vectors and resource fragment vectors, and merges them with cross-version mapping edges to generate a knowledge graph.
[0008] The beneficial effects of this invention are as follows: This invention quantifies the semantic coverage morphology and granularity differences of concepts through a multi-scale semantic coverage curvature index, providing a key basis for cross-version concept alignment and the extraction of relationships between concepts within the same version. Cross-version mapping combines semantic similarity and curvature consistency to reduce association bias caused by granularity mismatch. Same-version relationship prediction integrates semantic, granular, and morphological features, making the dependencies or associations between concepts more consistent with knowledge logic. The construction of evidence edges establishes a traceable link between concepts and original educational resources, enhancing the interpretability of knowledge. The resulting cross-version knowledge graph integrates scattered subject knowledge from multiple versions, with complete node attributes, clear relationships, and resource support. It adapts to the characteristics of multiple version updates of educational resources, providing a structurally sound and reusable knowledge system for applications such as intelligent teaching and personalized learning. Attached Figure Description
[0009] Figure 1 This is a flowchart of a knowledge graph construction method for subject-based educational resources according to the present invention; Figure 2 This is a schematic diagram of a knowledge graph construction system for subject-based educational resources according to the present invention; Figure 3 This is a schematic diagram of the computational logic of the present invention.
[0010] In the diagram: 201 Structural positioning information extraction module, 202 Resource fragment vector generation module, 203 Semantic coverage curvature index calculation module, 204 Concept node vector generation module, 205 Cross-version mapping edge construction module, 206 Knowledge graph generation module. Detailed Implementation
[0011] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.
[0012] It should be noted that, unless otherwise defined, the technical or scientific terms used in one or more embodiments of the present invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in one or more embodiments of the present invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" indicate that the element or object preceding the term encompasses the elements or objects listed following the term and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0013] like Figures 1-3 As shown, a method for constructing a knowledge graph for subject-based educational resources includes the following steps: Step S101: Perform neural sequence annotation on multi-source subject education resources, locate the text span of concept mentions, and simultaneously extract the structural positioning information of each text span in the resource chapter level; Step S102: Extract a multi-scale context view based on the structural positioning information, and encode the multi-scale context view into a multi-scale distributed vector set by combining the scale embedding vector. At the same time, encode the resource fragments to generate resource fragment vectors. Step S103: Calculate the average intrinsic dimension that varies with the neighborhood scale based on the nearest neighbor distance sequence of the vector set, and perform second-order difference summation on the average intrinsic dimension change curve to obtain the multi-scale semantic coverage curvature index. Step S104: Map the multi-scale semantic coverage curvature index to a query vector and input it into the curvature modulation aggregation network. Calculate the attention weights of each point in the vector set. Perform weighted aggregation on the vector set accordingly to generate concept node vectors and output granular scalars simultaneously. Step S105: Calculate the cosine similarity between concept node vectors of different versions, and subtract the absolute value of the difference between the corresponding multi-scale semantic coverage curvature exponents as curvature penalty terms. Establish cross-version mapping edges based on the calculation results. Step S106: The concept node vector, granularity scalar and multi-scale semantic coverage curvature index are concatenated into a comprehensive feature vector to predict and establish concept relationship edges. At the same time, evidence edges are established according to the matching degree between the concept node vector and the resource fragment vector, and merged with cross-version mapping edges to generate a knowledge graph.
[0014] In one embodiment of the present invention, neural sequence annotation is performed on multi-source subject-specific educational resources to locate the text span of concept references, and the structural positioning information of each text span in the resource chapter hierarchy is extracted simultaneously, including: For a single resource fragment text Constructing a lexical sequence And calculate the vector sequence: in The length of the word sequence. For neural network encoding functions, For the set of parameters of the neural network encoding function, For word sequence The corresponding vector sequence, For the first Vector representation of each word element; For any satisfying Text span Calculate the vector representation of the text span: in For text span The vector representation of , The vector representation of the word at the starting position. The vector representation of the term at the terminating position. For a vector sequence within the text span The average pooling result, This is a vector concatenation operation; Calculate the probability of mentioning a concept across text spans: in The concept of probability is mentioned. and For linear scoring parameters, It is a logical function; Constructing structural location information for text span: Depend on , , , , composition; in For structural positioning information, , , , , The respective text spans Chapter-level tags, section-level tags, subsection-level tags, paragraph-level tags, and sentence-level tags; Generate a collection of concept mention text spans: in The concept refers to a set of text spans, where each element contains both text span and text span. Concept mention probability With structural positioning information .
[0015] It should be noted that a single resource fragment text is a text fragment related to subject education, reflecting specific educational content; it can be collected through techniques such as crawling educational resource databases, user uploads, or obtaining data from official interfaces. A word sequence is the sequence of the smallest semantic units after splitting a single resource fragment text. A neural network encoding function is a function that performs vector transformation on the word sequence. The vector sequence is the output of the neural network encoding function, composed of vector representations of words. The vector representation of a word is a numerical vector form of a single word, reflecting the semantic information of that single word. The text span is a continuous word fragment in the word sequence, reflecting the text range that may belong to the concept mentioned. The starting position is the starting index of the text span in the word sequence, reflecting the starting boundary of the text span. The ending position is the ending index of the text span in the word sequence, reflecting the ending boundary of the text span. Linear scoring parameters are the weight vector and bias term used for scoring the text span vector. The logistic function is a function that maps the result of linear operations to probability values, i.e., the Sigmoid function. The vector representation of the text span is a concatenated vector of the word vector representation at the start position, the word vector representation at the end position, and the average pooling result.
[0016] It should be noted that chapter-level tags identify the chapter to which the text spans, reflecting the chapter-level position of the text span; they can be collected through techniques such as parsing the table of contents and formatting tags (e.g., chapter title tags) of educational resources. Section-level tags identify the section to which the text spans, reflecting the section-level position of the text span; they can be collected through techniques such as parsing the table of contents and formatting tags (e.g., section title tags) of educational resources. Subsection-level tags identify the subsection to which the text spans, reflecting the subsection-level position of the text span; they can be collected through techniques such as parsing the table of contents and formatting tags (e.g., subsection title tags) of educational resources. Paragraph-level tags identify the paragraph to which the text spans, reflecting the paragraph-level position of the text span; they can be collected through techniques such as parsing paragraph separators and formatting tags (e.g., paragraph tags) of educational resources. Sentence-level tags identify the sentence to which the text spans, reflecting the sentence-level position of the text span; they can be collected through techniques such as parsing sentence-ending punctuation (e.g., periods, question marks, exclamation marks) and formatting tags (e.g., sentence tags) of educational resources. Structural location information is a combination of chapter-level markers, section-level markers, subsection-level markers, paragraph-level markers, and sentence-level markers. The concept mention text span set is a collection that includes text span, concept mention probability, and structural location information.
[0017] It should be noted that the neural network encoding function preferentially adopts the BERT-type Transformer encoder, with a hidden layer dimension of 768, 12 layers, 12 attention heads, and a FeedForward network dimension of 3072. This structure is mature in semantic encoding tasks in the field of natural language processing, effectively capturing long-distance dependencies and contextual semantic relationships between words, and is suitable for the text semantic features of educational resources. The dimension of the weight vector is consistent with the dimension of the word vector representation (e.g., 768 dimensions), and it is initialized using a Xavier normal distribution; the initial value of the bias term is set to 0.1, with a value range of 0 to 0.5; this initialization method can avoid gradient vanishing or gradient exploding problems in the early stage of training, ensuring stable convergence of the model. The screening threshold for concept mention probability is set to 0.5. When the concept mention probability of a text span is greater than or equal to 0.5, the text span is included in the concept mention text span set; when the probability is less than 0.5, it is discarded. In addition, chapter level markers use two Arabic numerals (01 to 99), section level markers use two Arabic numerals (01 to 99), subsection level markers use two Arabic numerals (01 to 99), paragraph level markers use two Arabic numerals (01 to 99), and sentence level markers use two Arabic numerals (01 to 99). The combination format is chapter-section-subsection-paragraph-sentence, for example, 03-01-02-05-04 represents the fourth sentence of the fifth paragraph of the first section of the third chapter; for example, if a text spans the first chapter of a high school physics textbook, the second section of linear motion, the first subsection of uniform linear motion, the third paragraph, and the second sentence, its five-level markers can be combined as 01-02-01-03-02.
[0018] It should be noted that this invention uses neural sequence annotation technology to locate the text span of concept mentions in multi-source subject-specific educational resources, and simultaneously extracts five-level structural location information, laying the foundation for subsequent multi-scale context truncation and multi-scale semantic coverage curvature index calculation. The process involves splitting the text into a sequence of lexical units, encoding them into vectors via a neural network, and obtaining the probability of concept mentions through feature concatenation, linear scoring, and logical mapping, ultimately forming a candidate set containing text span, probability, and structural location information. This preserves the specific location information of concept mentions, avoids the cross-version granularity drift problem caused by directly defining knowledge graph nodes, provides accurate candidate data and location basis for subsequent steps, adapts to the hierarchical structure characteristics of subject-specific educational resources, and ensures the stability and reliability of subsequent knowledge graph construction.
[0019] In one embodiment of the present invention, a multi-scale context view is extracted based on structural positioning information, and the multi-scale context view is encoded into a multi-scale distributed vector set by combining scale embedding vectors. Simultaneously, resource fragments are encoded to generate resource fragment vectors, including: The set of text spans that mention the concept any element in Capture multi-scale contextual views: in For text span The sentence, For inclusion paragraphs, For inclusion The section, For inclusion In the chapter, the above four items are based on structural positioning information. get; For each scale corresponding to the multi-scale context view Calculate multi-scale vectors: in For multi-scale vectors, For neural network encoding functions, For the set of parameters of the neural network encoding function, For the corresponding scale The scale embedding vector, This is a vector 2-norm normalization operation; Aggregate all resource texts and the collection of text spans mentioned in the concept The elements in the dataset are used to generate a multi-scale distributed vector set: in It is a set of vectors distributed across multiple scales; For each resource text Calculate the resource fragment vector: ; in For resource fragment vectors, Consistent with the neural network encoding function in multi-scale vector computation, and constructing a set of resource fragment vectors. .
[0020] It should be noted that the multi-scale context view is a four-layer text set determined based on structural localization information. The scale embedding vector is a learnable vector that distinguishes different context scales. The multi-scale vector is the final vector after vector addition and normalization of the encoding result. The multi-scale distributed vector set is a set formed by aggregating multi-scale vectors from all versions and all concept mentions. The resource fragment vector is the vector representation of a single resource fragment text after neural network encoding, reflecting the overall semantic information of the resource fragment. The pair storage result is the associated storage data of a single resource fragment text and its corresponding resource fragment vector, reflecting the one-to-one mapping relationship between text and vector. This invention extracts a multi-scale context view based on structural localization information, generates multi-scale vectors through encoding, scale fusion, and normalization, and aggregates them to form a multi-scale distributed vector set. Simultaneously, the same encoder is reused to generate resource fragment vectors, providing basic data for subsequent multi-scale semantic coverage curvature index calculation and evidence edge construction. The process involves starting from the structural localization information of concept mentions, obtaining four layers of context text, encoding, fusing scale features and normalizing, aggregating multi-scale vectors from all versions, and simultaneously encoding and pairing resource fragments for storage. This approach leverages the inherent hierarchical structure of educational resources to construct a fixed multi-scale view, giving multi-scale features a clear physical meaning. The combination of scale embedding and normalization improves vector quality, and the reused encoder ensures semantic space consistency. The convergence of multi-scale vectors forms a set of data that supports curvature calculation, which will not be elaborated upon here.
[0021] It should be noted that the dimension of the scale embedding vector is consistent with that of the multi-scale vector (e.g., 768 dimensions); the initialization uses a normal distribution with a mean of 0 and a variance of 0.01; the learning update rule is to update it along with other parameters through the backpropagation algorithm during model training, with the update step size consistent with the parameter update step size of the neural network encoding function (e.g., 0.001); this setting ensures that the scale embedding vector can adaptively learn the feature differences at different scales. The encoding length limits for different scale text content in the multi-scale context view are as follows: sentence scale (scale 0) is limited to 512 tokens, paragraph scale (scale 1) is limited to 1024 tokens, section scale (scale 2) is limited to 2048 tokens, and chapter scale (scale 3) is limited to 4096 tokens; when the limit is exceeded, a tail truncation method is used; this limit adapts to the input requirements of mainstream encoders, while balancing encoding efficiency and semantic integrity. The paired storage uses a text identifier-vector data association format. The text identifier is a unique number for the resource fragment (composed of version number + fragment sequence number, such as V1-0035), and the vector data is stored as a floating-point array. The overall storage uses JSON format, with the text identifier as the key and the vector floating-point array as the value. The deduplication rule for the multi-scale distributed vector set uses a cosine similarity threshold. The cosine similarity between the vector to be added and the existing vectors in the set is calculated. If the maximum value is greater than 0.99, it is determined to be a duplicate vector and is not added; otherwise, the vector is retained. This rule can effectively remove redundant vectors, reduce computation, and preserve the diversity of conceptual semantics.
[0022] In one embodiment of the present invention, the average intrinsic dimension varying with neighborhood scale is calculated based on the nearest neighbor distance sequence of the vector set, and the average intrinsic dimension variation curve is summed by second-order difference to obtain the multi-scale semantic coverage curvature index, including: Let the set of vectors of the multi-scale distribution be . For any vector in the set With the remaining vectors Calculate the Euclidean distance: ; in The number of vectors, For the L2 norm, it will be targeted After sorting all distances, let's denote the distances with... The The distance between the nearest neighbors is ; Define the neighborhood scale set: ; in For floor operations, The number of vectors in a multi-scale distributed vector set; For any vector With any scale in the neighborhood scale set Calculate the local intrinsic dimension estimate: in For local intrinsic dimension estimation, For natural logarithm operations, For the first The distance between nearest neighbors. For the first The distance between nearest neighbors; Computational scale The average intrinsic dimension is: in The average intrinsic dimension; make Neighborhood-scale set Calculate the second difference for three adjacent scale values: And calculate the multi-scale semantic coverage curvature index: in It is a multi-scale semantic coverage curvature index.
[0023] It should be noted that Euclidean distance is the straight-line distance between two vectors in a multi-scale distributed vector set. The distance of the k-th nearest neighbor is the Euclidean distance between a given vector and the k-th nearest vector in the set after sorting, reflecting the local density characteristics of the vector at the k-neighbor scale. The neighborhood scale set is a numerical set that grows in powers of two and is limited by the square root of the number of vectors, reflecting the neighborhood range at different levels. The distance of the j-th nearest neighbor is the Euclidean distance between a given vector and the j-th nearest vector in the set after sorting, reflecting the distance distribution characteristics of the vector within a local range. The local intrinsic dimension estimate is the locally effective dimension of the vector calculated based on the logarithm of the nearest neighbor distance ratio, reflecting the complexity of the local region where the vector is located. The average intrinsic dimension is the arithmetic mean of the local intrinsic dimension estimates of all vectors in the set, reflecting the average complexity of the entire vector set at a specific neighborhood scale. The second difference is the quadratic difference between the average intrinsic dimensions corresponding to three adjacent scales, reflecting the bending strength of the average intrinsic dimension change curve. The absolute value of the second-order difference is its non-negative form, reflecting the absolute intensity of the curve's curvature. The multi-scale semantic coverage curvature index is the sum of the absolute values of all second-order differences, reflecting the overall curvature of the vector set as the neighborhood scale changes.
[0024] It should be noted that the minimum value of the neighborhood scale set is 2. A minimum scale of 2 avoids the influence of outliers caused by a single nearest neighbor and effectively captures local spatial features, ensuring the stability of the local intrinsic dimension estimation. Furthermore, starting from 2, powers of 2 are successively used until the maximum value does not exceed the square root of the number of vectors, thus adapting to the size of the vector set and avoiding the mixing of small-scale noise and large-scale global noise, ensuring the reasonableness of the intrinsic dimension estimation. The reasonable range for the local intrinsic dimension estimation is 0.1 to 20. When the local intrinsic dimension estimation is less than 0.1, it is corrected to 0.1; when it is greater than 20, it is corrected to 20. This range covers the common complexity range of educational resource vector sets. The correction rule avoids the interference of extreme outliers on the calculation of the average intrinsic dimension, ensuring the reliability of the subsequent curvature exponent. The reasonable range for the curvature index of multi-scale semantic coverage is 0 to 50. When the calculated result is greater than 50, it is determined to be an invalid value. In this case, the median of all valid curvature indices of the vector set is used as the replacement. This range is adapted to the complexity characteristics of the multi-scale vector set of educational resources. The invalid value replacement rule can handle abnormal results in extreme cases and ensure the availability of the curvature index.
[0025] It should be noted that the neighborhood scale set, which grows in powers of two and is limited by the square root of the number of vectors, includes: first, calculating the square root of the number of vectors in the multi-scale distributed vector set; then, taking the logarithm of this square root and rounding it down to the nearest integer; using 2 as the base and this integer as the exponent to obtain the maximum scale value; then, starting from 2, generating all values that do not exceed the maximum scale value in powers of two to form the neighborhood scale set; for example, if the number of vectors is 100, the square root is 10, the logarithm is rounded down to 3, the maximum scale value is 8, and the neighborhood scale set is 2, 4, 8. The process of obtaining the multi-scale semantic coverage curvature index by summing the second-order differences of the mean intrinsic dimension variation curve involves: first, calculating the mean intrinsic dimension corresponding to each neighborhood scale to form the mean intrinsic dimension variation curve; then, selecting the mean intrinsic dimension corresponding to three adjacent scale values in the curve, calculating the second-order difference and taking the absolute value; finally, summing all the absolute values to obtain the multi-scale semantic coverage curvature index. For example, if the neighborhood scale set is 2, 4, 8, and 16, the corresponding mean intrinsic dimensions are 1.2, 1.8, 2.1, and 2.0, the second-order differences are 0.3 and -0.5, respectively, and the sum of the absolute values is 0.8, that is, the multi-scale semantic coverage curvature index is 0.8.
[0026] In one embodiment of the present invention, a multi-scale semantic coverage curvature index is mapped to a query vector and input into a curvature modulation aggregation network. Attention weights are calculated for each point in the vector set. Based on this, a weighted aggregation is performed on the vector set to generate concept node vectors, and a granular scalar is output synchronously, including: Using multi-scale semantic coverage curvature index With a set of vectors distributed at multiple scales Calculate the query vector: ; in For query vector, and To generate parameters to be learned for the query, Let be the number of elements in a multi-scale distributed vector set. It is the natural logarithm of the number of elements. This is a vector concatenation operation. It is the hyperbolic tangent function; Calculate the first vector in the multi-scale distribution vector set. vectors Attention weights: in For attention weights, It is an exponential function. For query vector and the first The inner product of vectors, It is a scaling constant that is the square root of the vector dimension; Perform weighted aggregation on a set of vectors with multi-scale distributions: ; in For concept node vectors; Calculate particle size scalar: ; in For particle size scalar, and Generate the parameters to be learned at the granular level. For logical functions, concept node vectors With particle size scalar As input for subsequent steps.
[0027] It should be noted that the query-generated learning parameters are the weight matrix and bias terms used to generate the query vector. The query vector is a vector that integrates the multi-scale semantic coverage curvature exponent and the vector set size information, reflecting the comprehensive characteristics of concept complexity and sample size. The inner product is the result of multiplying the query vector and the multi-scale vectors element-wise and then summing them, reflecting the semantic similarity between the two vectors. The scaling result is the value obtained by dividing the inner product by the square root of the vector dimension, reflecting the normalized similarity characteristics. The exponential function value is the result of taking the exponent of the scaling result, reflecting the exponential amplification of similarity. The attention weight is the ratio of the exponential function value of a single vector to the sum of the exponential function values of all vectors, reflecting the importance of that vector in the aggregation. The concept node vector is the vector obtained by weighting and summing the multi-scale vectors with attention weights, reflecting the comprehensive semantic and geometric characteristics of the concept. The granularity-generated learning parameters are the weight vector and bias terms used to generate the granularity scalar. The granularity scalar is a value in the range of 0 to 1, reflecting the granularity tendency and coverage complexity characteristics of the concept.
[0028] It should be noted that the specific dimensions, initialization methods, and training update rules for generating the parameters to be learned from the query include: the weight matrix dimension is vector dimension (e.g., 768) × 2, and the bias term dimension is 768; during initialization, the weight matrix adopts a Xavier normal distribution, and the initial value of the bias term is set to 0.1; the training update rule is to update the parameters together with other learnable parameters of the model through the backpropagation algorithm, with the update step size consistent with the encoder parameters (e.g., 0.001). The specific dimensions and initialization methods for generating the parameters to be learned at the granular level include: the weight vector dimension is vector dimension (e.g., 768) + 1, and the bias term dimension is 1; during initialization, the weight vector adopts a Xavier normal distribution, and the initial value of the bias term is set to 0.1. The inner product of the query vector and the multi-scale vector is prone to excessively large values in high-dimensional space. Scaling using the square root of the vector dimension can reduce the numerical magnitude and avoid softmax function saturation. The process of generating a query vector by concatenating the multi-scale semantic coverage curvature index with the natural logarithm of the number of elements involves: first obtaining the multi-scale semantic coverage curvature index (scalar) and the natural logarithm of the number of elements in the multi-scale distributed vector set (scalar); concatenating the two scalars in sequence into a two-dimensional vector; then performing linear calculations with the query-generated learning parameters (weight matrix and bias term); and finally processing the query vector using the hyperbolic tangent function.
[0029] It should be noted that this invention uses the multi-scale semantic coverage curvature index as the core control variable to drive the curvature modulation aggregation network. By fusing concept complexity and sample size information, a query vector is generated. Adaptive attention weights are calculated, and the multi-scale vectors are weighted and aggregated to obtain concept node vectors. Simultaneously, a granular scalar is generated, providing a core representation for subsequent cross-version mapping and relational edge construction. The process involves concatenating the multi-scale semantic coverage curvature index and the natural logarithm of the number of elements to generate a query vector. After scaling and normalization, attention weights are obtained. After aggregating to generate node vectors, the multi-scale semantic coverage curvature index is concatenated again to generate a granular scalar. The attention weights can be dynamically adjusted according to the semantic coverage shape of the concept, enabling the concept node vector to capture both semantic and geometric features. The granular scalar accurately extracts the granularity tendency of the concept, avoiding the limitations of traditional fixed aggregation and improving the adaptability and stability of concept representation. Further details are omitted here.
[0030] In one embodiment of the present invention, the cosine similarity between concept node vectors of different versions is calculated, and the absolute value of the difference between the corresponding multi-scale semantic coverage curvature indices is subtracted as a curvature penalty term. Based on the calculation results, cross-version mapping edges are established, including: Set version The concept node vector set is And the corresponding set of multi-scale semantic coverage curvature indices is , set version The concept node vector set is And the corresponding set of multi-scale semantic coverage curvature indices is ,in and Each represents the quantity of a concept; For any pair Calculate the cosine similarity: in For cosine similarity, It is the vector norm 2; Calculate the curvature penalty term: ; in For curvature penalty terms; Calculate the mapping score: ; in Score the mapping; Calculate edge weights: in For border rights, It is an exponential function; Establish a cross-version mapping edge set: in This is a cross-version mapping edge set used to merge and generate a knowledge graph.
[0031] It should be noted that the first version's concept node vector set is the set of node vectors corresponding to all concepts in the first version, reflecting the comprehensive semantic and geometric features of the first version's concepts. The first version's multi-scale semantic coverage curvature index set is the set of multi-scale semantic coverage curvature indices corresponding to all concepts in the first version, reflecting the semantic coverage morphological features of the first version's concepts. The second version's concept node vector set is the set of node vectors corresponding to all concepts in the second version, reflecting the comprehensive semantic and geometric features of the second version's concepts. The second version's multi-scale semantic coverage curvature index set is the set of multi-scale semantic coverage curvature indices corresponding to all concepts in the second version, reflecting the semantic coverage morphological features of the second version's concepts. The second norm of the first version's concept node vectors is the square root of the sum of the squares of each element of the first version's concept node vector, reflecting the magnitude of the vector. The second norm of the second version's concept node vectors is the square root of the sum of the squares of each element of the second version's concept node vector, reflecting the magnitude of the vector. Cosine similarity is the ratio of the inner product to the product of the second norms of the two vectors, reflecting the semantic similarity between the two vectors. The curvature penalty term is the absolute value of the difference in the multi-scale semantic coverage curvature index between the two version concepts, reflecting the degree of difference in the semantic coverage morphology of cross-version concepts. The mapping score is the result of subtracting the curvature penalty term from the cosine similarity, reflecting the comprehensive matching degree of cross-version concepts. The exponential function value is the result of exponentially taking the mapping score, reflecting the exponential amplification feature of the comprehensive matching degree. The edge weight is the ratio of the exponential function value of a single mapping score to the sum of the exponential function values of all corresponding concepts in the second version, reflecting the strength of the cross-version mapping relationship. The starting point marker is a unique identifier identifying the starting point of the cross-version mapping edge (first version concept), reflecting the beginning of the mapping relationship; it can be generated using the combination rule of "first version number - concept index". The ending point marker is a unique identifier identifying the ending point of the cross-version mapping edge (second version concept), reflecting the end of the mapping relationship; it can be generated using the combination rule of "second version number - concept index". The cross-version mapping edge set is the set of all cross-version mapping edges, reflecting the complete mapping relationship between the two version concepts.
[0032] It should be noted that constructing the mapping score by combining cosine similarity and curvature penalty term involves: firstly, calculating the cosine similarity (semantic dimension) and curvature penalty term (morphological dimension) for cross-version concept pairs separately; then, subtracting the curvature penalty term from the cosine similarity to obtain the mapping score, achieving a dual consideration of semantic matching and morphological consistency. For example, if the multi-scale semantic coverage curvature index of concept A in version 1 is 0.6, and the multi-scale semantic coverage curvature index of concept B in version 2 is 0.5, with an absolute difference of 0.1 (curvature penalty term), their cosine similarity is 0.85, and the mapping score is 0.85 - 0.1 = 0.75; if concept C and A also have a cosine similarity of 0.85, but the absolute difference in their multi-scale semantic coverage curvature indices is 0.3, and the mapping score is 0.55, then B is preferentially chosen as the mapping object for A; thus avoiding the granularity mismatch problem caused by relying solely on semantic similarity. The weight coefficient of the curvature penalty term is 0.1, that is, mapping score = cosine similarity - 0.1 × curvature penalty term; the basis for this value is to balance the influence of semantic similarity and curvature consistency, to avoid the penalty term being too large to cover up semantic association, or too small to suppress granular mismatch, and to adapt to the semantic and morphological characteristics of the concept of educational resources.
[0033] It should be noted that the valid range for the mapping score is -1 to 1. A score less than -0.5 is considered an outlier; the outlier is corrected to -0.5 to avoid interference from extreme values in edge weight calculation. The edge weight filtering threshold is 0.1. Edges with a weight greater than or equal to 0.1 are retained; edges with a weight less than 0.1 are discarded to avoid redundancy in the graph structure caused by weakly associated mapping edges. The start and end markers use a combination format of "version number-concept index". The version number is a two-digit Arabic numeral (01 to 99), and the concept index is a four-digit Arabic numeral (0001 to 9999). For example, V1-0008 represents the 8th concept in the first version, and V2-0012 represents the 12th concept in the second version. The storage structure of the cross-version mapping edge set adopts JSON format. The top layer is the version pair identifier (such as V1-V2), and the lower layer is the list of starting point markers. Each starting point marker corresponds to a key-value pair of ending point marker and edge weight. The association query rules support querying by starting point marker, ending point marker, and edge weight range. For example, query all mapping edges corresponding to V1-0008, or mapping edges with edge weight greater than 0.3.
[0034] It should be noted that this invention constructs a comprehensive mapping score by combining the semantic similarity of cross-version concepts with the consistency of multi-scale semantic coverage curvature. The edge weights are then quantified, and the start and end markers are paired to form a cross-version mapping edge set. This achieves precise association between concepts from different versions, solving the problem of concept incomparability caused by cross-version granularity drift. The process involves collecting the concept node vector sets and multi-scale semantic coverage curvature index sets of two versions, calculating cosine similarity and curvature penalty terms, obtaining a mapping score, quantifying it into edge weights, and then pairing them to construct a mapping edge set. The dual-dimensional scoring mechanism improves the rationality of cross-version mapping, avoiding granularity mismatch caused by a single semantic dimension; the quantification of edge weights makes the mapping strength clearly distinguishable, facilitating the selection of effective mappings; and the structured mapping edge set is suitable for subsequent knowledge graph merging needs, conforming to the characteristics of multi-version updates of subject education resources, and providing a stable mapping foundation for constructing a unified cross-version knowledge graph.
[0035] In one embodiment of the present invention, concept node vectors, granularity scalars, and multi-scale semantic coverage curvature exponents are concatenated into a comprehensive feature vector to predict and establish concept relationship edges. Simultaneously, evidence edges are established based on the matching degree between concept node vectors and resource fragment vectors, and these are merged with cross-version mapping edges to generate a knowledge graph, including: For any ordered pair of concepts in the same version Construct a comprehensive feature vector: in For the comprehensive feature vector, and For concept node vectors, and For particle size scalar, and It is a multi-scale semantic coverage curvature index. This is a vector concatenation operation. This is for absolute value operations; For a set of relation types Each relation type in Calculate the unnormalized score and the relationship probability: in For unnormalized scores, and For linear transformation parameters, For relational probability, It is an exponential function; Form a set of concept relation edges: ; in For the set of concept relation edges; For any concept With any resource text Calculate the degree of matching: ; in For the degree of matching, For logical functions, For resource fragment vectors; Forming a set of evidence edges: ; in For the evidence edge set; With concept node set Cross-version mapping edge set For input, perform a merge operation: ; in For knowledge graphs.
[0036] It should be noted that the comprehensive feature vector is a high-dimensional vector formed by concatenating the concept node vector, granularity scalar, and absolute values of the differences between the two types of vectors, reflecting the comprehensive semantic, granular, and morphological features of the concept pair. The linear transformation parameters for a specific relation type are the weight vector and bias term used for relation prediction. The unnormalized score is the result of the operation between the comprehensive feature vector and the linear transformation parameters, reflecting the original tendency of the concept pair to belong to a certain type of relation. The relation probability is the result of the unnormalized score after softmax normalization, reflecting the likelihood of the concept pair belonging to a certain type of relation. A concept relation edge is a directed edge containing the relation type, start and end concepts, and relation probability, reflecting the semantic association between two concepts. The set of concept relation edges is the set of all valid concept relation edges, reflecting the complete semantic relation network between concepts within the same version. The matching degree is the result of processing the inner product of the concept node vector and the resource fragment vector using a logical function, reflecting the semantic compatibility between the two. Evidence edges are directed edges containing start and end identifiers and matching degrees, reflecting the evidence support relationship between the concept and the resource fragment. The set of evidence edges is the set of all valid evidence edges, reflecting the complete association network between the concept and the resource. A concept node set is a collection containing node information (node vectors, granularity scalars, and multi-scale semantic coverage curvature indices) for all concepts. The set merging operation integrates the concept node set with the three types of edge sets into a complete graph. A knowledge graph is a structured knowledge system containing both node and edge sets, reflecting the complete knowledge relationships within subject-specific educational resources.
[0037] It should be noted that the dimension of the weight vector of the linear transformation parameters for specific relation types is consistent with the dimension of the comprehensive feature vector (e.g., 1538 dimensions), and the dimension of the bias term is 1. During initialization, the weight vector adopts a Xavier normal distribution, and the initial value of the bias term is set to 0.1, which will not be elaborated upon here. The specific types of relations include four categories: prerequisite, containment, association, and causation, totaling four types. This can be expanded to six types according to subject requirements (adding equivalence and supplementation). This classification covers the core semantic relations of educational resource concepts and adapts to the knowledge organization characteristics of different subjects. The outlier criterion for unnormalized scores is a score less than -10 or greater than 10. The processing method is to correct scores less than -10 to -10 and scores greater than 10 to 10; this avoids extreme values causing an imbalance in weight distribution after softmax normalization and ensures the rationality of relation probabilities. The screening threshold for relation probabilities is 0.3. When the probability of a relation type is greater than or equal to 0.3, the relation edge is retained; when it is less than 0.3, it is discarded. The valid range for the matching degree is 0 to 1; the filtering condition is a matching degree greater than or equal to 0.2, and evidence edges that meet this condition are retained. The specific order of the set merging operation is: concept node set, cross-version mapping edge set, concept relationship edge set, and evidence edge set; nodes are integrated first, and then various types of edges are added in sequence to ensure that no association between edges and nodes is missed. In addition, the knowledge graph is stored in the attribute graph format. Node attributes include identifier, node vector, granularity scalar, and multi-scale semantic coverage curvature index, and edge attributes include identifier, weight, and relationship type (or matching degree); the data association rule is that the start and end identifiers of the edges correspond one-to-one with the node identifiers, and it supports association queries by node, edge type, and weight range.
[0038] It should be noted that this invention integrates the semantic, granular, and morphological features of concepts to construct a comprehensive feature vector for accurately predicting concept relationship edges. Simultaneously, it establishes evidence edges between concepts and resources, and then integrates all nodes and the three types of edge sets to generate a complete cross-version subject education knowledge graph, achieving the dual goals of knowledge association and evidence support. The process involves concatenating the comprehensive feature vector to predict relationship edges, calculating the matching degree to establish evidence edges, and merging all sets in a fixed order. The comprehensive feature vector improves the accuracy of concept relationship prediction, the evidence edges provide traceable resource support for concepts, and the cross-version mapping edges connect knowledge from different versions. The resulting knowledge graph contains a complete semantic relationship network and possesses resource evidence support, exhibiting a complete structure and high credibility, thus adapting to the knowledge organization and application needs of subject education resources.
[0039] In one embodiment of the present invention, such as Figure 2 As shown, a knowledge graph construction system for subject-based educational resources includes: The structural positioning information extraction module 201 performs neural sequence annotation on multi-source subject education resources, locates the text span of concept mentions, and simultaneously extracts the structural positioning information of each text span in the resource chapter level; The resource fragment vector generation module 202 extracts a multi-scale context view based on structural positioning information, and encodes the multi-scale context view into a multi-scale distributed vector set by combining the scale embedding vector. At the same time, it encodes the resource fragments to generate resource fragment vectors. The semantic coverage curvature index calculation module 203 calculates the average intrinsic dimension that varies with the neighborhood scale based on the nearest neighbor distance sequence of the vector set, and performs second-order difference summation on the average intrinsic dimension change curve to obtain the multi-scale semantic coverage curvature index. The concept node vector generation module 204 maps the multi-scale semantic coverage curvature index to a query vector and inputs it into the curvature modulation aggregation network. It calculates the attention weight of each point in the vector set and performs weighted aggregation on the vector set to generate concept node vectors, and outputs granular scalars simultaneously. The cross-version mapping edge construction module 205 calculates the cosine similarity between concept node vectors of different versions and subtracts the absolute value of the difference between the corresponding multi-scale semantic coverage curvature exponents as a curvature penalty term, and establishes cross-version mapping edges based on the calculation results. The knowledge graph generation module 206 concatenates concept node vectors, granular scalars, and multi-scale semantic coverage curvature indices into a comprehensive feature vector to predict and establish concept relationship edges. At the same time, it establishes evidence edges based on the matching degree between concept node vectors and resource fragment vectors, and merges them with cross-version mapping edges to generate a knowledge graph.
[0040] It should be noted that the multi-disciplinary educational resources collected during deployment include digitized texts of printed textbooks, electronic teaching materials, online question banks and solutions, experimental operation instructions, course handouts, etc., covering the core knowledge content of all subjects from elementary to high school. Specifically, compliant electronic resources can be obtained in batches through the official open interface of the educational resource platform. Printed textbooks and teaching materials are converted into text format using optical character recognition technology. User-uploaded compliant educational resource texts can also be received. Version information of the resources (such as textbook publication batch and teaching material revision number) is recorded simultaneously during the collection process. After collection, the format is first standardized, unifying text encoding and paragraph separators; then, deduplication is performed, deleting completely duplicate resource fragments; finally, the text is parsed according to the chapter-level structure, marking the boundaries of chapters, sections, subsections, paragraphs, and sentences, laying the foundation for subsequent structural location information extraction. The preprocessed resource text is split into fragments (each fragment not exceeding 500 words), input into the deployed neural network model (including encoding functions, curvature calculation modules, aggregation networks, etc.); then, all node and edge data are integrated according to the above modules to form a queryable and updatable structured knowledge graph.
[0041] It should be noted that the final output is a complete cross-version subject-specific educational knowledge graph, comprising four core components: a set of concept nodes, a set of cross-version mapping edges, a set of concept relationship edges, and a set of evidence edges. All data is stored in attribute graph format, supporting relational queries and dynamic updates. Specifically, taking a single node as an example: the node identifier is V2-0156, the corresponding concept is a quadratic equation, and the node attributes include a 768-dimensional node vector, a granularity scalar of 0.42, and a multi-scale semantic coverage curvature index of 0.68, fully carrying the semantic, granular, and morphological features of the concept. Specifically, taking a single mapping edge as an example: the starting point is identified as V1-0120 (old version concept quadratic equation), and the ending point is identified as V2-0156 (new version concept quadratic equation), with an edge weight of 0.87, reflecting a strong association between the same concept in the two versions; another mapping edge starts at V1-0121 (old version concept factorization method), and ends at V2-0157 (new version concept quadratic equation factorization solution method), with an edge weight of 0.32, reflecting a weak association after concept splitting in the version update. Specifically, taking a single relation edge as an example: the starting point is V2-0156 (quadratic equation), and the ending point is V2-0168 (discriminant of roots), with a relation type of preconditioning and a relation probability of 0.73, clearly indicating the semantic dependency between the two concepts; another relation edge starts at V2-0156 (quadratic equation), and ends at V2-0172 (quadratic function graph), with a relation type of association and a relation probability of 0.61, reflecting the correlation between the concepts. Specifically, taking a single evidence edge as an example: starting point V2-0156 (a quadratic equation in one variable), ending point U2-0342 (the second paragraph of the third section of the second chapter of the ninth-grade mathematics textbook), the matching degree is 0.79, providing clear resource-supporting evidence for the concept.
[0042] It should be noted that the interval and threshold sizes are set for ease of comparison. The size of the threshold depends on the amount of sample data and the base number set by those skilled in the art for each set of sample data, as long as it does not affect the proportional relationship between the parameter and the quantized value. Furthermore, the above formulas are all dimensionless calculations, and the formulas are derived from software simulations using a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0043] The embodiments of this example have been described above. However, this example is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of this example, and all of them are within the protection scope of this example.
Claims
1. A method for constructing knowledge graphs for subject-based educational resources, characterized in that, Includes the following steps: Step S101: Perform neural sequence annotation on multi-source subject education resources, locate the text span of concept mentions, and simultaneously extract the structural positioning information of each text span in the resource chapter level; Step S102: Extract a multi-scale context view based on the structural positioning information, and encode the multi-scale context view into a multi-scale distributed vector set by combining the scale embedding vector. At the same time, encode the resource fragments to generate resource fragment vectors. Step S103: Calculate the average intrinsic dimension that varies with the neighborhood scale based on the nearest neighbor distance sequence of the vector set, and perform second-order difference summation on the average intrinsic dimension variation curve to obtain the multi-scale semantic coverage curvature index. Step S104: Map the multi-scale semantic coverage curvature index to a query vector and input it into the curvature modulation aggregation network. Calculate the attention weights of each point in the vector set. Perform weighted aggregation on the vector set accordingly to generate concept node vectors and output granular scalars simultaneously. Step S105: Calculate the cosine similarity between concept node vectors of different versions, and subtract the absolute value of the difference between the corresponding multi-scale semantic coverage curvature exponents as curvature penalty terms. Establish cross-version mapping edges based on the calculation results. Step S106: The concept node vector, granularity scalar and multi-scale semantic coverage curvature index are concatenated into a comprehensive feature vector to predict and establish concept relationship edges. At the same time, evidence edges are established according to the matching degree between the concept node vector and the resource fragment vector, and merged with cross-version mapping edges to generate a knowledge graph.
2. The method for constructing a knowledge graph for subject-based educational resources according to claim 1, characterized in that, Neural sequence annotation is performed on multi-disciplinary educational resources to locate the text span of concept references and simultaneously extract the structural location information of each text span at the resource chapter level, including: A single resource fragment text is divided into a word sequence; a neural network encoding function is used to calculate the single resource fragment text to generate a vector sequence corresponding to the word sequence, and the vector sequence is composed of the vector representation of the word; For any text span in a word sequence where the starting position is less than the ending position, obtain the word vector representations at the starting and ending positions of the text span, and perform average pooling on all word vector representations covered by the text span to obtain the pooling result. The word vector representations at the starting and ending positions are concatenated with the pooling result to generate a vector representation of the text span. The vector representation of text span is processed with linear scoring parameters, and the result is mapped using a logical function to obtain the concept mention probability of text span. Based on the position of the text span in the resource chapter level, extract the chapter level markers, section level markers, subsection level markers, paragraph level markers, and sentence level markers covering the text span, and combine the chapter level markers, section level markers, subsection level markers, paragraph level markers, and sentence level markers into structural positioning information.
3. The method for constructing a knowledge graph for subject-based educational resources according to claim 2, characterized in that, Based on structural localization information, a multi-scale context view is extracted. Combined with scale embedding vectors, the multi-scale context view is encoded into a multi-scale distributed vector set. Simultaneously, resource fragments are encoded to generate resource fragment vectors, including: Combine text span, concept mention probability, and structural location information and store them in a concept mention text span set; For any element in the concept mention text span set, a multi-scale context view is determined based on structural positioning information. The multi-scale context view consists of sentences containing concept mention text spans, paragraphs containing sentences, sections containing paragraphs, and chapters containing sections. The neural network encoding function is used to calculate the text content in the multi-scale context view. The calculation result is then added to the scale embedding vector of the corresponding level, and the result is normalized by the vector L2 norm to generate a multi-scale vector. Iterate through the entire set of concept reference text spans across all resource versions, aggregate all generated multi-scale vectors, and construct a multi-scale distributed vector set; The neural network encoding function is used to calculate the individual resource fragment text, generate resource fragment vectors, and store the individual resource fragment text and resource fragment vectors as pairs.
4. The method for constructing a knowledge graph for subject-based educational resources according to claim 1, characterized in that, The average intrinsic dimension, varying with neighborhood scale, is calculated based on the nearest neighbor distance sequence of the vector set. A second-order difference summation is then performed on the average intrinsic dimension variation curve to obtain the multi-scale semantic coverage curvature index, including: For any vector in a multi-scale distributed vector set, calculate the Euclidean distance between the arbitrary vector and the other vectors in the set, and sort the Euclidean distances by numerical value to determine the distance of the k-th nearest neighbor of the arbitrary vector. Define a neighborhood scale set, which consists of values that increase in powers of two, and the maximum value is limited by the square root of the number of vectors in the multi-scale distributed vector set. For any vector and any scale value in the neighborhood scale set, calculate the natural logarithm of the ratio of the distance of the k-th nearest neighbor to the distance of the j-th nearest neighbor, sum the natural logarithms of all ratios and take the reciprocal to obtain the local intrinsic dimension estimate. The arithmetic mean of the local intrinsic dimension estimates of all vectors in a multi-scale distributed vector set is used to obtain the average intrinsic dimension of the corresponding scale value. Select three adjacent scale values from the neighborhood scale set, calculate the second difference of the average intrinsic dimension, and perform an accumulation operation on the absolute values of all second differences to obtain the multi-scale semantic coverage curvature index.
5. The method for constructing a knowledge graph for subject-based educational resources according to claim 1, characterized in that, The multi-scale semantic coverage curvature index is mapped to a query vector and input into a curvature modulation aggregation network. Attention weights are calculated for each point in the vector set, and weighted aggregation is performed on the vector set to generate concept node vectors. Simultaneously, granular scalars are output, including: Obtain the natural logarithm of the number of elements in a multi-scale distributed vector set, and perform vector concatenation operation on the multi-scale semantic coverage curvature index and the natural logarithm of the number of elements; The vector concatenation operation result is combined with the query-generated learning parameters to perform linear calculations, and the hyperbolic tangent function is used to process the linear calculation result to generate the query vector; Calculate the inner product of the query vector and any vector in the multi-scale distributed vector set, scale the inner product using the square root of the vector dimension, and calculate the exponential function value of the scaling result. The attention weight of any vector is obtained by dividing the value of the exponential function by the sum of the exponential function values of all vectors in the multi-scale distributed vector set. The attention weights are used to perform a weighted summation operation on a multi-scale distributed vector set to generate concept node vectors; The concept node vector and the multi-scale semantic coverage curvature exponent are concatenated. The result of the concatenation is then used to perform linear calculations with the granularity-generated learning parameters. The results of the linear calculations are processed using logical functions to output the granularity scalar.
6. The method for constructing a knowledge graph for subject-based educational resources according to claim 1, characterized in that, Calculate the cosine similarity between concept node vectors of different versions, and subtract the absolute value of the difference between the corresponding multi-scale semantic coverage curvature indices as a curvature penalty term. Based on the calculation results, establish cross-version mapping edges, including: Obtain the first version of the concept node vector set and the multi-scale semantic coverage curvature index set, as well as the second version of the concept node vector set and the multi-scale semantic coverage curvature index set; For the concept node vectors of the first version and the concept node vectors of the second version, calculate the inner product of the concept node vectors of the first version and the concept node vectors of the second version, and divide the inner product by the product of the L2 norm of the concept node vectors of the first version and the L2 norm of the concept node vectors of the second version to obtain the cosine similarity. For the first version of the multi-scale semantic coverage curvature index and the second version of the multi-scale semantic coverage curvature index, calculate the difference between the first version of the multi-scale semantic coverage curvature index and the second version of the multi-scale semantic coverage curvature index, and take the absolute value of the difference to obtain the curvature penalty term; Subtract the curvature penalty term from the cosine similarity to obtain the mapping score; Calculate the exponential function value of the mapping score, and divide the exponential function value of the mapping score by the sum of the exponential function values of the mapping scores of all concepts in the second version to obtain the edge weight; Combine the edge weights with the start and end markers to create cross-version mapping edges and include them in the cross-version mapping edge set.
7. The method for constructing a knowledge graph for subject-based educational resources according to claim 1, characterized in that, Concept node vectors, granularity scalars, and multi-scale semantic coverage curvature exponents are concatenated into a comprehensive feature vector to predict and establish concept relationship edges. Simultaneously, evidence edges are established based on the matching degree between concept node vectors and resource fragment vectors, and these are merged with cross-version mapping edges to generate a knowledge graph, including: For any ordered pair of concepts in the same version, obtain the concept node vector of the starting concept, the granularity scalar of the starting concept, the multi-scale semantic coverage curvature index of the starting concept, and the concept node vector of the terminating concept, the granularity scalar of the terminating concept, and the multi-scale semantic coverage curvature index of the terminating concept. Perform vector concatenation operations on the concept node vector of the starting concept, the concept node vector of the ending concept, the absolute value of the difference between the granularity scalar of the starting concept and the ending concept, and the absolute value of the difference between the multi-scale semantic coverage curvature exponent of the starting concept and the ending concept to generate a comprehensive feature vector. A linear calculation is performed on the integrated feature vector and the linear transformation parameters for a specific relation type to generate an unnormalized score; Calculate the exponential function value of the unnormalized score, divide the exponential function value of the unnormalized score by the sum of the exponential function values of the unnormalized scores of all relation types, and obtain the relation probability. Establish concept relationship edges based on relationship probabilities, and construct a set of concept relationship edges; Calculate the inner product of the concept node vector and the resource fragment vector, and process the inner product using a logical function to obtain the degree of matching between the concept node vector and the resource fragment vector; Based on the degree of matching, establish evidence edges and construct a set of evidence edges; Perform a set merging operation on the concept node set, cross-version mapping edge set, concept relationship edge set, and evidence edge set to output a knowledge graph.
8. A knowledge graph construction system for subject-based educational resources, characterized in that, The method for constructing a knowledge graph for subject-based educational resources as described in any one of claims 1 to 7 includes: The structural positioning information extraction module performs neural sequence annotation on multi-source subject education resources, locates the text span of concept mentions, and simultaneously extracts the structural positioning information of each text span in the resource chapter level; The resource fragment vector generation module extracts a multi-scale context view based on structural positioning information, and encodes the multi-scale context view into a multi-scale distributed vector set by combining the scale embedding vector. At the same time, it encodes the resource fragments to generate resource fragment vectors. The semantic coverage curvature index calculation module calculates the average intrinsic dimension that varies with the neighborhood scale based on the nearest neighbor distance sequence of the vector set, and performs second-order difference summation on the average intrinsic dimension change curve to obtain the multi-scale semantic coverage curvature index. The concept node vector generation module maps the multi-scale semantic coverage curvature index to query vectors and inputs them into the curvature modulation aggregation network. It calculates the attention weights of each point in the vector set and performs weighted aggregation on the vector set to generate concept node vectors, and outputs granular scalars simultaneously. The cross-version mapping edge construction module calculates the cosine similarity between concept node vectors of different versions and subtracts the absolute value of the difference between the corresponding multi-scale semantic coverage curvature exponents as a curvature penalty term, and establishes cross-version mapping edges based on the calculation results. The knowledge graph generation module concatenates concept node vectors, granular scalars, and multi-scale semantic coverage curvature indices into a comprehensive feature vector to predict and establish concept relationship edges. At the same time, it establishes evidence edges based on the matching degree between concept node vectors and resource fragment vectors, and merges them with cross-version mapping edges to generate a knowledge graph.