A four-dimensional index-based automatic sci-tech literature knowledge labeling method and system
By constructing a four-dimensional index architecture and a cross-level semantic association model, the problems of inaccurate knowledge recognition and insufficient hierarchical annotation in scientific and technological literature have been solved, achieving efficient and accurate literature processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-04-07
AI Technical Summary
The inaccurate identification of knowledge in existing scientific and technological literature and the lack of hierarchical annotation lead to low efficiency and accuracy in literature processing.
A four-dimensional index architecture is constructed, which includes a data input layer, a four-dimensional index construction layer, a knowledge enrichment area identification layer, a boundary detection and annotation layer, and a result output layer. Index parsing is performed through document-level indexing, paragraph-level indexing, sentence-level indexing, and entity-level indexing. Multi-level annotation is performed by combining semantic density calculation and cross-level semantic association model.
It enables accurate identification and multi-level annotation of scientific and technological literature knowledge, improving the efficiency and accuracy of literature processing.
Smart Images

Figure CN121303304B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing technology, specifically to a method and system for automatic annotation of scientific and technological literature knowledge based on four-dimensional indexing. Background Technology
[0002] In the field of scientific and technological literature processing, traditional knowledge annotation methods mostly rely on manual or single-level automated processing, which has obvious limitations: manual annotation is time-consuming and labor-intensive, and difficult to handle massive amounts of literature; automated methods often focus on single entity or sentence levels, lacking in-depth mining of cross-level semantic relationships between documents, paragraphs, sentences, and entities, resulting in insufficient accuracy in knowledge recognition and chaotic hierarchical annotation. At the same time, existing technologies have low efficiency in identifying knowledge-rich areas, easily missing key information, and the annotation results lack structured presentation, making it difficult to meet users' needs for quickly acquiring and utilizing the core knowledge of scientific and technological literature, seriously restricting the efficient dissemination and transformation of scientific and technological knowledge.
[0003] Existing technologies suffer from inaccurate knowledge recognition in scientific and technological literature and insufficient hierarchical labeling, leading to low efficiency and accuracy in literature processing. Summary of the Invention
[0004] This application provides a method and system for automatic annotation of scientific and technological literature knowledge based on four-dimensional indexing, which is used to address the technical problems of inaccurate identification of scientific and technological literature knowledge and insufficient hierarchical annotation in the prior art, resulting in low efficiency and accuracy of literature processing.
[0005] In view of the above problems, this application provides a method and system for automatic annotation of scientific and technological literature knowledge based on four-dimensional index.
[0006] The first aspect of this application provides a method for automatic annotation of scientific and technological literature knowledge based on four-dimensional indexing, the method comprising:
[0007] A four-dimensional index architecture is constructed, comprising a data input layer, a four-dimensional index construction layer, a knowledge enrichment region identification layer, a boundary detection and annotation layer, and a result output layer. The specific structure of the four-dimensional index construction layer includes a document-level index, a paragraph-level index, a sentence-level index, and an entity-level index. The data input layer receives multi-format scientific and technological documents, and based on the document-level index, paragraph-level index, sentence-level index, and entity-level index, the documents are sequentially indexed and parsed to obtain the document-level semantic structure, paragraph-level semantic relationship graph, sentence-level semantic relationship network, and entity-level semantic relationships. The knowledge-rich region identification layer performs semantic density calculation and knowledge-rich region identification on the document layer semantic structure, paragraph layer semantic relationship graph, sentence semantic relationship network and entity semantic relationship to obtain knowledge-rich region identification results; a cross-level semantic association model is established, and a recursive annotation architecture is designed based on the boundary detection annotation layer. The recursive annotation architecture and the cross-level semantic association model are used to predict entity boundaries and perform multi-level annotation on the knowledge-rich region identification results to obtain structured semantic annotation results. The structured semantic annotation results are then visualized and displayed through the result output layer.
[0008] A second aspect of this application provides an automatic knowledge annotation system for scientific and technological literature based on a four-dimensional index, the system comprising:
[0009] A four-dimensional index architecture construction module is used to construct a four-dimensional index architecture, which includes a data input layer, a four-dimensional index construction layer, a knowledge enrichment region identification layer, a boundary detection and annotation layer, and a result output layer. The specific structure of the four-dimensional index construction layer includes a document-level index, a paragraph-level index, a sentence-level index, and an entity-level index. An index parsing module is used to receive multi-format scientific and technological documents through the data input layer and sequentially parse the documents based on the document-level index, paragraph-level index, sentence-level index, and entity-level index to obtain the document-level semantic structure, paragraph-level semantic relationship graph, sentence-level semantic relationship network, and entity-level semantic relationship. The knowledge-rich region identification module is used to perform semantic density calculation and knowledge-rich region identification on the document-level semantic structure, paragraph-level semantic relationship graph, sentence semantic relationship network, and entity semantic relationship based on the knowledge-rich region identification layer, and obtain the knowledge-rich region identification result; the annotation result acquisition module is used to establish a cross-level semantic association model, design a recursive annotation architecture based on the boundary detection annotation layer, use the recursive annotation architecture and the cross-level semantic association model to perform entity boundary prediction and multi-level annotation on the knowledge-rich region identification result, obtain structured semantic annotation result, and visualize the structured semantic annotation result through the result output layer.
[0010] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0011] A four-dimensional indexing architecture is constructed, comprising a data input layer, a four-dimensional index construction layer, a knowledge-rich region identification layer, a boundary detection and annotation layer, and a result output layer. The data input layer receives multi-format scientific and technological documents, and sequentially parses these documents based on the document-level index, paragraph-level index, sentence-level index, and entity-level index to obtain the document-level semantic structure, paragraph-level semantic relationship graph, sentence-level semantic relationship network, and entity semantic relationships. Semantic density calculation and knowledge-rich region identification are performed to obtain the knowledge-rich region identification results. A cross-level semantic association model is established, and a recursive annotation architecture and the cross-level semantic association model are used to predict entity boundaries and perform multi-level annotation on the knowledge-rich region identification results, obtaining structured semantic annotation results. The result output layer visualizes and displays the structured semantic annotation results. This achieves the technical effect of accurate identification and multi-level annotation of scientific and technological document knowledge, improving the efficiency and accuracy of document processing. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A schematic flowchart of a method for automatically annotating scientific and technological literature knowledge based on four-dimensional indexing, provided for an embodiment of this application;
[0014] Figure 2 This is a schematic diagram of the structure of an automatic knowledge annotation system for scientific and technological literature based on a four-dimensional index, provided in an embodiment of this application.
[0015] Figure labeling: 10 for four-dimensional index architecture construction module, 20 for index parsing module, 30 for rich region identification module, and 40 for annotation result acquisition module. Detailed Implementation
[0016] This application provides a method and system for automatic annotation of scientific and technological literature knowledge based on four-dimensional indexing, which addresses the technical problems of inaccurate identification of scientific and technological literature knowledge and insufficient hierarchical annotation in the prior art, resulting in low efficiency and accuracy of literature processing.
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0018] Example 1, as Figure 1 As shown, this application provides a method for automatic annotation of scientific and technological literature knowledge based on four-dimensional indexing, the method comprising:
[0019] Step S100: Construct a four-dimensional index architecture, which includes a data input layer, a four-dimensional index construction layer, a knowledge enrichment region identification layer, a boundary detection and annotation layer, and a result output layer. The specific structure of the four-dimensional index construction layer includes a document layer index, a paragraph layer index, a sentence layer index, and an entity layer index.
[0020] Specifically, a four-dimensional index architecture is constructed, comprising a data input layer, a four-dimensional index construction layer, a knowledge enrichment region identification layer, a boundary detection and annotation layer, and a result output layer. The four-dimensional index construction layer is further divided into document-level indexing, paragraph-level indexing, sentence-level indexing, and entity-level indexing. The data input layer receives and parses scientific documents in multiple formats such as PDF, HTML, and XML. The four-dimensional index construction layer processes information at different granularities through four levels of indexing: the document-level index focuses on the overall document structure and thematic features; the paragraph-level index focuses on paragraph segmentation and inter-paragraph relationships; the sentence-level index emphasizes sentence semantic encoding and inter-sentence associations; and the entity-level index targets the identification and relationship extraction of multiple entity types. The knowledge enrichment region identification layer is used for semantic density calculation and important region identification. The boundary detection and annotation layer undertakes entity boundary prediction and multi-level annotation tasks. The result output layer is responsible for the visualization of structured annotation results. These layers work together to form a complete processing system from document input to annotation output.
[0021] Step S200: Receive multi-format scientific and technological documents through the data input layer, and sequentially perform index parsing on the multi-format scientific and technological documents based on the document layer index, paragraph layer index, sentence layer index, and entity layer index to obtain the document layer semantic structure, paragraph layer semantic relationship graph, sentence semantic relationship network, and entity semantic relationship.
[0022] Specifically, the system receives scientific and technological documents in multiple formats, including PDF, HTML, and XML, through the data input layer. Then, based on the four-dimensional index construction layer, it sequentially performs index parsing of the document layer index, paragraph layer index, sentence layer index, and entity layer index. Specifically, the document layer index performs document structure parsing (identifying titles, abstracts, chapters, etc.) and global feature extraction (such as subject area and research topic) to obtain the document layer semantic structure. The paragraph layer index performs intelligent paragraph segmentation (combining semantic coherence and physical segmentation) and inter-paragraph relationship modeling (constructing a graph structure including relationships such as sequence, citation, and topic inheritance) to build a paragraph layer semantic relationship graph. The sentence layer index performs sentence semantic encoding (e.g., using the SciBERT model to generate semantic vectors) and semantic relationship analysis (identifying causal, comparative, and other inter-sentence relationships) to generate a sentence semantic relationship network. Finally, the entity layer index performs multi-type entity recognition (including scientific concepts, research methods, experimental materials, etc.) and entity relationship extraction (mining synonymous, hierarchical, and other relationships between entities) to obtain entity semantic relationships.
[0023] Step S300: Based on the knowledge enrichment region identification layer, perform semantic density calculation and knowledge enrichment region identification on the document layer semantic structure, paragraph layer semantic relationship graph, sentence semantic relationship network and entity semantic relationship to obtain the knowledge enrichment region identification result.
[0024] Specifically, based on the knowledge-rich region identification layer, multi-dimensional semantic features are first extracted from the document-level semantic structure, paragraph-level semantic relationship graph, sentence semantic relationship network, and entity semantic relationship to obtain lexical-level feature sets (such as terminology density and concept complexity), syntactic-level feature sets (such as syntactic complexity and grammatical patterns), and semantic-level feature sets (such as concept relevance and knowledge level depth). Next, a density calculation model is designed, incorporating a multi-feature fusion model (using a multi-layer perceptron structure) and an attention weighting mechanism. Simultaneously, a multi-scale sliding window strategy covering word-level, sentence-level, and paragraph-level windows is designed. Then, this multi-scale sliding window strategy is applied, based on the density calculation model... The semantic density of the above semantic structure and relationships is calculated and Gaussian smoothed to obtain the semantic window density distribution curve. Then, based on the curve, the threshold of the knowledge-rich region is determined as T = μ + α × σ (where μ is the average semantic density of the document, σ is the standard deviation of the semantic density, and α is an adaptive parameter of 1.5-2.0), and a multi-level threshold strategy is set. Then, by calculating the density gradient and smoothing the gradient of the semantic window density distribution curve, boundary candidate points are detected and verified. The multi-level threshold strategy is used to segment and label the rich region of the curve to obtain the basic hierarchical knowledge-rich region. Finally, local search optimization and multi-scale boundary fusion are performed to determine the knowledge-rich region recognition result.
[0025] Step S400: Establish a cross-level semantic association model, design a recursive annotation architecture based on the boundary detection annotation layer, use the recursive annotation architecture and the cross-level semantic association model to perform entity boundary prediction and multi-level annotation on the knowledge-rich area identification results to obtain structured semantic annotation results, and visualize the structured semantic annotation results through the result output layer.
[0026] Specifically, a cross-level semantic association model is established. This model obtains hierarchical association modeling objectives at the document, paragraph, sentence, and entity levels, and performs hierarchical semantic representation on multi-format scientific documents to obtain semantic vectors for each level. This model is then constructed through cross-level association learning, global consistency constraints, and dynamic association weight learning. Subsequently, based on the boundary detection annotation layer, a recursive annotation architecture is designed, including document-level coarse-grained annotation, paragraph-level medium-grained annotation, sentence-level fine-grained annotation, and entity-level precise annotation. This recursive annotation architecture is then used to process the knowledge-rich region identification results, obtaining topic recognition through document-level annotation. The key area localization and global entity type prediction results are obtained. Paragraph-level annotation yields the results of topic refinement, entity density analysis, and pre-establishment of cross-paragraph relationships. Sentence-level annotation enables semantic role analysis, syntactic-guided entity boundary prediction, and context-aware entity classification. Entity-level annotation completes accurate boundary detection, nested entity processing, and standardized entity linking, thus obtaining annotation results at each level. After determining the initial semantic annotation results based on these results, global consistency verification and recursive feedback optimization are performed in conjunction with the cross-level semantic association model to obtain structured semantic annotation results. Finally, the results are visualized and displayed through the result output layer.
[0027] In one possible implementation, step S200 further includes:
[0028] Step S210: Based on the document layer index, perform document structure parsing and global feature extraction on the multi-format scientific and technological documents to obtain the document layer semantic structure.
[0029] Step S220: Based on the paragraph-level index, perform paragraph segmentation and inter-paragraph relationship modeling on the multi-format scientific and technological documents, and construct a paragraph-level semantic relationship graph.
[0030] Step S230: Use the sentence-level index to perform sentence semantic encoding and semantic relation analysis on the multi-format scientific and technological documents to generate a sentence semantic relation network.
[0031] Step S240: Based on the entity layer index, perform multi-type entity recognition and entity relation extraction on the multi-format scientific and technological documents to obtain entity semantic relations.
[0032] Specifically, based on document-level indexing, multi-format scientific and technological literature is processed. First, document structure parsing is performed to identify the standard academic literature structure, such as title, abstract, introduction, methods, results, discussion, and conclusion, clarifying the hierarchical relationship and content boundaries of each part. At the same time, global document features are extracted to obtain meta-information, including authors, journals, and publication time, as well as features reflecting the overall attributes of the literature, such as subject area, research topic, citation pattern, and keyword distribution. Then, a 768-dimensional document-level semantic vector is generated through models such as BERT. Finally, this information is integrated into a document-level semantic structure that includes metadata, structure, content, and semantic vector.
[0033] This paper processes multi-format scientific and technological documents based on paragraph-level indexing. First, intelligent paragraph segmentation is performed, which not only considers the physical segmentation of the documents, but also the semantic coherence, so that each paragraph maintains the relative semantic integrity. Then, the relationship model between paragraphs is performed to analyze the sequential relationship, citation relationship, topic inheritance relationship and other relationships between paragraphs. These complex relationships are learned through graph neural network, and each paragraph is represented as a structure that includes the document ID, the position information in the document, the paragraph topic vector, the list of entities in the paragraph and the semantic coherence score. Finally, a paragraph-level semantic relationship graph that can reflect the semantic relationship between paragraphs is constructed.
[0034] This paper employs sentence-level indexing to process multi-format scientific and technological documents. First, it uses a pre-trained model such as SciBERT to perform deep semantic encoding on each sentence, generating a 768-dimensional sentence semantic vector. Simultaneously, it parses the sentence's syntactic dependency tree, identifies entity mentions within the sentence, and calculates the sentence importance score. Subsequently, it performs semantic relationship analysis to identify causal, comparative, and complementary relationships between sentences. It uses a graph attention network (GAT) to learn these complex inter-sentence dependencies, representing each sentence as a structure containing its paragraph ID, sentence semantic vector, syntactic dependency tree, entity mentions, and importance score. Finally, it generates a sentence semantic relationship network that reflects the semantic connections between sentences.
[0035] This paper processes multi-format scientific and technological documents based on entity layer indexing. First, it performs multi-type entity recognition to identify entity types such as scientific concepts, research methods, experimental materials, numerical data, institutions, and citation information in the documents. Each entity is represented as a structure that includes entity text content, entity type label, four-dimensional location coordinates (document, paragraph, sentence, and entity position in the sentence), contextual semantic vector, and recognition confidence score. Then, entity relationship extraction is performed, which not only mines the relationships between entities within sentences, but also extracts long-distance entity relationships across sentences and paragraphs, and identifies synonyms, hierarchical relationships, and other associations between entities. Finally, it obtains entity semantic relationships that can reflect various types of entities and their interrelationships.
[0036] In one possible implementation, step S300 further includes:
[0037] Step S310: Based on the knowledge enrichment region identification layer, perform multi-dimensional semantic feature extraction on the document layer semantic structure, paragraph layer semantic relationship graph, sentence semantic relationship network and entity semantic relationship to obtain lexical feature set, syntactic feature set and semantic feature set.
[0038] Step S320: Design a density calculation model, which includes a multi-feature fusion model and an attention weight mechanism, wherein the multi-feature fusion model adopts a multilayer perceptron structure.
[0039] Step S330: Design a multi-scale sliding window strategy, which includes word-level window, sentence-level window and paragraph-level window.
[0040] Step S340: Using the multi-scale sliding window strategy, based on the density calculation model, perform semantic density calculation and knowledge enrichment region identification on the document layer semantic structure, paragraph layer semantic relationship graph, sentence semantic relationship network and entity semantic relationship to obtain the knowledge enrichment region identification result.
[0041] Specifically, based on the knowledge-rich region identification layer, multi-dimensional semantic features are extracted from the document-level semantic structure, paragraph-level semantic relationship graph, sentence semantic relationship network, and entity semantic relationship. Specifically, the lexical-level feature set is obtained by calculating the proportion of technical terms in a unit of text (technical term density), assessing concept complexity based on semantic depth and abstraction, and analyzing the novelty and importance of words (novelty index); the syntactic-level feature set is obtained by calculating syntactic complexity based on the depth and branching degree of the syntactic tree, identifying specific grammatical patterns in scientific writing (grammatical pattern analysis), and statistically analyzing the relationship density between modifiers and modified words (modification relationship density); and the semantic-level feature set is obtained by calculating the semantic association strength between concepts in the text (concept association degree), assessing the abstraction level and professional depth of knowledge (knowledge level depth), and calculating information entropy based on semantic distribution (information entropy calculation), ultimately yielding the corresponding three-level feature sets.
[0042] The designed density calculation model includes a multi-feature fusion model and an attention weighting mechanism. The multi-feature fusion model adopts a multilayer perceptron (MLP) structure, which concatenates the lexical feature set (32-dimensional vector), the syntactic feature set (24-dimensional vector), and the semantic feature set (64-dimensional vector) to form a comprehensive feature vector with dimensions of 32+24+64=120. This vector is then input into the MLP for nonlinear transformation and feature fusion, outputting a basic score of semantic density. The attention weighting mechanism calculates the product of the weight matrix and the feature vector and processes it through the softmax function to obtain the dynamic attention weight of each feature. This weight is then weighted and fused with the corresponding feature vector. Finally, combined with the output of the MLP, it achieves adaptive adjustment of the importance of different features, improving the accuracy of density calculation.
[0043] The designed multi-scale sliding window strategy includes word-level windows, sentence-level windows, and paragraph-level windows. The word-level window is set to a size of 5-10 words with a step size of 1 word, used for fine-grained lexical feature capture of the text. The sentence-level window is set to a size of 3-5 sentences with a step size of 1 sentence, used for medium-grained analysis of semantic connections and syntactic features between sentences. The paragraph-level window is set to a size of 2-3 paragraphs with a step size of 1 paragraph, used for coarse-grained grasp of thematic coherence and knowledge distribution between paragraphs. By combining and applying windows of different scales, comprehensive coverage and accurate extraction of semantic information from different levels, from micro to macro, of scientific and technological literature can be achieved.
[0044] A multi-scale sliding window strategy at the word, sentence, and paragraph levels is employed. Based on a multi-feature fusion model incorporating a multi-layer perceptron structure and a density calculation model with an attention weighting mechanism, the semantic structure of the document layer, the semantic relationship graph of the paragraph layer, the semantic relationship network of sentences, and the semantic relationship of entities are processed. The text is traversed by sliding windows at each scale. The density calculation model fuses lexical, syntactic, and semantic features and assigns dynamic weights. The semantic density of each window is calculated and Gaussian smoothed to obtain the semantic window density distribution curve. Based on this curve, the threshold for the knowledge-rich region is determined as T = μ + α × σ (μ is the average semantic density of the document, σ is the standard deviation of the semantic density, and α ranges from 1.5 to 2.0), and a multi-level threshold strategy is set. Density gradient calculation and gradient smoothing are performed on the curve. Boundary candidate points are detected and verified. The multi-level threshold strategy is used to segment and label the rich regions to obtain the basic hierarchical knowledge-rich regions. After local search optimization and multi-scale boundary fusion, the final knowledge-rich region identification result is obtained.
[0045] In one possible implementation, step S340 further includes:
[0046] Step S341: Using the multi-scale sliding window strategy, based on the density calculation model, perform semantic density calculation and density Gaussian smoothing on the document layer semantic structure, paragraph layer semantic relationship graph, sentence semantic relationship network and entity semantic relationship to obtain the semantic window density distribution curve.
[0047] Step S342: Based on the semantic window density distribution curve, determine the knowledge enrichment area threshold T = μ + α × σ, where μ is the document average semantic density, σ is the semantic density standard deviation, and α is an adaptive parameter, typically ranging from 1.5 to 2.0.
[0048] Step S343: Based on the knowledge enrichment area threshold, set a multi-level threshold strategy.
[0049] Step S344: Use the multi-level threshold strategy to identify the knowledge-rich region boundary and perform adaptive boundary optimization on the semantic window density distribution curve to determine the knowledge-rich region identification result.
[0050] Specifically, a multi-scale sliding window strategy is adopted, consisting of word-level windows (5-10 words in size, step size 1 word), sentence-level windows (3-5 sentences in size, step size 1 sentence), and paragraph-level windows (2-3 paragraphs in size, step size 1 paragraph). Based on a density calculation model that includes a multi-feature fusion model (using a multi-layer perceptron structure) and an attention weight mechanism, the semantic structure of the document layer, the semantic relationship graph of the paragraph layer, the semantic relationship network of sentences, and the semantic relationship of entities are processed. The semantic density is calculated by sliding each window, and then Gaussian smoothing is used to process the noise of the density curve (i.e., smoothed-density[i] = Σ(j = -k to k)gaussian-weight[j] × original-density[i+j]), finally obtaining the semantic window density distribution curve.
[0051] Based on the semantic window density distribution curve obtained through a multi-scale sliding window strategy and density calculation model, the average semantic density of the document as reflected by the curve is first calculated, i.e., the document average semantic density μ. Then, the dispersion of each semantic density value in the curve relative to the average value is calculated to obtain the semantic density standard deviation σ. Subsequently, the knowledge-rich area threshold T is determined. Its calculation formula is T equal to the document average semantic density μ plus the product of the adaptive parameter α and the semantic density standard deviation σ, where α is a parameter dynamically adjusted according to the subject domain characteristics, knowledge complexity and text structure of scientific and technological literature. It usually takes a value range of 1.5-2.0. This threshold can effectively distinguish between areas with high semantic density and ordinary areas in the document, providing a quantitative standard for the accurate identification of knowledge-rich areas in the future.
[0052] Based on the knowledge-rich region threshold, and combined with the overall semantic density distribution of the document, a multi-level threshold strategy is set. By calculating the document's average semantic density (μ) and semantic density standard deviation (σ), the knowledge-rich regions are divided into different importance levels. The first-level rich region corresponds to the region with a density greater than μ + 2σ, which is the core important region; the second-level rich region corresponds to the region with a density greater than μ + σ, which is the important region; and the third-level rich region corresponds to the region with a density greater than μ + 0.5σ, which is the generally important region. This achieves accurate differentiation of regions with different knowledge densities and provides a basis for the reasonable allocation of subsequent labeled resources.
[0053] Using the aforementioned multi-level threshold strategy, the knowledge-rich region boundaries are identified on the semantic window density distribution curve. By calculating the first-order gradient (reflecting the density change trend) and the second-order gradient (reflecting the density change acceleration), and combining the intervals in the curve where the density value exceeds each level of threshold, preliminary candidate boundary points are determined. Subsequently, adaptive boundary optimization is performed. Based on the characteristics of different levels of enriched regions, the window range is adjusted near the initial boundary through local search, so that the boundary conforms to the natural segmentation of the semantic topic. At the same time, the recognition results of multi-scale sliding windows are integrated to eliminate boundary deviation. Finally, the precise boundaries and levels of each knowledge-rich region are determined, forming the knowledge-rich region recognition results.
[0054] In one possible implementation, step S344 further includes:
[0055] Step S3441: Perform density gradient calculation and gradient smoothing based on the semantic window density distribution curve to obtain the semantic density gradient calculation result.
[0056] Step S3442: Perform boundary candidate point detection and boundary point verification and screening on the semantic density gradient calculation results to determine the boundary candidate points.
[0057] Step S3443: Using the multi-level threshold strategy, the semantic window density distribution curve is segmented and labeled based on the boundary candidate points to obtain the basic hierarchical knowledge enrichment region.
[0058] Step S3444: Perform local search optimization and multi-scale boundary fusion on the basic hierarchical knowledge enrichment region to determine the recognition result of the knowledge enrichment region.
[0059] Specifically, density gradients are calculated based on the semantic window density distribution curve. The first-order gradient reflects the density change trend and is calculated as gradient-1[i] = (density[i+1] - density[i-1]) / 2, meaning the first-order gradient value at position i is equal to the difference between the density values at positions i+1 and i-1 divided by 2. The second-order gradient reflects the acceleration of density change and is calculated as gradient-2[i] = density[i+1] - 2 × density[i] + density[i-1], meaning the second-order gradient value at position i is equal to the density value at position i+1 minus twice the density value at position i plus the density value at position i-1. The obtained first-order and second-order gradients are then smoothed using the same method as density Gaussian smoothing, i.e., weighted summation of gradient values using Gaussian weights to reduce noise interference, ultimately yielding the semantic density gradient calculation result.
[0060] When detecting boundary candidate points based on the semantic density gradient calculation results, the following steps are taken: First, zero-crossing detection is performed. If the previous value (gradient-1[i-1]) in the first-order gradient is greater than 0 and the current value (gradient-1[i]) is less than 0, or the previous value is less than 0 and the current value is greater than 0, the position i is added to the boundary candidate point set. Then, extreme point detection is performed. If the product of the previous value (gradient-2[i-1]) and the current value (gradient-2[i]) in the second-order gradient is less than 0, the position i is added to the boundary candidate point set. Subsequently, boundary point verification and screening are performed. The t-test is used to calculate the t-statistic of the semantic density difference between the regions before and after the candidate point. The t-statistic is calculated by subtracting the average semantic density of the region before the candidate point from the average semantic density of the region after the candidate point, and dividing the result by the square root of the product of the pooled variance and (the inverse of the sample size of the region before the candidate point plus the inverse of the sample size of the region after the candidate point). If the absolute value of the t-statistic is greater than the set threshold, the candidate point is included in the confirmed boundary candidate point set, and the boundary candidate point is finally determined.
[0061] Using the aforementioned multi-level threshold strategy (Level 1 threshold μ+2σ, Level 2 threshold μ+σ, Level 3 threshold μ+0.5σ), the semantic window density distribution curve is processed based on the determined boundary candidate points. First, the curve is divided into several continuous intervals using the boundary candidate points as segmentation markers, and the average semantic density within each interval is calculated. Then, the average semantic density of each interval is compared with the multi-level thresholds. If the average density exceeds the Level 1 threshold, it is marked as a Level 1 enriched region; if it exceeds the Level 2 threshold (but does not reach Level 1), it is marked as a Level 2 enriched region; if it exceeds the Level 3 threshold (but does not reach Level 2), it is marked as a Level 3 enriched region. Intervals that do not reach the Level 3 threshold are not marked as enriched regions. Finally, a basic hierarchical knowledge enriched region containing the location, level, and corresponding density features of each interval is formed.
[0062] Local search optimization is performed on the basic hierarchical knowledge-rich regions. Starting from the initial boundary of each enrichment region, the search slides within ±2 windows, calculating the semantic topic consistency score (measured by the mean cosine similarity of words within the window) corresponding to different boundary positions. The position with the highest score is selected as the optimized boundary. At the same time, multi-scale boundary fusion is performed, and the boundary information identified by word-level, sentence-level, and paragraph-level windows is weighted and fused (word-level weight 0.4, sentence-level weight 0.3, paragraph-level weight 0.3) to eliminate cross-scale boundary bias. Finally, the precise boundaries, ranges, and corresponding semantic feature labels of each level of knowledge-rich region are determined, forming the knowledge-rich region identification result.
[0063] In one possible implementation, step S400 further includes:
[0064] Step S410: Obtain the hierarchical association modeling target, which includes document-level association, paragraph-level association, sentence-level association, and entity-level association.
[0065] Step S420: Perform hierarchical semantic representation on the multi-format scientific and technological documents according to the hierarchical association modeling objective to obtain document-level semantic vectors, paragraph-level semantic vectors, sentence-level semantic vectors, and entity-level semantic vectors.
[0066] Step S430: Perform cross-level association learning on the document-level semantic vector, paragraph-level semantic vector, sentence-level semantic vector, and entity-level semantic vector to generate a basic semantic association model.
[0067] Step S440: Perform global consistency constraints and dynamic association weight learning on the basic semantic association model to establish the cross-level semantic association model.
[0068] Specifically, the goal of hierarchical association modeling is to obtain the following: document-level associations include thematic inheritance relationships between chapters (i.e., the continuation and derivation relationships of different chapters on the core theme), citation network relationships (association networks formed through mutual citations between documents), and global concept distribution patterns (the distribution patterns and association characteristics of concepts throughout the document); paragraph-level associations cover logical relationships between paragraphs (such as causal, comparative, and progressive relationships that reflect the logical connection between paragraphs), thematic coherence relationships (the consistency and coherence of paragraphs in terms of thematic content), and spatial position relationships (the association formed by the position of paragraphs in the document); sentence-level associations involve inter-sentence referential relationships (the referential connection between sentences achieved through pronouns, etc.), semantic similarity relationships (the degree of similarity between sentences in terms of semantic content), and argumentation structure relationships (the relationship between premises and conclusions formed by sentences in the argumentation process); entity-level associations include synonym relationships (the semantic equivalence between different entities), hierarchical relationships (the class and subordinate relationships between entities), and associated entity relationships (the association between entities based on specific semantics).
[0069] Based on the goal of association modeling at the document, paragraph, sentence, and entity levels, hierarchical semantic representation is performed on multi-format scientific and technological documents. For the sentence level, a pre-trained language model (such as BERT) is used to encode each sentence, generating a sentence-level semantic vector containing contextual information. For the paragraph level, based on the semantic relationships between sentences within a paragraph, an attention mechanism is used to weighted aggregate the semantic vectors of all sentences in the paragraph, resulting in a paragraph-level semantic vector that reflects the paragraph's theme and logic. For the document level, the semantic vectors of each paragraph within the document are fused by combining the theme inheritance relationship between chapters and the global concept distribution pattern, generating a document-level semantic vector that reflects the overall content of the document. For the entity level, an entity embedding model (such as TransE) is used to encode entities in the document, and the vector representation is optimized by combining synonym, hyponymy, and other association relationships between entities, resulting in an entity-level semantic vector.
[0070] By combining the hierarchical structure features of documents, cross-level association learning is performed on semantic vectors at the document, paragraph, sentence, and entity levels. A basic semantic association model is generated through a hierarchical propagation mechanism and an attention-based association network. Specifically, the hierarchical propagation mechanism employs a message propagation algorithm to achieve bidirectional flow of semantic information. Bottom-up message propagation means that the upward message of the i-th node is equal to the aggregation of semantic information of its child nodes; top-down message propagation means that the downward message of the i-th node is obtained by propagating the semantic information of its parent node. Then, the original vector is fused with the upward and downward messages to obtain the updated semantic representation of the i-th node. Meanwhile, the attention association network uses graph attention network to model complex relationships between levels. For the relationship between feature description and corresponding implementation content, the constraint relationship between entity and sentence, etc., the attention weight is calculated by the softmax function of (weight matrix multiplied by the concatenation of query vector and key vector plus bias term) (where query vector and key vector are taken from different level vectors, and weight matrix and bias term are learnable parameters). Then, the context vector is calculated by the weighted sum of attention weight and value vector (i.e., each value vector is multiplied by the corresponding attention weight and then summed) to aggregate key semantic information, and finally generate a basic semantic association model that can reflect the relationship between different levels of the document.
[0071] To establish a cross-level semantic association model, a global consistency constraint and dynamic association weight learning are applied to the basic semantic association model. The global consistency constraint is achieved through annotation consistency checks, ensuring compatibility between annotation results at different levels. Entity annotations must be semantically consistent with the sentence they belong to, sentence annotations must fit the theme of their paragraph, and paragraph annotations must be consistent with the overall document content. A conflict detection algorithm is also designed to automatically identify and resolve annotation conflicts. If a conflict is detected between an entity annotation and the sentence context, the conflict is resolved by incorporating the global context to obtain a reasonable annotation result. Based on this, dynamic association weight learning is performed. The association weight values are dynamically adjusted according to the tightness, frequency, and importance of semantic associations between levels, ultimately establishing a cross-level semantic association model that accurately reflects the semantic associations between different levels of the document.
[0072] In one possible implementation, step S400 further includes:
[0073] Step S450: The recursive annotation architecture specifically includes document-level coarse-grained annotation, paragraph-level medium-grained annotation, sentence-level fine-grained annotation, and entity-level precise annotation.
[0074] Specifically, the designed recursive annotation architecture is as follows: First, document-level coarse-grained annotation is performed. Based on the document-level semantic vector and global topic distribution, the overall technical field, core research topics, and main technical directions of the document are annotated to determine the macro-scope of the document. Next, paragraph-level medium-grained annotation is carried out. Combining the paragraph-level semantic vector and the logical relationship between paragraphs, the topic category and functional positioning of each paragraph in the document (such as introduction, core discussion, conclusion, etc.) are annotated. Then, sentence-level fine-grained annotation is performed. Based on the sentence-level semantic vector and the relationship between sentences, the semantic role of sentences (such as premise, conclusion, method description, result statement, etc.) and key information types (such as technical features, parameter data, experimental phenomena, etc.) are annotated. Finally, entity-level precise annotation is implemented. Based on the entity-level semantic vector and the relationship between entities, the technical terms, variables, device components, and other entities in the document are annotated with type (such as equipment, algorithm, material, etc.) and attribute (such as entity function, parameters, state, etc.). Through recursive annotation from coarse to fine, a multi-level precise characterization of the document content is achieved.
[0075] In one possible implementation, step S400 further includes:
[0076] Step S460: Use the recursive annotation architecture to perform entity boundary prediction and multi-level annotation on the knowledge-rich region identification results to obtain document-level annotation results, paragraph-level annotation results, sentence-level annotation results and entity-level annotation results.
[0077] Step S470: Determine the initial semantic annotation results based on the document-level annotation results, paragraph-level annotation results, sentence-level annotation results, and entity-level annotation results.
[0078] Step S480: Based on the cross-level semantic association model, perform global consistency verification and recursive feedback optimization on the initial semantic annotation results to obtain the structured semantic annotation results.
[0079] Specifically, a recursive annotation architecture is used to process the knowledge-rich region identification results. Document-level coarse-grained annotation is used to identify document topics, clarifying the core topic category to which the knowledge-rich region belongs. Simultaneously, key areas are pre-located to pinpoint the distribution range of key content, and global entity type prediction is implemented to determine the approximate category of all entities, thus obtaining document-level annotation results. Based on paragraph-level medium-grained annotation, the knowledge-rich region identification results are refined into paragraph topics, accurately extracting the specific topics of each paragraph. Entity density analysis within paragraphs is conducted to understand entity distribution, and cross-paragraph relationships are pre-established to grasp inter-paragraph relationships. The process involves several steps: First, paragraph-level annotations are obtained. Second, sentence-level fine-grained annotations are used to analyze the semantic roles of sentences in the knowledge-rich region identification results to clarify their semantic expression. Syntactic-guided entity boundary prediction is used to determine the boundary range of entities. Context-aware entity classification is then combined to accurately classify entity categories, resulting in sentence-level annotations. Finally, based on precise entity-level annotations, accurate boundary detection is performed on the knowledge-rich region identification results to determine the accurate boundaries of entities. Nested entities are processed to clarify the inclusion relationships between entities, and standardized entity links are performed to associate them with a standard terminology system, resulting in entity-level annotations.
[0080] Based on document-level, paragraph-level, sentence-level, and entity-level annotation results, a multi-dimensional fusion algorithm is designed to determine the initial semantic annotation results. First, a weighted fusion strategy is used (assigning weights according to the importance of each level of annotation in semantic expression, such as entity-level annotations having higher weights than sentence-level annotations) to remove duplicate annotation information at the same level. Then, association rule mining (such as identifying the inherent mapping relationship between entities and sentences, and paragraphs and documents) is used to supplement the missing association annotations between levels. Finally, an entity linking tool is used to associate and bind the scattered entity annotations with sentence and paragraph annotations to form the initial semantic annotation results containing semantic associations between levels.
[0081] Based on the cross-level semantic association model, the following methods are used to process the initial semantic annotation results to obtain structured semantic annotation results. Global consistency verification involves calculating the association scores between the annotation results at each level and the output of the cross-level semantic association model (e.g., using cosine similarity algorithm to compare the matching degree between entity annotations and sentence semantic vectors, and sentence annotations and paragraph topic vectors), setting thresholds to filter out inconsistent annotations. Recursive feedback optimization involves constructing a feedback iteration mechanism, using the verified inconsistent annotations as samples to input into the annotation model for parameter adjustment (e.g., using gradient descent to optimize model weights), and regenerating the annotation results. This verification and optimization process is repeated until global consistency is achieved, ultimately forming structured semantic annotation results.
[0082] In one possible implementation, step S460 further includes:
[0083] Step S461: Use the document-level coarse-grained annotation to perform document topic identification, key area pre-location, and global entity type prediction on the knowledge-rich area identification results to obtain document-level annotation results.
[0084] Step S462: Based on the granular annotation at the paragraph level, refine the paragraph topic, analyze the entity density within the paragraph, and pre-establish cross-paragraph relationships based on the identification results of the knowledge-rich area to obtain the paragraph-level annotation results.
[0085] Step S463: Use the sentence-level fine-grained annotation to perform sentence semantic role analysis, syntactic-guided entity boundary prediction, and context-aware entity classification on the knowledge-rich region identification results to obtain sentence-level annotation results.
[0086] Step S464: Based on the entity-level precise annotation, perform precise boundary detection, nested entity processing, and entity standardization linking on the knowledge-rich region identification results to obtain entity-level annotation results.
[0087] Specifically, when processing the identification results of knowledge-rich areas using document-level coarse-grained annotation, topic models (such as LDA models) are used to cluster the text into topics, extract high-frequency keywords, and match them with a pre-defined topic thesaurus to complete document topic identification. Text density algorithms (such as calculating word frequency density and semantic similarity) are combined to locate chapters or segments with high knowledge density, thereby achieving pre-positioning of key areas. Rule-based and statistical entity classifiers (integrating domain dictionary matching and part-of-speech feature analysis) are used to perform preliminary type classification of the identified entities, complete global entity type prediction, and finally integrate the above results to obtain document-level annotation results.
[0088] When processing the knowledge-rich region identification results based on paragraph-level granular annotation, a paragraph vector model (such as Doc2Vec) is used to convert each paragraph into a vector representation. The paragraph vectors are then clustered using the K-means clustering algorithm, and paragraph topics are refined by combining them with a pre-defined subdivided topic tag library. Entity counting tools are used to count the frequency and distribution density of entities in each paragraph, calculate the entity density value (number of entities / paragraph length), and classify the density levels to achieve entity density analysis within paragraphs. By calculating the cosine similarity between paragraphs, combined with paragraph position information and conjunction recognition (such as "in addition" and "on the contrary"), a paragraph association matrix is constructed to pre-establish semantic relationships across paragraphs. Finally, the above results are integrated to obtain the paragraph-level annotation results.
[0089] When processing the knowledge-rich region identification results using sentence-level fine-grained annotation, a semantic role annotation model (such as a BERT-based SRL model) is used to identify predicates and related arguments (such as agent, patient, time, and place) in the sentence, clarifying the semantic roles of each component of the sentence to complete the sentence semantic role analysis. Syntactic analysis tools (such as a dependency parser) are used to parse the syntactic structure of the sentence, locating the possible boundary range of entities based on syntactic relations such as subject-predicate and verb-object. Syntactic-guided entity boundary prediction is achieved by combining the part-of-speech features of the entity's start and end words (such as nouns and noun phrases). Based on the semantic vectors of words within the context window (generated through a pre-trained language model), the semantic similarity between the entity to be classified and the candidate category labels is calculated. Context-aware entity classification is completed by combining the collocation relationships of the entities in the sentence. Finally, the above results are integrated to obtain the sentence-level annotation results.
[0090] When processing knowledge-rich region identification results based on entity-level precise annotation, precise boundary detection employs a boundary detection algorithm incorporating an attention mechanism. This algorithm focuses on key features through attention weight allocation, simultaneously inputting grammatical features (such as part-of-speech tags and dependency syntax relations), semantic features (such as word vectors generated by a pre-trained language model), and prosodic features (such as pause positions and rhythmic patterns within sentences). The model calculates and outputs the probability distribution of the entity's start and end positions to determine the precise boundary. For nested entity processing, a hierarchical nested entity recognition algorithm is designed to address the nested structures in scientific literature. This algorithm employs a multi-label annotation strategy, using stacking... The neural network first identifies outer entities, then uses fine-grained features (such as the entity's internal syntactic dependency tree) to identify inner entities within their text scope, while simultaneously recording the nested hierarchical relationships of entities. Entity standardization and linking are achieved by matching the identified entities with standard knowledge bases (such as domain ontology and knowledge graph) using an entity linking tool. The similarity between entity name strings and semantic vector cosine similarity is calculated, and the standard entity with the highest matching degree is selected to establish a link, thus achieving standardized representation. At the same time, related knowledge is extracted from the knowledge base to supplement the entity information, thus completing knowledge enhancement. Finally, the results of the three steps are integrated to obtain the entity-level annotation results.
[0091] Example 2 is based on the same inventive concept as the four-dimensional index-based automatic annotation method for scientific and technological literature knowledge in the previous examples, such as... Figure 2 As shown, this application provides an automatic annotation system for scientific and technological literature based on a four-dimensional index. The system and method embodiments in this application are based on the same inventive concept. The system includes:
[0092] The four-dimensional index architecture construction module 10 is used to construct a four-dimensional index architecture, which includes a data input layer, a four-dimensional index construction layer, a knowledge enrichment area identification layer, a boundary detection and annotation layer, and a result output layer. The specific structure of the four-dimensional index construction layer includes a document layer index, a paragraph layer index, a sentence layer index, and an entity layer index.
[0093] The index parsing module 20 is used to receive multi-format scientific and technological documents through the data input layer, and to perform index parsing on the multi-format scientific and technological documents in sequence based on the document layer index, paragraph layer index, sentence layer index and entity layer index to obtain the document layer semantic structure, paragraph layer semantic relationship graph, sentence semantic relationship network and entity semantic relationship.
[0094] The knowledge enrichment region identification module 30 is used to perform semantic density calculation and knowledge enrichment region identification on the document layer semantic structure, paragraph layer semantic relationship graph, sentence semantic relationship network and entity semantic relationship based on the knowledge enrichment region identification layer, so as to obtain the knowledge enrichment region identification result.
[0095] The annotation result acquisition module 40 is used to establish a cross-level semantic association model, design a recursive annotation architecture based on the boundary detection annotation layer, use the recursive annotation architecture and the cross-level semantic association model to perform entity boundary prediction and multi-level annotation on the knowledge-rich area identification results, obtain structured semantic annotation results, and visualize the structured semantic annotation results through the result output layer.
[0096] Furthermore, the system is also used to implement the following functions:
[0097] Based on the document-level index, document structure parsing and global feature extraction are performed on the multi-format scientific and technological documents to obtain the document-level semantic structure; based on the paragraph-level index, paragraph segmentation and inter-paragraph relationship modeling are performed on the multi-format scientific and technological documents to construct a paragraph-level semantic relationship graph; the sentence-level index is used to perform sentence semantic encoding and semantic relationship analysis on the multi-format scientific and technological documents to generate a sentence semantic relationship network; based on the entity-level index, multi-type entity recognition and entity relationship extraction are performed on the multi-format scientific and technological documents to obtain entity semantic relationships.
[0098] Furthermore, the system is also used to implement the following functions:
[0099] Based on the knowledge-rich region identification layer, multi-dimensional semantic features are extracted from the document-level semantic structure, paragraph-level semantic relationship graph, sentence semantic relationship network, and entity semantic relationship to obtain lexical-level feature sets, syntactic-level feature sets, and semantic-level feature sets. A density calculation model is designed, which includes a multi-feature fusion model and an attention weight mechanism. The multi-feature fusion model adopts a multilayer perceptron structure. A multi-scale sliding window strategy is designed, which includes word-level windows, sentence-level windows, and paragraph-level windows. Using the multi-scale sliding window strategy, semantic density calculation and knowledge-rich region identification are performed on the document-level semantic structure, paragraph-level semantic relationship graph, sentence semantic relationship network, and entity semantic relationship based on the density calculation model to obtain the knowledge-rich region identification results.
[0100] Furthermore, the system is also used to implement the following functions:
[0101] The multi-scale sliding window strategy is used to calculate semantic density and perform density Gaussian smoothing on the document-level semantic structure, paragraph-level semantic relationship graph, sentence semantic relationship network, and entity semantic relationship based on the density calculation model, obtaining a semantic window density distribution curve. Based on the semantic window density distribution curve, the knowledge-rich region threshold T = μ + α × σ is determined, where μ is the document's average semantic density, σ is the semantic density standard deviation, and α is an adaptive parameter, typically ranging from 1.5 to 2.0. Based on the knowledge-rich region threshold, a multi-level threshold strategy is set. The multi-level threshold strategy is then used to identify the knowledge-rich region boundary and perform adaptive boundary optimization on the semantic window density distribution curve, determining the knowledge-rich region identification result.
[0102] Furthermore, the system is also used to implement the following functions:
[0103] Density gradient calculation and gradient smoothing are performed based on the semantic window density distribution curve to obtain the semantic density gradient calculation result. Boundary candidate points are detected and verified based on the semantic density gradient calculation result to determine the boundary candidate points. The multi-level threshold strategy is used to segment and label the enriched region of the semantic window density distribution curve based on the boundary candidate points to obtain the basic hierarchical knowledge enriched region. Local search optimization and multi-scale boundary fusion are performed on the basic hierarchical knowledge enriched region to determine the knowledge enriched region identification result.
[0104] Furthermore, the system is also used to implement the following functions:
[0105] Obtain hierarchical association modeling objectives, including document-level association, paragraph-level association, sentence-level association, and entity-level association; perform hierarchical semantic representation on the multi-format scientific and technological documents according to the hierarchical association modeling objectives to obtain document-level semantic vectors, paragraph-level semantic vectors, sentence-level semantic vectors, and entity-level semantic vectors; perform cross-hierarchical association learning on the document-level semantic vectors, paragraph-level semantic vectors, sentence-level semantic vectors, and entity-level semantic vectors to generate a basic semantic association model; apply global consistency constraints and dynamic association weight learning to the basic semantic association model to establish the cross-hierarchical semantic association model.
[0106] Furthermore, the system is also used to implement the following functions:
[0107] The recursive annotation architecture specifically includes document-level coarse-grained annotation, paragraph-level medium-grained annotation, sentence-level fine-grained annotation, and entity-level precise annotation.
[0108] Furthermore, the system is also used to implement the following functions:
[0109] The recursive annotation architecture is used to predict entity boundaries and perform multi-level annotation on the knowledge-rich region identification results, obtaining document-level annotation results, paragraph-level annotation results, sentence-level annotation results, and entity-level annotation results. Based on the document-level annotation results, paragraph-level annotation results, sentence-level annotation results, and entity-level annotation results, initial semantic annotation results are determined. Based on the cross-level semantic association model, the initial semantic annotation results are subjected to global consistency verification and recursive feedback optimization to obtain the structured semantic annotation results.
[0110] Furthermore, the system is also used to implement the following functions:
[0111] The document-level coarse-grained annotation is used to perform document topic identification, key region pre-location, and global entity type prediction on the knowledge-rich region identification results to obtain document-level annotation results. Based on the paragraph-level medium-grained annotation, the paragraph topic is refined, intra-paragraph entity density is analyzed, and cross-paragraph relationship is pre-established on the knowledge-rich region identification results to obtain paragraph-level annotation results. The sentence-level fine-grained annotation is used to perform sentence semantic role analysis, syntactic-guided entity boundary prediction, and context-aware entity classification on the knowledge-rich region identification results to obtain sentence-level annotation results. Based on the entity-level precise annotation, the knowledge-rich region identification results are subjected to precise boundary detection, nested entity processing, and entity standardization linking to obtain entity-level annotation results.
[0112] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0113] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0114] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.
Claims
1. A method for automatic annotation of scientific and technological literature knowledge based on four-dimensional indexing, characterized in that, The method includes: A four-dimensional index architecture is constructed, which includes a data input layer, a four-dimensional index construction layer, a knowledge enrichment region identification layer, a boundary detection and annotation layer, and a result output layer. The specific structure of the four-dimensional index construction layer includes document layer index, paragraph layer index, sentence layer index, and entity layer index. The data input layer receives multi-format scientific and technological documents, and the multi-format scientific and technological documents are indexed and parsed sequentially based on the document layer index, paragraph layer index, sentence layer index and entity layer index to obtain the document layer semantic structure, paragraph layer semantic relationship graph, sentence semantic relationship network and entity semantic relationship; Based on the knowledge enrichment region identification layer, semantic density calculation and knowledge enrichment region identification are performed on the semantic structure of the document layer, the semantic relationship graph of the paragraph layer, the semantic relationship network of the sentence and the semantic relationship of the entity, to obtain the knowledge enrichment region identification result; A cross-level semantic association model is established. Based on the boundary detection annotation layer, a recursive annotation architecture is designed. The recursive annotation architecture and the cross-level semantic association model are used to predict entity boundaries and perform multi-level annotation on the knowledge-rich area identification results to obtain structured semantic annotation results. The structured semantic annotation results are then visualized and displayed through the result output layer. The obtained knowledge-rich region identification results include: Based on the knowledge enrichment region identification layer, multi-dimensional semantic features are extracted from the document layer semantic structure, paragraph layer semantic relationship graph, sentence semantic relationship network and entity semantic relationship to obtain lexical feature set, syntactic feature set and semantic feature set; Design a density calculation model, which includes a multi-feature fusion model and an attention weight mechanism, wherein the multi-feature fusion model adopts a multilayer perceptron structure; Design a multi-scale sliding window strategy, which includes word-level window, sentence-level window and paragraph-level window; Using the multi-scale sliding window strategy, based on the density calculation model, semantic density calculation and knowledge enrichment region identification are performed on the document layer semantic structure, paragraph layer semantic relationship graph, sentence semantic relationship network and entity semantic relationship to obtain knowledge enrichment region identification results; The obtained knowledge-rich region identification results include: Using the multi-scale sliding window strategy, semantic density calculation and density Gaussian smoothing are performed on the document layer semantic structure, paragraph layer semantic relationship graph, sentence semantic relationship network and entity semantic relationship based on the density calculation model to obtain the semantic window density distribution curve. Based on the semantic window density distribution curve, the knowledge enrichment area threshold T is determined as T = μ + α × σ, where μ is the average semantic density of the document, σ is the standard deviation of the semantic density, and α is an adaptive parameter with a value of 1.5-2.
0. Based on the knowledge enrichment area threshold, a multi-level threshold strategy is set; The multi-level threshold strategy is used to identify the knowledge-rich region boundary and perform adaptive boundary optimization on the semantic window density distribution curve to determine the knowledge-rich region identification result.
2. The method for automatic annotation of scientific and technological literature knowledge based on four-dimensional indexing as described in claim 1, characterized in that, The process of obtaining the document-level semantic structure, paragraph-level semantic relationship graph, sentence semantic relationship network, and entity semantic relationships includes: Based on the document layer index, the document structure of the multi-format scientific and technological documents is parsed and global features are extracted to obtain the document layer semantic structure. Based on the paragraph-level index, paragraph segmentation and inter-paragraph relationship modeling are performed on the multi-format scientific and technological documents to construct a paragraph-level semantic relationship graph. The sentence-level index is used to perform sentence semantic encoding and semantic relation analysis on the multi-format scientific and technological documents to generate a sentence semantic relation network. Based on the entity layer index, multi-type entity recognition and entity relation extraction are performed on the multi-format scientific and technological documents to obtain entity semantic relations.
3. The method for automatic annotation of scientific and technological literature knowledge based on four-dimensional indexing as described in claim 1, characterized in that, The determination of the knowledge-rich region identification result includes: Based on the semantic window density distribution curve, density gradient calculation and gradient smoothing are performed to obtain the semantic density gradient calculation result. The semantic density gradient calculation results are used to detect and verify boundary candidate points to determine the boundary candidate points. The multi-level threshold strategy is used to segment and label the semantic window density distribution curve based on the boundary candidate points to obtain the basic hierarchical knowledge enrichment region. The basic hierarchical knowledge-rich region is subjected to local search optimization and multi-scale boundary fusion to determine the identification result of the knowledge-rich region.
4. The method for automatic annotation of scientific and technological literature knowledge based on four-dimensional indexing as described in claim 1, characterized in that, The establishment of a cross-level semantic association model includes: Obtain the hierarchical association modeling objectives, which include document-level association, paragraph-level association, sentence-level association, and entity-level association; Based on the hierarchical association modeling objective, the multi-format scientific and technological documents are subjected to hierarchical semantic representation to obtain document-level semantic vectors, paragraph-level semantic vectors, sentence-level semantic vectors, and entity-level semantic vectors. Cross-level association learning is performed on the document-level semantic vectors, paragraph-level semantic vectors, sentence-level semantic vectors, and entity-level semantic vectors to generate a basic semantic association model; The basic semantic association model is subjected to global consistency constraints and dynamic association weight learning to establish the cross-level semantic association model.
5. The method for automatic annotation of scientific and technological literature knowledge based on four-dimensional indexing as described in claim 1, characterized in that, The recursive annotation architecture specifically includes document-level coarse-grained annotation, paragraph-level medium-grained annotation, sentence-level fine-grained annotation, and entity-level precise annotation.
6. The method for automatic annotation of scientific and technological literature knowledge based on four-dimensional indexing as described in claim 5, characterized in that, The obtained structured semantic annotation results include: The recursive annotation architecture is used to perform entity boundary prediction and multi-level annotation on the knowledge-rich region identification results to obtain document-level annotation results, paragraph-level annotation results, sentence-level annotation results and entity-level annotation results. Based on the document-level annotation results, paragraph-level annotation results, sentence-level annotation results, and entity-level annotation results, determine the initial semantic annotation results; Based on the cross-level semantic association model, the initial semantic annotation results are subjected to global consistency verification and recursive feedback optimization to obtain the structured semantic annotation results.
7. The method for automatic annotation of scientific and technological literature knowledge based on four-dimensional indexing as described in claim 6, characterized in that, The process of obtaining document-level annotation results, paragraph-level annotation results, sentence-level annotation results, and entity-level annotation results includes: The document-level coarse-grained annotation is used to perform document topic identification, key region pre-location, and global entity type prediction on the knowledge-rich area identification results to obtain document-level annotation results; Based on the paragraph-level granular annotation, the knowledge-rich area identification results are refined into paragraph topics, analyzed for entity density within paragraphs, and pre-established for cross-paragraph relationships to obtain paragraph-level annotation results. The sentence-level fine-grained annotation is used to perform sentence semantic role analysis, syntactic-guided entity boundary prediction, and context-aware entity classification on the knowledge-rich region identification results to obtain sentence-level annotation results; Based on the entity-level precise annotation, the identification results of the knowledge-rich region are subjected to precise boundary detection, nested entity processing, and entity standardization linking to obtain entity-level annotation results.
8. An automatic knowledge annotation system for scientific and technological literature based on four-dimensional indexing, characterized in that, The system is used to implement the automatic annotation method for scientific and technological literature knowledge based on four-dimensional indexing as described in any one of claims 1-7, and the system comprises: The four-dimensional index architecture construction module is used to construct a four-dimensional index architecture. The four-dimensional index architecture includes a data input layer, a four-dimensional index construction layer, a knowledge enrichment area identification layer, a boundary detection and annotation layer, and a result output layer. The specific structure of the four-dimensional index construction layer includes document layer index, paragraph layer index, sentence layer index, and entity layer index. The index parsing module is used to receive multi-format scientific and technological documents through the data input layer, and to perform index parsing on the multi-format scientific and technological documents in sequence based on the document layer index, paragraph layer index, sentence layer index and entity layer index to obtain the document layer semantic structure, paragraph layer semantic relationship graph, sentence semantic relationship network and entity semantic relationship; The knowledge enrichment region identification module is used to perform semantic density calculation and knowledge enrichment region identification on the document layer semantic structure, paragraph layer semantic relationship graph, sentence semantic relationship network and entity semantic relationship based on the knowledge enrichment region identification layer, so as to obtain the knowledge enrichment region identification result. The annotation result acquisition module is used to establish a cross-level semantic association model, design a recursive annotation architecture based on the boundary detection annotation layer, use the recursive annotation architecture and the cross-level semantic association model to perform entity boundary prediction and multi-level annotation on the knowledge-rich area identification results to obtain structured semantic annotation results, and visualize the structured semantic annotation results through the result output layer.
Citation Information
Patent Citations
Segmented semantic annotation method in weak annotation environment
CN110888991A
Financial document level event extraction method and system based on multi-semantic enhancement
CN119808793A