Artificial intelligence-based literature structured extraction method and system

Through the preprocessing and feature extraction model of the original document data, a structured element collection is generated, which solves the problem of incomplete extraction of information within the document, and efficient and accurate document structured processing is achieved, improving the efficiency and quality of document reading and intelligent application.

CN120336416AActive Publication Date: 2025-07-18JIEHELIX (SHANGHAI) MEDICAL TECH CO LTD

Patent Information

Application Number
CN202510820361.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing literature processing technologies are difficult to deeply explore structured information within the literature, resulting in incomplete and inaccurate information extraction, affecting literature understanding and intelligent application.

Method used

By obtaining the original literature data, pre-processing, the pre-trained literature feature extraction model is called, content features and structural features are extracted, and structured analysis and integration optimization are carried out to generate a structured feature collection.

Benefits of technology

It realizes the accurate extraction of the subject elements, logical relationship elements and key information elements of the literature, builds their complex relationships, improves the automation and intelligence level of literature processing, and improves the efficiency and quality of literature reading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336416A_ABST
    Figure CN120336416A_ABST
Patent Text Reader

Abstract

The invention provides a literature structured extraction method and system based on artificial intelligence, and aims to solve the problem that information extraction is incomplete and inaccurate in an existing literature processing technology. Firstly, an original literature data set comprising a plurality of literature units is obtained and preprocessed to obtain a standardized literature data set; calling a pre-trained literature feature extraction model to process the standardized literature data, extracting content features and structural features of the literature units, based on the content features and the structural features of the literature units, executing structured analysis processing to generate a structured element set containing theme elements, logic relation elements and key information elements, and finally, storing the structured element set. And integrating and optimizing the structured element set to generate a final structured literature result containing the element association relationship, thereby realizing automatic and structured extraction of the literature information, improving the literature processing efficiency and accuracy, and providing a new way for intelligent management and application of the literature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology. Specifically, it relates to a method and system for literature structured extraction based on artificial intelligence. Background Art

[0002] In today's information age, the number of literatures in fields such as academic research, technical reports, and business analysis has increased explosively. These literatures not only cover a wide range of topics, but also their content structures have become increasingly complex, including various forms such as chapter division, paragraph organization, chart embedding, and citation marking. For scientific researchers, enterprise decision-makers, and ordinary readers, how to quickly and accurately extract valuable information from a vast amount of literatures and understand the internal logical structure and key points of the literatures has become an urgent problem to be solved.

[0003] Traditional literature processing methods mainly rely on manual reading and collation. This method is not only time-consuming and laborious, with low efficiency, but also easily affected by subjective factors such as personal knowledge background and reading habits, resulting in difficulties in ensuring the accuracy and comprehensiveness of information extraction. Although natural language processing and machine learning technologies have made certain progress in the field of literature processing in recent years, most of the existing literature processing technologies focus on the surface-level extraction of literature content, such as keyword extraction, abstract generation, etc., and fail to deeply explore the structured information inside the literatures. These structured information, such as chapter titles, logical relationships between paragraphs, positioning and association of key information points, etc., play a crucial role in deeply understanding literature content, constructing knowledge graphs, and realizing intelligent applications of literatures. Summary of the Invention

[0004] In view of the problems mentioned above, in combination with the first aspect of the present invention, embodiments of the present invention provide a method for literature structured extraction based on artificial intelligence. The method includes: Obtain a set of original literature data, where the set of original literature data contains multiple literature units, and each literature unit consists of text content and meta-information; Preprocess the set of original literature data to obtain a set of standardized literature data, where the set of standardized literature data contains text paragraphs and meta-information entries in a unified format; Call a pre-trained literature feature extraction model to process the set of standardized literature data to obtain the content features and structural features of the literature units; Perform structured parsing processing based on the content features and the structural features to generate a set of structured elements of the literature units, where the set of structured elements contains topic elements, logical relationship elements, and key information elements; Integrate and optimize the set of structured elements to generate a final structured literature result containing element association relationships.

[0005] On the other hand, an embodiment of the present invention further provides a literature structured extraction system based on artificial intelligence, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0006] Based on the above aspects, the embodiment of the present invention realizes the in-depth mining and efficient utilization of the original literature data by comprehensively applying advanced technical means such as preprocessing, feature extraction, structured parsing and integration optimization. It can automatically identify and accurately extract the theme elements in the literature to ensure the accurate grasp of the core content of the literature. At the same time, by deeply analyzing the logical structure of the literature, the logical relationship elements are extracted to reveal the logical context and argumentation process inside the literature. In addition, it can effectively capture the key information elements in the literature. More importantly, it can build the complex association relationships between these key information elements to form a complete structured literature result, enabling users to intuitively and comprehensively understand the literature content, improving the efficiency and quality of literature reading, and thus significantly improving the automation and intelligence level of literature processing. Description of the Drawings

[0007] Figure 1 is a schematic execution flow diagram of the literature structured extraction method based on artificial intelligence provided by an embodiment of the present invention.

[0008] Figure 2 is a schematic diagram of exemplary hardware and software components of the literature structured extraction system based on artificial intelligence provided by an embodiment of the present invention. Detailed Embodiments

[0009] The present invention will be specifically described below with reference to the accompanying drawings of the specification. Figure 1 is a schematic flow diagram of the literature structured extraction method based on artificial intelligence provided by an embodiment of the present invention. The literature structured extraction method based on artificial intelligence will be introduced in detail below.

[0010] Step S110: Obtain a set of original literature data, which includes multiple literature units, and each literature unit is composed of text content and meta-information.

[0011] In this embodiment, in order to obtain the original literature data set, diversified data collection strategies can be adopted. For example, literature can be collected from multiple authoritative and widely used academic resource channels, including but not limited to well-known comprehensive academic databases, specific databases in professional fields, open access academic platforms, and internal knowledge bases of universities and research institutions. For different data sources, appropriate collection methods can be used. For example, for databases that support API interfaces, specific program code can be written to send requests to obtain literature data according to the parameters and format requirements specified in the API documentation. This method can efficiently and accurately obtain the required literature information, including text content and meta-information. For websites without API interfaces, web crawler technology can be used. The web crawler program will locate the position of the literature data in the web page according to the preset rules, parse the HTML or XML structure, and extract the text content and meta-information. During the collection process, in order to ensure the legality and compliance of the data, the usage terms of each data source and relevant laws and regulations will be followed.

[0012] When obtaining literature data, it usually involves various formats of literature, such as PDF, DOC, TXT, etc. Different formats of literature have different storage structures and encoding methods. For PDF format literature, since it usually contains complex layout information and possible encryption protection, professional PDF parsing libraries can be used to convert its content into a processable text form. For DOC format literature, the automated interfaces of office software or specialized document parsing tools can be used to extract the text content and meta-information. During the collection process, the literature can be preliminarily screened to eliminate those that are incomplete, damaged, or do not meet the requirements of the research field, ensuring that the literature units in the original literature data set have a certain quality and relevance.

[0013] The text content of each literature unit is the core information carrier of the literature, which may include parts such as research background, experimental methods, result analysis, and conclusions. The meta-information provides additional descriptions and identifications for the literature, such as author names, author affiliated institutions, publication dates, keywords, literature types (such as journal papers, conference papers, research reports, etc.). These meta-information helps to classify, retrieve, and analyze the literature in the follow-up. Through the above comprehensive and detailed collection and collation process, the literature data from different channels and different formats are integrated together to form an original literature data set containing multiple literature units.

[0014] Step S120: Preprocess the original literature data set to obtain a standardized literature data set, which contains text paragraphs and meta-information entries in a unified format.

[0015] In this embodiment, due to the diversity of sources and formats of the documents in the original document data set, there are problems such as inconsistent formats, redundant content, missing or scattered meta-information. Therefore, preprocessing is required to obtain a standardized document data set. The preprocessing process mainly includes steps such as format conversion, denoising, segmentation, meta-information field alignment, missing value filling, and association binding.

[0016] Step S121: Perform format conversion processing on each document unit in the original document data set, convert the text content from different sources into a unified plain text format, and obtain the converted text content.

[0017] In this embodiment, documents in different formats have significant differences in storage and presentation methods. To facilitate subsequent processing and analysis, they need to be converted into a unified plain text format. For PDF format documents, a dedicated PDF text extraction library is used. Its working principle is to parse the internal structure of the PDF file, identify the text objects therein, and convert them into plain text. During the parsing process, possible character encoding problems will be handled to ensure the accuracy of the extracted text content. For DOC format documents, the COM interface of office software or an open-source document parsing library can be used to extract their content as plain text. During the extraction process, the format settings in the document, such as font, font size, color, paragraph format, etc., can be removed, and only the text information is retained. For documents in other formats, corresponding conversion methods will also be adopted to ensure that the text content of all documents is converted into a unified plain text format. The converted text content will serve as the basis for subsequent processing, facilitating operations such as denoising and segmentation.

[0018] Step S122: Perform denoising processing on the converted text content, remove redundant symbols, duplicate paragraphs, and additional information irrelevant to the core content of the document, and obtain the denoised text content.

[0019] In this embodiment, the converted text content may contain a large amount of redundant information, such as non-text symbols, duplicate paragraphs, reference lists, appendix descriptions, copyright statements, etc. These information will interfere with subsequent analysis and processing, so denoising processing is required. The denoising processing includes the following sub-steps: Step S1221: Identify the symbol types in the converted text content, and filter out non-text symbols and hyperlink markers.

[0020] In this embodiment, regular expressions are used to identify and filter non-text symbols and hyperlink tags. Regular expressions are a powerful text matching tool. By defining specific patterns, various symbols and tags in the text can be accurately identified. For example, define a regular expression pattern to match common punctuation marks, special characters, and HTML tags and other non-text symbols. For hyperlink tags, a pattern starting with "http" or "https" can be defined for matching. When traversing the text content, once a matching symbol or tag is found, it is deleted from the text. To ensure the accuracy and integrity of the filtering, different types of documents can be tested and the regular expression patterns can be optimized to adapt to various possible situations.

[0021] Step S1222: Extract the repeated character sequences in the text content and calculate the occurrence frequency of the repeated character sequences in the continuous text.

[0022] In this embodiment, a sliding window method is adopted to extract the repeated character sequences in the text content. Set an appropriate window size, which will be adjusted according to the characteristics of the document and the processing requirements. The window slides from left to right in the text, moving one character position at a time, and intercepting the character sequence within the window. For each intercepted character sequence, it can be searched throughout the text to count the number of times it appears. Then, according to the total length of the text and the number of times the character sequence appears, its occurrence frequency is calculated. To improve the calculation efficiency, data structures such as hash tables can be used to store the already counted character sequences and their occurrence times to avoid repeated calculations.

[0023] Step S1223: Based on the occurrence frequency, perform redundancy judgment on the repeated character sequences, and delete the repeated paragraphs whose occurrence frequency exceeds the preset threshold.

[0024] In this embodiment, a preset occurrence frequency threshold is set, which is determined based on the statistical analysis of a large amount of document data and actual processing experience. For the calculated occurrence frequency of each repeated character sequence, it is compared with the preset threshold. If the occurrence frequency of a certain repeated character sequence exceeds the threshold, the paragraph where the character sequence is located is considered redundant and is deleted from the text. During the judgment process, the length of the character sequence and context information can be considered to avoid misdeleting important content. For example, if a short character sequence appears frequently in the text but it is a key expression of an important concept, the paragraph where it is located will not be deleted.

[0025] Step S1224: Detect the additional information area in the text content, and the additional information area includes a reference list, appendix instructions, and copyright statements.

[0026] In this embodiment, the additional information area is detected by analyzing the structure and features of the text. For a reference list, there is usually a specific title identifier, such as "References". Using regular expressions or string matching algorithms, these title identifiers are searched in the text to determine the starting position of the reference list. For appendix descriptions and copyright statements, corresponding iconic words or format features are also searched, such as "Appendix", "Copyright", etc. Once these iconic information are found, the scope of the additional information area can be determined.

[0027] Step S1225: extracting the starting position and the ending position of the additional information area, and cutting off the additional information located before the starting position and after the ending position in the text content.

[0028] In this embodiment, after determining the starting position and the ending position of the additional information area, the additional information located before the starting position and after the ending position in the text content is truncated by string truncation. For example, using the string slicing operation in the programming language, the core text content to be retained is extracted according to the starting position and the ending position, and the additional information is discarded. In the truncation process, the integrity and coherence of the retained text content can be ensured, and the loss of important information can be avoided.

[0029] Step S1226: performing space normalization processing on the truncated text content, replacing multiple consecutive space characters with a single space character, and generating denoised text content.

[0030] In this embodiment, a regular expression or a simple string replacement algorithm is used to normalize the space of the truncated text content. The regular expression can define a pattern that matches multiple consecutive space characters. When matching consecutive spaces are found in the text, they are replaced with a single space character. This processing can make the text content more standardized and neat, which is convenient for subsequent text analysis and processing.

[0031] Step S123: segment the denoised text content according to semantic delimiters to generate text paragraphs with logical coherence.

[0032] In this embodiment, the semantic delimiter is determined according to the semantic and logical structure of the text. Common semantic delimiters include punctuation marks such as full stops, exclamation marks, question marks, and some specific paragraph delimiter marks. By traversing the denoised text content, the positions of the semantic delimiters are identified. During the identification process, the context of the punctuation marks and the integrity of the sentences can be considered to avoid incorrect segmentation. For example, in some abbreviations or specific expressions, a full stop may not indicate the end of a sentence and requires special processing. According to the positions of the semantic delimiters, the text content is segmented into multiple paragraphs. After segmentation, each paragraph can be checked to ensure that the content within the paragraph is logically coherent, avoiding situations where paragraphs are too long or too short or logically chaotic. If a problem is found in a certain paragraph, it can be adjusted and merged according to the context information.

[0033] Step S124: Perform field alignment processing on the meta-information of the literature unit, and map the scattered meta-information entries to the preset standard meta-information fields.

[0034] In this embodiment, the meta-information of different literature units may be stored with different field names and formats. For the convenience of subsequent unified processing and analysis, field alignment processing is required. First, a set of preset standard meta-information fields are defined, and these fields include author, publication date, keywords, literature type, etc. For the meta-information of each literature unit, analyze its field name and content, and map it to the standard meta-information fields. For example, if the meta-information of a certain literature unit contains an "Author" field, and the standard meta-information field is "author", then map the value of the "Author" field to the "author" field. During the mapping process, the semantic similarity of the field names and the consistency of the content can be considered. For fields that cannot be directly mapped, manual intervention or classification and mapping can be performed through machine learning algorithms.

[0035] Step S125: Perform missing value filling processing on the meta-information after field alignment, and complete the missing meta-information entries based on the statistical rules of the meta-information of literature units of the same type.

[0036] In this embodiment, there may be missing values in the meta-information after field alignment. To ensure the integrity of the meta-information, missing value filling processing is required. First, classify the literature units according to the type of literature, such as journal papers, conference papers, research reports, etc. For each type of literature unit, count the value distribution and statistical rules of each field in its meta-information. For example, count the common value range and distribution frequency of the publication date in a certain type of literature unit. For the missing meta-information entries, fill them according to the statistical rules of this type of literature unit. If the publication date of a certain literature unit is missing, it can be estimated and filled according to the average year or common year range of the publication date of this type of literature unit. For some missing values that cannot be filled by statistical rules, further information retrieval or manual supplementation can be carried out.

[0037] Step S126: Perform an association binding process on the text paragraphs with logical coherence and the meta-information entries to generate a standardized literature data set containing a text paragraph set and a meta-information entry set.

[0038] In this embodiment, by establishing a unique identifier, the text paragraphs with logical coherence and the meta-information entries are associated and bound. Assign a unique identifier to each literature unit, and this identifier can be the literature number, DOI (Digital Object Identifier), etc. For the text paragraph set and the meta-information entry set, associate each text paragraph and the corresponding meta-information entry with this unique identifier. For example, store the text paragraph set and the meta-information entry set of a certain literature unit in a data structure, using the unique identifier as the key and the text paragraph set and the meta-information entry set as the values. Through the above association binding process, a standardized literature data set containing a text paragraph set and a meta-information entry set is formed.

[0039] Step S130: Invoke a pre-trained literature feature extraction model to process the standardized literature data set to obtain the content features and structural features of the literature unit.

[0040] In this embodiment, the pre-trained literature feature extraction model is trained based on a large amount of annotated literature data, and it can effectively extract the content features and structural features of the literature unit from the standardized literature data set. The processing process includes the following steps: Step S131: Input the text paragraphs in the standardized literature data set into the text encoding module of the literature feature extraction model, and generate paragraph semantic vectors through word embedding and context encoding operations.

[0041] In this embodiment, the text encoding module is an important part of the document feature extraction model, which is mainly responsible for converting text paragraphs into vector representations. First, a word embedding operation is performed to map each word in the text paragraph into a low-dimensional vector space. There are various methods for word embedding, such as Word2Vec, GloVe, etc. These methods can capture the semantic relationships between words through learning a large amount of text data, making words with similar semantics closer in the vector space. During the word embedding process, each word can be converted into a corresponding vector according to the word sequence in the text paragraph. Then, a context encoding operation is carried out, considering the context information of the word in the paragraph to further enrich the representation of the word vector. Context encoding can use models such as recurrent neural network (RNN), long short-term memory network (LSTM), or gated recurrent unit (GRU), or models based on the Transformer architecture, such as BERT, etc. These models can capture the long-distance dependencies between words and generate word vectors containing context information. Finally, the vectors of all words in the paragraph are aggregated to generate a paragraph semantic vector. The aggregation method can be average, maximum, or weighted summation, etc., which is selected according to specific tasks and model requirements.

[0042] Step S132: Input the meta-information entries in the standardized document data set into the meta-information encoding module of the document feature extraction model, and generate meta-information feature vectors through category feature conversion and numerical feature normalization operations.

[0043] In this embodiment, the meta-information entries include various types of information, such as category features (document type, author's affiliated institution, etc.) and numerical features (publication date, citation count, etc.). The meta-information encoding module will process different types of meta-information differently. For category features, methods such as one-hot encoding or label encoding are used to convert them into vector representations. One-hot encoding creates a binary vector for each category, with only the position corresponding to the category being 1 and the rest being 0. Label encoding assigns a unique integer label to each category. For numerical features, in order to eliminate the dimensional differences between different features, a normalization operation is performed. The normalization method can be min-max normalization or z-score normalization, etc. Min-max normalization maps the values of numerical features to a fixed interval, such as [0, 1]. Z-score normalization converts it into a standard normal distribution according to the mean and standard deviation of the feature. Through category feature conversion and numerical feature normalization operations, the meta-information entries are converted into meta-information feature vectors.

[0044] Step S133: Perform cross-dimensional correlation analysis on the paragraph semantic vector and the meta-information feature vector to generate an association feature vector containing the association relationship between the text and the meta-information.

[0045] In this embodiment, the paragraph semantic vector and the meta-information feature vector respectively represent the text content and meta-information of the document, but there may be a potential association relationship between them. To mine this association relationship, cross-dimensional correlation analysis is performed. The attention mechanism can be used to calculate the attention weights between the paragraph semantic vector and the meta-information feature vector. The attention weights represent the importance degree of each vector element in the association relationship. By performing weighted summation on the attention weights, the paragraph semantic vector and the meta-information feature vector are fused to generate an association feature vector containing the association relationship between the text and the meta-information, which can better reflect the overall features of the document.

[0046] Step S134: Input the association feature vector into the content feature extraction layer of the document feature extraction model, and extract the content features reflecting the core content of the document through feature screening and dimensionality reduction operations.

[0047] In this embodiment, the main task of the content feature extraction layer is to extract the features reflecting the core content of the document from the association feature vector. First, a feature screening operation is performed. By calculating the importance scores of each element in the association feature vector, the elements with higher importance scores are selected as candidate features. The calculation of importance scores can use feature selection algorithms such as chi-square test and information gain. Then, a dimensionality reduction operation is performed to compress the high-dimensional candidate feature vector to a lower dimension to reduce the computational amount and storage cost. Methods for dimensionality reduction include principal component analysis (PCA), linear discriminant analysis (LDA), etc. Principal component analysis projects the data onto the principal component directions found, achieving dimensionality reduction. Through feature screening and dimensionality reduction operations, the content features reflecting the core content of the document are extracted, which can thus summarize the main themes and key information of the document.

[0048] Step S135: Input the paragraph semantic vector into the structure feature extraction layer of the document feature extraction model, and extract the structure features reflecting the organization mode of the document through paragraph order modeling and logical relationship recognition operations.

[0049] In this embodiment, the structure feature extraction layer is mainly responsible for extracting the structure features reflecting the organization mode of the document from the paragraph semantic vector. The processing process includes the following sub-steps: Step S1351: Construct an order index of the paragraph semantic vector, which is used to represent the original arrangement order of the text paragraphs in the document.

[0050] In this embodiment, in order to preserve the original arrangement order information of text paragraphs, a sequential index of paragraph semantic vectors is constructed. In the standardized literature data set, each text paragraph has a specific position in the literature. By recording the paragraph position corresponding to each paragraph semantic vector, a sequential index is constructed. The sequential index can be a simple integer sequence, where each integer corresponds to a paragraph semantic vector, indicating the order of the paragraph semantic vector in the literature.

[0051] Step S1352: Perform joint encoding processing on the paragraph semantic vectors and the sequential index to generate a temporal feature vector containing order information.

[0052] In this embodiment, in order to incorporate the order information of paragraphs into the paragraph semantic vectors, joint encoding processing is performed. A method based on a sequence model, such as LSTM or GRU, can be used to take the paragraph semantic vectors and the sequential index as inputs for encoding. During the encoding process, the model will consider the order information of paragraphs and learn the temporal relationship between paragraphs. Through joint encoding processing, a temporal feature vector containing order information is generated, and this vector can better reflect the structural features of the literature.

[0053] Step S1353: Extract the cosine similarity values between adjacent paragraph semantic vectors and calculate the change gradient of the cosine similarity values in consecutive paragraphs.

[0054] In this embodiment, the cosine similarity is a commonly used metric to measure the similarity between two vectors. For adjacent paragraph semantic vectors, calculate the cosine similarity values between them. The calculation of the cosine similarity value is based on the inner product of the vectors and the vector norms, which reflects the similarity degree of the two vectors in terms of direction. After calculating the cosine similarity values of adjacent paragraph semantic vectors, further calculate the change gradient of these cosine similarity values in consecutive paragraphs. The change gradient represents the speed and direction of the change of the cosine similarity value in the paragraph sequence. By analyzing the change gradient, the change situation of semantic connection between paragraphs can be captured.

[0055] Specifically, in order to calculate the change gradient of the cosine similarity value, the sequence of cosine similarity values of adjacent paragraphs can be traversed. For each cosine similarity value in the sequence of cosine similarity values, calculate the difference between it and the previous cosine similarity value, and this difference is the change gradient at that position. The calculation of the change gradient can help identify the mutation points of semantic connection between paragraphs, and these mutation points may correspond to chapter transitions or theme switches in the literature.

[0056] Step S1354: Identify the logical transition points between paragraphs based on the change gradient, and these logical transition points are used to demarcate the chapter boundary information in the literature.

[0057] In this embodiment, the logical transition points between paragraphs are identified by analyzing the change gradient. A change gradient threshold is set. When the change gradient at a certain position exceeds this change gradient threshold, this position is considered a logical transition point. The setting of this change gradient threshold is determined based on a large amount of literature data statistics and practical application experience, which can ensure the accuracy of chapter division while avoiding the problems of over-division or under-division.

[0058] The identification of logical transition points is very important for dividing the chapter boundary information in the literature. In the literature, different chapters usually have different themes and content structures, and the transition between chapters is often accompanied by a large change in semantics, which will be reflected in the change gradient of the cosine similarity value. By identifying logical transition points, the literature can be divided into different chapters.

[0059] Step S1355: Perform attention mechanism processing on the time series feature vector to generate an attention weight distribution reflecting the importance of paragraphs.

[0060] In this embodiment, the attention mechanism is a technology that can automatically focus on the important parts in the sequence. The time series feature vector is input into the attention mechanism module, and this module will calculate the attention weights of each element of the time series feature vector. The attention weights represent the importance of each paragraph to the overall literature structure and content.

[0061] The calculation process of the attention mechanism usually includes calculating the similarity between the query vector, key vector, and value vector, and then converting the similarity into a probability distribution through the softmax function. This probability distribution is the attention weight. By performing attention mechanism processing on the time series feature vector, important paragraphs can be highlighted and unimportant paragraphs can be suppressed, thus generating an attention weight distribution reflecting the importance of paragraphs.

[0062] Step S1356: Construct a literature structure tree based on the attention weight distribution and chapter boundary information. The node features of this literature structure tree are used as structure features reflecting the organization method of the literature.

[0063] In this embodiment, the literature structure tree is constructed by combining the attention weight distribution and chapter boundary information. First, the literature is divided into different chapters according to the chapter boundary information, and each chapter serves as a node of the literature structure tree. Then, corresponding weights are assigned to each node according to the attention weight distribution, and the weights represent the importance of this chapter in the literature.

[0064] During the process of constructing the literature structure tree, the hierarchical relationship between chapters is also considered. For chapters with a subordinate relationship, they are constructed into a parent-child node relationship to form a tree structure. The node features of the literature structure tree include the weight of the node, the semantic features of the chapter represented by the node, etc. These node features can be used as structural features reflecting the literature organization method for subsequent structured parsing and analysis.

[0065] Step S136: Standardize the content features and structural features to generate a set of literature features with a unified dimensional representation.

[0066] In this embodiment, since the content features and structural features are extracted by different methods and modules, in order to facilitate subsequent processing and analysis, it is necessary to standardize the content features and structural features.

[0067] Through the standardization process, a set of literature features with a unified dimensional representation is generated. This set of literature features contains the content features and structural features of the literature and can more comprehensively reflect the overall features of the literature.

[0068] Step S140: Perform structured parsing processing based on the content features and structural features to generate a set of structured elements of the literature unit. This set of structured elements includes topic elements, logical relationship elements, and key information elements.

[0069] In this embodiment, the purpose of the structured parsing processing is to extract valuable structured information from the content features and structural features of the literature to form a set of structured elements. It specifically includes the following steps: Step S141: Input the content features into the topic recognition module, and generate topic elements through operations such as topic word extraction and topic distribution probability calculation. The topic elements include core topic words and descriptions of the topic coverage range.

[0070] In this embodiment, the topic recognition module receives the content features as input and generates topic elements through a series of operations. First, topic word extraction is performed. Topic word extraction can adopt a statistical method, such as the term frequency-inverse document frequency (TF-IDF) algorithm. This algorithm determines the importance of a word by calculating the frequency of the word in the document and the inverse frequency in the entire document set, and the words with higher importance are selected as topic words. It can also adopt a deep learning method, such as using a pre-trained language model for topic word extraction.

[0071] After extracting the topic words, further calculate the topic distribution probability. The topic distribution probability represents the probability of each topic appearing in the literature, which can be calculated by clustering the content features or using a probabilistic topic model. Through the operations of topic word extraction and topic distribution probability calculation, topic elements containing core topic words and descriptions of the topic coverage are generated. The core topic words summarize the main topics of the literature, and the topic coverage description explains the distribution of each topic in the literature.

[0072] Step S142: Input the structural features into the logical parsing module, and generate logical relationship elements through paragraph association analysis and chapter level identification operations. The logical relationship elements include the derivation relationship between paragraphs and the subordinate relationship between chapters.

[0073] In this embodiment, the logical parsing module receives the structural features as input and generates logical relationship elements through paragraph association analysis and chapter level identification operations. The specific steps are as follows: Step S1421: Extract the paragraph association features in the structural features. The paragraph association features include the semantic cohesion strength and content continuity index of adjacent paragraphs.

[0074] In this embodiment, the structural features contain the association information between paragraphs, and the paragraph association features are extracted by analyzing this information. The paragraph association features can be obtained by calculating the semantic similarity between adjacent paragraphs, the citation relationship between paragraphs, etc. The semantic cohesion strength reflects the closeness of the semantics between adjacent paragraphs, and the content continuity index represents the coherence of the content between adjacent paragraphs. By extracting the paragraph association features, the logical relationship between paragraphs can be better understood.

[0075] Step S1422: Construct a paragraph relationship graph based on the paragraph association features. The edge weights of the paragraph relationship graph represent the degree of logical association between paragraphs.

[0076] In this embodiment, a paragraph relationship graph is constructed according to the extracted paragraph association features. The paragraph relationship graph is a graph structure, where the nodes in the graph represent paragraphs and the edges represent the logical association between paragraphs. The weights of the edges are determined by factors such as the semantic cohesion strength and content continuity index in the paragraph association features. A large edge weight indicates a high degree of logical association between paragraphs.

[0077] The method of graph theory can be used to construct the paragraph relationship graph. For example, an adjacency matrix can be used to represent the structure of the graph. Through the paragraph relationship graph, the logical relationship between paragraphs can be intuitively displayed.

[0078] Step S1423: Perform community discovery processing on the paragraph relationship graph to identify the paragraph communities with significant association relationships as the literature chapters.

[0079] In this embodiment, community detection is a technique for identifying sets of closely connected nodes in a graph structure. Community detection is performed on the paragraph relationship graph, and by calculating the connection strength between nodes and the cohesion within communities, paragraph communities with significant correlation relationships are identified. These paragraph communities can be regarded as chapters in the literature because the paragraphs within them have strong logical connections, while the logical connections with paragraphs in other communities are relatively weak.

[0080] Multiple algorithms can be used for community detection, such as the Louvain algorithm, spectral clustering algorithm, etc. These algorithms can divide paragraphs into different communities according to the structure and edge weights of the paragraph relationship graph, thereby identifying the chapters of the literature.

[0081] Step S1424: Extract the chapter-level features in the structural features, and the chapter-level features include the font size information and indentation depth information of the chapter titles.

[0082] In this embodiment, the structural features also include the hierarchical information of the chapters, and by extracting this information, the chapter-level features can be obtained. The chapter-level features can be obtained by analyzing the text structure and format information of the literature, such as the font size and indentation depth of the chapter titles. The font size and indentation depth are usually related to the hierarchy of the chapters. Larger fonts and smaller indentation depths usually indicate higher-level chapters.

[0083] By extracting the chapter-level features, the chapter structure and hierarchical relationship of the literature can be understood, providing a basis for determining the subordinate relationship between chapters.

[0084] Step S1425: Determine the subordinate relationship between chapters based on the chapter-level features, and the subordinate relationship includes the corresponding relationship between the parent chapter and the child chapter.

[0085] In this embodiment, the subordinate relationship between chapters is determined according to the extracted chapter-level features. By comparing the font size and indentation depth of the chapter titles, the hierarchical relationship between chapters can be judged. If the title font of a chapter is smaller and the indentation depth is larger, then it is very likely to be a sub-chapter of another chapter.

[0086] Determining the subordinate relationship between chapters can construct a chapter-level tree. The nodes in the tree represent chapters, and the edges represent the subordinate relationship between chapters. Through the chapter-level tree, the chapter structure and hierarchical relationship of the literature can be clearly displayed.

[0087] Step S1426: Perform matching verification processing on the paragraph communities and chapter-level features in the paragraph relationship graph to generate logical relationship elements including the derivation relationship between paragraphs and the subordinate relationship between chapters.

[0088] In this embodiment, in order to ensure the consistency of the paragraph community and the chapter-level features, matching verification processing is performed on them. The paragraph communities identified in the paragraph relationship diagram are compared with the chapters determined according to the chapter-level features to check whether they correspond and are consistent.

[0089] If the paragraph community and the chapter-level features are inconsistent, adjustment and correction are required. For example, if a paragraph community contains chapter contents at multiple different levels, it is necessary to further analyze the logical relationships between the paragraphs and re-divide the paragraph community. Through the matching verification processing, logical relationship elements including the derivation relationships between paragraphs and the subordination relationships between chapters are generated, and these logical relationship elements can accurately reflect the logical structure of the literature.

[0090] Step S143: Input the content features and structure features into the key information extraction module, and generate key information elements through entity recognition and key sentence screening operations. The key information elements include the core research conclusions and key experimental conditions.

[0091] In this embodiment, the key information extraction module receives the content features and structure features as inputs and generates key information elements through entity recognition and key sentence screening operations. Entity recognition can adopt the named entity recognition (NER) technology, which can identify entities in the text, such as person names, place names, organization names, experimental equipment names, etc. Key sentence screening can adopt machine learning-based methods, such as support vector machines (SVM), decision trees, etc., or can also adopt deep learning-based methods, such as long short-term memory networks (LSTM), convolutional neural networks (CNN), etc.

[0092] When performing key sentence screening, the importance of the sentence and its relevance to the theme can be considered. Sentences with higher importance and stronger relevance to the theme will be selected as key sentences, and the information contained in the key sentences can be used as key information elements, such as core research conclusions and key experimental conditions, etc.

[0093] Step S144: Perform conflict detection processing on the theme elements, logical relationship elements, and key information elements to identify the semantic contradiction points between different elements.

[0094] In this embodiment, since the theme elements, logical relationship elements, and key information elements are generated through different methods and modules, there may be semantic contradictions between them. In order to ensure the consistency of the structured element set, conflict detection processing is required.

[0095] Conflict detection processing can be achieved by comparing the semantic information between different elements. For example, if the theme mentioned in the theme element is inconsistent with the core research conclusion mentioned in the key information element, or the chapter relationship described in the logical relationship element does not match the distribution position of the key experimental conditions in the key information element, it is considered that there is a semantic contradiction point.

[0096] Step S145: Based on the semantic contradiction points, correct the theme elements, logical relationship elements, and key information elements to generate a set of structured elements that pass the consistency check.

[0097] In this embodiment, after identifying the semantic contradiction points, it is necessary to correct the theme elements, logical relationship elements, and key information elements. The correction process can be adjusted according to the specific contradiction situation. For example, if there is a contradiction between the theme elements and the key information elements, the results of theme word extraction and key information extraction can be re-evaluated, and the content of the theme elements and key information elements can be adjusted.

[0098] If there is a contradiction between the logical relationship elements and the key information elements, the logical relationship between paragraphs and the subordinate relationship between chapters can be re-analyzed, and the content of the logical relationship elements can be adjusted. Through the correction process, it is ensured that there is no semantic contradiction between the theme elements, logical relationship elements, and key information elements in the set of structured elements, and a set of structured elements that pass the consistency check is generated.

[0099] Step S150: Integrate and optimize the set of structured elements to generate a final structured literature result that includes the element association relationship.

[0100] In this embodiment, the purpose of the integration and optimization process is to associate and optimize each element in the set of structured elements to form a final structured literature result that includes the element association relationship. Specifically, it includes the following steps: Step S151: Construct a feature index for each element in the set of structured elements, which is used to record the unique identifiers of the theme elements, logical relationship elements, and key information elements.

[0101] In this embodiment, in order to facilitate the management and association of each element in the set of structured elements, a feature index for each element is constructed. The feature index can adopt data structures such as hash tables, and store the unique identifiers of the theme elements, logical relationship elements, and key information elements as keys, and the specific content of the elements as values.

[0102] The unique identifier can be the name, number, etc. of the element. Through the feature index, each element in the set of structured elements can be quickly located and accessed.

[0103] Step S152: Perform an association analysis process on the feature index to identify the semantic mapping relationship between the theme elements and the key information elements, and this semantic mapping relationship includes the corresponding relationship between the theme words and the key experimental conditions.

[0104] In this embodiment, by performing correlation analysis on the feature index, the semantic mapping relationship between the theme elements and the key information elements is identified. The corresponding relationship between them can be found by comparing the core theme words in the theme elements and the key experimental conditions and other information in the key information elements.

[0105] For example, if a certain theme word in the theme element is semantically related to a certain key experimental condition in the key information element, then the mapping relationship between them can be established. The semantic mapping relationship can help to understand the internal connection between the theme and the key information of the literature.

[0106] Step S153: Identify the structural dependence relationship between the logical relationship elements and the key information elements, where the structural dependence relationship includes the chapter subordination relationship and the distribution position relationship of the core research conclusions.

[0107] In this embodiment, the structural dependence relationship between the logical relationship elements and the key information elements is analyzed. By observing the chapter subordination relationship in the logical relationship elements and the distribution position of the core research conclusions in the key information elements, the association between them is found.

[0108] For example, if a certain core research conclusion always appears in the sub-chapter under a certain parent chapter, then the chapter subordination relationship and the distribution position relationship of the core research conclusions can be established. The structural dependence relationship can reflect the logical structure of the literature and the distribution law of the key information.

[0109] Step S154: Construct an element association graph based on the semantic mapping relationship and the structural dependence relationship. The nodes of the element association graph are structured elements, and the edges are the association relationships between the elements.

[0110] In this embodiment, an element association graph is constructed according to the identified semantic mapping relationship and structural dependence relationship. The element association graph is a graph structure. The nodes in the graph represent structured elements, such as theme elements, logical relationship elements, and key information elements, and the edges represent the association relationships between the elements, such as semantic mapping relationships and structural dependence relationships.

[0111] When constructing the element association graph, it is also necessary to set a weight attribute for the edges. The weight attribute represents the strength of the association between the elements. The calculation of the weight attribute can be determined according to the importance and stability of the association relationship. For example, the weight of the semantic mapping relationship can be determined by calculating the co-occurrence frequency of the theme word and the key information element, and the weight of the structural dependence relationship can be determined by the appearance consistency of the logical relationship element in different literature units.

[0112] Step S1541: Extract the association strength value in the semantic mapping relationship, which is calculated by the co-occurrence frequency of the theme word and the key information element.

[0113] In this embodiment, in order to determine the association strength of the semantic mapping relationship, the association strength value in the semantic mapping relationship is extracted. The association strength value can be obtained by calculating the co-occurrence frequency of the subject word and the key information element. The co-occurrence frequency represents the ratio of the number of times the subject word and the key information element appear simultaneously in the literature to the number of times they each appear.

[0114] By counting the co-occurrence of the subject word and the key information element in the literature, the association strength value can be calculated. The higher the association strength value, the closer the association between the subject word and the key information element.

[0115] Step S1542: Extract the association stability value in the structural dependency relationship. This association stability value is calculated by the appearance consistency of the logical relationship elements in different literature units.

[0116] In this embodiment, in order to determine the association stability of the structural dependency relationship, the association stability value in the structural dependency relationship is extracted. The association stability value can be obtained by calculating the appearance consistency of the logical relationship elements in different literature units.

[0117] Specifically, multiple literature units can be analyzed to count the occurrences of the logical relationship elements in these literature units. If the appearance patterns of the logical relationship elements in different literature units are relatively consistent, then the association stability value is higher; otherwise, it is lower. The association stability value can reflect the reliability of the structural dependency relationship.

[0118] Step S1543: Set a weight attribute for the edges of the element association graph. This weight attribute is obtained by weighted summation of the association strength value and the association stability value.

[0119] In this embodiment, in order to comprehensively consider the association strength of the semantic mapping relationship and the association stability of the structural dependency relationship, a weight attribute is set for the edges of the element association graph. The weight attribute can be obtained by weighted summation of the association strength value and the association stability value.

[0120] The weights for weighted summation can be adjusted according to specific application scenarios and requirements. For example, if more emphasis is placed on the association strength of the semantic mapping relationship, the weight of the association strength value can be appropriately increased; if more emphasis is placed on the association stability of the structural dependency relationship, the weight of the association stability value can be appropriately increased.

[0121] Step S1544: Perform a merging process on the duplicate elements in the structured element set, and merge the subject elements, logical relationship elements, or key information elements with the same semantics into a single node.

[0122] In this embodiment, to simplify the structure of the element association graph, duplicate elements in the structured element set are merged. If there are theme elements, logical relationship elements, or key information elements with the same semantics, they are merged into a single node.

[0123] During the merging process, all attribute information of the elements can be retained and integrated into the merged single node. By merging duplicate elements, the number of nodes in the element association graph can be reduced, improving the readability and processing efficiency of the graph.

[0124] Step S1545: Perform attribute supplementation processing on the merged element nodes, and integrate element attribute information from different sources into the same node.

[0125] In this embodiment, after merging the element nodes, there may be a situation where the element attribute information is scattered. To make the attribute information of the element nodes more complete, attribute supplementation processing is performed on the merged element nodes. Integrate element attribute information from different sources into the same node. For example, if a certain theme element obtains different attribute descriptions in different analysis processes, such as determining the coverage of the theme in the theme recognition module and discovering the associated attributes with other elements in the conflict detection process, then these attribute information from different sources need to be merged into the node corresponding to the theme element.

[0126] When performing attribute supplementation, the compatibility and consistency of the attributes need to be considered. For the same type of attribute, if there are multiple values, they need to be merged or the most appropriate value is selected. For example, if both sources provide descriptions of the theme coverage, but the expressions are slightly different, they can be merged into a more accurate and complete description through semantic analysis and other methods. For different types of attributes, they are directly added to the attribute list of the node.

[0127] Step S1546: Generate a complete element association graph based on the weight attribute and the attribute information of the element nodes. This element association graph is used to represent the multi-dimensional association relationship between structured elements.

[0128] In this embodiment, after completing the weight setting of the edges, the merging of the element nodes, and the attribute supplementation, a complete element association graph is generated based on this information. The element association graph is a graph structure that comprehensively represents the multi-dimensional association relationship between structured elements. It not only includes the semantic mapping relationship and structural dependency relationship between the elements, but also reflects the strength and stability of the association through the weight of the edges, and details the characteristics of each element through the attribute information of the nodes.

[0129] The process of generating the element association graph includes determining the positions of the nodes and the connection methods of the edges. The positions of the nodes can be arranged according to the degree of association tightness between the elements. Nodes with a closer association can be placed closer to visually display the relationship between them. The connection methods of the edges are determined according to the previously identified semantic mapping relationships and structural dependency relationships to ensure that the edges in the graph accurately represent the associations between the elements.

[0130] Step S155: Perform a topological sorting process on the element association graph to generate a structured arrangement scheme reflecting the importance order of the elements.

[0131] In this embodiment, in order to further clarify the importance order between the structured elements, a topological sorting process is performed on the element association graph. Topological sorting is an algorithm for sorting a directed acyclic graph, which can arrange the nodes into a linear sequence according to the dependency relationships between the nodes in the graph, such that for any directed edge, the starting node always appears before the ending node in the sequence.

[0132] In the element association graph, the dependency relationships between the nodes reflect the logical order and importance between the elements. Through topological sorting, a structured arrangement scheme reflecting the importance order of the elements can be generated. During the sorting process, the weights of the edges can be considered. Nodes connected by edges with larger weights will be more prioritized in the sorting because this indicates that their association is closer and their importance is higher.

[0133] Step S156: Reorganize the structured element set according to the structured arrangement scheme to generate a final structured document result containing the element association relationships.

[0134] In this embodiment, the structured element set is reorganized according to the generated structured arrangement scheme. The structured elements are rearranged in the order of the arrangement scheme, and combined with the association relationships in the element association graph, the association information between the elements is also incorporated into the final result.

[0135] During the reorganization process, the detailed information and attributes of each element can be retained, and at the same time, their association relationships are presented in a clear and readable manner. For example, the theme elements, logical relationship elements, and key information elements can be listed in order of importance, and other elements associated with each element can be marked behind it. The final structured document result generated in this way can comprehensively and accurately reflect the theme, logical structure, and key information of the document, as well as the relationships between them, facilitating subsequent applications such as document analysis, knowledge mining, and information retrieval.

[0136] Furthermore, the pre-trained document feature extraction model is constructed through the following steps: In this embodiment, constructing a pre-trained literature feature extraction model is one of the key steps in the entire literature structured extraction method. Through learning a large amount of labeled literature data, the model can automatically extract valuable content features and structural features from the literature. The construction process includes the following steps: Step S211: Obtain a training data set containing labeled literature, where the labeled literature includes manually labeled content feature tags and structural feature tags.

[0137] In this embodiment, in order to train the literature feature extraction model, it is necessary to obtain a training data set containing labeled literature. These labeled literatures have been carefully analyzed and labeled manually, and include content feature tags and structural feature tags. Content feature tags can be the core themes, key concepts, etc. of the literature, and structural feature tags can be the chapter structures of the literature, the logical relationships between paragraphs, etc.

[0138] The way to obtain the training data set can be to collect literatures from multiple academic resource platforms and organize professional annotators for annotation. During the annotation process, detailed annotation specifications and guidelines need to be formulated to ensure the accuracy and consistency of annotation. For example, for content feature tags, clearly stipulate how to determine the core theme and key concepts; for structural feature tags, stipulate how to divide chapters and identify the logical relationships between paragraphs.

[0139] Step S212: Input the training data set into the initial feature extraction model, and generate predicted content features and predicted structural features through forward propagation operations.

[0140] In this embodiment, the obtained training data set is input into the initial feature extraction model. The initial feature extraction model is a pre-designed neural network model, which includes multiple modules and layers, such as a text encoding module, a meta-information encoding module, a content feature extraction layer, a structural feature extraction layer, etc.

[0141] After inputting the training data, the model will perform forward propagation operations. Forward propagation refers to the process in which data starts from the input layer, passes through the processing of each module and layer in turn, and finally generates predicted content features and predicted structural features. In each module and layer, the data will undergo a series of calculations and transformations, such as word embedding, context encoding, feature screening, dimensionality compression, etc. These operations are all to extract meaningful features from the input literature data.

[0142] Step S213: Calculate the first loss value between the predicted content features and the content feature tags, and this first loss value is calculated through a cross-entropy loss function.

[0143] In this embodiment, in order to measure the difference between the predicted content features generated by the model and the content feature labels manually annotated, the first loss value is calculated. The cross-entropy loss function is a commonly used loss function, which can measure the difference between two probability distributions. In this model, the predicted content features and the content feature labels can be regarded as two probability distributions, and their difference is calculated through the cross-entropy loss function.

[0144] Specifically, the cross-entropy loss function compares the predicted content features and the content feature labels. For each feature element, it calculates the difference between its predicted value and the true value, and then sums up these differences after weighting. The weights for the weighted sum can be set according to the importance of the features. Feature elements with higher importance will account for a larger proportion in the loss calculation.

[0145] Step S214: Calculate the second loss value between the predicted structural features and the structural feature labels. This second loss value is calculated through the mean squared error loss function.

[0146] In this embodiment, similarly, in order to measure the difference between the predicted structural features generated by the model and the structural feature labels manually annotated, the second loss value is calculated. The mean squared error loss function is a function used to measure the difference between two numerical sequences. It calculates the square of the difference between each corresponding element, and then sums up these squared values and takes the average.

[0147] When calculating the second loss value between the predicted structural features and the structural feature labels, the mean squared error loss function compares the predicted value and the true value of each structural feature element, calculates the square of the difference between them, and then sums up and averages the squared differences of all elements. The mean squared error loss function can effectively reflect the overall difference between the predicted structural features and the true structural features.

[0148] Step S215: Combine the first loss value and the second loss value to generate a joint loss function, and optimize the network parameters of the initial feature extraction model through backpropagation operations.

[0149] In this embodiment, in order to comprehensively consider the accuracy of the predicted content features and the predicted structural features, the first loss value and the second loss value are combined to generate a joint loss function. The joint loss function can combine the first loss value and the second loss value through weighted summation, and the weights can be adjusted according to specific application requirements. For example, if more attention is paid to the accuracy of the content features, the weight of the first loss value can be appropriately increased; if more attention is paid to the accuracy of the structural features, the weight of the second loss value can be appropriately increased.

[0150] After obtaining the combined loss function, the backpropagation algorithm is used to optimize the network parameters of the initial feature extraction model. The backpropagation algorithm is an optimization algorithm based on gradient descent. It calculates the gradient of the combined loss function with respect to each parameter in the model, and then adjusts the values of the parameters according to the direction and magnitude of the gradient, so that the value of the combined loss function gradually decreases. In each backpropagation process, the model updates the parameters to improve the fitting ability for the training data.

[0151] Step S216: When the convergence value of the combined loss function satisfies a preset threshold, stop the training and determine the corresponding initial feature extraction model as the pre-trained literature feature extraction model.

[0152] In this embodiment, during the training process, the convergence situation of the combined loss function can be continuously monitored. A convergence threshold is preset. When the value of the combined loss function drops below this threshold, it is considered that the model has converged and the training process can be stopped. At this time, the current initial feature extraction model is determined as the pre-trained literature feature extraction model.

[0153] In actual training, there may be situations where the combined loss function no longer decreases or decreases very slowly within a certain period of time, which can also be used as a basis for judging model convergence. By continuously adjusting the model parameters and training strategies, ensure that the model can accurately extract content features and structural features from the literature when reaching the convergence condition.

[0154] During the entire literature structure extraction process, attention also needs to be paid to data privacy protection and anti-leakage issues. In the data collection stage, for literature involving privacy-sensitive data, such as literature containing personal identity information, business secrets, etc., strict screening and processing are required. Data desensitization technology can be used to replace, delete, or encrypt sensitive information to ensure the security of data during collection and storage.

[0155] During the model training and use process, corresponding privacy protection measures should also be taken. For example, the training data of the model can be encrypted and stored, and only authorized personnel can access it. In the model inference stage, privacy protection processing should also be carried out on the input literature data to avoid leaking sensitive information. At the same time, a perfect security management system should be established to regularly conduct security checks and vulnerability repairs on the data and the model to ensure the security and reliability of the entire literature structure extraction system.

[0156] Figure 2 FIG. shows a schematic diagram of exemplary hardware and software components of an artificial intelligence-based literature structure extraction system 100 that can implement the ideas of the present application provided by some embodiments of the present application. For example, the processor 120 can be used on the artificial intelligence-based literature structure extraction system 100 and is used to execute the functions in the present application.

[0157] The artificial intelligence-based literature structured extraction system 100 can be a general-purpose server or a special-purpose server, both of which can be used to implement the artificial intelligence-based literature structured extraction method of this application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0158] For example, the artificial intelligence-based literature structured extraction system 100 can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROM, or RAM, or any combination thereof. Exemplarily, the artificial intelligence-based literature structured extraction system 100 can also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of this application can be implemented according to these program instructions. The artificial intelligence-based literature structured extraction system 100 also includes an input / output I / O interface 150 between the computer and other input / output devices.

[0159] For ease of explanation, only one processor is described in the artificial intelligence-based literature structured extraction system 100. However, it should be noted that the artificial intelligence-based literature structured extraction system 100 in this application can also include multiple processors. Therefore, the steps performed by one processor described in this application can also be jointly executed or separately executed by multiple processors. For example, if the processor of the artificial intelligence-based literature structured extraction system 100 executes step A and step B, it should be understood that step A and step B can also be jointly executed by two different processors or separately executed in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.

[0160] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above-mentioned artificial intelligence-based literature structured extraction method is implemented.

[0161] It should be noted that, in order to simplify the expression of the disclosure of the present invention and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are merged into one embodiment, drawing, or description thereof.

Claims

1. A method for structured extraction of literature based on artificial intelligence, characterized in that, The method includes: Obtain a set of original literature data, where the set of original literature data contains multiple literature units, and each literature unit consists of text content and meta-information; Preprocess the set of original literature data to obtain a set of standardized literature data, where the set of standardized literature data contains text paragraphs and meta-information entries in a unified format; Call a pre-trained literature feature extraction model to process the set of standardized literature data to obtain the content features and structural features of the literature units; Perform structured parsing processing based on the content features and the structural features to generate a set of structured elements of the literature units, where the set of structured elements contains topic elements, logical relationship elements, and key information elements; Perform integration and optimization processing on the set of structured elements to generate a final structured literature result containing element association relationships.

2. The method for document structured extraction based on artificial intelligence according to claim 1, wherein The preprocessing of the set of original literature data to obtain a set of standardized literature data, where the set of standardized literature data contains text paragraphs and meta-information entries in a unified format, includes: Perform format conversion processing on each literature unit in the set of original literature data to convert text content from different sources into a unified plain text format to obtain the converted text content; Denoise the converted text content to remove redundant symbols, duplicate paragraphs, and additional information unrelated to the core content of the literature to obtain the denoised text content; Segment the denoised text content according to semantic delimiters to generate text paragraphs with logical coherence; Perform field alignment processing on the meta-information of the literature unit to map the scattered meta-information entries into a preset standard meta-information field; Perform missing value filling processing on the meta-information after field alignment, and complete the missing meta-information entries based on the meta-information statistical rules of literature units of the same type; Perform association binding processing on the text paragraphs with logical coherence and the meta-information entries to generate a set of standardized literature data containing a set of text paragraphs and a set of meta-information entries.

3. The method for literature structured extraction based on artificial intelligence according to claim 2, wherein The denoising processing of the converted text content to remove redundant symbols, duplicate paragraphs, and additional information unrelated to the core content of the literature to obtain the denoised text content includes: Identify the symbol types in the converted text content and filter out non-text symbols and hyperlink markers; Extract the repeated character sequences in the text content and calculate the occurrence frequency of the repeated character sequences in the continuous text; Perform redundancy judgment on the repeated character sequences based on the occurrence frequency, and delete the duplicate paragraphs with an occurrence frequency exceeding a preset threshold; Detect the additional information areas in the text content, where the additional information areas include reference lists, appendix descriptions, and copyright statements; Extract the start position and end position of the additional information area, and truncate the additional information in the text content before the start position and after the end position; Perform space normalization processing on the truncated text content to replace multiple consecutive space characters with a single space character to generate the denoised text content.

4. The method for literature structured extraction based on artificial intelligence according to claim 1, wherein Invoking the pre-trained literature feature extraction model to process the standardized literature data set to obtain the content features and structural features of the literature unit, including: Inputting the text paragraphs in the standardized literature data set into the text encoding module of the literature feature extraction model, and generating paragraph semantic vectors through word embedding and context encoding operations; Inputting the meta-information entries in the standardized literature data set into the meta-information encoding module of the literature feature extraction model, and generating meta-information feature vectors through category feature conversion and numerical feature normalization operations; Performing cross-dimensional correlation analysis on the paragraph semantic vectors and the meta-information feature vectors to generate correlation feature vectors containing the correlation relationship between text and meta-information; Inputting the correlation feature vectors into the content feature extraction layer of the literature feature extraction model, and extracting content features reflecting the core content of the literature through feature screening and dimension compression operations; Inputting the paragraph semantic vectors into the structural feature extraction layer of the literature feature extraction model, and extracting structural features reflecting the organization mode of the literature through paragraph order modeling and logical relationship recognition operations; Performing standardization processing on the content features and the structural features to generate a literature feature set with a unified dimension representation.

5. The method for literature structured extraction based on artificial intelligence according to claim 4, wherein The step of inputting the paragraph semantic vectors into the structural feature extraction layer of the literature feature extraction model, and extracting structural features reflecting the organization mode of the literature through paragraph order modeling and logical relationship recognition operations includes: Constructing a sequential index of the paragraph semantic vectors, where the sequential index is used to represent the original arrangement order of text paragraphs in the literature; Performing joint encoding on the paragraph semantic vectors and the sequential index to generate time-series feature vectors containing sequential information; Extracting the cosine similarity values between adjacent paragraph semantic vectors, and calculating the change gradient of the cosine similarity values in consecutive paragraphs; Identifying logical transition points between paragraphs based on the change gradient, where the logical transition points are used to divide the chapter boundary information in the literature; Performing an attention mechanism on the time-series feature vectors to generate an attention weight distribution reflecting the importance of paragraphs; Constructing a literature structure tree based on the attention weight distribution and the chapter boundary information, and using the node features of the literature structure tree as structural features reflecting the organization mode of the literature.

6. The method for literature structured extraction based on artificial intelligence according to claim 1, wherein Performing structured parsing processing based on the content features and the structural features to generate a set of structured elements of the literature unit, where the set of structured elements includes theme elements, logical relationship elements, and key information elements, including: Inputting the content features into a theme recognition module, and generating theme elements through theme word extraction and theme distribution probability calculation operations, where the theme elements include core theme words and theme coverage descriptions; Inputting the structural features into a logical parsing module, and generating logical relationship elements through paragraph correlation analysis and chapter hierarchy recognition operations, where the logical relationship elements include derivation relationships between paragraphs and subordination relationships between chapters; Input the content features and the structural features into the key information extraction module, and generate key information elements through entity recognition and key sentence screening operations. The key information elements include the core research conclusions and the key experimental conditions; Perform conflict detection processing on the theme elements, the logical relationship elements, and the key information elements to identify the semantic contradiction points between different elements; Based on the semantic contradiction points, perform correction processing on the theme elements, the logical relationship elements, and the key information elements to generate a set of structured elements that pass the consistency verification.

7. The method for literature structured extraction based on artificial intelligence according to claim 6, wherein Input the structural features into the logical analysis module, and generate logical relationship elements through paragraph association analysis and chapter level identification operations. The logical relationship elements include the derivation relationship between paragraphs and the subordination relationship between chapters, including: Extract the paragraph association features in the structural features. The paragraph association features include the semantic connection strength and the content continuity index of adjacent paragraphs; Construct a paragraph relationship graph based on the paragraph association features. The edge weights of the paragraph relationship graph represent the logical association degree between paragraphs; Perform community discovery processing on the paragraph relationship graph to identify the paragraph communities with significant association relationships as the literature chapters; Extract the chapter level features in the structural features. The chapter level features include the font size information and the indentation depth information of the chapter titles; Determine the subordination relationship between chapters based on the chapter level features. The subordination relationship includes the corresponding relationship between the parent chapter and the subchapter; Perform matching verification processing on the paragraph communities in the paragraph relationship graph and the chapter level features to generate logical relationship elements that include the derivation relationship between paragraphs and the subordination relationship between chapters.

8. The method for literature structured extraction based on artificial intelligence according to claim 1, wherein Perform integration and optimization processing on the set of structured elements to generate a final structured literature result that includes the association relationships between elements, including: Construct a feature index for each element in the set of structured elements. The feature index is used to record the unique identifiers of the theme elements, the logical relationship elements, and the key information elements; Perform association analysis processing on the feature index to identify the semantic mapping relationship between the theme elements and the key information elements. The semantic mapping relationship includes the corresponding relationship between the theme words and the key experimental conditions; Identify the structural dependence relationship between the logical relationship elements and the key information elements. The structural dependence relationship includes the distribution position relationship between the chapter subordination relationship and the core research conclusions; Construct an element association graph based on the semantic mapping relationship and the structural dependence relationship. The nodes of the element association graph are the structured elements, and the edges are the association relationships between elements; Perform topological sorting processing on the element association graph to generate a structured arrangement scheme that reflects the importance order of elements; Reorganize the set of structured elements according to the structured arrangement scheme to generate a final structured literature result that includes the association relationships between elements; Among them, constructing the element association graph based on the semantic mapping relationship and the structural dependence relationship, where the nodes of the element association graph are the structured elements and the edges are the association relationships between elements, includes: Extract the association strength value in the semantic mapping relationship, where the association strength value is calculated by the co-occurrence frequency of the subject term and the key information element; Extract the association stability value in the structural dependency relationship, where the association stability value is calculated by the occurrence consistency of the logical relationship element in different document units; Set a weight attribute for the edges of the element association graph, where the weight attribute is obtained by weighted summation of the association strength value and the association stability value; Perform a merging process on the duplicate elements in the structured element set, and merge the subject elements, logical relationship elements, or key information elements with the same semantics into a single node; Perform an attribute supplementation process on the merged element nodes, and integrate the element attribute information from different sources into the same node; Generate a complete element association graph based on the weight attribute and the attribute information of the element nodes, where the element association graph is used to represent the multi-dimensional association relationship between the structured elements.

9. The method for literature structured extraction based on artificial intelligence according to claim 1, wherein The pre-trained document feature extraction model is constructed through the following steps: Obtain a training data set containing annotated documents, where the annotated documents contain manually annotated content feature labels and structural feature labels; Input the training data set into the initial feature extraction model, and generate predicted content features and predicted structural features through forward propagation operations; Calculate the first loss value between the predicted content features and the content feature labels, where the first loss value is calculated by the cross-entropy loss function; Calculate the second loss value between the predicted structural features and the structural feature labels, where the second loss value is calculated by the mean squared error loss function; Fuse the first loss value and the second loss value to generate a joint loss function, and optimize the network parameters of the initial feature extraction model through backpropagation operations; When the convergence value of the joint loss function satisfies a preset threshold, stop training and determine the corresponding initial feature extraction model as the pre-trained document feature extraction model.

10. An artificial intelligence-based literature structured extraction system, characterized in that, It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the artificial intelligence-based document structured extraction method according to any one of claims 1-9 above.

Citation Information

Patent Citations

  • Literature information pushing method, device and system and storage medium

    CN117527888A

  • Scientific and technical element-conversion-structure multi-dimensional coupling calculation method

    CN118586790A

  • Text structure element automatic identification method based on hierarchical neural network

    CN119514521A

  • Method and system for analyzing and predicting theme trend of scientific and technical literature

    CN120068882A

Cited By

  • Report generation method and system based on multi-agent architecture

    CN120745571A

  • Composite knowledge chain construction method for scientific research logic representation

    CN120930751A

  • A composite knowledge chain construction method for scientific research logic representation

    CN120930751B

  • Home disabled elderly safety care education content system construction method and system

    CN121093939A

  • Scientific and technological intelligence deep analysis method and system based on cross-modal semantic enhancement

    CN121303139A