Academic viewpoint extraction method and system applied to academic literature
By performing structural standardization of academic literature and calling pre-trained models for semantic feature extraction, the problem of difficulty in automatically extracting views in the existing technology is solved, efficient and accurate academic views extraction is achieved, and the efficiency and quality of academic research is improved.
Patent Information
- Application Number
- CN202510820344.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The prior art is difficult to extract point-of-view information from academic literature automatically and accurately, resulting in inefficiency and susceptibility to subjective factors.
By obtaining the original literature data, performing structural normalization preprocessing, calling the pre-trained view feature extraction model for semantic feature extraction and view detection, and generating view detection results, including text position markers and view content summary.
It realizes the automation and accurate extraction of academic views, improves the efficiency and quality of academic research, and provides intuitive and comprehensive academic views information.
Smart Images

Figure CN120337937A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology. Specifically, it relates to a method and system for extracting academic viewpoints applied to academic literature. Background Art
[0002] In the field of academic research, as an important carrier for knowledge dissemination and innovation, the quantity and complexity of academic literature are constantly increasing. When conducting academic research, researchers often need to extract key information from a large number of literatures, especially the academic viewpoints in the literatures, in order to deeply understand the development trends in the research field and grasp the research trends. However, traditional methods for reading and analyzing academic literature mainly rely on manual reading and note-taking. This way is not only inefficient but also easily affected by subjective factors, resulting in incomplete and inaccurate information extraction.
[0003] With the rapid development of natural language processing and machine learning technologies, although some automated literature analysis tools have emerged, most of these tools focus on the superficial analysis of literatures, such as keyword extraction, text classification, etc., and fail to deeply mine the academic viewpoints in the literatures and their specific positions in the text. Summary of the Invention
[0004] In view of the problems mentioned above, in combination with the first aspect of the present invention, embodiments of the present invention provide a method for extracting academic viewpoints applied to academic literature. The method includes: Obtain the original literature data set of the academic literature to be analyzed, where the original literature data set contains multiple literature text units with chapter structures; Perform preprocessing on the original literature data set to obtain a preprocessed literature data set with standardized structures, where the preprocessed literature data set contains literature text segments divided by paragraphs and corresponding chapter category tags; Call a pre-trained viewpoint feature extraction model to perform semantic feature extraction processing on the preprocessed literature data set to obtain the literature semantic features of each literature text unit, where the literature semantic features include the semantic correlation degree between text segments and the core discussion direction features; Perform viewpoint detection processing on the literature semantic features through the viewpoint feature extraction model to generate the viewpoint detection results of each literature text unit, where the viewpoint detection results include the text position tags of the academic viewpoints to be extracted and the summary of the viewpoint content; Generate an academic viewpoint extraction result including the summary of the viewpoint content and the corresponding text position tags based on the viewpoint detection results, and output the academic viewpoint extraction result to a target analysis terminal to support academic content analysis operations.
[0005] In another aspect, an academic view extraction system for academic literature provided by an embodiment of the present invention includes a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions, or codes, and the processor is used to execute the programs, instructions, or codes in the machine-readable storage medium to implement the above method.
[0006] Based on the above aspects, the embodiment of the present invention realizes the automatic and accurate extraction of view information in academic literature by comprehensively applying technical means such as document preprocessing, semantic feature extraction, and view detection. First, the original document data is preprocessed with structural standardization to provide a unified and standardized text basis for subsequent view extraction; then, a pre-trained view feature extraction model is called to extract semantic features from the preprocessed document data, effectively capturing the semantic correlation degree and core discussion direction features between text segments; then, through view detection processing, text position markers and view content summaries containing the academic views to be extracted are accurately generated; finally, an academic view extraction result containing the view content summary and the corresponding text position markers is generated, providing intuitive and comprehensive academic view information for scientific researchers. Generally speaking, the technical solution of this application significantly improves the automation level and accuracy of academic view extraction, helps scientific researchers quickly grasp the core views in the literature, improves the efficiency and quality of academic research, and promotes the dissemination and innovation of academic knowledge. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 is a schematic flowchart of the execution process of the academic view extraction method for academic literature provided by an embodiment of the present invention.
[0008] Figure 2 is a schematic diagram of exemplary hardware and software components of the academic view extraction system for academic literature provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0009] The present invention will be specifically described below with reference to the accompanying drawings of the specification. Figure 1 is a schematic flowchart of the academic view extraction method for academic literature provided by an embodiment of the present invention. The academic view extraction method for academic literature will be introduced in detail below.
[0010] Step S110: Obtain an original document data set of the academic literature to be analyzed, where the original document data set includes multiple document text units with a chapter structure.
[0011] In the scenario of academic research and analysis, in order to comprehensively and accurately extract academic viewpoints, it is necessary to first obtain the original literature data set of the academic literature to be analyzed. These literatures come from a wide range of sources, such as professional academic databases, well-known academic journal websites, etc. Taking the medical research field as an example, relevant academic literatures can be obtained from sources such as PubMed and the Chinese Biomedical Literature Database. Each literature has its specific chapter structure, such as title, abstract, introduction, methods, results, discussion, conclusion, etc. The title summarizes the core of the research, the abstract extracts the key points, the introduction elaborates the background and purpose, the methods describe the research means, the results present the research findings, the discussion analyzes the results, and the conclusion summarizes the research value. The original literature data set is composed of numerous such literature text units with chapter structures.
[0012] When collecting academic literature data for training and analysis, it is necessary to strictly abide by laws and regulations to ensure that the data collection obtains legal authorization. For example, when obtaining literature data from a professional academic database, establish a legal cooperation relationship with the operator of the database. By signing a formal data use agreement, clarify key information such as the scope, purpose, and term of data use. For example, when cooperating with a well-known academic journal database, the agreement will stipulate that the obtained data can only be used for research on academic viewpoint extraction and cannot be used for commercial profit or other unauthorized purposes. At the same time, pay the corresponding fees according to the requirements of the agreement to obtain the legal right to use the data. For some literatures not included in the database or requiring additional permissions, directly contact the authors or copyright owners of the literatures and request authorization to use their literature data. In the communication, clearly explain the purpose, method, and expected results of data use to obtain the understanding and support of the authors. Communication can be carried out through emails, official letters, etc., and retain relevant communication records as evidence of authorization. Even for publicly available academic literature data, it is necessary to carefully study its terms of use and copyright statements. Some publicly available data may have specific usage restrictions, such as requiring appropriate citation and annotation when using. When collecting and using these data, strictly abide by its regulations to ensure that the data use complies with legal requirements.
[0013] Step S120: Preprocess the original literature data set to obtain a preprocessed literature data set with a standardized structure. The preprocessed literature data set contains literature text fragments divided by paragraphs and corresponding chapter category tags.
[0014] Since the literatures in the original literature data set come from diverse sources and have large differences in format and structure, it is not conducive to subsequent analysis. Therefore, it is necessary to preprocess it to standardize the literature structure. The preprocessed literature data set contains literature text fragments divided by paragraphs and corresponding chapter category tags, which helps to clarify the chapter attribution of the text fragments in the literature.
[0015] Step S121: Uniformly process the format of each document text unit in the original document data set to eliminate the typesetting differences between different documents, and obtain initial document texts with consistent formats.
[0016] The file format types of different documents are diverse, including plain text format, rich text format, PDF format, etc. For these different formats, different methods should be adopted to eliminate typesetting differences to obtain initial document texts with consistent formats.
[0017] Step S1211: Identify the file format type of the document text unit in the original document data set, and the file format type includes at least one of plain text format, rich text format, and PDF format.
[0018] First, it is necessary to identify the file format type of the document text unit in the original document data set. It can be initially judged by the file suffix. For example, a file ending with.txt is in plain text format, a file ending with.docx may be in rich text format, and a file ending with.pdf is in PDF format. More accurate identification can also be carried out by reading the header information of the file, etc., to ensure that the correct processing method is adopted for documents of different formats.
[0019] Step S1212: Perform mark clearing processing on the document text unit in rich text format, remove the format marks of bold, italic, and underline, and retain the text content.
[0020] For the document text unit in rich text format, the format marks such as bold, italic, and underline it contains are mainly used for format settings during document display and have no actual semantic value for academic view extraction. Parsing tools can be used to parse the markup language of rich text files to identify and remove these format marks. For example, for rich text files based on HTML format, use regular expressions to match and delete 、 、 Such tags to obtain pure text content.
[0021] Step S1213: Perform text extraction processing on the PDF-formatted literature text unit, identify the text content in the scanned PDF through optical character recognition technology, and generate an editable text string.
[0022] If the literature text unit is in PDF format and is a scanned PDF, the text in it exists in the form of images and cannot be directly processed. At this time, it is necessary to rely on optical character recognition (OCR) technology, select a suitable OCR engine, such as Tesseract OCR, and adjust relevant parameters according to the language type, resolution, etc. of the document to recognize the text in the image and generate an editable text string.
[0023] Step S1214: Perform line break normalization processing on the extracted text string, merge consecutive line breaks into a single line break, and eliminate text breaks caused by forced line breaks within a paragraph.
[0024] The extracted text string may have the problem of too many consecutive line breaks, resulting in text breaks within a paragraph and affecting semantic understanding. Traverse the text string programmatically, merge consecutive line breaks into a single line break, and make the paragraph structure of the text clearer. For example, in Python, the replacement method of the string can be used to replace multiple consecutive line breaks with a single line break.
[0025] Step S1215: Perform unified encoding format processing on all literature text units, and convert texts in different encoding formats to a unified Unicode encoding format.
[0026] Different literatures may adopt different encoding formats, such as UTF-8, GBK, etc. To avoid encoding errors in subsequent processing, it is necessary to unify the encoding formats of all literature text units to the Unicode encoding format. It can be achieved using the encoding conversion functions in programming languages. For example, in Python, the encode() and decode() methods are used for encoding conversion.
[0027] Step S1216: Perform content alignment processing on the processed literature text units in various formats to generate an initial literature text with a consistent format.
[0028] Perform content alignment processing on the literature text units in different formats after the above processing, including setting unified font, font size, line spacing, etc., so that all literatures have a unified style in appearance, which is convenient for subsequent processing and analysis, and finally generate an initial literature text with a consistent format.
[0029] Step S122: Perform chapter recognition processing on the initial document text, extract the text content of the document title, abstract, introduction, method, result, discussion, and conclusion, and determine the chapter boundaries of each part.
[0030] In order to accurately extract the text content of each important chapter in the document, it is necessary to perform chapter recognition processing on the initial document text and determine the chapter boundaries of each part.
[0031] Step S1221: Identify the first line of text content of the initial document text, extract the text containing the title keywords as the document title, and record the starting position and ending position of the title.
[0032] The first line of the initial document text is identified through the preset title keyword rules. Title keywords are usually words that can summarize the core content of the document. Regular expressions are used to match whether the first line of text contains these keywords. If the match is successful, the text is determined as the document title, and its starting and ending character positions in the initial document text are recorded.
[0033] Step S1222: Search for a paragraph containing the summary keyword in the text after the title, determine the starting position of the summary part, and extract the continuous text to the position where the next chapter keyword appears as the end position of the summary part.
[0034] After determining the title, search for a paragraph containing the abstract keyword (such as "Abstract") in the text after the title to determine the starting position of the abstract. Then extract the continuous text until the next chapter keyword (such as "Introduction") is encountered, and this position is used as the end position of the abstract.
[0035] Step S1223: Search the text after the abstract for paragraphs containing the introduction keyword, method keyword, result keyword, discussion keyword, and conclusion keyword in turn, and determine the starting position of each chapter.
[0036] In the text after the abstract, use regular expressions to search for paragraphs containing introduction keywords (such as "introduction" "research background"), method keywords (such as "method" "experimental design"), result keywords (such as "result" "experimental results"), discussion keywords (such as "discussion" "analysis"), and conclusion keywords (such as "conclusion" "research conclusion"), so as to determine the starting position of each chapter.
[0037] Step S1224: perform boundary verification processing on the text content after each chapter keyword to check whether the subsequent text contains feature words related to the chapter theme, and the feature words include at least one of research background, experimental design, and data analysis.
[0038] To ensure the accuracy of chapter division, boundary verification is performed on the text content after the keywords of each chapter. Check whether the subsequent text contains characteristic words related to the chapter theme. For example, in the introduction chapter, check whether it contains characteristic words such as "research status" and "problem formulation", and in the method chapter, check whether it contains characteristic words such as "experimental method" and "sample collection". By judging whether these characteristic words are included in the subsequent text, the rationality of the chapter boundary is verified.
[0039] Step S1225: Record the starting character position and ending character position of each chapter to generate a chapter boundary marker set.
[0040] After determining the starting position of each chapter and boundary verification, record the starting character position and ending character position of each chapter. From this, the starting character positions and ending character positions of each chapter can be organized into a chapter boundary marker set.
[0041] Step S1226: Extract the text content of each chapter according to the chapter boundary marker set to generate a chapter content set including the literature title, abstract, introduction, method, results, discussion, and conclusion parts.
[0042] According to the chapter boundary marker set, extract the text content of each chapter from the initial literature text, combine these contents together to generate a chapter content set including the literature title, abstract, introduction, method, results, discussion, and conclusion parts.
[0043] Step S123: Perform paragraph segmentation processing on the initial literature text based on the chapter boundary, divide the continuous text within each chapter into multiple paragraph units, and obtain a literature text fragment divided by paragraphs.
[0044] After determining the boundaries and contents of each chapter, perform paragraph segmentation processing on the continuous text within each chapter according to features such as line breaks and punctuation marks in the text. When a line break is encountered and the text in the next line is semantically independent from the previous line, or according to punctuation marks indicating the end of a sentence such as a period, exclamation mark, or question mark, it can be judged as the start of a new paragraph. Through paragraph segmentation processing, the chapter text is divided into multiple paragraph units to obtain a literature text fragment divided by paragraphs. These fragments are the basic text units for further analysis in the follow-up.
[0045] Step S124: Perform punctuation standardization processing on each paragraph unit, unify the full-width and half-width symbol formats and correct missing punctuation to generate a paragraph text with standardized punctuation.
[0046] Due to different document sources and input methods, there may be inconsistent formatting and missing punctuation marks in the paragraph units. For the problem of inconsistent full-width and half-width symbol formatting, the full-width punctuation marks are converted to half-width punctuation marks through character encoding conversion. For the case of missing punctuation marks, they are judged and supplemented according to the semantic and grammatical rules of the text. For example, in a long sentence, a full stop is added at an appropriate position according to the logical structure and expression intention of the sentence. Through the standardization process of punctuation marks, paragraph texts with standardized punctuation marks are generated, making the semantic expression of the text clearer and more accurate.
[0047] Step S125: Perform category marking processing on the content of each chapter obtained by the chapter recognition processing, and add corresponding chapter category marks to each paragraph unit. The chapter category marks include at least one of the title category, abstract category, method category, and conclusion category.
[0048] After completing the paragraph segmentation and punctuation standardization processing, according to the content of each chapter obtained by the previous chapter recognition processing, add corresponding chapter category marks to each paragraph unit. The paragraph units located in the title chapter are marked as the title category, the paragraph units located in the abstract chapter are marked as the abstract category, the paragraph units located in the method chapter are marked as the method category, and the paragraph units located in the conclusion chapter are marked as the conclusion category. For the paragraph units in chapters such as introduction, results, and discussion, corresponding category marks can also be added according to specific research needs and analysis purposes, such as introduction - background category, introduction - purpose category, etc. By adding chapter category marks, the chapter attribution and semantic function of each paragraph unit in the document are clarified.
[0049] Step S126: Integrate the document text fragments divided by paragraphs and the corresponding chapter category marks to generate a preprocessed document data set with standardized structure.
[0050] After completing the processing of paragraph segmentation, punctuation standardization, and chapter category marking, integrate the document text fragments divided by paragraphs and the corresponding chapter category marks. A dictionary data structure can be used to store the paragraph text as the key and the chapter category mark as the value in a dictionary; or a database table form can be used, with the paragraph text and chapter category mark as different fields respectively, and each record corresponding to a paragraph unit. Through data integration, all the paragraph text fragments and chapter category marks are combined together to generate a preprocessed document data set with standardized structure, which is convenient for subsequent semantic feature extraction and opinion detection operations.
[0051] Step S130: Call a pre-trained opinion feature extraction model to perform semantic feature extraction processing on the preprocessed document data set to obtain the document semantic features of each document text unit. The document semantic features include the semantic correlation degree and the core discussion direction features among text fragments.
[0052] After obtaining the preprocessed literature data set with normalized structure, call the pre-trained opinion feature extraction model to perform semantic feature extraction on it. The pre-trained opinion feature extraction model is trained based on a large amount of academic literature data and can effectively capture the semantic information in the text. By processing the preprocessed literature data set through this model, the literature semantic features of each literature text unit are obtained, including the semantic correlation degree between text segments and the core discussion direction features.
[0053] Step S131: Input the literature text segments in the preprocessed literature data set into the text encoding layer of the opinion feature extraction model, perform tokenization on each paragraph unit, and generate token embedding vectors.
[0054] Input the literature text segments in the preprocessed literature data set into the text encoding layer of the opinion feature extraction model, and perform tokenization on each paragraph unit in this layer to generate token embedding vectors.
[0055] Step S1311: Perform word segmentation on each paragraph unit to split the continuous text into token units with semantic meanings, and the token units include Chinese words and English words.
[0056] For Chinese paragraph units, use a word segmentation tool such as the Jieba word segmentation tool to split the text into appropriate words according to the Chinese language rules; for English paragraph units, use the split() method in Python to split the text into words according to spaces, and split the continuous text into token units with semantic meanings.
[0057] Step S1312: Perform lowercase conversion on the token units after word segmentation to unify the case format of English words.
[0058] To eliminate the influence of the case of English words on subsequent processing, perform lowercase conversion on the English words in the token units after word segmentation, convert all English words to lowercase format, and make the representation of English words more unified.
[0059] Step S1313: Remove the meaningless symbols in the token units, and the meaningless symbols include at least one of special punctuation marks, mathematical symbols, and hyperlinks, and retain the tokens with semantic information.
[0060] In the token units, there may be meaningless symbols such as special punctuation marks, mathematical symbols, and hyperlinks. These symbols are not helpful for semantic understanding and need to be removed. These meaningless symbols can be matched through regular expressions and deleted, and only the tokens with semantic information are retained.
[0061] Step S1314: Input the processed token units into the vocabulary mapping module of the text encoding layer, and map each token unit to its corresponding token index according to the predefined vocabulary.
[0062] The predefined vocabulary contains a large number of common token units. Input the processed token units into the vocabulary mapping module of the text encoding layer, and find the corresponding token index for each token unit according to the vocabulary. If the token unit exists in the vocabulary, map it to the corresponding index; if not, a special index (such as an unknown word index) can be used to represent it.
[0063] Step S1315: Invoke the embedding matrix of the text encoding layer to perform vector transformation processing on the token index, and generate a token embedding vector for each token unit. The dimension of the token embedding vector is consistent with the number of columns of the embedding matrix.
[0064] The embedding matrix of the text encoding layer is a pre-trained matrix. Its number of rows is equal to the size of the vocabulary, and the number of columns represents the dimension of the token embedding vector. Invoke this embedding matrix, use the token index as the row index of the matrix, and extract the corresponding row vector from the embedding matrix, which is the token embedding vector for each token unit. The dimension of the token embedding vector is consistent with the number of columns of the embedding matrix.
[0065] Step S1316: Perform positional encoding processing on the token embedding vectors of each paragraph unit, add a positional encoding vector representing its position in the paragraph to each token embedding vector, and generate a set of token embedding vectors containing position information.
[0066] In order to enable the model to perceive the position information of tokens in a paragraph, perform positional encoding processing on the token embedding vectors of each paragraph unit. According to the position of the token in the paragraph, generate the corresponding positional encoding vector, and add the positional encoding vector to the token embedding vector to obtain a set of token embedding vectors containing position information.
[0067] Step S132: Perform context correlation modeling processing on the token embedding vectors through the self-attention mechanism module of the text encoding layer, and generate local context feature vectors containing semantic relationships between words within the paragraph.
[0068] After obtaining the set of token embedding vectors containing position information, use the self-attention mechanism module of the text encoding layer to perform context correlation modeling processing on the token embedding vectors. The self-attention mechanism module can calculate the correlation between tokens, and generate local context feature vectors containing semantic relationships between words within the paragraph through operations such as weighted summation of token embedding vectors. This local context feature vector can reflect the semantic association between words within the paragraph.
[0069] Step S133: Input the local context feature vector into the cross-paragraph association layer of the opinion feature extraction model to analyze the semantic coherence between different paragraph units, and calculate the semantic overlap degree and the theme continuity parameter of adjacent paragraphs.
[0070] Input the local context feature vector into the cross-paragraph association layer of the opinion feature extraction model. The main function of this layer is to analyze the semantic coherence between different paragraph units. By comparing the local context feature vectors of adjacent paragraphs, calculate the semantic overlap degree and the theme continuity parameter between them. The semantic overlap degree can be measured by calculating the similarity between vectors (such as cosine similarity), and the theme continuity parameter can be determined according to the occurrence and distribution of theme keywords in the paragraphs to judge the semantic association degree between adjacent paragraphs.
[0071] Step S134: Build an inter-paragraph semantic association model based on the semantic overlap degree and the theme continuity parameter, and generate a semantic association degree descriptor between text fragments, where the semantic association degree descriptor is used to represent the cohesion strength of paragraph units in the discourse logic.
[0072] Build an inter-paragraph semantic association model according to the calculated semantic overlap degree and the theme continuity parameter. This model can generate a semantic association degree descriptor between text fragments by means of weighted combination of the semantic overlap degree and the theme continuity parameter. The semantic association degree descriptor is a multi-dimensional vector, which can represent the cohesion strength of paragraph units in the discourse logic.
[0073] Step S135: Conduct a joint analysis and processing on the local context feature vector and the inter-paragraph semantic association degree descriptor to extract the core discourse theme repeatedly emphasized in the literature text unit, and generate a core discourse direction feature, where the core discourse direction feature includes a set of theme keywords and a theme occurrence frequency parameter.
[0074] Conduct a joint analysis and processing on the local context feature vector and the inter-paragraph semantic association degree descriptor. By mining the theme of the local context feature vector and combining the inter-paragraph semantic association degree descriptor to judge the continuity of the theme between different paragraphs, extract the core discourse theme repeatedly emphasized in the literature text unit. The core discourse direction feature includes a set of theme keywords and a theme occurrence frequency parameter. The set of theme keywords represents the key expressions of the core discourse theme, and the theme occurrence frequency parameter reflects the frequency of each theme appearing in the literature. Through this information, the core discourse direction of the literature can be clarified.
[0075] Step S136: Conduct a feature fusion processing on the semantic association degree descriptor between text fragments and the core discourse direction feature to generate a literature semantic feature for each literature text unit.
[0076] Perform feature fusion processing on the semantic correlation degree descriptor between text segments and the core discussion direction feature. These two features can be combined by concatenation to generate the document semantic feature of each document text unit. The document semantic feature is a comprehensive feature vector that contains the semantic correlation information and the core discussion direction information between text segments.
[0077] Step S140: Perform opinion detection processing on the document semantic feature through the opinion feature extraction model to generate the opinion detection result of each document text unit. The opinion detection result includes the text position marker of the academic opinion to be extracted and the summary of the opinion content.
[0078] After obtaining the document semantic feature of each document text unit, it is necessary to carry out opinion detection processing with the help of the opinion feature extraction model to generate the opinion detection result. This result includes the text position marker of the academic opinion to be extracted and the summary of the opinion content. By deeply analyzing the document semantic feature, the position of the academic opinion in the document can be accurately located and its core content can be refined.
[0079] Step S141: Input the document semantic feature into the opinion positioning layer of the opinion feature extraction model, analyze the distribution density of the set of topic keywords in the core discussion direction feature in the paragraph unit, and identify the candidate paragraphs where the topic keywords gather.
[0080] Input the document semantic feature into the opinion positioning layer of the opinion feature extraction model. The main task of this layer is to analyze the distribution density of the set of topic keywords in the core discussion direction feature in the paragraph unit. The set of topic keywords reflects the core discussion topic of the document, and its distribution density in the paragraph can reflect the degree of association between the paragraph and the core discussion. The distribution density is determined by calculating the frequency of the topic keywords appearing in each paragraph. When the distribution density of the topic keywords in a certain paragraph is high, it means that this paragraph may contain important academic opinions and is identified as a candidate paragraph. For example, in a document about climate change, if topic keywords such as "greenhouse gas emissions" and "global warming" frequently appear in a certain paragraph, then this paragraph may be a candidate paragraph.
[0081] Step S142: Perform sentiment analysis processing on the local context feature vector of the candidate paragraph, extract the sentiment tendency parameters representing affirmative, negative, or innovative discussions, and screen out the target paragraphs whose sentiment tendency parameters exceed the preset threshold.
[0082] Perform sentiment analysis processing on the local context feature vector of the candidate paragraph to extract the sentiment tendency parameters representing affirmative, negative, or innovative discussions. The sentiment tendency parameters can reflect the sentiment attitude of the discussions in the paragraph and are crucial for judging the nature of the academic opinion.
[0083] Step S1421: Input the local context feature vector of the candidate paragraph into the sentiment analysis sub-module of the opinion feature extraction model. The sentiment analysis sub-module includes an affirmative classifier, a negative classifier, and an innovative classifier.
[0084] Input the local context feature vector of the candidate paragraph into the sentiment analysis sub-module of the opinion feature extraction model. This sub-module includes an affirmative classifier, a negative classifier, and an innovative classifier. These three classifiers are respectively used to judge whether the sentiment tendency of the paragraph is affirmative, negative, or innovative. Each classifier has been trained with a large amount of data and can accurately classify according to the input local context feature vector.
[0085] Step S1422: Classify the local context feature vector through the affirmative classifier to generate an affirmative probability value indicating that the paragraph content has an affirmative attitude towards a certain claim.
[0086] The affirmative classifier classifies the local context feature vector and generates an affirmative probability value indicating that the paragraph content has an affirmative attitude towards a certain claim according to its internal classification rules and training model. This probability value reflects the likelihood of expressing an affirmative view in the paragraph. For example, in a paragraph discussing the effectiveness of a certain treatment method, the affirmative classifier will analyze the local context feature vector to judge the probability that the paragraph has an affirmative attitude towards this treatment method.
[0087] Step S1423: Classify the local context feature vector through the negative classifier to generate a negative probability value indicating that the paragraph content has a negative attitude towards a certain claim.
[0088] The negative classifier classifies the local context feature vector in a similar way to generate a negative probability value indicating that the paragraph content has a negative attitude towards a certain claim. This probability value reflects the likelihood of expressing a negative view in the paragraph. For example, in a literature evaluating the effect of a certain policy, the negative classifier will analyze the local context feature vector of the relevant paragraph to determine the probability that it has a negative attitude towards this policy.
[0089] Step S1424: Classify the local context feature vector through the innovative classifier to generate an innovative probability value indicating that the paragraph content proposes a new theory or a new method.
[0090] The innovative classifier processes the local context feature vector to generate an innovative probability value indicating that the paragraph content proposes a new theory or a new method. This innovative probability value is used to measure the degree of innovation of the paragraph. In some literatures on frontier research, the innovative classifier can help identify paragraphs that propose novel views or methods.
[0091] Step S1425: Combine the affirmative probability value, the negative probability value, and the innovation probability value into a sentiment tendency parameter set.
[0092] Combine the affirmative probability value, the negative probability value, and the innovation probability value into a sentiment tendency parameter set, which comprehensively reflects the sentiment tendency characteristics of the candidate paragraph.
[0093] Step S1426: Screen out the candidate paragraphs in the sentiment tendency parameter set where at least one probability value exceeds a preset threshold as target paragraphs, and the preset threshold is used to distinguish the statements with clear sentiment tendency from the neutral descriptions.
[0094] Set a preset threshold to distinguish the statements with clear sentiment tendency from the neutral descriptions. Screen out the candidate paragraphs in the sentiment tendency parameter set where at least one probability value exceeds the preset threshold as target paragraphs. For example, if the affirmative probability value exceeds the preset threshold, it indicates that the paragraph has a strong affirmative attitude towards a certain claim; if the innovation probability value exceeds the preset threshold, it means that the paragraph may put forward a new theory or method. Through the above screening method, it is possible to focus on the paragraphs with clear view expressions.
[0095] Step S143: In the target paragraphs, locate the key nodes of the argument logic based on the semantic association degree descriptors between text fragments, and the key nodes include the text positions of putting forward claims, providing evidence, and drawing conclusions.
[0096] In the target paragraphs, locate the key nodes of the argument logic according to the semantic association degree descriptors between text fragments. The semantic association degree descriptors reflect the cohesion strength of paragraph units in the argument logic. By analyzing its numerical value and change trend, the text positions of key argument links such as putting forward claims, providing evidence, and drawing conclusions can be found. For example, when the semantic association degree descriptor shows that the connection between two text fragments is close and there is an obvious causal relationship, it may indicate the positions of a claim and the corresponding evidence; when a text fragment is closely related to the previous content and is summary in nature, it may be the position of drawing a conclusion.
[0097] Step S144: Perform summary generation processing on the text content of the key nodes, and extract short sentences containing the core claim or conclusion as the summary of the view content.
[0098] Perform abstract generation processing on the text content of key nodes to extract short sentences containing core claims or conclusions as the summary of view content. The abstract generation process takes into account the semantic information and importance of the text, removes redundant information, and retains the most core content. For example, in a key node text presenting a certain theory, the abstract generation processing will extract the core expression of the theory; in a key node text reaching a conclusion, a summary statement can be extracted to form the summary of view content.
[0099] Step S145: Record the starting character position and ending character position of the key node in the literature text unit, and generate corresponding text position markers.
[0100] Record the starting character position and ending character position of the key node in the literature text unit, and generate corresponding text position markers based on this. The text position markers can accurately identify the specific position of the view in the literature, facilitating subsequent viewing and citation. By recording the starting and ending character positions, the text range of the key node can be accurately located, providing convenience for the analysis and organization of views.
[0101] Step S146: Perform an association matching process on the summary of view content and the corresponding text position markers to generate the view detection results for each literature text unit.
[0102] Perform an association matching process on the summary of view content and the corresponding text position markers, so that each summary of view content corresponds to an accurate text position marker. Through the above association matching, the view detection results for each literature text unit are generated. The view detection results present the text position and core content of the academic views to be extracted in a structured manner.
[0103] Step S150: Generate an academic view extraction result containing the summary of view content and the corresponding text position markers based on the view detection results, and output the academic view extraction result to the target analysis terminal to support academic content analysis operations.
[0104] After obtaining the view detection results for each literature text unit, it is necessary to generate an academic view extraction result containing the summary of view content and the corresponding text position markers based on these results, and output it to the target analysis terminal to support various academic content analysis operations.
[0105] Step S151: Perform redundancy elimination processing on the summary of view content in the view detection results, merge the view content expressing the same or similar claims, and generate a compressed view content set.
[0106] Perform redundancy elimination on the view content summaries in the view detection results. Since there may be multiple view content summaries in the literature expressing the same or similar claims, these redundant information will increase the complexity of subsequent analysis, so merging processing is required.
[0107] For example, step S1511: Calculate the semantic similarity between any two view contents in the view content summary. The semantic similarity is determined by comparing the keyword overlap rate and the matching degree of the core argument direction features of the view content.
[0108] Calculate the semantic similarity between any two view contents in the view content summary. This process is achieved by comparing the keyword overlap rate and the matching degree of the core argument direction features of the view content. The keyword overlap rate can be determined by calculating the ratio of the number of identical keywords in two view contents to the total number of keywords; the matching degree of the core argument direction features is comprehensively judged based on the set of topic keywords and the topic occurrence frequency parameter in the core argument direction features. For example, if most of the keywords in two view contents are the same and the core argument direction features are also highly matched, then their semantic similarity is relatively high.
[0109] Step S1512: Construct a similarity matrix of view contents. The elements in the similarity matrix represent the semantic similarity degree between two view contents.
[0110] Construct a similarity matrix of view contents based on the calculated semantic similarity. The similarity matrix is a two-dimensional matrix, and each element in the matrix represents the semantic similarity degree between two view contents. The rows and columns of the matrix correspond to different view contents respectively. The larger the value of the element, the more similar the corresponding two view contents are semantically.
[0111] Step S1513: Perform hierarchical clustering analysis based on the similarity matrix, and divide the view contents with semantic similarity exceeding a preset similarity threshold into the same clustering group.
[0112] Perform hierarchical clustering analysis based on the similarity matrix. Preset a similarity threshold, and divide the view contents with semantic similarity exceeding this threshold into the same clustering group. Hierarchical clustering analysis will gradually merge similar view contents into clustering groups according to the element values in the similarity matrix, forming a hierarchical clustering structure. In the above way, the view contents expressing the same or similar claims can be grouped together.
[0113] Step S1514: Perform content fusion processing on the view contents in each clustering group, and extract the common core claim of each view content as the merged view content.
[0114] Perform content fusion processing on the opinion content in each clustering group, analyze the semantic information of each opinion content in the clustering group, and extract their common core claims as the merged opinion content. This process requires comprehensive consideration of the expression methods and key points of each opinion content, removing the different parts and retaining the common core opinions. For example, in a clustering group about a certain technology application, extract the common expressions about the advantages of the technology application in each opinion content to form the merged opinion content.
[0115] Step S1515: Record the number of original opinion contents corresponding to each merged opinion content, and generate an opinion content record containing the merging information.
[0116] Record the number of original opinion contents corresponding to each merged opinion content, save this information together with the merged opinion content, and generate an opinion content record containing the merging information. This record can reflect the information of the merging process, such as how many original opinion contents a certain merged opinion content is merged from.
[0117] Step S1516: Organize the merged opinion content and the corresponding merging information to generate a compressed opinion content set, where the semantic similarity of different opinion contents in the compressed opinion content set is lower than the preset similarity threshold.
[0118] Organize the merged opinion content and the corresponding merging information to form a compressed opinion content set. The semantic similarity of different opinion contents in this set is lower than the preset similarity threshold, that is, the opinion contents in the set have high independence and difference, removing redundant information and making the opinion contents more refined.
[0119] Step S152: Perform topic grouping processing on the opinion content set according to the topic relevance of each opinion content in the opinion content set, and generate multiple topic-related opinion subsets.
[0120] Perform topic grouping processing according to the topic relevance of each opinion content in the compressed opinion content set. Analyze the topic keywords and core discussion directions of each opinion content, and divide the topic-related opinion contents into the same opinion subset. For example, in a literature about energy research, divide the opinion contents related to solar energy utilization into one subset, and divide the opinion contents related to wind energy utilization into another subset, forming multiple topic-related opinion subsets, which is convenient for subsequent in-depth analysis of opinions on different topics.
[0121] Step S153: Extract the text position markers corresponding to the opinion content summaries in each opinion subset, and count the distribution of each text position marker in the literature text unit to generate an opinion distribution density map.
[0122] Extract the text position markers corresponding to the summary of the view content in each view subset, and count the distribution of these text position markers in the literature text unit. By analyzing the distribution of the text position markers, the distribution density of different views in the literature can be understood. For example, count the number of text position markers containing a certain view subset in each chapter, and generate a view distribution density map according to the statistical results. The view distribution density map visually shows the distribution of different views in the literature in a graphical way, which helps to quickly grasp the distribution law of views in the literature.
[0123] Step S154: Add a topic label to each view subset, and the topic label is generated based on the set of topic keywords in the core discussion direction feature.
[0124] Add a topic label to each view subset, and the topic label is generated based on the set of topic keywords in the core discussion direction feature. Extract the keywords related to the topic of the view subset from the core discussion direction feature and combine them into a topic label. For example, for a view subset on biodiversity conservation, extract keywords such as "biodiversity" and "conservation measures" from the core discussion direction feature to form a topic label of "biodiversity conservation". The topic label can concisely summarize the topic content of the view subset, facilitating the identification and classification of the view subset.
[0125] Step S155: Integrate the topic label, the view content subset, and the corresponding view distribution density map to generate an academic view extraction result including a structured hierarchy.
[0126] Integrate the topic label, the view content subset, and the corresponding view distribution density map to form an academic view extraction result including a structured hierarchy. It can be integrated in ways such as a tree structure or a list structure, with the topic label as the upper-level node, the view content subset as the middle-level node, and the view distribution density map as the relevant information of the lower-level node. Through the above structured integration method, the academic view extraction result has a clear hierarchical structure, facilitating understanding and analysis.
[0127] Step S156: Transmit the academic view extraction result to the target analysis terminal through a data interface, and the target analysis terminal is used to perform operations such as academic view comparison, research trend analysis, or literature quality assessment.
[0128] The academic view extraction results are transmitted to the target analysis terminal through a data interface. The target analysis terminal can be a dedicated data analysis software or hardware device, which can perform various academic content analysis operations using the received academic view extraction results, such as academic view comparison, research trend analysis, or literature quality assessment. For example, in the academic view comparison operation, the target analysis terminal can compare the view contents of the same topic in different literatures to find the differences and commonalities of the views; in the research trend analysis, it can analyze the research trends in a certain field based on the view distribution density map and the changes of topic tags; in the literature quality assessment, it can evaluate the literature quality based on aspects such as the innovation and rationality of the views.
[0129] In the above embodiments, the view feature extraction model mainly consists of a text encoding layer, a cross-paragraph association layer, a view localization layer, and a sentiment analysis sub-module, etc. These modules and levels cooperate with each other to achieve the effective extraction of views in academic literature.
[0130] The text encoding layer is the basic processing layer of the view feature extraction model, mainly responsible for performing tokenization on the input literature text fragments and generating token embedding vectors. It includes a vocabulary mapping module and an embedding matrix. The vocabulary mapping module maps token units to corresponding token indices according to a predefined vocabulary, and the embedding matrix converts the token indices into token embedding vectors. In addition, positional encoding processing is also performed on the token embedding vectors to retain the position information of the tokens in the paragraph.
[0131] The cross-paragraph association layer receives the local context feature vectors output by the text encoding layer, and its main function is to analyze the semantic coherence between different paragraph units. By calculating the semantic overlap degree and the topic continuity parameter of adjacent paragraphs, a semantic association model between paragraphs is constructed to generate a semantic association degree descriptor between text fragments.
[0132] The view localization layer receives the literature semantic features. By analyzing the distribution density of the set of topic keywords in the core discussion direction features in paragraph units, it identifies the candidate paragraphs where the topic keywords gather, providing localization information for subsequent view detection.
[0133] The sentiment analysis sub-module includes a positive classifier, a negative classifier, and an innovation classifier, which are used to perform sentiment tendency analysis on the local context feature vectors of the candidate paragraphs to generate sentiment tendency parameters representing positive, negative, or innovative discussions.
[0134] There are clear connection relationships among these modules and levels. The output of the text encoding layer serves as the input of the cross-paragraph association layer, and the output of the cross-paragraph association layer, together with the local context feature vectors of the text encoding layer, is used to generate literature semantic features. The literature semantic features are input into the opinion location layer, and the local context feature vectors of the candidate paragraphs screened by the opinion location layer are input into the sentiment analysis sub-module.
[0135] The quality and diversity of the training data are crucial for the performance of the model. A large number of academic literatures are collected as training data, and these literatures should cover different fields, different research directions, and different writing styles to ensure the wide applicability of the model.
[0136] The collected academic literatures are manually annotated, and the annotation content includes the text location markers of each academic opinion in the literature and the summary of the opinion content. The annotation process needs to follow unified standards and specifications to ensure the accuracy and consistency of the annotation. For example, when annotating a medical research literature, clearly indicate the text location where each research conclusion is located, and extract the core content of the conclusion as the summary of the opinion content.
[0137] The annotated data is divided into a training set, a validation set, and a test set. The training set is used for parameter learning of the model, the validation set is used to adjust the hyperparameters of the model during the training process, and the test set is used to evaluate the final performance of the model. The division ratio can be adjusted according to the actual situation, but generally follows certain principles to ensure that each data set can fully play its role.
[0138] Before training the opinion feature extraction model, a series of training parameters need to be set, and these parameters will affect the training process and final performance of the opinion feature extraction model.
[0139] The learning rate controls the step size of the opinion feature extraction model during each parameter update. If the learning rate is too large, the opinion feature extraction model may skip the optimal solution; if the learning rate is too small, the convergence speed of the opinion feature extraction model will be very slow. A dynamic learning rate strategy can be adopted, using a larger learning rate at the beginning of training and gradually decreasing the learning rate as training progresses to balance the convergence speed and accuracy of the opinion feature extraction model.
[0140] The batch size refers to the number of data samples input into the opinion feature extraction model during each training. A larger batch size can improve the stability and efficiency of training, but may occupy more memory; a smaller batch size can increase the randomness of the opinion feature extraction model, helping to jump out of local optimal solutions. It is necessary to select an appropriate batch size according to the computing resources and the complexity of the opinion feature extraction model.
[0141] The number of training epochs represents the number of times the entire training dataset is traversed by the model. If the number of training epochs is too small, the model may not be able to fully learn the features in the data; if the number of training epochs is too large, the model may overfit. The appropriate number of training epochs can be determined by observing the performance of the validation set. Stop training when the performance of the validation set no longer improves.
[0142] The training process of the opinion feature extraction model is an iterative optimization process. By continuously adjusting the model's parameters, the output result of the model is made as close as possible to the labeled true result.
[0143] For example, input the literature text fragments in the training set into the opinion feature extraction model and process them sequentially according to the module and hierarchical structure of the opinion feature extraction model. The text encoding layer tokenizes and embeds the text, the cross-paragraph association layer analyzes the semantic associations between paragraphs, the opinion location layer locates the candidate paragraphs, the sentiment analysis sub-module conducts sentiment tendency analysis, and finally outputs the predicted opinion detection results. Compare the predicted results of the opinion feature extraction model with the labeled true results and calculate the loss value. The loss function can be selected as the cross-entropy loss function, etc. It measures the degree of difference between the model's predicted results and the true results. The smaller the loss value, the more accurate the model's prediction. According to the calculated loss value, use the backpropagation algorithm to calculate the gradients of the model's various parameters. The gradient represents the rate of change of the loss function with respect to the parameters, and the direction of parameter update can be determined through the gradient.
[0144] According to the calculated gradients and the preset learning rate, update the model's parameters. The updated parameters will make the model output predictions closer to the true results in the next forward propagation.
[0145] Continuously repeat the processes of forward propagation, loss calculation, backpropagation, and parameter update until the preset number of training epochs is reached or the performance of the validation set no longer improves.
[0146] After completing the model training, it is necessary to evaluate the performance of the model using the test set. For example, select appropriate evaluation metrics to measure the performance of the model, such as accuracy, recall rate, F1 value, etc. Accuracy represents the proportion of samples correctly predicted by the model in the total samples, recall rate represents the proportion of positive samples correctly predicted by the model in the actual positive samples, and the F1 value is the harmonic mean of accuracy and recall rate. By comprehensively considering these metrics, the performance of the model can be evaluated comprehensively. If the performance of the opinion feature extraction model is not ideal, the model needs to be optimized. One can start from aspects such as adjusting training parameters, increasing training data, and improving the model structure. For example, if it is found that the model has overfitting, one can try to increase the regularization term or reduce the complexity of the model; if the accuracy of the model is low, one can consider increasing the diversity of training data or adjusting parameters such as the learning rate.
[0147] Figure 2 FIG. 1 shows a schematic diagram of exemplary hardware and software components of an academic view extraction system 100 for academic literature that can implement the ideas of the present application provided by some embodiments of the present application. For example, a processor 120 can be used on the academic view extraction system 100 for academic literature and is used to execute the functions in the present application.
[0148] The academic view extraction system 100 for academic literature can be a general-purpose server or a special-purpose server, both of which can be used to implement the academic view extraction method for academic literature of the present application. Although only one server is shown in the present application, for convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0149] For example, the academic view extraction system 100 for academic literature can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROM, or RAM, or any combination thereof. Exemplarily, the academic view extraction system 100 for academic literature can also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The academic view extraction system 100 for academic literature also includes an I / O interface 150 between the computer and other input / output devices.
[0150] For ease of explanation, only one processor is described in the academic view extraction system 100 for academic literature. However, it should be noted that the academic view extraction system 100 for academic literature in the present application can also include multiple processors. Therefore, the steps performed by one processor described in the present application can also be jointly executed or separately executed by multiple processors. For example, if the processor of the academic view extraction system 100 for academic literature executes steps A and B, it should be understood that steps A and B can also be jointly executed by two different processors or separately executed in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.
[0151] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above-mentioned academic view extraction method for academic literature is implemented.
[0152] It should be noted that, in order to simplify the description of the present invention disclosure and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are incorporated into one embodiment, drawing or description thereof.
Claims
1. A method for extracting academic viewpoints applied to academic literature, characterized in that, The method includes: Obtaining an original literature data set of the academic literature to be analyzed, where the original literature data set contains multiple literature text units with a chapter structure; Performing preprocessing on the original literature data set to obtain a preprocessed literature data set with a standardized structure, where the preprocessed literature data set contains literature text fragments divided by paragraphs and corresponding chapter category tags; Invoking a pre-trained viewpoint feature extraction model to perform semantic feature extraction processing on the preprocessed literature data set to obtain the literature semantic features of each literature text unit, where the literature semantic features include the semantic correlation degree between text fragments and the core discussion direction features; Performing viewpoint detection processing on the literature semantic features through the viewpoint feature extraction model to generate a viewpoint detection result for each literature text unit, where the viewpoint detection result includes the text position tag of the academic viewpoint to be extracted and the summary of the viewpoint content; Generating an academic viewpoint extraction result including the summary of the viewpoint content and the corresponding text position tag based on the viewpoint detection result, and outputting the academic viewpoint extraction result to the target analysis terminal to support academic content analysis operations.
2. The academic view extraction method applied to academic literature according to claim 1, characterized in that, The performing preprocessing on the original literature data set to obtain a preprocessed literature data set with a standardized structure, where the preprocessed literature data set contains literature text fragments divided by paragraphs and corresponding chapter category tags, includes: Performing format unification processing on each literature text unit in the original literature data set to eliminate the typesetting differences of different literatures and obtain an initial literature text with a consistent format; Performing chapter recognition processing on the initial literature text, extracting the text content of the literature title, abstract, introduction, method, result, discussion, and conclusion parts, and determining the chapter boundaries of each part; Performing paragraph segmentation processing on the initial literature text based on the chapter boundaries, dividing the continuous text within each chapter into multiple paragraph units, and obtaining literature text fragments divided by paragraphs; Performing punctuation standardization processing on each paragraph unit, unifying the full-width and half-width symbol formats and correcting missing punctuation to generate paragraph text with standardized punctuation; Performing category tagging processing on the content of each chapter obtained through the chapter recognition processing, adding corresponding chapter category tags to each paragraph unit, where the chapter category tags include at least one of the title category, abstract category, method category, and conclusion category; Integrating the literature text fragments divided by paragraphs and the corresponding chapter category tags to generate a preprocessed literature data set with a standardized structure.
3. The academic view extraction method applied to academic literature according to claim 2, characterized in that, The performing format unification processing on each literature text unit in the original literature data set to eliminate the typesetting differences of different literatures and obtain an initial literature text with a consistent format, includes: Identifying the file format type of the literature text unit in the original literature data set, where the file format type includes at least one of the plain text format, rich text format, and PDF format; Performing tag removal processing on the literature text unit in the rich text format, removing the format tags of bold, italic, and underline, and retaining the text content; Perform text extraction processing on the literature text units in PDF format, identify the text content in scanned PDF through optical character recognition technology, and generate editable text strings; Perform line break normalization processing on the extracted text strings, merge consecutive line breaks into a single line break, and eliminate text breaks caused by forced line breaks within paragraphs; Perform unified encoding format processing on all literature text units, and convert text in different encoding formats into a unified Unicode encoding format; Perform content alignment processing on the processed literature text units in various formats to generate initial literature texts with consistent formats.
4. The academic view extraction method applied to academic literature according to claim 2, characterized in that, Perform chapter recognition processing on the initial literature text, extract the text content of the literature title, abstract, introduction, method, result, discussion, and conclusion parts, and determine the chapter boundaries of each part, including: Identify the first-line text content of the initial literature text, extract the text containing the title keywords as the literature title, and record the starting and ending positions of the title; Search for paragraphs containing abstract keywords in the text after the title, determine the starting position of the abstract part, and extract the consecutive text until the position where the next chapter keyword appears as the ending position of the abstract part; Search for paragraphs containing introduction keywords, method keywords, result keywords, discussion keywords, and conclusion keywords in the text after the abstract part in sequence, and determine the starting positions of each chapter respectively; Perform boundary verification processing on the text content after each chapter keyword, and check whether the subsequent text contains feature words related to the chapter theme, and the feature words include at least one of research background, experimental design, and data analysis; Record the starting and ending character positions of each chapter to generate a chapter boundary marker set; Extract the text content of each chapter according to the chapter boundary marker set to generate a chapter content set containing the literature title, abstract, introduction, method, result, discussion, and conclusion parts.
5. The academic view extraction method applied to academic literature according to claim 1, characterized in that Call the pre-trained opinion feature extraction model to perform semantic feature extraction processing on the preprocessed literature data set to obtain the literature semantic features of each literature text unit, and the literature semantic features include the semantic correlation degree and the core discussion direction feature between text segments, including: Input the literature text segments in the preprocessed literature data set into the text encoding layer of the opinion feature extraction model, perform tokenization processing on each paragraph unit and generate token embedding vectors; Perform context correlation modeling processing on the token embedding vectors through the self-attention mechanism module of the text encoding layer to generate local context feature vectors containing the semantic relationships between words within the paragraph; Input the local context feature vectors into the cross-paragraph association layer of the opinion feature extraction model, analyze the semantic coherence between different paragraph units, and calculate the semantic overlap degree and the theme continuity parameter of adjacent paragraphs; Construct a semantic association model between paragraphs based on the semantic overlap degree and the theme continuity parameter, and generate a semantic association degree descriptor between text segments, and the semantic association degree descriptor is used to represent the cohesion strength of paragraph units in the discussion logic; Jointly analyze and process the local context feature vectors and the inter-paragraph semantic relevance descriptors to extract the core discussion topics repeatedly emphasized in the literature text units, and generate core discussion direction features, where the core discussion direction features include a set of topic keywords and a topic frequency parameter; Perform feature fusion processing on the semantic relevance descriptors between the text segments and the core discussion direction features to generate the literature semantic features of each literature text unit.
6. The academic view extraction method applied to academic literature according to claim 5, characterized in that, Inputting the literature text segments in the preprocessed literature data set into the text encoding layer of the view feature extraction model, and performing tokenization processing on each paragraph unit and generating token embedding vectors, including: Perform word segmentation on each paragraph unit to split the continuous text into token units with semantic meanings, where the token units include Chinese words and English words; Perform lowercase conversion processing on the token units after word segmentation to unify the case format of English words; Remove meaningless symbols in the token units, where the meaningless symbols include at least one of special punctuation marks, mathematical symbols, and hyperlinks, and retain the token units with semantic information; Input the processed token units into the vocabulary mapping module of the text encoding layer, and map each token unit to a corresponding token index according to the predefined vocabulary; Call the embedding matrix of the text encoding layer to perform vector conversion processing on the token index to generate the token embedding vector of each token unit, where the dimension of the token embedding vector is consistent with the number of columns of the embedding matrix; Perform position encoding processing on the token embedding vectors of each paragraph unit, and add a position encoding vector representing its position in the paragraph to each token embedding vector to generate a set of token embedding vectors containing position information.
7. The method for extracting academic viewpoints applied to academic literature according to claim 1, characterized in that, Performing view detection processing on the literature semantic features through the view feature extraction model to generate the view detection results of each literature text unit, where the view detection results include the text position markers of the academic views to be extracted and the view content summaries, including: Input the literature semantic features into the view positioning layer of the view feature extraction model, analyze the distribution density of the set of topic keywords in the core discussion direction features in the paragraph units, and identify the candidate paragraphs where the topic keywords gather; Perform sentiment analysis processing on the local context feature vectors of the candidate paragraphs to extract sentiment tendency parameters representing affirmative, negative, or innovative discussions, and screen the target paragraphs whose sentiment tendency parameters exceed the preset threshold; In the target paragraphs, locate the key nodes of the discussion logic based on the semantic relevance descriptors between the text segments, where the key nodes include the text positions of putting forward claims, providing evidence, and drawing conclusions; Perform summary generation processing on the text content of the key nodes, and extract short sentences containing the core claims or conclusions as the view content summaries; Record the starting character position and the ending character position of the key nodes in the literature text unit to generate corresponding text position markers; Perform associated matching processing on the view content summaries and the corresponding text position markers to generate the view detection results of each literature text unit.
8. The academic view extraction method applied to academic literature according to claim 7, characterized in that, Performing sentiment analysis processing on the local context feature vectors of the candidate paragraphs, extracting sentiment tendency parameters representing affirmative, negative, or innovative discourses, and screening target paragraphs with sentiment tendency parameters exceeding a preset threshold, including: Inputting the local context feature vectors of the candidate paragraphs into the sentiment analysis sub-module of the opinion feature extraction model, where the sentiment analysis sub-module includes an affirmative classifier, a negative classifier, and an innovative classifier; Classifying the local context feature vectors through the affirmative classifier to generate an affirmative probability value indicating that the paragraph content holds an affirmative attitude towards a certain claim; Classifying the local context feature vectors through the negative classifier to generate a negative probability value indicating that the paragraph content holds a negative attitude towards a certain claim; Classifying the local context feature vectors through the innovative classifier to generate an innovative probability value indicating that the paragraph content proposes a new theory or method; Combining the affirmative probability value, negative probability value, and innovative probability value into a sentiment tendency parameter set; Screening candidate paragraphs with at least one probability value exceeding the preset threshold in the sentiment tendency parameter set as target paragraphs, where the preset threshold is used to distinguish discourses with clear sentiment tendencies from neutral descriptions.
9. The academic view extraction method applied to academic literature according to claim 1, characterized in that, Generating an academic opinion extraction result including an opinion content summary and corresponding text position markers based on the opinion detection result, and outputting the academic opinion extraction result to a target analysis terminal to support academic content analysis operations, including: Performing redundancy elimination processing on the opinion content summaries in the opinion detection result, merging opinion contents expressing the same or similar claims, and generating a compressed opinion content set; Performing topic grouping processing on the opinion content set according to the topic relevance of each opinion content in the opinion content set, and generating multiple opinion subsets related to topics; Extracting the text position markers corresponding to the opinion content summaries in each opinion subset, counting the distribution of each text position marker in the literature text unit, and generating an opinion distribution density map; Adding a topic label to each opinion subset, where the topic label is generated based on the set of topic keywords in the core discourse direction feature; Integrating the topic label, opinion content subset, and the corresponding opinion distribution density map to generate an academic opinion extraction result including a structured hierarchy; Transmitting the academic opinion extraction result to the target analysis terminal through a data interface, where the target analysis terminal is used to perform operations such as academic opinion comparison, research trend analysis, or literature quality assessment.
10. An academic view extraction system applied to academic literature, characterized in that, Including a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions, or codes, and the processor is used to execute the programs, instructions, or codes in the memory to implement the academic opinion extraction method for academic literature according to any one of claims 1-9 above.
Citation Information
Patent Citations
Information processor, feature extraction method, recording medium, and program
CN101031919A
Feeling detection method, feeling detection device, feeling detection program containing the method, and recording medium containing the program
CN101506874A
Scholars viewpoint extraction method based on webpage texts
CN110263319A
Text data viewpoint abstract generation method and system based on sentence meaning structure model
CN110889292A
Net citizen viewpoint analysis method based on large language model and topic model
CN117688182A
Cited By
Biomedical literature elaborated content generation system and generation method
CN122153053A
Intelligent document batch processing system and method based on large language model
CN122173458A