Academic viewpoint labeling and analyzing method and system

By annotating and analyzing academic literature and generating snapshots of academic views, the problems of low efficiency and high cost of computing resources in the existing technology are solved, and efficient and accurate recognition and analysis of academic views are achieved, and reading experience is improved.

CN120068800AActive Publication Date: 2025-05-30BEIJING BZD TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510153270.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-30
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

The prior art is inefficient and consumes computational resources when identifying and mining academic views in academic literature, resulting in readers facing huge reading burdens.

Method used

By searching and classifying historical academic literature, labeling analysis conditions are formulated, and the XML format data to be labeled and analyzed using these conditions are labeled and analyzed, and a snapshot of academic opinions is generated, including location annotation, categories and content.

Benefits of technology

It improves the efficiency and accuracy of finding valuable academic opinions in a large number of academic literature, reduces the consumption of computing resources, provides a better reading experience, and lays a foundation for the identification of view plagiarism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068800A_ABST
    Figure CN120068800A_ABST
Patent Text Reader

Abstract

The invention discloses an academic viewpoint annotation and analysis method and system, and belongs to the technical field of academic viewpoint recognition, and the method comprises the steps: obtaining historical academic literatures, classifying the historical academic literatures, obtaining annotation analysis conditions of academic viewpoints, obtaining XML format data of the academic literatures, and preprocessing the XML format data. The method comprises the following steps: taking an academic viewpoint as to-be-annotated analysis text data, annotating and analyzing the to-be-annotated analysis text data by using an annotation analysis condition to obtain an annotation result and an analysis result, and obtaining an academic viewpoint snapshot according to the annotation result and the analysis result. The processing efficiency and the recognition accuracy are improved, and foundation and bedding work is laid for identification of the plagiarism of the viewpoints according to the academic viewpoint snapshots.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of academic opinion recognition. Specifically, it relates to a method and system for academic opinion annotation and analysis. Background Art

[0002] The objects for academic opinion mining generally include: user evaluations on social platforms and e-commerce platforms; news opinion sentences in Internet news; subjective comments on various things published by users on blogs and BBS forums, taking the subjective comments as the objects for opinion extraction; opinions expressed by experts and scholars on social hot events on the Internet, lacking the annotation, analysis, and mining of academic opinions in academic literature.

[0003] Among the current methods for extracting academic opinions from academic literature, many are based on word analysis. For example, feature analysis based on sentiment words, polarity words, negation words, phrases, etc. is used to identify opinion sentences. When building a self-training model under the condition of limited annotation quantity or without manual annotation, the model can learn the rules from the data and accurately identify and mine the sentence segments of academic opinions. However, when applied to academic literature with a large amount of academic text content, to efficiently identify the sentences containing academic opinions from the academic text, a large amount of computing resources are required, and it is time-consuming and laborious to find valuable academic opinions, bringing a huge reading burden to readers. Summary of the Invention

[0004] To solve the above problems and technical deficiencies, the embodiments of this application adopt the following technical solutions. An academic opinion annotation and analysis method includes the following steps:

[0005] Step 1: Retrieve and download historical academic literature that has completed academic annotation and analysis, classify the historical academic literature, and formulate the annotation and analysis conditions for academic opinions;

[0006] The annotation and analysis conditions include: the last chapter part of the academic literature, the last part of the paragraph unit, and the end of the natural paragraph;

[0007] Step 2: Receive the XML format data of the academic literature to be annotated and analyzed, preprocess the XML format data, and use it as the text data to be annotated and analyzed;

[0008] Step 3: Use the annotation and analysis conditions to annotate and analyze the text data to be annotated and analyzed, obtain the annotation result and the analysis result, and generate an academic opinion snapshot according to the annotation result and the analysis result;

[0009] Among them, the annotation result includes: the position annotation of the academic opinion sentence in the paragraph unit and the position annotation of the paragraph unit in the full text of the text;

[0010] Among them, the analysis result includes: the category and content of the academic opinion;

[0011] Among them, the content of academic viewpoints includes: main conclusions, core viewpoints, arguments, topic sentences, and conclusions;

[0012] Among them, conclusions include: major conclusions obtained by analyzing the last chapter of academic literature, medium conclusions obtained by analyzing the last part of paragraph units, and minor conclusions obtained by analyzing the end of natural paragraphs.

[0013] Preferably, the process of formulating the annotation and analysis conditions is as follows:

[0014] Retrieve and download historical academic literature that has completed academic annotation and analysis, and conduct the first viewpoint classification on the academic viewpoints of the historical academic literature;

[0015] Judge the importance of each academic viewpoint category in the results of the first viewpoint classification, and conduct a defined expression and the second viewpoint classification on the academic viewpoints of the historical academic literature according to the importance judgment results;

[0016] Formulate conditions for the annotation and analysis of academic viewpoints according to the defined expression and the results of the second viewpoint classification, and obtain the annotation and analysis conditions.

[0017] Clean the XML format data of academic literature;

[0018] Segment text sentences, and then segment each part of the paragraph unit;

[0019] According to the IMRaD structure, divide the full text of the text into five parts: introduction, discussion, methods, results, and conclusions in terms of chapters, for annotating and analyzing the positions of academic viewpoint sentences in paragraphs or paragraph units, as well as the positions of paragraphs or paragraph units in the full text.

[0020] Furthermore, the academic viewpoint snapshot also includes: metadata attributes;

[0021] Among them, metadata attributes include: author, institution, year, article title, source journal.

[0022] An academic viewpoint annotation and analysis system includes:

[0023] A text acquisition module, configured to acquire XML format data of academic literature, preprocess the XML format data, and use it as text data to be annotated and analyzed;

[0024] A condition formulation module, configured to acquire historical academic literature, classify the historical academic literature, and obtain the annotation and analysis conditions of academic viewpoints;

[0025] The annotation and analysis conditions include: the last chapter of academic literature, the last part of paragraph units, and the end of natural paragraphs;

[0026] The annotation analysis module is used to perform annotation and analysis on the text data to be annotated and analyzed using annotation analysis conditions, obtain annotation results and analysis results, and obtain an academic view snapshot based on the annotation results and analysis results;

[0027] Among them, the annotation results include: the position annotation of the academic view sentence in the paragraph unit and the position annotation of the paragraph unit in the full text of the text;

[0028] Among them, the analysis results include: the category and content of the academic view;

[0029] Among them, the content of the academic view includes: main conclusions, core views, arguments, topic sentences, and conclusions;

[0030] Among them, the conclusions include: major conclusions obtained by analyzing the last chapter part of the academic literature, medium conclusions obtained by analyzing the last part of the paragraph unit, and minor conclusions obtained by analyzing the end of the natural paragraph.

[0031] Furthermore, the academic view snapshot further includes: metadata attributes;

[0032] Among them, the metadata attributes include: author, institution, year, article title, source journal;

[0033] The summarization and generalization priority of the major conclusion is higher than that of the medium conclusion;

[0034] The summarization and generalization priority of the medium conclusion is higher than that of the minor conclusion.

[0035] Preferably, the conclusion further includes a commentary. The commentary is the content of the author's commentary on the cited views in the text. It is necessary to identify and annotate the commentary. By identifying and annotating the academic view sentences cited by the author in the text and the corresponding author commentary content, the identification and annotation of the commentary are completed.

[0036] Furthermore, there are four forms of commentary manifestations;

[0037] The first form of commentary manifestation is that the cited view is in the front and the commentary content is in the back, and the cited view and the commentary share one sentence;

[0038] The second form of commentary manifestation is that the cited view is in the front and the commentary content is in the back, and the cited view and the commentary are divided into multiple sentences;

[0039] The third form of commentary manifestation is that there is no commentary content, only the cited view, and there is a logical relationship between the sentence before the cited view sentence, the cited view sentence, and the sentence after the cited view sentence;

[0040] The fourth form of review presentation is a view without citation, only the content of the review. The content of the review is a paragraph unit that summarizes the existing research results. There is only the content of the review in the paragraph unit and no corresponding cited view.

[0041] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the academic view annotation and analysis method described above are implemented.

[0042] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the academic view annotation and analysis method described above are implemented.

[0043] Compared with the prior art, the beneficial effects of the embodiments of the present application are as follows:

[0044] By clarifying the categories of various academic views in academic literature, formulating conditions for the annotation and analysis of academic views according to the importance of academic views, and performing annotation and analysis on the text data to be annotated and analyzed according to the annotation analysis conditions to generate an academic view snapshot, the present application can improve the document processing efficiency and recognition accuracy when searching for valuable academic views in a large number of academic literatures, lay a foundation and make preparations for the identification of view plagiarism based on the academic view snapshot, and provide a better reading experience for readers. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In the drawings:

[0046] Figure 1 is a schematic diagram of the method steps of the embodiment of the present application;

[0047] Figure 2 is a schematic diagram of the method flow of the embodiment of the present application;

[0048] Figure 3 is a schematic diagram of the system structure of the embodiment of the present application;

[0049] Figure 4 is a schematic diagram of the device structure of the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and illustrated in the drawings here can be arranged and designed in various different configurations.

[0051] Embodiment 1:

[0052] As Figure 1 andFigure 2 As shown in Figure 2 , a method for annotating and analyzing academic viewpoints includes the following steps:

[0053] Retrieve and download historical academic documents that have completed academic annotation and analysis, classify the historical academic documents, and formulate the annotation and analysis conditions for academic viewpoints;

[0054] The annotation and analysis conditions include: the last chapter part of the academic document, the last part of the paragraph unit, and the end of the natural paragraph;

[0055] The process of formulating the annotation and analysis conditions is as follows:

[0056] Obtain historical academic documents and conduct the first viewpoint classification on the academic viewpoints of the historical academic documents;

[0057] Judge the importance of each academic viewpoint category in the results of the first viewpoint classification, and conduct a defined expression and a second viewpoint classification on the academic viewpoints of the historical academic documents according to the importance judgment results;

[0058] Formulate the conditions for the annotation and analysis of academic viewpoints according to the defined expression and the results of the second viewpoint classification, and obtain the annotation and analysis conditions;

[0059] Among them, the annotation results include: the position annotation of the academic viewpoint sentence in the paragraph unit and the position annotation of the paragraph unit in the full text of the text;

[0060] Among them, the analysis results include: the category and content of the academic viewpoint;

[0061] Among them, the content of the academic viewpoint includes: main conclusions, core viewpoints, arguments, topic sentences, and conclusions;

[0062] Among them, the conclusions include: the major conclusions obtained by analyzing the last chapter part of the academic document, the medium conclusions obtained by analyzing the last part of the paragraph unit, and the minor conclusions obtained by analyzing the end of the natural paragraph.

[0063] Receive the XML format data of the academic document to be annotated and analyzed, preprocess the XML format data, and use it as the text data to be annotated and analyzed;

[0064] The preprocessing process is as follows:

[0065] Clean the data of the XML format data of the academic document;

[0066] Segment the text sentences, and then segment each part of the paragraph unit;

[0067] The full text is segmented into five parts: Introduction, Discussion, Methods, Results, and Conclusion, based on the IMRaD structure, to mark and analyze the positions of academic opinion sentences in paragraphs or paragraph units, as well as the positions of paragraphs or paragraph units in the full text.

[0068] Use the annotation analysis conditions to annotate and analyze the text data to be annotated and analyzed, obtain the annotation results and analysis results, and generate an academic opinion snapshot based on the annotation results and analysis results;

[0069] The academic opinion snapshot also includes: metadata attributes;

[0070] Among them, the metadata attributes include: author, institution, year, title of the article, source journal.

[0071] Academic opinions are the core content of academic literature, an important part of academic information exchange. They are the views and understandings held by the opinion holders on the issues, things, or phenomena studied in their own research fields. The formation of academic opinions is the result of scholars' thinking activities and plays an important role in their scientific research activities and cognitive structures. For the same research question, different scholars or schools of thought will form completely different academic opinions due to differences in discourse systems, research perspectives, and prior knowledge.

[0072] Based on the important role of academic opinions in academic communication, in order to protect the originality of academic opinions and at the same time reveal academic misconduct, in the industry standard of the press and publication in China, "Academic Publishing Norms - Definition of Academic Misconduct in Journals", the manifestations of such misconduct as "opinion plagiarism" are summarized as the following 5 types:

[0073] a) Directly using the arguments, opinions, conclusions, etc. in the published literature of others without citation;

[0074] b) Using the arguments, opinions, conclusions, etc. of others without citation after paraphrasing them without changing their original meaning;

[0075] c) Using the arguments, opinions, conclusions, etc. of others after deleting some content without citation;

[0076] d) Using the arguments, opinions, conclusions, etc. of others after splitting or reorganizing them without citation;

[0077] e) Using the arguments, opinions, conclusions, etc. of others after adding some content without citation.

[0078] From the above manifestations of "opinion plagiarism", it can be seen that the substantial content of the main objects of opinion plagiarism is the "arguments, opinions, conclusions", etc. original to the author.

[0079] Therefore, the scope of academic opinions is the original arguments, opinions, conclusions, etc. of the author in academic literature.

[0080] In order to accurately identify and extract the arguments, viewpoints, and conclusions in academic literature, based on the analysis of the text structures of a large number of academic literature, for the summary viewpoints and paragraphs originally created by the authors in academic literature, category expansion is carried out on the basis of arguments, viewpoints, and conclusions. The expanded categories include: main conclusions and core viewpoints, arguments and topic sentences, and detailed elaborations of conclusions.

[0081] At the same time, according to the importance of academic viewpoints, classification is carried out at the full-text level and paragraph-unit level of the academic viewpoints, definitions and differentiations are made for various types of academic viewpoints, and conditions are formulated for the annotation and analysis of academic viewpoints:

[0082] Main conclusions and core viewpoints are the most basic views of the authors on the research questions discussed in academic literature, the most important ideological viewpoints put forward in the text, and also the high-level summary and concentration of all the minor conclusions and sub-arguments in the text. Main conclusions and core viewpoints belong to the category of sentences and paragraphs of the authors' original viewpoints at the full-text level of the whole article and are generally located in the "Conclusion" part of the position in the text.

[0083] Arguments and topic sentences are a topic or hypothesis put forward for a thing or phenomenon, etc., and then based on this topic or hypothesis, evidence to support this topic or hypothesis is sought. Arguments and topic sentences belong to the category of sentences and paragraphs of the authors' original viewpoints at the paragraph-unit level and are generally located at the beginning of the paragraph unit. Among them, sub-arguments are several ideological viewpoints subordinate to the central argument and serving to elaborate the central argument.

[0084] Conclusions are conclusive contents deduced from facts and logic through investigations and logical analyses of things or phenomena, and thus have a certain degree of authority, rigor, and practicality. Conclusions can be divided into major conclusions, medium conclusions, and minor conclusions;

[0085] When the last chapter part of the thesis is the major conclusion, this part is the summary and generalization of the full text, that is, the above-mentioned "main conclusion", belonging to the conclusion at the full-text level of the whole article;

[0086] Minor conclusions are the results written in standard paragraphs. When the end of a natural paragraph is a minor conclusion, this part is the summary of this natural paragraph and belongs to the conclusion at the natural paragraph level;

[0087] When the last part at the paragraph-unit level is the medium conclusion, this part is the summary and generalization of the paragraph unit, and the degree of summary and generalization is between that of the major conclusion and the minor conclusion;

[0088] That is, the degree of summary and generalization of the minor conclusion < the degree of summary and generalization of the medium conclusion < the degree of summary and generalization of the main conclusion.

[0089] In addition to the above-mentioned opinion sentence segments, there is another type of opinion sentence segment called "commentary", which is the author's commentary on the cited opinions in the text. By marking and identifying the academic opinion sentences cited by the author in the text and matching them with the corresponding author's commentary content, the commentary content is an opinion sentence segment of equal importance to the above-mentioned various types in academic literature and is also one of the objects of key research, marking, and excavation.

[0090] For example, in Text Example 1:

[0091] The identification of entity boundaries is relatively difficult and has always been a difficult point in named entity recognition. In recent years, deep learning methods with deep neural network structures have been continuously developed and have shown better effects than traditional methods in many named entity recognition tasks [6-7]; however, as a specific type of named entity, there are still few relevant research results on the extraction of theoretical terms at present; the extraction of theoretical terms is different from the tasks of other types of named entity recognition (such as person name recognition, place name recognition, etc.). At present, there is no large-scale mature corpus, and the complex composition and cross-disciplinary distribution characteristics of theoretical terms also increase the difficulty of their extraction; therefore, how to apply deep learning methods to the extraction of theoretical terms, improve the extraction performance and reduce the manual annotation cost of the training set is a problem that needs to be studied currently.

[0092] Text Example 1 can clearly express four different types of opinion sentences: "argument / thesis, cited opinion, author's commentary on the cited opinion, small conclusion".

[0093] Among them, the argument / thesis sentence of Text Example 1 is "The identification of entity boundaries is relatively difficult and has always been a difficult point in named entity recognition", which is located at the beginning of the paragraph, is the argument put forward for the content of this paragraph, and is also the thesis sentence of this paragraph;

[0094] The cited opinion sentence is "In recent years, deep learning methods with deep neural network structures have been continuously developed and have shown better effects than traditional methods in many named entity recognition tasks [6-7]". This sentence follows the argument and thesis sentence closely, and the author uses the opinion of other authors as evidence.

[0095] The author's commentary on the cited opinion is "However, as a specific type of named entity, there are still few relevant research results on the extraction of theoretical terms at present; the extraction of theoretical terms is different from the tasks of other types of named entity recognition (such as person name recognition, place name recognition, etc.). At present, there is no large-scale mature corpus, and the complex composition and cross-disciplinary distribution characteristics of theoretical terms also increase the difficulty of their extraction". These two sentences of "commentary" follow the author's sentence of citing the opinion of other authors closely, are the content of the author's commentary on the cited opinion, and are also the original personal opinions of the author.

[0096] The last sentence is a summary made by the author for this paragraph, and its content is "Therefore, how to apply deep learning methods to theoretical term extraction, improve extraction performance and reduce the manual annotation cost of the training set is an issue that needs to be studied currently". Since it is only a conclusive statement for this paragraph, it is marked as "small conclusion"; when this sentence is a summary statement made by the author for the content at the paragraph unit level, it is marked as "medium conclusion"; when this sentence is a summary statement made by the author for the content at the whole text level, it is marked as "big conclusion", that is, "main conclusion" or "core view".

[0097] For example, in Text Example 2:

[0098] Conclusion:

[0099] The extraction of theoretical terms is the basis for large-scale literature content analysis and in-depth revelation of interdisciplinary knowledge transfer; this paper proposes a method for theoretical term extraction based on deep learning. Based on the widely used Bi-LSTM-CRF deep learning model as the basic framework, a theoretical term extraction model is constructed, and the feature construction and annotation method of theoretical terms in this model are studied. The extraction performance of this model under different feature trainings is experimentally compared and analyzed; the experimental results show the effectiveness of the theoretical term end-word features, context features, word-formation features, negative example features, word vector features, and part-of-speech features constructed in this paper for theoretical term extraction. The model incorporating the above features is significantly higher than the traditional method under different training set ratios and has good generality and scalability; this paper also proposes a self-training algorithm, enabling the model to perform weakly supervised learning on unlabeled data, further improving the extraction performance while reducing the manual annotation cost; however, there are also problems such as the inability to recall some theoretical terms and extraction errors. Next, it will be studied to construct character- and word-level convolutional layers in the model so that the model can more effectively learn the extraction features of theoretical terms, and the application of co-training algorithms in the self-training of the model will be studied to weaken the influence of noise; in addition, how to reasonably select incremental self-training batches and automatically adjust thresholds unsupervised is also an issue to be further studied and solved in the future; it can be said that this paper not only provides a more effective general method for theoretical term extraction but also provides a method reference for the research on other types of named entity recognition.

[0100] Text Example 2 is marked as "big conclusion", that is, "main conclusion" or "core view".

[0101] There are generally four forms of presentation for comments in academic literature:

[0102] Form of presentation of comments 1: Cited view + comment = one sentence, with the cited view in front and the comment behind, combined into one sentence.

[0103] For example, in Text Example 3:

[0104] Term citation can effectively measure interdisciplinary research. Analyzing the innovation and borrowing of theoretical terms helps to reveal the evolution and cross-penetration of core knowledge in different disciplines.

[0105] In this example sentence, the cited view is that "term citation can effectively measure interdisciplinary research".

[0106] The content of the commentary is that "analyzing the innovation and borrowing of theoretical terms helps to reveal the evolution and cross-penetration of core knowledge in different disciplines".

[0107] Commentary form two: The cited view comes first, and the commentary comes later, each being a single sentence, for a total of two or three sentences.

[0108] For example, in text example four:

[0109] In recent years, deep learning methods with deep neural network structures have been continuously developed and have shown better performance than traditional methods in many named entity recognition tasks [6-7]; as a specific type of named entity, there are still relatively few research results on the extraction of theoretical terms; the extraction of theoretical terms is different from other types of named entity recognition tasks (such as person name recognition, place name recognition, etc.), and there is currently no large-scale mature corpus, and the complex composition and full-discipline distribution characteristics of theoretical terms also make their extraction more difficult.

[0110] In this example sentence, the cited view is that "in recent years, deep learning methods with deep neural network structures have been continuously developed and have shown better performance than traditional methods in many named entity recognition tasks [6-7]".

[0111] The first commentary content in text example four is that "as a specific type of named entity, there are still relatively few research results on the extraction of theoretical terms".

[0112] The second commentary content in text example four is that "the extraction of theoretical terms is different from other types of named entity recognition tasks (such as person name recognition, place name recognition, etc.), and there is currently no large-scale mature corpus, and the complex composition and full-discipline distribution characteristics of theoretical terms also make their extraction more difficult".

[0113] Commentary form three: There is no commentary. Although the author cites the academic views of other authors, there is no commentary content. This situation is also common in academic literature. Therefore, it is necessary to clarify that "cited view" and "commentary" do not necessarily appear in the text at the same time. Based on this, it is possible to avoid marking and identifying commentary content that does not correspond to the cited view.

[0114] For example, in text example five:

[0115] While the structure of the hidden layer and output layer of the model remains unchanged, the input features of the input layer have an important impact on the learning performance of the model [6, 12 - 13, 19].

[0116] The whole sentence of this example is a cited view. The sentence before this example is "As can be seen from the above analysis, the Bi - LSTM - CRF model learns the deep features of the data from the features of the input sequence dataset, and realizes the representation of this type of data for sequence labeling prediction"; the sentence after this example is "This paper intends to use word vector (word embedding), part - of - speech (POS) and entity (ENT) features of theoretical terms as the input of the Bi - LSTM - CRF model, and study the influence of each feature on the effectiveness of theoretical term extraction, in order to achieve a more effective theoretical term extraction model".

[0117] From the relationship among the cited view sentence, the sentence before it and the sentence after it, it can be seen that the above example sentence is a cited view sentence, but there is no corresponding commentary content.

[0118] Therefore, although this sentence is a cited view, from its relationship with the previous sentence, it can be seen that the more important nature of this sentence is similar to a commentary sentence. It is both a summary sentence of the previous sentence and a commentary on a view sentence. That is to say, sometimes, a summary sentence can be regarded as a view sentence, and a cited view sentence can be regarded as a commentary sentence, which is determined by its content nature.

[0119] Therefore, although the next sentence follows the cited view sentence closely, it is not marked as "commentary" because its nature is not the author's commentary on the cited view.

[0120] Commentary manifestation form four: Due to the special nature of the paragraph unit, for example, in the paragraph unit that summarizes the previous research results, there is a situation where there is only commentary without corresponding cited views.

[0121] For example, in Text Example Six:

[0122] 2. Related research:

[0123] 2.1. Theoretical term extraction:

[0124] Regarding the extraction of theoretical terms, scholars have conducted certain research. Wang Fang et al. [2] used manual annotation to identify 586 theories in 1,822 articles published in the Journal of Intelligence Science from 2000 to 2013. Chen Feng et al. [3] further used a machine learning method based on CRF to study the automatic extraction method of theoretical terms in academic journals. They conducted extraction experiments on the titles and abstracts of the above 1,822 papers. When the proportion of the training set was increased to 80%, a relatively good accuracy rate (93.33%) was achieved, but the recall rate was still not high (56%). In addition, progress has also been made in the extraction of method terms, which is similar to the extraction of theories. Hua Berlin [8] used a rule-based method to extract intelligence method terms from Chinese academic literature. He conducted experimental tests on the full-text data of 46 papers in the 1st to 4th issues of the Journal of Intelligence Science in 2013, and both the recall rate (94%) and the accuracy rate (92%) were relatively high.

[0125] The above-mentioned research all adopted manual rules or traditional machine learning methods to extract for one discipline and one type of literature, and their generality, scalability, and machine learning performance still need to be greatly improved.

[0126] The chapter particularity in Text Example 6 leads to the particularity of the paragraph nature. "2. Related Research" is a review. Therefore, the relationships between paragraphs and sentences in this chapter need to pay attention to their particularity. This section is "2.1. Extraction of Theoretical Terms". The entire first paragraph quotes the research results of predecessors without viewpoints, so it is not marked. The second paragraph, which is the last paragraph of this section, should be considered as the author's commentary on the above-cited content.

[0127] Example 2:

[0128] As Figure 3 shown, an academic viewpoint annotation and analysis system includes:

[0129] A text acquisition module, used to acquire XML format data of academic literature, preprocess the XML format data, and use it as the text data to be annotated and analyzed;

[0130] The preprocessing process is as follows:

[0131] Clean the XML format data of academic literature;

[0132] According to the IMRaD structure, divide the full text of the text into five parts: introduction, discussion, method, result, and conclusion in units of chapters;

[0133] Divide the paragraph units of each part, and then divide the text sentences of each paragraph unit.

[0134] A condition formulation module, configured to obtain historical academic documents, classify the historical academic documents, and obtain the annotation analysis conditions for academic viewpoints;

[0135] The annotation analysis conditions include: the last chapter part of the academic document, the last part of the paragraph unit, and the end of the natural paragraph;

[0136] The process of formulating the annotation analysis conditions is as follows:

[0137] Obtain historical academic documents and conduct a first viewpoint classification on the academic viewpoints of the historical academic documents;

[0138] Judge the importance of each academic viewpoint category in the results of the first viewpoint classification, and define and express the academic viewpoints of the historical academic documents and conduct a second viewpoint classification based on the importance judgment results;

[0139] Formulate conditions for the annotation and analysis of academic viewpoints based on the definition expression and the results of the second viewpoint classification, and obtain the annotation analysis conditions.

[0140] An annotation analysis module, configured to use the annotation analysis conditions to annotate and analyze the text data to be annotated and analyzed, obtain the annotation results and analysis results, and obtain an academic viewpoint snapshot based on the annotation results and analysis results;

[0141] Among them, the annotation results include: the position annotation of the academic viewpoint sentence in the paragraph unit and the position annotation of the paragraph unit in the full text of the text;

[0142] Among them, the analysis results include: the category and content of the academic viewpoint;

[0143] Among them, the content of the academic viewpoint includes: main conclusions, core viewpoints, arguments, topic sentences, and conclusions;

[0144] Among them, the conclusions include: major conclusions obtained by analyzing the last chapter part of the academic document, medium conclusions obtained by analyzing the last part of the paragraph unit, and minor conclusions obtained by analyzing the end of the natural paragraph.

[0145] The academic viewpoint snapshot also includes: metadata attributes;

[0146] Among them, the metadata attributes include: author, institution, year, title of the article, source journal;

[0147] The summarization and generalization priority of the major conclusions is higher than that of the medium conclusions;

[0148] The summarization and generalization priority of the medium conclusions is higher than that of the minor conclusions.

[0149] The conclusion also includes comments, which are the author's comments on the cited views in the text. It is necessary to identify and label the comments. By identifying and labeling the academic view sentences cited by the author in the text and the corresponding author's comment content, the identification and labeling of the comments are completed.

[0150] There are four forms of comment expressions;

[0151] The first form of comment expression is that the cited view comes first and the comment content comes later, and the cited view and the comment share one sentence;

[0152] The second form of comment expression is that the cited view comes first and the comment content comes later, and the cited view and the comment are divided into multiple sentences;

[0153] The third form of comment expression is that there is no comment content, only the cited view, and there is a logical relationship between the sentence before the cited view sentence, the cited view sentence, and the sentence after the cited view sentence;

[0154] The fourth form of comment expression is that there is no cited view, only the comment content, and the comment content is a paragraph unit summarizing the existing research results. There is only comment content in the paragraph unit and no corresponding cited view.

[0155] Example 3:

[0156] As Figure 4 shown, from the hardware level, the present application provides an embodiment of an electronic device for all or part of the content of the academic view annotation and analysis method. The electronic device specifically includes the following content:

[0157] Step 1, obtain historical academic documents, classify the historical academic documents, and obtain the annotation and analysis conditions of academic views;

[0158] The annotation and analysis conditions include: the last chapter part of the academic document, the last part of the paragraph unit, and the end of the natural paragraph;

[0159] Step 2, obtain the XML format data of the academic document, preprocess the XML format data, and use it as the text data to be annotated and analyzed;

[0160] Step 3, use the annotation and analysis conditions to annotate and analyze the text data to be annotated and analyzed, obtain the annotation result and the analysis result, and obtain the academic view snapshot according to the annotation result and the analysis result;

[0161] Among them, the annotation result includes: the position annotation of the academic view sentence in the paragraph unit and the position annotation of the paragraph unit in the full text of the text;

[0162] Among them, the analysis result includes: the category and content of the academic view;

[0163] Among them, the content of academic viewpoints includes: main conclusions, core viewpoints, arguments, topic sentences, and conclusions;

[0164] Among them, conclusions include: major conclusions obtained by analyzing the last chapter of academic literature, medium conclusions obtained by analyzing the last part of paragraph units, and minor conclusions obtained by analyzing the end of natural paragraphs.

[0165] At the hardware level, in order to effectively improve the flexibility, generality, and acquisition efficiency of data acquisition, the present application provides an embodiment of an electronic device for all or part of the content in an academic viewpoint annotation and analysis method. The electronic device specifically includes the following:

[0166] A processor, a memory, a communications interface, and a bus; wherein, the processor, the memory, and the communications interface complete mutual communication through the bus; the communications interface is used to implement information transmission between a data acquisition device based on a distributed aggregation model and related devices such as a core business system, a user terminal, and a related database, etc. The logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., and this embodiment is not limited thereto. In this embodiment, the logic controller can be implemented with reference to the embodiments of the academic viewpoint annotation and analysis method and the embodiments of the data acquisition device based on the distributed aggregation model, and the content thereof is incorporated herein, and the repeated parts will not be described again.

[0167] It can be understood that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.

[0168] In practical applications, part of the academic viewpoint annotation and analysis method can be executed on the electronic device side as described above, or all operations can be completed in the client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. The present application does not make any limitations thereto. If all operations are completed in the client device, the client device may further include a processor.

[0169] The above-mentioned client device may have a communication module (i.e., communication unit), which can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.

[0170] Embodiment 4:

[0171] An embodiment of the present application also provides a computer-readable storage medium that can implement all steps of the academic view annotation and analysis method with the execution subject being a server or a client in the above embodiments. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, all steps of the academic view annotation and analysis method with the execution subject being a server or a client in the above embodiments are implemented. For example, when the processor executes the computer program, the following steps are implemented:

[0172] Step 1: Obtain historical academic documents, classify the historical academic documents, and obtain the annotation and analysis conditions of academic views;

[0173] The annotation and analysis conditions include: the last chapter part of the academic document, the last part of the paragraph unit, and the end of the natural paragraph;

[0174] Step 2: Obtain the XML format data of the academic document, preprocess the XML format data, and use it as the text data to be annotated and analyzed;

[0175] Step 3: Use the annotation and analysis conditions to annotate and analyze the text data to be annotated and analyzed, obtain the annotation result and the analysis result, and obtain the academic view snapshot according to the annotation result and the analysis result;

[0176] Among them, the annotation result includes: the position annotation of the academic view sentence in the paragraph unit and the position annotation of the paragraph unit in the full text of the text;

[0177] Among them, the analysis result includes: the category and content of the academic view;

[0178] Among them, the content of the academic view includes: the main conclusion, the core view, the argument, the topic sentence, and the conclusion;

[0179] Among them, the conclusion includes: the major conclusion obtained by analyzing the last chapter part of the academic document, the medium conclusion obtained by analyzing the last part of the paragraph unit, and the minor conclusion obtained by analyzing the end of the natural paragraph.

[0180] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, apparatus, or computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0181] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (apparatuses), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0182] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0183] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0184] The above-described embodiments merely represent the preferred embodiments of the present application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications, improvements, and substitutions can be made, and these all fall within the protection scope of the present application.

Claims

1. A method for marking and analyzing academic opinions, characterized in that: The following steps are involved: Step 1: Retrieve and download historical academic documents that have been academically annotated and analyzed, classify the historical academic documents, and formulate the annotation and analysis conditions of academic viewpoints; The conditions for marking analysis include: the last chapter of academic literature, the last part of a paragraph unit, and the end of a natural paragraph; Step 2: receiving XML format data of the academic document to be annotated and analyzed, preprocessing the XML format data, and using it as text data to be annotated and analyzed; Step 3: Use the annotation analysis conditions to annotate and analyze the text data to be annotated and analyzed, obtain the annotation results and analysis results, and generate a snapshot of academic opinions based on the annotation results and analysis results; The annotation results include: the position annotation of the academic opinion sentence in the paragraph unit and the position annotation of the paragraph unit in the full text; Among them, the analysis results include: the categories and contents of academic opinions; Among them, the contents of academic viewpoints include: main conclusions, core ideas, arguments, topic sentences and conclusions; Among them, the conclusions include: the main conclusion obtained by analyzing the last chapter of the academic literature, the middle conclusion obtained by analyzing the last part of the paragraph unit, and the small conclusion obtained by analyzing the end of the natural paragraph.

2. According to the method for marking and analyzing academic opinions in claim 1, it is characterized in that: The formulation process of the annotation analysis conditions is as follows: Retrieve and download historical academic documents that have been academically annotated and analyzed, and conduct the first classification of academic viewpoints in historical academic documents; Determine the importance of each academic opinion category in the first opinion classification results, and define and express the academic opinions of historical academic documents and classify the opinions for the second time based on the importance judgment results; According to the definition statement and the results of the second opinion classification, conditions are formulated for the annotation and analysis of academic opinions to obtain the annotation and analysis conditions.

3. The method for marking and analyzing academic viewpoints according to claim 1, characterized in that: The pre-processing process is as follows: Perform data cleaning on XML format data of academic literature; Segment the text sentences, and then segment the paragraph units of each part; According to the IMRaD structure, the full text is divided into five parts: introduction, discussion, methods, results and conclusion. It is used to mark and analyze the position of academic opinion sentences in paragraphs or paragraph units, as well as the position of paragraphs or paragraph units in the full text.

4. The method for marking and analyzing academic viewpoints according to claim 1, characterized in that: The scholarly viewpoint snapshot also includes: metadata attributes; Among them, metadata attributes include: author, institution, year, title, and source journal.

5. An academic opinion annotation and analysis system, characterized in that: include: The text acquisition module is used to acquire XML format data of academic documents, pre-process the XML format data, and use it as text data to be annotated and analyzed; The condition formulation module is used to retrieve and download historical academic documents that have been academically annotated and analyzed, classify historical academic documents, and formulate the annotation and analysis conditions of academic viewpoints; The conditions for marking analysis include: the last chapter of academic literature, the last part of a paragraph unit, and the end of a natural paragraph; An annotation analysis module, used to annotate and analyze text data to be annotated and analyzed using annotation analysis conditions, obtain annotation results and analysis results, and generate academic opinion snapshots based on the annotation results and analysis results; The annotation results include: the position annotation of the academic opinion sentence in the paragraph unit and the position annotation of the paragraph unit in the full text; Among them, the analysis results include: the categories and contents of academic opinions; Among them, the contents of academic viewpoints include: main conclusions, core ideas, arguments, topic sentences and conclusions; Among them, the conclusions include: the main conclusion obtained by analyzing the last chapter of the academic literature, the middle conclusion obtained by analyzing the last part of the paragraph unit, and the small conclusion obtained by analyzing the end of the natural paragraph. The summary and generalization of the main conclusion has a higher priority than the summary and generalization of the middle conclusion, and the summary and generalization of the middle conclusion has a higher priority than the summary and generalization of the small conclusion.

6. The academic opinion annotation and analysis system according to claim 5, characterized in that: The scholarly viewpoint snapshot also includes: metadata attributes; Among them, metadata attributes include: author, institution, year, title, and source journal.

7. The academic opinion annotation and analysis system according to claim 5, characterized in that: The content of the academic opinion also includes comments, which are the author's comments on the opinions cited in the text. By identifying and marking the comments, the academic opinion sentences cited by the author in the text and the corresponding author's comments are identified and marked, and the identification and marking of the comments are completed.

8. The academic opinion annotation and analysis system according to claim 7, characterized in that: There are four forms of review expressions, and the specific forms of the four forms of review expressions are as follows: The first form of commentary is to put the quoted viewpoint first and the commentary content second, with the quoted viewpoint and the commentary sharing the same sentence; The second form of commentary is to put the quoted viewpoint first and the commentary content second, with the quoted viewpoint and commentary divided into multiple sentences; The third type of commentary expression is that there is only a quoted opinion, and there is a logical relationship between the sentence before the quoted opinion, the sentence after the quoted opinion, and the sentence after the quoted opinion. The fourth form of review is to have only review content, which is a paragraph unit summarizing existing research results. The paragraph unit only contains review content and no corresponding cited opinions.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the academic opinion annotation and analysis method described in claim 1 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the academic opinion annotation and analysis method described in claim 1 are implemented.

Citation Information

Patent Citations

  • Method, device and system for processing electronic book comments

    CN103631782A

  • Evaluation method and device based on argumentation text structure

    CN108595407A

  • Scientific and technical literature review automatic generation method and device

    CN118278365A

  • Scientific research auxiliary system and method based on large language model

    CN118940833A