A method and system for marking and analyzing academic opinions

By annotating and analyzing academic literature and generating snapshots of academic opinions, the problems of high computing resource consumption and low efficiency in existing technologies are solved, and efficient identification and analysis of academic opinions are achieved, improving recognition accuracy and reading experience.

CN120068800BActive Publication Date: 2025-09-12BEIJING BZD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510153270.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-09-12
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

Existing technologies consume a lot of computing resources and take a long time to identify and mine academic opinions in academic literature, which puts a heavy burden on readers and lacks efficient annotation and analysis methods.

Method used

By retrieving and downloading historical academic documents, formulating annotation and analysis conditions, and using XML format data for preprocessing, academic opinions are annotated and analyzed, and a snapshot of academic opinions is generated, including location annotation and content analysis, to identify the main conclusions, core ideas, arguments, topic sentences and conclusions.

Benefits of technology

It improves the efficiency and accuracy of identifying valuable opinions in academic literature, reduces computing resource consumption, provides a better reading experience, and lays a foundation for identifying plagiarism of opinions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068800B_ABST
    Figure CN120068800B_ABST
Patent Text Reader

Abstract

The present application discloses a method and system for annotating and analyzing academic opinions, belonging to the technical field of academic opinion recognition. The method and system include obtaining historical academic documents, classifying the historical academic documents, obtaining annotation and analysis conditions for academic opinions, obtaining XML format data of the academic documents, pre-processing the XML format data and treating it as text data to be annotated and analyzed, annotating and analyzing the text data to be annotated and analyzed using the annotation and analysis conditions, obtaining annotation results and analysis results, and obtaining academic opinion snapshots based on the annotation and analysis results. The method can improve processing efficiency and recognition accuracy when searching for valuable academic opinions in a large number of academic documents, and lay a foundation and preparatory work for the identification of opinion plagiarism based on the academic opinion snapshots.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of academic opinion identification, and specifically relates to a method and system for marking and analyzing academic opinions. Background Art

[0002] The objects of academic opinion mining generally include: user reviews on social platforms and e-commerce platforms; news opinion sentences in Internet news; subjective comments on various things posted by users on blogs and BBS forums, with subjective comments as the objects of opinion extraction; opinions expressed by experts and scholars on social hot topics on the Internet. There is a lack of annotation, analysis and mining of academic opinions in academic literature.

[0003] Many current methods for extracting academic opinions from academic documents are based on word analysis. For example, they identify opinion sentences based on feature analysis of sentiment words, polarity words, negative words, and phrases. When the number of annotations is limited or no manual annotation is required, a self-training model is constructed to learn patterns from the data. The model can accurately identify and mine sentences containing academic opinions. However, when applied to academic documents with a large amount of academic text content, a large amount of computing resources is required to efficiently identify sentences containing academic opinions from academic texts. This is time-consuming and labor-intensive to find valuable academic opinions, and places a huge reading burden on readers. Summary of the Invention

[0004] To solve the above problems and technical defects, the present application embodiment adopts the following technical solution, a method for marking and analyzing academic opinions, comprising the following steps:

[0005] Step 1: Search and download historical academic documents that have been annotated and analyzed, classify them, and formulate annotation and analysis conditions for academic viewpoints;

[0006] The marking analysis conditions include: the last chapter of academic documents, the last part of paragraph units and the end of natural paragraphs;

[0007] Step 2: receiving XML format data of the academic document to be annotated and analyzed, pre-processing the XML format data, and using it as text data to be annotated and analyzed;

[0008] Step 3: Use the annotation analysis conditions to annotate and analyze the text data to be annotated and analyzed, obtain the annotation results and analysis results, and generate an academic opinion snapshot based on the annotation results and analysis results;

[0009] The annotation results include: the position annotation of the academic opinion sentence in the paragraph unit and the position annotation of the paragraph unit in the full text;

[0010] Among them, the analysis results include: the categories and contents of academic opinions;

[0011] Among them, the content of academic viewpoints includes: main conclusions, core ideas, arguments, topic sentences and conclusions;

[0012] Among them, the conclusions include: the main conclusions obtained by analyzing the last chapter of the academic literature, the middle conclusions obtained by analyzing the last part of the paragraph unit, and the small conclusions obtained by analyzing the end of the natural paragraph.

[0013] Preferably, the process of formulating the labeling analysis conditions is as follows:

[0014] Retrieve and download historical academic documents that have been academically annotated and analyzed, and conduct a first-level classification of academic viewpoints in historical academic documents;

[0015] Determine the importance of each academic opinion category in the first opinion classification results, and define and express the academic opinions of historical academic documents and conduct a second opinion classification based on the importance judgment results;

[0016] According to the definition statement and the second opinion classification results, conditions are formulated for the annotation and analysis of academic opinions to obtain the annotation and analysis conditions.

[0017] Perform data cleaning on XML format data of academic literature;

[0018] Segment the text sentences, and then segment the paragraph units of each part;

[0019] According to the IMRaD structure, the full text is divided into five parts: introduction, discussion, methods, results and conclusion. It is used to mark and analyze the position of academic opinion sentences in paragraphs or paragraph units, as well as the position of paragraphs or paragraph units in the full text.

[0020] Furthermore, the academic opinion snapshot also includes: metadata attributes;

[0021] Among them, metadata attributes include: author, institution, year, title, and source journal.

[0022] An academic opinion annotation and analysis system, comprising:

[0023] The text acquisition module is used to obtain XML format data of academic documents, pre-process the XML format data, and use it as text data to be annotated and analyzed;

[0024] The condition formulation module is used to obtain historical academic documents, classify historical academic documents, and obtain the annotation and analysis conditions of academic viewpoints;

[0025] The marking analysis conditions include: the last chapter of academic documents, the last part of paragraph units and the end of natural paragraphs;

[0026] An annotation analysis module is used to annotate and analyze text data to be annotated using annotation analysis conditions, obtain annotation results and analysis results, and obtain a snapshot of academic opinions based on the annotation results and analysis results;

[0027] The annotation results include: the position annotation of the academic opinion sentence in the paragraph unit and the position annotation of the paragraph unit in the full text;

[0028] Among them, the analysis results include: the categories and contents of academic opinions;

[0029] Among them, the content of academic viewpoints includes: main conclusions, core ideas, arguments, topic sentences and conclusions;

[0030] Among them, the conclusions include: the main conclusions obtained by analyzing the last chapter of the academic literature, the middle conclusions obtained by analyzing the last part of the paragraph unit, and the small conclusions obtained by analyzing the end of the natural paragraph.

[0031] Furthermore, the academic opinion snapshot also includes: metadata attributes;

[0032] Among them, metadata attributes include: author, institution, year, title, and source journal;

[0033] The summary and generalization of the major conclusion takes precedence over the summary and generalization of the medium conclusion;

[0034] The summary and generalization of the middle conclusion takes precedence over the summary and generalization of the small conclusion.

[0035] Preferably, the conclusion also includes comments, which are the author's comments on the viewpoints cited in the text. The comments need to be identified and annotated. The identification and annotation of the comments are completed by identifying and annotating the academic viewpoint sentences cited by the author in the text and the corresponding author comments.

[0036] Furthermore, the review has four review expression forms;

[0037] The first form of commentary is to put the quoted viewpoint first and the commentary content second, with the quoted viewpoint and the commentary sharing the same sentence.

[0038] The second form of commentary is to put the quoted viewpoint first and the commentary content second, with the quoted viewpoint and commentary divided into multiple sentences;

[0039] The third type of commentary is a statement with no commentary content but only quoted opinions. There is a logical relationship between the sentence before the quoted opinion, the sentence after the quoted opinion, and the sentence after the quoted opinion.

[0040] The fourth form of review is to have no cited opinions, only review content. The review content is a paragraph unit that summarizes existing research results. The paragraph unit only contains review content and no corresponding cited opinions.

[0041] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned academic opinion annotation and analysis method are implemented.

[0042] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the academic opinion annotation and analysis method described above.

[0043] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0044] This application clarifies the categories of various academic viewpoints in academic documents, formulates conditions for the annotation and analysis of academic viewpoints according to their importance, annotates and analyzes the text data to be annotated and analyzed according to the annotation and analysis conditions, and generates a snapshot of academic viewpoints. This can improve the document processing efficiency and recognition accuracy when searching for valuable academic viewpoints in a large number of academic documents, and lays the foundation and preparatory work for the identification of plagiarism of viewpoints based on the academic viewpoint snapshots, so that readers have a better reading experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In the attached figure:

[0046] Figure 1 A schematic diagram of the method steps of an embodiment of the present application;

[0047] Figure 2 This is a schematic diagram of the method flow of an embodiment of the present application;

[0048] Figure 3 A schematic diagram of the system structure of an embodiment of the present application;

[0049] Figure 4 This is a schematic diagram of the device structure of an embodiment of the present application. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.

[0051] Example 1:

[0052] like Figure 1 and Figure 2 As shown, a method for marking and analyzing academic opinions includes the following steps:

[0053] Retrieve and download historical academic documents that have been annotated and analyzed, classify them, and formulate annotation and analysis conditions for academic viewpoints;

[0054] The marking analysis conditions include: the last chapter of academic documents, the last part of paragraph units and the end of natural paragraphs;

[0055] The process of formulating the annotation analysis conditions is as follows:

[0056] Obtain historical academic documents and conduct the first classification of academic viewpoints in the historical academic documents;

[0057] Determine the importance of each academic opinion category in the first opinion classification results, and define and express the academic opinions of historical academic documents and conduct a second opinion classification based on the importance judgment results;

[0058] According to the definition and the second opinion classification results, conditions are formulated for the annotation and analysis of academic opinions to obtain the annotation and analysis conditions;

[0059] The annotation results include: the position annotation of the academic opinion sentence in the paragraph unit and the position annotation of the paragraph unit in the full text;

[0060] Among them, the analysis results include: the categories and contents of academic opinions;

[0061] Among them, the content of academic viewpoints includes: main conclusions, core ideas, arguments, topic sentences and conclusions;

[0062] Among them, the conclusions include: the main conclusions obtained by analyzing the last chapter of the academic literature, the middle conclusions obtained by analyzing the last part of the paragraph unit, and the small conclusions obtained by analyzing the end of the natural paragraph.

[0063] Receive XML format data of academic documents to be annotated and analyzed, pre-process the XML format data, and use it as text data to be annotated and analyzed;

[0064] The preprocessing process is as follows:

[0065] Perform data cleaning on XML format data of academic literature;

[0066] Segment the text sentences, and then segment the paragraph units of each part;

[0067] According to the IMRaD structure, the full text is divided into five parts: introduction, discussion, methods, results and conclusion. It is used to mark and analyze the position of academic opinion sentences in paragraphs or paragraph units, as well as the position of paragraphs or paragraph units in the full text.

[0068] Use the annotation analysis conditions to annotate and analyze the text data to be annotated and analyzed, obtain the annotation results and analysis results, and generate a snapshot of academic opinions based on the annotation results and analysis results;

[0069] The Scholarly Viewpoint Snapshot also includes: metadata attributes;

[0070] Among them, metadata attributes include: author, institution, year, title, and source journal.

[0071] Academic opinions are the core content of academic literature and an important part of academic information exchange. They are the views and understandings held by the opinion holders on the issues, things or phenomena being studied in their own research fields. The formation of academic opinions is the result of scholars' thinking activities and plays an important role in their scientific research activities and cognitive structure. For the same research problem, different scholars or schools will form completely different academic opinions due to differences in discourse systems, research perspectives, and prior knowledge.

[0072] Given the importance of academic viewpoints in academic exchanges, and in order to protect their originality and expose academic misconduct, my country's press and publishing industry standard, "Definition of Academic Misconduct in Academic Publishing Standards and Journals," summarizes the manifestations of misconduct such as "plagiarism" into the following five categories:

[0073] a) Directly using arguments, viewpoints, conclusions, etc. from others' published literature without citation;

[0074] b) Paraphrasing others' arguments, viewpoints, conclusions, etc. without changing their original meaning and using them without citing them;

[0075] c) Using others' arguments, viewpoints, conclusions, etc. without citing them after deleting parts of them;

[0076] d) Disassemble or reorganize others' arguments, viewpoints, conclusions, etc. and use them without citation;

[0077] e) Adding some content to other people's arguments, opinions, conclusions, etc. and using them without citation.

[0078] From the above-mentioned manifestations of "plagiarism of ideas", it can be seen that the essential content of the main object of plagiarism of ideas is the author's original "arguments, viewpoints, conclusions", etc.

[0079] Therefore, the scope of academic views includes the author's original arguments, viewpoints, conclusions, etc. in academic literature.

[0080] In order to accurately identify and discover the arguments, viewpoints and conclusions in academic literature, based on the analysis of the chapter structure of a large number of academic documents, the categories of the original summary opinion sentences of the authors in the academic literature are expanded on the basis of the arguments, viewpoints and conclusions. The expanded categories include: main conclusions and core viewpoints, arguments and theme sentences, and detailed elaboration of conclusions.

[0081] At the same time, according to the importance of academic viewpoints, academic viewpoints are classified at the chapter and paragraph levels, various academic viewpoints are defined, expressed and distinguished, and conditions are formulated for the annotation and analysis of academic viewpoints:

[0082] The main conclusions and core viewpoints are the author's most basic views on the research issues discussed in the academic literature. They are the most important ideas and viewpoints put forward in the text. They are also a high-level summary and concentration of all the small conclusions and sub-arguments in the text. The main conclusions and core viewpoints belong to the author's original viewpoints at the level of the entire text, and are generally located in the "Conclusion" part of the text.

[0083] The thesis and topic sentence are a topic or hypothesis raised about a thing or phenomenon, and then they are expanded based on this topic or hypothesis to find evidence to support this topic or hypothesis. The thesis and topic sentence are the author's original opinion sentences at the paragraph unit level, and are generally located at the beginning of the paragraph unit. The sub-arguments are several ideas and viewpoints that are subordinate to the central argument and serve to explain the central argument.

[0084] Conclusion is the conclusive content deduced from facts and logic based on investigation and logical analysis of things or phenomena. Therefore, it has a certain degree of authority, rigor and practicality. Conclusions can be divided into major conclusions, medium conclusions and minor conclusions.

[0085] When the last chapter of the paper is the conclusion, it is a summary and generalization of the entire paper, that is, the "main conclusion" mentioned above, which belongs to the conclusion of the entire paper;

[0086] The small conclusion is the result of writing according to the standard paragraph. When a small conclusion is written at the end of a paragraph, it is a summary of the paragraph and belongs to the conclusion at the paragraph level.

[0087] When the last part of the paragraph unit is the middle conclusion, it is a summary and generalization of the paragraph unit, and the summary and generalization are between the big conclusion and the small conclusion;

[0088] That is, the summary and generalization of the small conclusion is less than the summary and generalization of the middle conclusion, which is less than the summary and generalization of the main conclusion.

[0089] In addition to the above-mentioned opinion-type sentences, there is another type of opinion-type sentence called "Commentary", which is the author's commentary on the opinions cited in the text. By marking and identifying the academic opinion-type sentences cited by the author in the text and pairing them with the corresponding author's commentary content, the commentary content is an opinion-type sentence in academic literature that is equally important as the above-mentioned types, and is also one of the key objects of research, marking and mining.

[0090] For example, in text example 1:

[0091] Identifying entity boundaries is difficult and has always been a challenge in named entity recognition. In recent years, deep learning methods with deep neural network structures have continued to develop, showing better results than traditional methods on many named entity recognition tasks [6-7]. However, as a specific type of named entity, there are currently few research results related to theoretical term extraction. The extraction of theoretical terms is different from other types of named entity recognition tasks (such as name recognition and place name recognition). Currently, there is no large-scale mature corpus. In addition, the complex structure and cross-disciplinary distribution of theoretical terms also make their extraction more difficult. Therefore, how to apply deep learning methods to theoretical term extraction to improve extraction performance and reduce the cost of manual annotation of training sets is a problem that needs to be studied.

[0092] Text example 1 can clearly express four different types of opinion sentences: "argument / theme, quoted opinion, author's comments on the quoted opinion, and small conclusion"

[0093] The thesis / topic sentence of text example 1 is "The identification of entity boundaries is relatively difficult, which has always been a difficulty in named entity recognition." It is located at the beginning of the paragraph and is the thesis for the content of this paragraph, and is also the topic sentence of this paragraph.

[0094] The opinion quote sentence is "In recent years, deep learning methods with deep neural network structures have continued to develop and have shown better results than traditional methods in many named entity recognition tasks [6-7]." This sentence follows the thesis and topic sentence, and the author uses the opinions of other authors as evidence;

[0095] The author's comment on the cited opinion is "However, as a specific type of named entity, there are currently few relevant research results on theoretical term extraction; the extraction of theoretical terms is different from other types of named entity recognition tasks (such as personal name recognition, place name recognition, etc.), and there is currently no large-scale mature corpus. Moreover, the complex structure and cross-disciplinary distribution of theoretical terms also make their extraction more difficult." These two "comments" follow the author's quoting of other authors' opinions. They are the author's comments on the cited opinions and are also the author's original personal opinions.

[0096] The last sentence is the author's summary of this paragraph, which reads "Therefore, how to apply deep learning methods to theoretical terminology extraction, improve extraction performance and reduce the cost of manual annotation of training sets is a problem that needs to be studied at present." Because it is only a conclusion statement of this paragraph, it is marked as "small conclusion"; when this sentence is a summary statement made by the author on the content of a paragraph unit, it is marked as "middle conclusion"; when this sentence is a summary statement made by the author on the content of the entire text, it is marked as "big conclusion", that is, "main conclusion" or "core point of view".

[0097] For example, in text example 2:

[0098] in conclusion:

[0099] The extraction of theoretical terms is the basis for large-scale literature content analysis and in-depth revelation of interdisciplinary knowledge transfer. This paper proposes a theoretical term extraction method based on deep learning. Taking the currently widely used Bi-LSTM-CRF deep learning model as the basic framework, this paper constructs a theoretical term extraction model and studies the feature construction and annotation method of theoretical terms in this model. It also conducts an experimental comparative analysis of the extraction performance of the model under different feature training. The experimental results show the effectiveness of the theoretical term tail word features, context features, word formation features, negative example features, word vector features, and part-of-speech features constructed in this paper for theoretical term extraction. The model incorporating the above features is significantly superior to the traditional method at different training set ratios, and has good versatility and Scalability; This paper also proposes a self-training algorithm that enables the model to perform weakly supervised learning on unlabeled data, further improving the extraction performance while reducing the cost of manual labeling; However, there are also problems with some theoretical terms that cannot be recalled and extracted incorrectly. The next step will be to study the construction of character and word-level convolutional layers in the model so that the model can more effectively learn the extraction features of theoretical terms, and study the application of collaborative training algorithms in model self-training to weaken the influence of noise; In addition, how to reasonably select incremental self-training batches and automatically adjust thresholds in an unsupervised manner is also a problem that needs to be further studied and solved in the future; It can be said that this paper not only provides a more effective general method for theoretical term extraction, but also provides a method reference for other types of named entity recognition research.

[0100] Text example 2 is marked as "main conclusion", which means "main conclusion" or "core point".

[0101] Commentary generally takes four forms in academic literature:

[0102] The first form of commentary: quoted opinion + comment = one sentence, with the quoted opinion coming first and the comment coming second, combining the two into one in one sentence.

[0103] For example, in text example three:

[0104] Term citation can effectively measure interdisciplinary studies[5]. Analyzing the innovation and borrowing of theoretical terms can help reveal the evolution and cross-penetration of core knowledge in different disciplines.

[0105] In this example sentence, the cited viewpoint is “Term citation can effectively measure interdisciplinary studies”;

[0106] The commentary states that "the innovation and borrowing of analytical theoretical terms helps reveal the evolution and cross-penetration of core knowledge in different disciplines."

[0107] The second form of commentary: the quoted opinion comes first, followed by the commentary, each one a sentence, for a total of two or three sentences.

[0108] For example, in Text Example 4:

[0109] In recent years, deep learning methods with deep neural network structures have continued to develop and have shown better results than traditional methods in many named entity recognition tasks [6-7]. As a specific type of named entity, there are currently few research results on theoretical term extraction. The extraction of theoretical terms is different from other types of named entity recognition tasks (such as person name recognition, place name recognition, etc.). There is currently no large-scale mature corpus, and the complex structure and cross-disciplinary distribution of theoretical terms also make their extraction more difficult.

[0110] In this example, the cited viewpoint is "In recent years, deep learning methods with deep neural network structures have continued to develop and have shown better results than traditional methods in many named entity recognition tasks [6-7]."

[0111] The first comment in text example 4 is “As a specific type of named entity, there are currently few relevant research results on theoretical term extraction”;

[0112] The second comment in text example 4 is, "The extraction of theoretical terms is different from other types of named entity recognition tasks (such as person name recognition, place name recognition, etc.). There is currently no large-scale mature corpus. In addition, the complex structure and cross-disciplinary distribution of theoretical terms also make their extraction more difficult."

[0113] The third form of commentary is no commentary. Although the author cites the academic opinions of other authors, there is no commentary content. This situation is also common in academic literature. Therefore, it is necessary to clarify that "cited opinions" and "commentary" do not necessarily appear in the text at the same time. Based on this, it can avoid marking and identifying commentary content that does not correspond to the cited opinions.

[0114] For example, in Text Example 5:

[0115] When the hidden layer and output layer structures of the model remain unchanged, the input features of the input layer have an important impact on the learning performance of the model [6,12-13,19].

[0116] The entire sentence in this example is a quoted opinion. The first sentence of the example sentence is "From the above analysis, we can see that the Bi-LSTM-CRF model learns the deep features of the data from the characteristics of the input sequence dataset, represents this type of data, and thus performs sequence labeling prediction." The second sentence of the example sentence is "This paper intends to use word embedding, part of speech (POS), and entity (ENT) features of theoretical terms as input to the Bi-LSTM-CRF model, and study the impact of each feature on the effectiveness of theoretical term extraction, in order to achieve a more effective theoretical term extraction model."

[0117] From the relationship between the quoted opinion sentence, the sentence before the quoted opinion sentence and the sentence after the quoted opinion sentence, it can be seen that the above example sentence is a quoted opinion sentence, but there is no corresponding commentary content.

[0118] Therefore, although this sentence is a quoted opinion, it can be seen from its relationship with the previous sentence that the more important nature of this sentence is similar to a commentary sentence. It is both a summary sentence of the previous sentence and a commentary on an opinion sentence. In other words, sometimes, a summary sentence can be regarded as an opinion sentence, and a quoted opinion sentence can be regarded as a commentary sentence, all depending on the nature of its content.

[0119] Therefore, although the next sentence follows the quoted opinion sentence, it is not marked as "comment" because its nature is not the author's comment on the quoted opinion.

[0120] The fourth form of commentary: Due to the special nature of the paragraph unit, for example, in a paragraph unit that reviews the research results of predecessors, there may be only comments but no corresponding cited opinions.

[0121] For example, in Text Example 6:

[0122] 2. Related research:

[0123] 2.1、Theoretical term extraction:

[0124] Scholars have conducted some research on the extraction of theoretical terms. Wang Fang et al. [2] used manual annotation methods to identify 586 theories in 1,822 articles published in the Journal of the Chinese Society for Information Science from 2000 to 2013. Chen Feng et al. [3] further used a CRF-based machine learning method to study the automatic extraction of theoretical terms in academic journals. They conducted extraction experiments on the titles and abstracts of the above 1,822 papers. When the proportion of the training set increased to 80%, a good accuracy rate (93.33%) was achieved, but the recall rate was still not high (56%). In addition, term extraction, a method similar to theoretical extraction, has also made progress. Hua Bolin [8] used a rule-based method to study the extraction of intelligence method terms from Chinese academic literature. The full-text data of 46 papers in the Journal of the Chinese Society for Information Science, No. 1 to No. 4, 2013, were used for experimental testing. The recall rate (94%) and accuracy rate (92%) were both high.

[0125] The above studies all use manual rules or traditional machine learning methods to extract data from one discipline and one document type. Their versatility, scalability and machine learning performance still need to be greatly improved.

[0126] The special nature of the chapters in Text Example 6 leads to the special nature of the paragraphs. "2. Related Research" is a review, so the paragraphs and the relationship between sentences in this chapter need to pay attention to their special nature. This section is "2.1. Theoretical Terminology Extraction". The entire first paragraph quotes the research results of predecessors, but there is no opinion, so it is not marked. The second paragraph, which is the last paragraph of this section, should be considered to be marked as the author's commentary on the above-mentioned quoted content.

[0127] Example 2:

[0128] like Figure 3 As shown, an academic opinion annotation and analysis system includes:

[0129] The text acquisition module is used to obtain XML format data of academic documents, pre-process the XML format data, and use it as text data to be annotated and analyzed;

[0130] The preprocessing process is as follows:

[0131] Perform data cleaning on XML format data of academic literature;

[0132] According to the IMRaD structure, the full text is divided into five parts: introduction, discussion, methods, results and conclusions.

[0133] Segment the paragraph units of each part, and then segment the text sentences of each paragraph unit.

[0134] The condition formulation module is used to obtain historical academic documents, classify historical academic documents, and obtain the annotation and analysis conditions of academic viewpoints;

[0135] The marking analysis conditions include: the last chapter of academic documents, the last part of paragraph units and the end of natural paragraphs;

[0136] The process of formulating the annotation analysis conditions is as follows:

[0137] Obtain historical academic documents and conduct the first classification of academic viewpoints in the historical academic documents;

[0138] Determine the importance of each academic opinion category in the first opinion classification results, and define and express the academic opinions of historical academic documents and conduct a second opinion classification based on the importance judgment results;

[0139] According to the definition statement and the second opinion classification results, conditions are formulated for the annotation and analysis of academic opinions to obtain the annotation and analysis conditions.

[0140] An annotation analysis module is used to annotate and analyze text data to be annotated using annotation analysis conditions, obtain annotation results and analysis results, and obtain a snapshot of academic opinions based on the annotation results and analysis results;

[0141] The annotation results include: the position annotation of the academic opinion sentence in the paragraph unit and the position annotation of the paragraph unit in the full text;

[0142] Among them, the analysis results include: the categories and contents of academic opinions;

[0143] Among them, the content of academic viewpoints includes: main conclusions, core ideas, arguments, topic sentences and conclusions;

[0144] Among them, the conclusions include: the main conclusions obtained by analyzing the last chapter of the academic literature, the middle conclusions obtained by analyzing the last part of the paragraph unit, and the small conclusions obtained by analyzing the end of the natural paragraph.

[0145] The Scholarly Viewpoint Snapshot also includes: metadata attributes;

[0146] Among them, metadata attributes include: author, institution, year, title, and source journal;

[0147] The summary and generalization of the major conclusion takes precedence over the summary and generalization of the medium conclusion;

[0148] The summary and generalization of the middle conclusion takes precedence over the summary and generalization of the small conclusion.

[0149] The conclusion also includes comments, which are the author's comments on the views cited in the text. The comments need to be identified and annotated. The identification and annotation of the comments are completed by identifying and annotating the academic viewpoint sentences cited by the author in the text and the corresponding author comments.

[0150] There are four types of commentary expressions;

[0151] The first form of commentary is to put the quoted viewpoint first and the commentary content second, with the quoted viewpoint and the commentary sharing the same sentence.

[0152] The second form of commentary is to put the quoted viewpoint first and the commentary content second, with the quoted viewpoint and commentary divided into multiple sentences;

[0153] The third type of commentary is a statement with no commentary content but only quoted opinions. There is a logical relationship between the sentence before the quoted opinion, the sentence after the quoted opinion, and the sentence after the quoted opinion.

[0154] The fourth form of review is to have no cited opinions, only review content. The review content is a paragraph unit that summarizes existing research results. The paragraph unit only contains review content and no corresponding cited opinions.

[0155] Example 3:

[0156] like Figure 4 As shown, from a hardware perspective, this application provides an embodiment of an electronic device that includes all or part of the content of the academic opinion annotation and analysis method, and the electronic device specifically includes the following content:

[0157] Step 1: Obtain historical academic documents, classify them, and obtain the annotation and analysis conditions of academic viewpoints;

[0158] The marking analysis conditions include: the last chapter of academic documents, the last part of paragraph units and the end of natural paragraphs;

[0159] Step 2: Obtain XML format data of academic documents, pre-process the XML format data, and use it as text data to be annotated and analyzed;

[0160] Step 3: Use the annotation analysis conditions to annotate and analyze the text data to be annotated and analyzed, obtain the annotation results and analysis results, and obtain a snapshot of academic opinions based on the annotation results and analysis results;

[0161] The annotation results include: the position annotation of the academic opinion sentence in the paragraph unit and the position annotation of the paragraph unit in the full text;

[0162] Among them, the analysis results include: the categories and contents of academic opinions;

[0163] Among them, the content of academic viewpoints includes: main conclusions, core ideas, arguments, topic sentences and conclusions;

[0164] Among them, the conclusions include: the main conclusions obtained by analyzing the last chapter of the academic literature, the middle conclusions obtained by analyzing the last part of the paragraph unit, and the small conclusions obtained by analyzing the end of the natural paragraph.

[0165] From a hardware perspective, in order to effectively improve the flexibility, versatility, and efficiency of data collection, this application provides an embodiment of an electronic device that implements all or part of the content of the academic opinion annotation and analysis method. The electronic device specifically includes the following content:

[0166] A processor, a memory, a communication interface, and a bus; wherein the processor, the memory, and the communication interface communicate with each other via the bus; the communication interface is used to implement information transmission between a data acquisition device based on a distributed model and related devices such as a core business system, a user terminal, and related databases; the logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., but this embodiment is not limited thereto. In this embodiment, the logic controller can be implemented with reference to the embodiment of the academic opinion annotation and analysis method in the embodiment, as well as the embodiment of the data acquisition device based on the distributed model, the contents of which are incorporated herein and repeated parts are not repeated.

[0167] It is understandable that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.

[0168] In practical applications, portions of the academic opinion annotation and analysis method may be executed on the electronic device side as described above, or all operations may be performed on the client device. The specific selection may be based on the processing capabilities of the client device and the limitations of the user's usage scenario, and this application does not impose any restrictions on this. If all operations are performed on the client device, the client device may also include a processor.

[0169] The above-mentioned client device may have a communication module (i.e., a communication unit), which can communicate with a remote server to realize data transmission with the server. The server may include a server on the task scheduling center side, and other implementation scenarios may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster consisting of multiple servers, or a server structure of a distributed device.

[0170] Example 4:

[0171] The embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the academic opinion annotation and analysis method in the above-mentioned embodiment, where the execution subject is a server or a client. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the computer program implements all steps of the academic opinion annotation and analysis method in the above-mentioned embodiment, where the execution subject is a server or a client. For example, when the processor executes the computer program, the following steps are implemented:

[0172] Step 1: Obtain historical academic documents, classify them, and obtain the annotation and analysis conditions of academic viewpoints;

[0173] The marking analysis conditions include: the last chapter of academic documents, the last part of paragraph units and the end of natural paragraphs;

[0174] Step 2: Obtain XML format data of academic documents, pre-process the XML format data, and use it as text data to be annotated and analyzed;

[0175] Step 3: Use the annotation analysis conditions to annotate and analyze the text data to be annotated and analyzed, obtain the annotation results and analysis results, and obtain a snapshot of academic opinions based on the annotation results and analysis results;

[0176] The annotation results include: the position annotation of the academic opinion sentence in the paragraph unit and the position annotation of the paragraph unit in the full text;

[0177] Among them, the analysis results include: the categories and contents of academic opinions;

[0178] Among them, the content of academic viewpoints includes: main conclusions, core ideas, arguments, topic sentences and conclusions;

[0179] Among them, the conclusions include: the main conclusions obtained by analyzing the last chapter of the academic literature, the middle conclusions obtained by analyzing the last part of the paragraph unit, and the small conclusions obtained by analyzing the end of the natural paragraph.

[0180] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, apparatuses, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0181] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (apparatus), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0182] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0183] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0184] The above-described embodiments merely represent preferred implementations of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person skilled in the art would be able to make numerous variations, improvements, and substitutions without departing from the spirit of the present application, all of which fall within the scope of protection of the present application.

Claims

1. A method for marking and analyzing academic opinions, characterized in that: The following steps are involved: Step 1: Search and download historical academic documents that have been annotated and analyzed, classify them, and formulate annotation and analysis conditions for academic viewpoints; The marking analysis conditions include: the last chapter of academic documents, the last part of paragraph units and the end of natural paragraphs; Step 2: receiving XML format data of the academic document to be annotated and analyzed, pre-processing the XML format data, and using it as text data to be annotated and analyzed; Step 3: Use the annotation analysis conditions to annotate and analyze the text data to be annotated and analyzed, obtain the annotation results and analysis results, and generate an academic opinion snapshot based on the annotation results and analysis results; The annotation results include: the position annotation of the academic opinion sentence in the paragraph unit and the position annotation of the paragraph unit in the full text; Among them, the analysis results include: the categories and contents of academic opinions; Among them, the content of academic viewpoints includes: core ideas, arguments, topic sentences and conclusions; The conclusions include: the main conclusions obtained by analyzing the last chapter of the academic literature, the middle conclusions obtained by analyzing the last part of the paragraph unit, and the small conclusions obtained by analyzing the end of the natural paragraph; The content of academic opinions also includes comments. Comments are the author's comments on the opinions cited in the article. By identifying and marking the comments, the academic opinion sentences cited by the author in the article and the corresponding author comments are identified and marked to complete the identification and marking of comments. There are four forms of commentary, and the specific forms of the four forms of commentary are as follows: The first form of commentary is to put the quoted viewpoint first and the commentary content second, with the quoted viewpoint and the commentary sharing the same sentence. The second form of commentary is to put the quoted viewpoint first and the commentary content second, with the quoted viewpoint and commentary divided into multiple sentences; The third form of commentary is to only quote opinions, and there is a logical relationship between the sentence before the quoted opinion, the sentence after the quoted opinion, and the sentence after the quoted opinion. The fourth form of review is to have only review content, which is a paragraph unit summarizing existing research results. The paragraph unit only contains review content and no corresponding cited opinions.

2. The method for marking and analyzing academic opinions according to claim 1, characterized in that: The process of formulating the labeling analysis conditions is as follows: Retrieve and download historical academic documents that have been academically annotated and analyzed, and conduct a first-level classification of academic viewpoints in historical academic documents; Determine the importance of each academic opinion category in the first opinion classification results, and define and express the academic opinions of historical academic documents and conduct a second opinion classification based on the importance judgment results; According to the definition statement and the second opinion classification results, conditions are formulated for the annotation and analysis of academic opinions to obtain the annotation and analysis conditions.

3. The method for marking and analyzing academic opinions according to claim 1, characterized in that: The pre-processing process is as follows: Perform data cleaning on XML format data of academic literature; Segment the text sentences, and then segment the paragraph units of each part; According to the IMRaD structure, the full text is divided into five parts: introduction, discussion, methods, results and conclusion. It is used to mark and analyze the position of academic opinion sentences in paragraphs or paragraph units, as well as the position of paragraphs or paragraph units in the full text.

4. The method for marking and analyzing academic opinions according to claim 1, characterized in that: The scholarly viewpoint snapshot also includes: metadata attributes; Among them, metadata attributes include: author, institution, year, title, and source journal.

5. An academic opinion annotation and analysis system, characterized by: include: The text acquisition module is used to obtain XML format data of academic documents, pre-process the XML format data, and use it as text data to be annotated and analyzed; The condition formulation module is used to retrieve and download historical academic documents that have been academically annotated and analyzed, classify historical academic documents, and formulate the annotation and analysis conditions of academic viewpoints; The marking analysis conditions include: the last chapter of academic documents, the last part of paragraph units and the end of natural paragraphs; An annotation analysis module is used to annotate and analyze text data to be annotated using annotation analysis conditions, obtain annotation results and analysis results, and generate academic opinion snapshots based on the annotation results and analysis results; The annotation results include: the position annotation of the academic opinion sentence in the paragraph unit and the position annotation of the paragraph unit in the full text; Among them, the analysis results include: the categories and contents of academic opinions; Among them, the content of academic viewpoints includes: core ideas, arguments, topic sentences and conclusions; The conclusions include: the main conclusion obtained by analyzing the last chapter of the academic literature, the middle conclusion obtained by analyzing the last part of the paragraph unit, and the small conclusion obtained by analyzing the end of the natural paragraph. The summary and generalization of the main conclusion takes precedence over the summary and generalization of the middle conclusion, and the summary and generalization of the middle conclusion takes precedence over the summary and generalization of the small conclusion. The content of academic opinions also includes comments. Comments are the author's comments on the opinions cited in the article. By identifying and marking the comments, the academic opinion sentences cited by the author in the article and the corresponding author comments are identified and marked to complete the identification and marking of comments. There are four forms of commentary, and the specific forms of the four forms of commentary are as follows: The first form of commentary is to put the quoted viewpoint first and the commentary content second, with the quoted viewpoint and the commentary sharing the same sentence. The second form of commentary is to put the quoted viewpoint first and the commentary content second, with the quoted viewpoint and commentary divided into multiple sentences; The third form of commentary is to only quote opinions, and there is a logical relationship between the sentence before the quoted opinion, the sentence after the quoted opinion, and the sentence after the quoted opinion. The fourth form of review is to have only review content, which is a paragraph unit summarizing existing research results. The paragraph unit only contains review content and no corresponding cited opinions.

6. The academic opinion annotation and analysis system according to claim 5, characterized in that: The scholarly viewpoint snapshot also includes: metadata attributes; Among them, metadata attributes include: author, institution, year, title, and source journal.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the academic opinion annotation and analysis method according to claim 1 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the academic opinion annotation and analysis method according to claim 1 are implemented.

Citation Information

Patent Citations

  • Scientific and technical literature review automatic generation method and device

    CN118278365A

  • Scientific research auxiliary system and method based on large language model

    CN118940833A