A method for analyzing and scoring a dissertation

By using machine learning and multiple linear regression modeling to perform structural analysis and scoring of dissertations, this method solves the problem of assessing the structural quality of dissertations, achieves accurate and objective scoring and improves efficiency, and is applicable to the quality assessment of various types of dissertations.

CN119180274BActive Publication Date: 2026-02-10FUJIAN NORMAL UNIV
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202411231509.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-04
Publication Date
2026-02-10
Estimated Expiration
2044-09-04

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively assess and improve the structural quality of dissertations, making it difficult to measure student engagement, and manual grading is subject to subjective bias and inefficiency.

Method used

Machine learning and analytical methods were employed to conduct structural analysis and scoring of dissertations through multiple linear regression modeling and various algorithms, including multidimensional evaluation of text, images, and references, and a structural analysis and scoring model was constructed.

Benefits of technology

It achieves accurate feedback and objective scoring of thesis structure, reduces human scoring bias, significantly improves work efficiency, is applicable to the quality assessment of various thesis types, and supports the standardization and quality improvement of academic research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119180274B_ABST
    Figure CN119180274B_ABST
Patent Text Reader

Abstract

The application discloses a kind of structure analysis and scoring method of dissertation, it is related to educational evaluation technical field.The method comprises the following steps: collecting completed degree authorization dissertation;Obtain the attribute data and derivative data of the collected dissertation, constitute data set A;Data set A is carried out to be missing and characteristic scaling, and data training set B is constructed;Multiple linear regression modeling is carried out to data training set B, and dissertation structure analysis and scoring model M are obtained;The attribute data and derivative data of the measured dissertation are obtained, and constitute measured data set C;Measured data set C is carried out to be missing and characteristic scaling, and measured evaluation data set D is constructed;Measured evaluation data set D is calculated by applying dissertation structure analysis and scoring model M, and the structure score of measured dissertation is obtained, and each structure is given rationalization suggestion.The application realizes the structure score of measured dissertation and the rationalization suggestion of each structure by modeling to the completed dissertation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of educational evaluation technology, and in particular to a method for structural analysis and scoring of dissertations. Background Technology

[0002] A thesis is a research report or scientific paper written by an author to obtain a specific degree. It not only represents different academic levels but is also an important source of literature. Theses are typically not published in academic journals and are mainly obtained through degree-granting institutions, specific collections, and private channels. The thesis is a core component of higher education, bearing the important task of cultivating students' innovative abilities ("from 0 to 1"), problem-solving skills, critical thinking abilities, and collaborative communication skills. It is the most crucial link in realizing the transformation from "learning knowledge" to "developing strong abilities." Furthermore, the thesis is also an important observation point for higher education evaluation, and its quality reflects, to some extent, the overall quality of graduates. To this end, my country has specifically issued a series of standards to regulate theses, such as the Rules for Writing Theses (GB / T7713.1-2006), the Writing Formats for Scientific and Technological Reports, Theses, and Academic Papers (GB / T7713-1987), and the Rules for Bibliographic References (GB / T 7714-2005). In January 2014, the Ministry of Education issued the "Measures for Random Inspection of Doctoral and Master's Dissertations" (Degree

[2014] No. 5), and in January 2021, the Ministry of Education issued the "Measures for Random Inspection of Undergraduate Graduation Theses (Designs) (Trial Implementation)" (Education Supervision

[2020] No. 5). The random inspection of undergraduate graduation theses focuses on assessing the "qualification" of the topic selection, writing arrangement, logical construction, professional ability, and academic norms, with a particular emphasis on the undergraduate's basic academic norms and basic academic literacy. The intrinsic quality of a thesis includes academic innovation, the significance of the thesis, and the logicality of the thesis, while extrinsic quality includes formal structure and format standardization. The evaluation of a thesis is a complex, multi-dimensional, and multi-step process, typically involving the supervisor, peer scholars, and a committee composed of several experts, who conduct quality evaluations from aspects such as academic value, methodology, and innovation.

[0003] Standardization, as an external criterion for evaluating the quality of dissertations, leads professors to spend considerable time checking student papers for formatting, resulting in an emphasis on format over content. In October 2010, the journal *Computer Engineering and Design* published "Research and Application of Java-based Document Format Checking Technology," which researched the use of Java for dissertation format checking. This technology enables the detection and matching of dissertation formats with standard documents, generating detailed format check reports and improving the efficiency of dissertation format checking. Document CN110069785A discloses a component-based dissertation standardization control and analysis platform and system. This system can automatically process dissertations according to the format requirements of different universities, disciplines, and majors using component-based conventions, solving the problems of low efficiency, high arbitrariness, and high error rate associated with traditional manual editing. In 2016, the journal *Higher Education Exploration* (S-level) published an article titled "Construction of an Undergraduate Graduation Thesis Quality Evaluation System Based on Academic Norms." Addressing the lack of consideration for academic irregularities in previous evaluation studies, the article introduced academic norm indicators and constructed a three-tiered evaluation method for undergraduate theses: 1. Goal Layer (Highest Level): This refers to the predetermined goals of the problem; 2. Criterion Layer (Middle Level): This refers to the criteria influencing the achievement of the goals, comprising five indicators: thesis topic selection and review, basic knowledge and theoretical application, ability level, overall thesis quality, and academic norms; 3. Measures Layer (Lowest Level): This refers to the measures to promote the achievement of the goals, comprising 19 secondary indicators, including alignment with professional training direction, appropriateness and operability of the topic selection, and basic knowledge. The second issue of the *Journal of Harbin Vocational and Technical College* in 2019 published an article titled "Evaluation of Undergraduate Graduation Thesis Quality and Its Improvement Path," analyzing data on the thesis quality of 297 graduates from the School of Economics at Anhui University of Finance and Economics between 2007 and 2017. The article concluded that the main reason for the low quality of undergraduate graduation theses was the failure of general education to play its due role, and recommended: formulating scientific general education training objectives and rationally arranging general education courses; improving students' logical thinking ability through general education; and enhancing students' innovative ability through general education. The March 2019 issue of *Beijing Social Sciences* published an article titled "Standardization Assessment of Highly Cited Doctoral Dissertations in Higher Education in China," pointing out prominent problems such as dissertations with unbalanced structural proportions in the main body, insufficient awareness of problems in the literature review, poor research design and methodology, confusion between the significance and problems of the research, and a lack of research ethics norms. The article suggested strengthening academic norms education, further improving the curriculum of research methods, strengthening training in educational research methods, formulating writing standards for doctoral dissertations in higher education, and improving the mechanism for doctoral students' research ethics education.The 5th issue of *Heilongjiang Higher Education Research* in 2019 published "Identification of the Main Characteristics of Poorly Reviewed Graduate Theses," which selected blind reviewer opinions from 411 poorly reviewed graduate theses from S University's 2017 graduating class. The study found specific deficiencies in the content of theses, including thesis themes, logical structure, argumentation and analysis, conclusions and suggestions, writing attitude, writing norms, and innovativeness. These characteristics are closely related and overlap. The *Journal of Editing* in June 2018 published "A Discussion on the Methods of Reviewing and Processing Figures and Tables in Scientific Journal Papers," pointing out that figures and tables play an indispensable role in papers, possessing advantages such as large data (or phenomena) capacity, high accuracy, and strong comparativeness. The aim should be to present data (or phenomena) analysis results in a concise, intuitive, and aesthetically pleasing manner. The review and processing of figures and tables should be based on the needs of the paper's content, following a logical order, considering the appropriateness of the selected figure and table types, the scientific validity, standardization, and page layout efficiency.

[0004] The quality of dissertations is closely related to the time and effort students invest in their graduation projects, but measuring student engagement is difficult due to a lack of effective data. Document CN107239900A discloses a method for evaluating the quality of undergraduate dissertations based on an extension cloud model, including the following steps: 1) Establishing an undergraduate dissertation quality evaluation index system; 2) Determining the classification of level standards and the level cloud model: dividing the evaluation indicators into 5 levels and providing the corresponding cloud model; 3) Obtaining evaluation values; 4) Determining the combined weights of the indicators: applying the order relation analysis method and the entropy weight method to determine the subjective and objective weights of the indicators, and determining the combined weights of the indicators based on the principle of minimum discriminative information; 5) Determining expert weights: applying the grey relational analysis method to determine expert weights; 6) Establishing an extension cloud correlation matrix; 7) Determining the evaluation levels, making the evaluation of undergraduate dissertation quality more reasonable, scientific, and stable. Literature CN108122180A discloses a real-time generation method for autonomous learning engagement based on online learning behavior. It generates engagement parameter Es from teaching video playback behavior, video viewing duration, and concurrent learning behavior data, automating and simplifying the acquisition of online learning engagement data. This makes it possible to acquire engagement data in real time and conduct process analysis, thereby completely changing the current thinking and methods for analyzing online learning engagement and making data-supported guidance and analysis of the online learning process possible. Literature CN109858769A provides a method for evaluating the quality of undergraduate graduation project guidance. It collects the percentage grades obtained by students in each required course before the graduation project; normalizes the grades of each course using the average grade and standard deviation of all students in each course to obtain a relative grade; obtains the average grade and standard deviation of each student's relative grades in each course; normalizes the graduation project grade using the average grade and standard deviation of all students in the graduation project to obtain a relative graduation project grade; and compares each student's relative graduation project grade with the average relative grade and standard deviation of their relative grades in the required courses to obtain an evaluation index for the quality of undergraduate graduation project guidance. CN109045664B discloses a deep learning-based diving scoring method, server, and system. It trains a diving scoring model using a known diving video dataset and corresponding diving score dataset; inputs the diving videos into the trained diving scoring model and outputs diving scores; this avoids human interference in diving scores and improves the accuracy of diving scores.CN110298038A discloses a text scoring method and apparatus. The method includes: segmenting the text to be scored into words to obtain word segments; determining the semantic vectors and part-of-speech vectors of the segmented words to obtain word vectors composed of the semantic vectors and part-of-speech vectors; inputting the word vectors into a pre-trained sequence encoder to obtain the output of the sequence encoder as the sequence encoding vector of the text to be scored; inputting the word vectors into a pre-trained tree encoder to obtain the output of the tree encoder as the tree encoding vector of the text to be scored; fusing the sequence encoding vector and the tree encoding vector of the text to be scored to obtain a fused encoding vector of the text to be scored; and determining the credibility of the text to be scored as a specified type of text based on the fused encoding vector, which is used as the score of the text to be scored. This method can improve the accuracy of scoring. CN110069785A discloses a paper standardization control and analysis platform and system based on component conventions. According to the style conventions for components provided by universities and majors, it converts student (writer) input data into XML components with relevant styles and compresses them into text documents. Simultaneously, it performs real-time plagiarism checks on the student's writing process and records historical writing information. This solves the problems of low efficiency, high arbitrariness, and high error rate of traditional manual editing. Based on historical tracking and analysis, it can monitor plagiarism behavior of students (writers) in real time during the paper writing process, effectively controlling academic misconduct. Based on deep learning-based text analysis methods, it can uncover potential relationships between research trends and entities in the paper's field and track the overall progress of the paper work. CN110069768A provides an automatic scoring method for English argumentative essays based on text structure. The method includes an automatic text component identification module and an automatic text structure scoring module. It comprehensively considers the impact of text component identification results and global structural features between paragraphs on the scoring task, maximizing the effectiveness of text structure scoring for argumentative essays. To address the current lack of a complete structural system for argumentative essays, document CN112214988A discloses an argumentative essay structure analysis method based on a combination of deep learning and rules. This method can automatically analyze the structure of argumentative essays without manual processing, thus accelerating the analysis process and saving labor costs. Document CN108595407A discloses an evaluation method and apparatus based on the structure of argumentative essays. This method obtains multiple discourse elements of the essay by identifying paragraph and sentence types. It then constructs sequential, planar, or hierarchical features of the essay's structure using these elements. Finally, it obtains the evaluation result of the essay based on these sequential, planar, or hierarchical features using a pre-defined feature model.In the process of doctoral education, the dissertation is a comprehensive reflection of a doctoral student's research ability, innovation ability, knowledge mastery and application ability, and written expression ability. Its quality not only reflects the research ability and academic level of the doctoral applicant but also, to a certain extent, reflects the level of graduate education in the training institution. This paper aims to find a method to quantitatively and systematically analyze the influencing factors of doctoral dissertation quality and provide suggestions for improving the quality of doctoral training. Literature CN111027868A discloses a method for evaluating the influencing factors of dissertation quality based on structural equation modeling. It collects author data after completing the dissertation, calculates initial variables, uses structural equation modeling to calculate the model fit index, iteratively corrects the structural equation model, and obtains the evaluation results of each variable on the quality of the dissertation.

[0005] CN110287319B invented a student evaluation text analysis method based on sentiment analysis technology, which analyzes and processes student evaluation texts and can solve the defect of incorrect classification of suggestive comments. In July 2021, the *Journal of Xi'an Aeronautical University*, Volume 39, Issue 4, published "Analysis of Research Hotspots and Trends in Undergraduate Graduation Thesis Quality Evaluation Based on Scientific Knowledge Graphs." Using literature related to undergraduate graduation thesis quality evaluation indexed by CNKI from 2000 to 2020 as a sample, the paper uses knowledge graphs and bibliometric methods to review the current research status of undergraduate graduation thesis quality evaluation by domestic scholars. Chronologically, it is divided into two stages: before 2005, the number of publications was relatively small, and related research mainly summarized the authors' work or management experience; after 2005, the number of publications increased rapidly, and the depth and breadth of research expanded accordingly. First, scholars are employing methods from various disciplines, including statistics, systems engineering, and even computer science, in their research. Keywords such as AHP (Analytic Hierarchy Process), fuzzy comprehensive evaluation, and the Delphi method are frequently encountered. Second, the scope of evaluation applications is expanding. More and more scholars are selectively choosing evaluation indicators based on the characteristics of their teaching or management disciplines and majors, constructing more suitable and targeted evaluation indicator systems. On this basis, they select mature evaluation methods and conduct further evaluations. In the coming years, in terms of research depth, scholars will focus more on the quantification and subdivision of undergraduate thesis quality evaluation indicators, making the methods and conclusions of quality evaluation more scientific, reasonable, applicable, and operable. At the same time, relevant theoretical knowledge of computer science and topology will be more deeply integrated into research related to undergraduate thesis quality evaluation. On the other hand, in terms of research breadth, in conjunction with the 14th Five-Year Plan, building a high-quality education system and standardizing undergraduate thesis quality evaluation methods are the overall research direction and trend for undergraduate thesis quality evaluation in the coming years.

[0006] For thesis plagiarism detection, for example, CNKI offers a series of products including the CNKI Undergraduate Thesis Detection System, the Postgraduate Thesis Academic Misconduct Detection System, and the CNKI Undergraduate Graduation Project (Thesis) Management System. The plagiarism detection of theses by relevant higher-level departments and university functional departments often relies on third-party providers such as Wanfang's search and plagiarism detection services. In June 2019, the journal *Research on Higher Financial Education* published an article titled "Plagiarism Rate, Supervisors, and the Quality of Undergraduate Graduation Theses," which pointed out that after the implementation of the plagiarism detection mechanism, students' main energy was focused on revising to reduce the plagiarism rate and pass the detection, making originally concise language verbose and original fluent sentences incoherent. Such revisions are meaningless and waste the energy of both students and teachers. Although universities have adopted measures such as undergraduate mentorship systems and the writing of term papers in recent years, the implementation results have been unsatisfactory. The September 2019 issue of *Computer Applications Research* published an article titled "A Comprehensive Method for Paper Plagiarism Detection and Evaluation." Addressing the overly extreme nature of internet-based plagiarism detection systems, the article proposes a novel, comprehensive method. This method detects anomalies and then requires manual review. The goal is to reduce false positives, allowing papers that were initially flagged as plagiarized to have a chance for review. When the same paper shows significantly different similarity rates on different plagiarism detection websites, this new comprehensive method incorporates human judgment into the final plagiarism assessment, reducing the influence of website-controlled similarity rates. This dual website / human hybrid detection compensates for the impact of website database issues on plagiarism results, improving the accuracy and reliability of the plagiarism detection results. The 2020 issue of the *Journal of Tianjin Normal University (Social Sciences Edition)*, Volume 2, published an article titled "Universities Urgently Need Methods for Detecting Knowledge Integrity to Address the Side Effects of Plagiarism Detection," which pointed out that plagiarism—the dishonest behavior of knowledge—targeted by plagiarism detection has not been eliminated; instead, it has added side effects that contradict the fundamental purpose of universities in cultivating morality and talent. In recent years, with the support of the internet, students' information application abilities have significantly improved, and their ability to acquire resources has been greatly enhanced. Universities' management systems for graduation theses have become increasingly standardized and diverse, such as self-checking and external checks. However, degree theses have not shown a corresponding improvement trend; instead, problems such as logical inconsistencies, weak writing, and empty content have gradually increased, with obvious instances of fabricated data, fabricated analyses and discussions, and fabricated citations.

[0007] Writing a dissertation is undoubtedly a challenging and personality-shaping process. In recent years, the rapid development of artificial intelligence (AI) technology has brought disruptive changes to academia. AI tools such as ChatGPT, Midjourney, Sora, and Suno have garnered significant attention for their exceptional capabilities in processing complex text and visual information. They not only perform translation and proofreading tasks more efficiently and accurately, but also facilitate content creation, generating high-quality articles, news reports, and even novels. These technologies have transformed the way information and knowledge are produced, becoming a crucial engine of social change and reshaping the interaction between humans and technology. The convergence of AI technology and dissertations will have a profound impact on both students and faculty, leading to a better understanding and fulfillment of their teaching needs, and freeing up more time and energy for both to focus on innovative research. The "Teachers' Digital Literacy" standard (JY / T0646—2022), released on November 30, 2022, includes digital application as one of the five primary dimensions of the teachers' digital literacy framework. Digital application is further subdivided into four secondary dimensions: digital instructional design, digital instructional implementation, digital academic assessment, and digital collaborative education. Specifically, digital instructional implementation requires: utilizing digital technology resources to conduct individualized guidance, being able to use digital technology resources to identify student learning differences and provide targeted guidance; digital academic assessment requires: ① selecting and using assessment data collection tools, being able to reasonably select and use digital tools to collect multimodal academic assessment data; ② applying data analysis models to conduct academic data analysis, being able to select and apply appropriate data analysis models to conduct academic data analysis; ③ achieving visualization and interpretation of academic data, being able to use digital tools to visualize and interpret the results of academic data analysis.

[0008] In terms of research depth, scholars will focus more on the quantification and subdivision of undergraduate thesis quality evaluation indicators to make the evaluation methods and conclusions more scientific and reasonable, thereby enhancing their applicability and operability. Simultaneously, relevant theoretical knowledge from computer science and topology will be more deeply integrated into undergraduate thesis quality evaluation research. On the other hand, in terms of research breadth, constructing a high-quality education system and standardizing undergraduate thesis quality evaluation methods are the overall research direction and trend for undergraduate thesis quality evaluation in the coming years.

[0009] The intrinsic value of a thesis lies in the rigor of its logical structure, the efficiency of information transmission, and the breadth and depth of its content. A current challenge is whether to provide faculty and students with an efficient thesis structure analysis tool that analyzes and scores theses based on their numerical characteristics (such as the total number and distribution of words, the number and distribution of citations and figures, paragraph structure, etc.) and their implicit characteristic factors.

[0010] Therefore, proposing a structural analysis and scoring method for dissertations to address the difficulties of existing technologies is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0011] In view of this, the present invention provides a method for structural analysis and scoring of dissertations. By performing machine learning and analysis on dissertations, a structural analysis and evaluation model for dissertations can be established, thereby realizing the structural analysis and evaluation of dissertations and promoting the improvement of the structural quality of dissertations.

[0012] To achieve the above objectives, the present invention adopts the following technical solution:

[0013] A method for structural analysis and scoring of dissertations, comprising the following steps:

[0014] S1. Collect dissertations that have completed degree authorization;

[0015] S2. Obtain the attribute data and derived data of the collected dissertations to form dataset A;

[0016] S3. Perform gap detection and feature scaling on dataset A to construct training set B;

[0017] S4. Perform multiple linear regression modeling on the training data set B to obtain the thesis structure analysis and scoring model M.

[0018] S5. Obtain the attribute data and derived data of the thesis to be tested, and form the dataset C to be tested;

[0019] S6. Perform missing data inspection and feature scaling on the dataset C to be tested, and construct the evaluation dataset D to be tested;

[0020] S7. Apply the thesis structure analysis and scoring model M to calculate the evaluation dataset D, obtain the thesis structure score, and provide rational suggestions for each structure.

[0021] Optionally, the above method may include a thesis data table, Table_Dissertation, which consists of the following fields: ID, Thesis Type, Enrollment Time, Completion Time, Degree Granting Institution, Major Field, Author Information, Supervisor, Chinese Abstract Word Count, English Abstract Word Count, Number of Chapters, Word Count and Distribution, References and Distribution, Figures and Distribution, Tables and Distribution, Acknowledgments, Thesis Grade, Third-Party Rating, Chinese Abstract Rating, English Abstract Rating, Acknowledgments Rating, Reference Citation Reasonableness Rating, and Figure Rating.

[0022] The attribute data of the thesis is stored in the thesis data table Table_Dissertation, which contains the following fields: thesis type, enrollment time, completion time, degree-granting institution, major field, author information, supervisor, Chinese abstract word count, English abstract word count, number of chapters, word count and distribution, references and distribution, figures and distribution, tables and distribution, acknowledgments, thesis grade, and third-party rating.

[0023] The derived data of the thesis are stored in the thesis data table Table_Dissertation, which includes the following fields: Chinese abstract score, English abstract score, acknowledgments score, reference citation rationality score, and graph score.

[0024] The fields for word count and distribution, references and distribution, figures and distribution, and tables and distribution are composed of strings consisting of numbers, the @ symbol, and the period symbol. The specific format is: number1@number2.number3.number4, where number1 represents the number of words / references / figures / tables, @ represents the position, number2 represents the first-level heading (chapter), number3 represents the second-level heading (section), and number4 represents the third-level heading (item). Multiple strings are separated by commas (,).

[0025] Optionally, in the above method, in S2, the attribute data and derived data of the collected dissertations are obtained to form dataset A, wherein...

[0026] The collected attribute data for dissertations include: dissertation type, enrollment time, completion time, degree-granting institution, major field, author information, supervisor, Chinese abstract word count, English abstract word count, number of chapters, word count and distribution, references and distribution, figures and distribution, tables and distribution, acknowledgments, dissertation grade, and third-party rating;

[0027] The collected data derived from dissertations include: Chinese abstract scores, English abstract scores, acknowledgments scores, reference citation validity scores, and figure scores;

[0028] The collected attribute data and derived data of dissertations constitute the dissertation dataset A.

[0029] Optionally, the above method can be used to obtain the derived data of the collected dissertations. Specifically, algorithm analysis is performed on the Chinese abstract, English abstract, acknowledgments, and references of the collected dissertations. Algorithm analysis is also performed on each figure in the collected dissertations to obtain a score, and the scores are combined to obtain the derived data of the dissertations.

[0030] Optionally, the scoring of the Chinese abstract and the English abstract are obtained by performing text analysis on the collected Chinese and English abstracts of the dissertations using algorithms, including: support vector machine, Naive Bayes, K-means clustering and neural network.

[0031] The acknowledgments score is obtained by analyzing the text of the acknowledgments in the collected dissertations using an algorithm. The algorithm used is the Python third-party library SnowNLP.

[0032] The reference citation rationality score is calculated by applying text analysis algorithms to score the reference citations of each collected dissertation, and then summing the scores to obtain the average value. The text analysis algorithms include: text similarity algorithm, semantic analysis algorithm, and graph algorithm.

[0033] The image scoring is achieved by applying image analysis algorithms to score each image in the collected dissertations one by one, and then summing the scores to calculate the average value. The image analysis algorithms include image quality analysis, image content understanding, and image correlation analysis.

[0034] As can be seen from the above technical solution, compared with the prior art, the present invention provides a method for structural analysis and scoring of dissertations, which has the following beneficial effects: 1) It applies multiple algorithms and technologies to perform multi-dimensional analysis of the dissertation structure, including text, images, references, etc., which is comprehensive and systematic; 2) It provides accurate feedback on the dissertation structure based on the analysis results, helping authors to identify and improve the shortcomings of the dissertation, thereby improving the quality of the dissertation structure; 3) It adopts automated analysis algorithms, which significantly reduces the subjective bias of manual scoring and ensures the objectivity and consistency of the scoring results; 4) It greatly improves work efficiency, especially when processing a large number of dissertations, significantly reducing the time and labor intensity of manual review; 5) This method has good scalability and can be adjusted according to the needs of different disciplines, and is applicable to the quality assessment of various types of dissertations; 6) The present invention not only significantly improves the scientificity and efficiency of dissertation review, but also provides strong support for the standardization and quality improvement of academic research. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0036] Figure 1 A flowchart of a method for structural analysis and scoring of dissertations provided by this invention. Detailed Implementation

[0037] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0038] In this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0039] Reference Figure 1 As shown, this invention discloses a method for structural analysis and scoring of dissertations, including the following steps:

[0040] S1. Collect dissertations that have completed degree authorization;

[0041] S2. Obtain the attribute data and derived data of the collected dissertations to form dataset A;

[0042] S3. Perform missing value detection (e.g., deleting missing values) and feature scaling (e.g., standardization, normalization, maximum absolute value scaling) algorithms on dataset A to construct data training set B;

[0043] S4. Model the training data B using a multiple linear regression algorithm (such as stepwise regression or Bayesian linear regression) to obtain the thesis structure analysis and scoring model M.

[0044] S5. Obtain the attribute data and derived data of the thesis to be tested, and form the dataset C to be tested;

[0045] S6. The dataset C to be tested is processed using the same missing detection and feature scaling algorithm as in step S3 to construct the evaluation dataset D to be tested.

[0046] S7. Apply the thesis structure analysis and scoring model M to calculate the evaluation dataset D, obtain the thesis structure score, and provide rational suggestions for each structure.

[0047] Furthermore, as shown in Table 1, there exists a thesis data table Table_Dissertation, which consists of the following fields: ID, Thesis Type, Enrollment Time, Completion Time, Degree Granting Institution, Major Field, Author Information, Supervisor, Chinese Abstract Word Count, English Abstract Word Count, Number of Chapters, Word Count and Distribution, References and Distribution, Graphs and Distribution, Tables and Distribution, Acknowledgments, Thesis Grade, Third-Party Rating, Chinese Abstract Rating, English Abstract Rating, Acknowledgments Rating, Reference Citation Reasonableness Rating, and Graph Rating.

[0048] Furthermore, the attribute data of the thesis is stored in the thesis data table Table_Dissertation, which contains the following fields: thesis type, enrollment time, completion time, degree-granting institution, major field, author information, supervisor, Chinese abstract word count, English abstract word count, number of chapters, word count and distribution, references and distribution, figures and distribution, tables and distribution, acknowledgments, thesis grade, and third-party rating.

[0049] The derived data of the thesis are stored in the thesis data table Table_Dissertation, which includes the following fields: Chinese abstract score, English abstract score, acknowledgments score, reference citation rationality score, and graph score.

[0050] The fields for word count and distribution, references and distribution, figures and distribution, and tables and distribution are composed of strings consisting of numbers, the @ symbol, and the period symbol. The specific format is: number1@number2.number3.number4, where number1 represents the number of words / references / figures / tables, @ represents the position, number2 represents the first-level heading (chapter), number3 represents the second-level heading (section), and number4 represents the third-level heading (item). Multiple strings are separated by commas (,).

[0051] Table 1. Dissertation Data Table

[0052]

[0053]

[0054] Furthermore, in S2, the attribute data and derived data of the collected dissertations are obtained to form dataset A.

[0055] Furthermore, the collected attribute data for dissertations include: dissertation type, enrollment time, completion time, degree-granting institution, field of study, author information, supervisor, Chinese abstract word count, English abstract word count, number of chapters, word count and distribution, references and distribution, figures and distribution, tables and distribution, acknowledgments, dissertation grade, and third-party rating;

[0056] The collected data derived from dissertations include: Chinese abstract scores, English abstract scores, acknowledgments scores, reference citation validity scores, and figure scores;

[0057] The collected attribute data and derived data of dissertations constitute the dissertation dataset A.

[0058] Furthermore, the collected thesis-derived data is obtained. Specifically, algorithmic analysis is performed on the Chinese abstracts, English abstracts, acknowledgments, and references of the collected theses. Algorithmic analysis is also performed on each figure in the collected theses to obtain a score, and the scores are combined to obtain the thesis-derived data.

[0059] Furthermore, the Chinese abstract score and the English abstract score are obtained by using algorithms to analyze the text of the collected dissertations' Chinese and English abstracts, respectively. The algorithms use common text analysis and natural language processing algorithms and tools, including: support vector machine, Naive Bayes, K-means clustering and neural network.

[0060] The acknowledgments score is obtained by analyzing the text of the acknowledgments in the collected dissertations using an algorithm. The algorithm uses common text analysis and natural language processing algorithms and tools, such as the Python third-party library SnowNLP.

[0061] The reference citation rationality score is calculated by applying text analysis algorithms to score the reference citations of each collected dissertation, and then summing the scores to obtain the average value. The text analysis algorithms use common text analysis and natural language processing algorithms and tools, including: text similarity algorithms (cosine similarity, word frequency-inverse document frequency), semantic analysis algorithms (word vector model, BERT), and graph algorithms (citation network analysis PageRank).

[0062] The image scoring is achieved by applying image analysis algorithms to score each image in the collected dissertations one by one, and then summing the scores to calculate the average. The image analysis algorithms use common image processing algorithms and tools, including image quality analysis (signal-to-noise ratio, contrast assessment, sharpness assessment, image compression rate), image content understanding (edge ​​detection, YOLO), and image correlation analysis (ResNet, VGG).

[0063] Furthermore, S5, acquire the attribute data and derived data of the thesis to be tested, forming the dataset C to be tested. Further, the attribute data of the thesis to be tested includes: thesis type, enrollment time, completion time, degree-granting institution, major field, author information, supervisor, Chinese abstract word count, English abstract word count, number of chapters, word count and distribution, references and distribution, figures and distribution, tables and distribution, acknowledgments, thesis grade, and third-party rating.

[0064] The derived data of the dissertation to be tested include: Chinese abstract score, English abstract score, acknowledgments score, reference citation rationality score, and figure score;

[0065] The attribute data and derived data of the dissertation to be tested constitute the dataset C.

[0066] Furthermore, the derived data of the thesis to be tested is obtained. Specifically, the Chinese abstract, English abstract, acknowledgments, and references of the thesis to be tested are analyzed by algorithms, and each figure in the thesis to be tested is analyzed by algorithms to obtain a score. The average score is calculated to obtain the derived data of the thesis to be tested.

[0067] The scores for the Chinese abstract and the English abstract are obtained by using algorithms such as support vector machine, naive Bayes, K-means clustering, and neural network to perform text analysis on the Chinese and English abstracts of the thesis to be tested, respectively.

[0068] The acknowledgments score is obtained by using algorithms such as the Python third-party library SnowNLP to analyze the text of the acknowledgments section of the thesis under test.

[0069] The reference citation rationality score is calculated by applying text analysis algorithms to score the reference citations of each collected dissertation, and then summing the scores to obtain the average value. The text analysis algorithms include: text similarity algorithms (cosine similarity, word frequency-inverse document frequency), semantic analysis algorithms (word vector model, BERT), and graph algorithms (citation network analysis PageRank).

[0070] The image scoring is achieved by applying image analysis algorithms to score each image in the collected dissertations one by one, and then summing the scores to calculate the average value. The image analysis algorithms include image quality analysis (signal-to-noise ratio, contrast assessment, sharpness assessment, image compression rate), image content understanding (edge ​​detection, YOLO), and image correlation analysis (ResNet, VGG).

[0071] In one specific embodiment, taking four collected dissertations to be tested as an example,

[0072] Collect dissertations that have completed degree authorization (self-review, external review, and defense); extract attribute data from these dissertations, analyze to obtain derived data, and construct a dissertation dataset A; perform gap checking and feature scaling on dissertation dataset A to ensure consistent importance of all features, and construct a dissertation training set B; perform multiple linear regression modeling on dissertation training set B to obtain a dissertation structure analysis and scoring model; conduct structure analysis and scoring of dissertations to be tested: extract attribute data from the dissertations to be tested, calculate and analyze to obtain derived data, and form a dissertation dataset C; perform gap checking and feature scaling on dissertation dataset C to construct a dissertation evaluation dataset D; apply the dissertation structure analysis and scoring model to calculate the dissertation evaluation dataset D, obtain the dissertation structure score, and provide rationalization suggestions for each structure; apply text analysis algorithms to score the Chinese abstract section; apply text analysis algorithms to score the English abstract section; apply text sentiment analysis algorithms to score the acknowledgments section; apply text analysis model algorithms to score the rationality of reference citations; apply graph analysis algorithms to score the quality of graphs. The following table shows the attribute data of four dissertations to be tested and the derived data obtained through calculation:

[0073] Attribute data and derived data tables of 4 dissertations to be tested

[0074]

[0075]

[0076]

[0077] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0078] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for structural analysis and scoring of dissertations, characterized in that, Includes the following steps: S1. Collect dissertations that have completed degree authorization; S2. Obtain the attribute data and derived data of the collected dissertations to form dataset A; The attribute data of a thesis includes: thesis type, enrollment time, completion time, degree-granting institution, major field, author information, supervisor, Chinese abstract word count, English abstract word count, number of chapters, word count and distribution, references and distribution, figures and distribution, tables and distribution, acknowledgments, thesis grade, and third-party rating; Algorithm analysis was performed on the Chinese abstracts, English abstracts, acknowledgments, and references of the collected dissertations. Algorithm analysis was also performed on each figure in the collected dissertations to obtain a score. The scores were then combined to obtain the dissertation-derived data. Derivative data for dissertations include: Chinese abstract score, English abstract score, acknowledgments score, reference citation validity score, and figure score; The attribute data and derived data of the thesis constitute the thesis dataset A; S3. Perform gap detection and feature scaling on dataset A to construct training set B; S4. Perform multiple linear regression modeling on the training data set B to obtain the thesis structure analysis and scoring model M. The scores for the Chinese and English abstracts are obtained by analyzing the Chinese and English abstracts of the collected dissertations using algorithms, including Support Vector Machine, Naive Bayes, K-means clustering, and neural networks. The acknowledgments score is obtained by analyzing the text of the acknowledgments in the collected dissertations using an algorithm. The algorithm used is the Python third-party library SnowNLP. The reference citation rationality score is calculated by applying text analysis algorithms to score the reference citations of each collected dissertation, and then summing the scores to obtain the average value. The text analysis algorithms include: text similarity algorithm, semantic analysis algorithm, and graph algorithm. The image scoring is achieved by applying image analysis algorithms to score each image in the collected dissertations one by one, and then summing up the scores to calculate the average value. The image analysis algorithms include image quality analysis, image content understanding, and image correlation analysis. S5. Obtain the attribute data and derived data of the thesis to be tested, and form the dataset C to be tested; S6. Perform gap detection and feature scaling on the dataset C to be tested, and construct the evaluation dataset D to be tested; S7. Apply the thesis structure analysis and scoring model M to calculate the evaluation dataset D, obtain the thesis structure score, and provide rational suggestions for each structure.

2. The structural analysis and scoring method for a dissertation according to claim 1, characterized in that, There exists a thesis data table Table_Dissertation, which consists of the following fields: ID, Thesis Type, Enrollment Time, Completion Time, Degree Granting Institution, Major Field, Author Information, Supervisor, Chinese Abstract Word Count, English Abstract Word Count, Number of Chapters, Word Count and Distribution, References and Distribution, Figures and Distribution, Tables and Distribution, Acknowledgments, Thesis Grade, Third-Party Rating, Chinese Abstract Rating, English Abstract Rating, Acknowledgments Rating, Reference Citation Reasonableness Rating, and Figure Rating. The attribute data of the thesis is stored in the thesis data table Table_Dissertation, which contains the following fields: thesis type, enrollment time, completion time, degree-granting institution, major field, author information, supervisor, Chinese abstract word count, English abstract word count, number of chapters, word count and distribution, references and distribution, figures and distribution, tables and distribution, acknowledgments, thesis grade, and third-party rating. The derived data of the thesis are stored in the thesis data table Table_Dissertation, which includes the following fields: Chinese abstract score, English abstract score, acknowledgments score, reference citation rationality score, and graph score. The fields for word count and distribution, references and distribution, figures and distribution, and tables and distribution are strings composed of numbers, the @ symbol, and the period ".". The specific format is: number1@number2.number3.number4, where number1 represents the number of words / references / figures / tables, @ represents the position, number2 represents the first-level heading "chapter", number3 represents the second-level heading "section", and number4 represents the third-level heading "item". Multiple strings are separated by commas ".".

Citation Information

Patent Citations

  • Extension cloud model-based undergraduate thesis quality evaluation method

    CN107239900A

  • Autonomous learning concentration degree real-time generating method based on online learning behavior

    CN108122180A

  • Deep learning-based diving scoring method, server, and system

    CN109045664B

  • Method for evaluating undergraduate thesis project guidance quality

    CN109858769A

  • Chapter structure-based English argumentation automatic scoring method

    CN110069768A