A method and system for intelligent examination of contents of graduation defense slides

By intelligently checking the slides used in university graduation defenses, the problems of fragmented logical expression, unclear focus, and insufficient interactivity in existing technologies have been solved. This has resulted in slide content that is logically coherent and has an appropriate information density, thereby improving the efficiency and quality of the defense.

CN120747995BActive Publication Date: 2025-11-28WEIKE ZHIJIAN (FOSHAN) TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511150228.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-28
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

In existing technologies, the PPT documents for university graduation defenses lack in-depth guidance, resulting in fragmented logical expression, blurred key points, information overload, and insufficient interactivity. It is difficult to clearly convey the core innovation points and argumentation of the research within a limited time, and general academic templates cannot solve the differentiated needs of disciplines.

Method used

This paper provides an intelligent method and system for checking the content of college graduation defense slides. By processing the slide documents in a structured manner, it performs multiple checks such as topic coverage, information density, logical coherence, relevance, and readability, and generates a detection report, providing optimization suggestions to help users improve the slide content.

Benefits of technology

It improved the quality of the slide content and the efficiency of the defense, helped students identify and solve problems, ensured that the slide content was logically coherent, had an appropriate information density, highlighted key points, and enhanced the interactivity and readability of the defense.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747995B_ABST
    Figure CN120747995B_ABST
Patent Text Reader

Abstract

The application discloses a kind of high school graduation defense slide content intelligent inspection method and system, it is related to intelligent detection technical field, including steps: slide document structured processing;For the text in each slide, divide title text and non-title text;Execute the following inspection items, obtain the inspection result or calculation result corresponding to each inspection item: check whether the content of each slide covers preset theme;Slide and preset theme mapping association;Calculate the information density score of non-title text of slide;Calculate the logical coherence score of non-title text in slide, judge whether logical coherence;Calculate the degree of cutting theme of non-title text and preset theme / title text in slide, judge whether text is cutting theme;Calculate the readability score of text in slide;Check whether there is forbidden word in the text of slide, and statistics forbidden word usage degree;Generation inspection report.Finally assist student to perfect / improve the quality of slide.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent detection, in particular to a method and system for intelligently checking the content of a graduation presentation slide in a university. BACKGROUND

[0002] Graduation thesis and defense are the core evaluation links of talent training in universities, and they constitute a closed-loop system of "written research" and "oral examination". Graduation thesis requires students to complete systematic academic research through steps such as topic selection, literature review, method design and argumentation, and its content needs to meet academic standards and be logically rigorous; the defense is a dynamic evaluation of the graduation thesis, which shows the PPT document, and combines the PPT on-site presentation and the question and answer test to test the students' understanding depth, logical consistency and innovation value of the research content. However, the existing teaching system pays more attention to guiding students in writing papers, but lacks guidance on graduation thesis defense PPT documents. Students usually make deletional modifications based on the structure and content of the thesis and directly fill them into the PPT document. However, due to the lack of in-depth understanding of the structure of the PPT document, the quality of the final displayed document content is poor, mainly existing the following problems:

[0003] (1) Logical expression is fragmented: there is an essential difference between the linear narrative of the thesis and the visualization and modularization requirements of the defense PPT. The PPT generated by simply deleting the content of the thesis by students often presents "large text stacking" or "logical jumping", making it difficult to clearly convey the core innovation points and argumentation context of the research within a limited time.

[0004] (2) Key points are fuzzy and information is overloaded: the defense PPT needs to highlight the research problem, method innovation, experimental verification and conclusion value, but under the existing way, students often fall into the "all-inclusive" trap, over-retaining details of the thesis (such as literature review, data preprocessing, etc.), which weakens the presentation of key achievements, making it difficult for the defense committee to quickly focus on the core contributions.

[0005] (3) Lack of interactivity and question and answer prediction: the defense PPT is not only a one-way display tool, but also needs to provide logical support for the question and answer session. The existing PPT making method lacks the design of potential problem prediction (such as visualization of key conclusions and comparison of controversial data), which makes students react passively in the question and answer session and makes it difficult to systematically respond to the questions.

[0006] In addition, although there are more general PPT templates circulating on the network, students can fill in information according to the preset panels. However, the general academic templates are not deeply adapted to the subject characteristics and research types (such as empirical research / theoretical research), and cannot solve the differences in content logic and expression form between different disciplines. SUMMARY

[0007] In view of the problems in the prior art, the application provides a high school graduation defense slide content intelligent inspection method and system, which can detect problems in the content of a defense slide document, generate a detection report and provide optimization suggestions, and finally assist users in finding and solving problems in the slide document.

[0008] The technical scheme of the application is implemented as follows:

[0009] A high school graduation defense slide content intelligent inspection method comprises the following steps:

[0010] S1, structurally process a slide document; the slide document comprises a plurality of slides; for the text in each slide, divide the text into title text and non-title text respectively.

[0011] S2, perform the following inspection items to obtain the inspection results or calculation results corresponding to each inspection item:

[0012] S2-1, check whether the content of each slide covers a preset theme; the preset theme comprises a plurality of themes; map and associate the slide and the preset theme one by one; display the preset theme that is not covered; the preset theme is, for example, "table of contents", "background", "meaning", "content scheme", "technical route", "algorithm design" and the like, which is set in the system by a user or a developer in advance, and whether all the titles and main content of the entire slide document contain the set theme is judged during detection; the mapping processing record is performed on the associated slide and preset theme;

[0013] Further, the preset theme has a plurality of sets, which correspond to different fields or majors, such as computer majors and mathematics majors, and theme settings of PPT documents of legal majors are different; before starting the system inspection, a user selects a set of preset themes in advance.

[0014] S2-2, calculate the information density score of the non-title text of the slide;

[0015] S2-3, calculate the logical coherence score of the non-title text in the slide; whether the logical coherence score indicates logical coherence is determined;

[0016] S2-4, calculate the topic-cutting degree of the non-title text and the title text in the slide; whether the text cuts the topic is determined according to the topic-cutting degree; the topic-cutting degree refers to the degree of fit between the text content and the topic.

[0017] S2-5, calculate the readability score of the text in the slide;

[0018] S2-6, check whether there is a forbidden word in a preset forbidden word table in the text of the slide, and count the use degree of the forbidden word;

[0019] S3, generating an inspection report showing the inspection results and the calculation results.

[0020] The inspection report, i.e., the presentation slide, is associated with a preset theme, a presentation information density score, a logical coherence score, and whether the text is logically coherent, on topic, and readable.

[0021] Further, optimization suggestions are generated based on the inspection results and the calculation results.

[0022] As a further optimization of the above scheme, the inspection items further include:

[0023] S2-7, identifying whether the number of text words in the slide meets a preset standard;

[0024] S2-8, identifying whether the slide contains a page number annotation;

[0025] S2-9, identifying whether the text contains non-standard use of punctuation marks; non-standard use of punctuation marks includes multiple consecutive identical punctuation marks (such as multiple consecutive commas), mixed use of Chinese and English punctuation marks, incomplete paired symbols (such as parentheses and quotation marks), and extra spaces. Punctuation standardization can be achieved through regular expressions.

[0026] S2-10, identifying whether the text contains errors and outputting the errors and the corresponding corrections.

[0027] As a further optimization of the above scheme, the calculation of the information density score is:

[0028] For a slide, the non-title text is processed for word segmentation and sentence segmentation, and the words and sentences are extracted;

[0029] The words are compared with a preset professional field library to determine whether the words are professional words, and the proportion of the professional words in the non-title text is calculated, denoted as A1;

[0030] It is determined whether the words appear only in the non-title text; the words that appear only in the non-title text are denoted as unique words, and the proportion of the unique words in the non-title text is calculated, denoted as A2;

[0031] The number of mathematical symbols in the non-title text is extracted, and the proportion of the number of mathematical symbols in the non-title text is calculated, denoted as A3;

[0032] The number of words is calculated, and then divided by a preset maximum number of words to obtain A4;

[0033] determine whether the word is a stop word, calculate the proportion of the number of words of the stop word in the non-title text, denoted as A5; stop word refers to a virtual word, pronoun, preposition, conjunction and the like in the text with extremely high frequency (such as "of", "is", "in"), but cannot express specific meaning when existing alone.

[0034] determine whether the word is repeated; the repeatedly appearing word is denoted as a repeated word; calculate the proportion of the number of words of the repeated word in the non-title text, denoted as A6;

[0035] weight and add A1, A2, A3, A4, A5, A6 to obtain the information density score; wherein A1, A2, A3, A4 are respectively added with positive weight, and A5, A6 are respectively added with negative weight, and the value range of the information density score is [0, 1].

[0036] Information density refers to the content of effective information in a unit page, and too sparse or redundant will lead to ambiguity of focus and affect the audience to grasp. The higher the information density score is, the greater the information density is. The weight value used for weighting is defined by the user or the developer.

[0037] As a further optimization of the above scheme, one of the slides includes one title text and multiple non-title texts.

[0038] Vectorize multiple non-title texts of one of the slides to obtain multiple vectors and form a similarity matrix; in the similarity matrix, calculate the similarity value of any two adjacent vectors, which is the logical coherence score.

[0039] By specifying the score range, the logical coherence score is combined to evaluate whether the text is logically smooth.

[0040] Further, in the similarity matrix, for any two adjacent vectors, the similarity value is calculated respectively, and the logical coherence score is calculated according to multiple similarity values;

[0041] The multiple vectors are represented as ; wherein n represents the number of non-title texts under one slide; represents the i-th vector;

[0042] The similarity value of two adjacent vectors is represented as .

[0043] The logical coherence score is represented as .

[0044] Further, if the logical coherence score is not in the specified score range, a false judgment program is started:

[0045] For each non-title text, use regular expression matching to find out whether there is any number sequence at the beginning of the non-title text; if there is, it is a false judgment, and the logical coherence score is invalid; otherwise, the logical coherence score is valid.

[0046] As a further optimization of the above scheme, in S2-7, the ratio of the number of text words of the slide to the number of text words of the slide document is calculated, denoted as B1; the preset standard is that B1 is within a preset range.

[0047] The proportion of the text of a theme in the total text word count affects the speed of the defense speech and the proportion of the speech content. Proper proportion can help the speaker control the speech speed and interaction time of different blocks.

[0048] Further, in a set of preset themes, each theme is assigned a preset range, expressed in percentage; for example, the word count of "background questions" accounts for 10-15%, "methods and results", "core data display and argumentation process" accounts for 60-75%, and "conclusions" accounts for 15-20%. The preset range of each set of preset themes is different and can be customized.

[0049] As a further optimization of the above scheme, one of the preset themes corresponds to a keyword dictionary; the keyword dictionary stores one or more theme keywords;

[0050] The theme keywords are matched with the title text; if any of the theme keywords matches the title text successfully, the preset theme is mapped and associated with the title text.

[0051] The title text is completed and associated with the slide. The application of the preset theme and the keyword dictionary is as follows: "Catalogue: [catalog, catalog]", "Content: [project content, content, research content]", where "catalog" and "content" are preset themes, and "[catalog, catalog]" and "[project content, content, research content]" are keyword dictionaries.

[0052] When matching, if there is an intersection between a keyword dictionary and a title text, it is considered to be a successful match.

[0053] In the keyword dictionary, one of the theme keywords is the preset theme itself.

[0054] Further, the theme matching method includes: exact match, fuzzy match and semantic match, and the priority gradually decreases from left to right.

[0055] The exact match is that the title text and one of the theme keywords are exactly the same;

[0056] The fuzzy matching is that a third-party tool is used to calculate the similarity between the theme keyword and the title text, and only when there is a similarity exceeding a preset threshold, the theme keyword with the highest similarity is matched with the title text successfully; the third-party tool can be a fuzzywuzzy library of Python.

[0057] The semantic matching is that a word embedding model is used to encode the title text and the theme keyword, and then the cosine similarity between the encodings is calculated; only when there is a cosine similarity exceeding a preset threshold, the theme keyword with the highest cosine similarity is matched with the title text successfully; the word embedding model can be BGE or MiniLM, etc.

[0058] As a further optimization of the above scheme, the non-title text is denoted as C1, and the title text is denoted as C2; the text vector similarity between C1 and C2 is calculated, which is the on-topic degree.

[0059] Further, the calculation of the text vector similarity is that C1 and C2 are respectively converted into vectors C5 and C6, and then the similarity value of C5 and C6 is calculated.

[0060] It is judged whether the on-topic degree is within a specified numerical range, if within the range, it is indicated that the on-topic degree is high; otherwise, it is indicated that the on-topic degree is low.

[0061] As a further optimization of the above scheme, the calculation of the readability score is:

[0062] For a slide, the non-title text is subjected to word segmentation processing and sentence segmentation processing, the words and sentences are extracted, and the words are subjected to part-of-speech tagging to obtain a tagged type; the tagged type includes adverbs, conjunctions and others.

[0063] The words are matched with a preset general Chinese word library to judge whether the words are rare words;

[0064] According to the number of characters of the words, it is judged whether the words are complex words; the complex words are words with more than 3 characters;

[0065] According to the tagged type, it is judged whether the words are functional words; the functional words include adverbs and conjunctions.

[0066] The proportions of the rare words, the complex words and the functional words in the non-title text are calculated and denoted as D1, D2 and D3 respectively; the number of the words is calculated and then divided by a preset maximum word value to obtain D4;

[0067] D1, D2, D3 and D4 are converted into percentages and then multiplied by 100; D4 is subjected to upper limit truncation processing so that the maximum value of D4 does not exceed a preset truncation value.

[0068] D1, D2, D3 and D4 are weighted and added to obtain a readability penalty score;

[0069] The readability score = 100 - the readability penalty score.

[0070] The higher the readability score, the higher the readability, and vice versa.

[0071] The application also provides a high school graduation defense slide content intelligent inspection system, which applies the high school graduation defense slide content intelligent inspection method.

[0072] The slide information extraction module is used to extract the text of each slide in the slide document, and the text includes title text and non-title text.

[0073] The theme coverage checking module is used to check whether the content of each slide covers the preset theme; the preset theme includes multiple; the slide and the preset theme are one-to-one mapped and associated; the preset theme is, for example, "catalog", "background", "meaning", "content scheme", "technology route", "algorithm design", etc., which is set in the system by the user or the developer in advance, and the judgment of whether all the titles and main contents of the entire slide document contain the set theme is made during detection; the associated slide and the preset theme are mapped and processed for record.

[0074] The text information density calculation module is used to calculate the information density score of the non-title text of the slide.

[0075] The text logical relationship checking module is used to calculate the logical coherence score of the non-title text in the slide; and whether the logical coherence is determined according to the logical coherence score.

[0076] The text topic cutting degree checking module is used to calculate the topic cutting degree of the non-title text and the title text in the slide; and whether the text is topic cutting is determined according to the topic cutting degree.

[0077] The text readability checking module is used to evaluate the readability of the text in the slide to obtain a readability score.

[0078] The forbidden word checking module is used to check whether there is a forbidden word in the preset forbidden word table in the text, and to count the use degree of the forbidden word.

[0079] The report generation module is used to generate an inspection report according to the checking results and calculation results of each module.

[0080] The inspection report displays whether the slide show is associated with a preset theme, a presentation information density score, a logical coherence score, and whether the text is logically coherent, on topic, and readable.

[0081] The modules are decoupled from each other, and a user can configure and select any number of modules to run during detection.

[0082] Further, the report generation module generates rebuttal questions based on the text; the report generation module is also provided with a general question template library; and the inspection report displays the rebuttal questions and general questions.

[0083] The pre-prepared rebuttal questions and general questions can help students rehearse the key points of the rebuttal and guide the review to focus on the preset topic, thereby improving the efficiency of the rebuttal in both directions.

[0084] As a further optimization of the above scheme, the following is further included:

[0085] A text length structure inspection module is configured to identify whether the number of characters in the text of the slide show meets a preset standard.

[0086] A page number annotation inspection module is configured to identify whether the slide show contains a page number annotation; specifically, a regular expression is used to detect whether the slide show contains a page number annotation.

[0087] A punctuation specification inspection module is configured to identify whether the text contains non-standardly used punctuation marks; non-standardly used punctuation marks include multiple identical punctuation marks in succession (such as multiple commas in succession), mixed use of Chinese and English punctuation marks, incomplete pairs of symbols (such as parentheses and quotation marks), and extra spaces.

[0088] A misspelled word inspection module; the misspelled word inspection module includes a misspelled word recognition model and a whitelist; the misspelled word recognition model is configured to identify whether the text contains misspelled words; and the whitelist stores custom words / characters to avoid false positives from the misspelled word recognition model.

[0089] Further, the whitelist also interacts with a forbidden word inspection module, and the forbidden word inspection module ignores words in the whitelist during inspection.

[0090] The inspection report also displays whether the number of characters meets the requirements, page number annotation vulnerabilities, incorrectly used punctuation marks, and misspelled words.

[0091] Compared with the prior art, the present application has the following beneficial effects:

[0092] The application provides a high school graduation defense slide content intelligent inspection method and system, aiming at the scene of efficient graduation defense, providing multiple selectable inspection items, evaluating the quality of slide content text from multiple angles such as information density score, logical coherence score, and topic relevance, and combining format specification review to assist college students in finding and correcting errors in the form of scores and error prompts, thereby improving the quality of slides. BRIEF DESCRIPTION OF DRAWINGS

[0093] Figure 1 is a flowchart of a high school graduation defense slide content intelligent inspection method provided by an embodiment of the application. DETAILED DESCRIPTION

[0094] In order to make the purpose, technical solutions and advantages of the application clearer, the technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only a part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.

[0095] As shown in Figure 1 , the embodiment provides a high school graduation defense slide content intelligent inspection method, including the following steps:

[0096] S1, structurally process the slide document; the slide document includes multiple slides; for the text in each slide, divide it into title text and non-title text respectively; in this embodiment, there is one title text and multiple non-title texts in a slide.

[0097] S2, perform the following inspection items to obtain the inspection results or calculation results corresponding to each inspection item:

[0098] S2-1, check whether the content of each slide covers the preset theme, that is, the preset theme covering and association check; the preset theme includes multiple; the preset theme is set by the user or the developer in the system in advance, such as "table of contents", "background", "meaning", "content scheme", "technical route", "algorithm design", etc., and the system judges whether all the titles and main contents of the entire slide document contain the set theme; the mapping processing record is performed on the associated slide and the preset theme; the preset theme has multiple sets, corresponding to different fields or professions, such as computer and mathematics, and the themes of PPT documents of legal professionals are different; the user selects a set of preset themes in advance before starting the system inspection.

[0099] In the embodiment, one preset theme corresponds to one keyword dictionary; the keyword dictionary stores one or more theme keywords; the application of the preset theme and the keyword dictionary is as follows: "Catalog: [catalog, catalog]" and "Content: [project content, content, research content]", wherein "catalog" and "content" are preset themes, and "[catalog, catalog]" and "[project content, content, research content]" are keyword dictionaries. In the keyword dictionary, one theme keyword is the preset theme itself.

[0100] The theme keywords are matched with the title text; if any one of the theme keywords is successfully matched with the title text, the preset theme is mapped and associated with the title text. The title text is completed association, and the slide is completed association.

[0101] During the matching, if there is an intersection between one keyword dictionary and the title text, it is considered as a successful match.

[0102] Specifically, the theme matching method includes: exact matching, fuzzy matching and semantic matching, and the priority gradually decreases from left to right. If the exact matching fails, fuzzy matching is tried; if the fuzzy matching fails, semantic matching is tried.

[0103] The exact matching is that the title text and one of the theme keywords are completely the same;

[0104] The fuzzy matching is that the similarity between the theme keywords and the title text is calculated by using a third-party tool, and only when there is a similarity exceeding a preset threshold, the theme keyword with the highest similarity is successfully matched with the title text; the third-party tool can be a fuzzywuzzy library of Python.

[0105] The semantic matching is that the title text and the theme keywords are encoded by using a word embedding model, and the cosine similarity between the encodings is calculated; only when there is a cosine similarity exceeding a preset threshold, the theme keyword with the highest cosine similarity is successfully matched with the title text; the word embedding model can be BGE or MiniLM, etc.

[0106] The preset theme not covered is displayed;

[0107] S2-2, calculate the information density score of the non-title text of the slide; in the embodiment, the calculation of the information density score is as follows:

[0108] For a slide, the non-title text is processed by word segmentation and sentence segmentation, and the words and sentences are extracted;

[0109] The words are compared with a preset professional field library to determine whether the words are professional words, and the proportion of the professional words in the non-title text is calculated, which is denoted as A1;

[0110] determining whether the vocabulary appears only in the non-title text; the vocabulary appearing only is recorded as a unique word; calculating the proportion of the unique word in the non-title text, recorded as A2;

[0111] extracting the number of mathematical symbols in the non-title text, calculating the proportion of the number of words of the mathematical symbols in the non-title text, recorded as A3;

[0112] calculating the number of words, and then dividing by the preset maximum number of words to obtain A4;

[0113] determining whether the vocabulary is a stop word, calculating the proportion of the number of words of the stop word in the non-title text, recorded as A5; the stop word refers to a virtual word, pronoun, preposition, conjunction and the like appearing in the text with a very high frequency (such as "of", "is", "in"), but unable to express specific meaning when existing alone.

[0114] determining whether the vocabulary appears repeatedly; the vocabulary appearing repeatedly is recorded as a repeated word; calculating the proportion of the number of words of the repeated word in the non-title text, recorded as A6;

[0115] weighting and adding A1, A2, A3, A4, A5 and A6 to obtain an information density score; wherein A1, A2, A3 and A4 are respectively added with positive weights, and A5 and A6 are respectively added with negative weights, and the value range of the information density score is [0, 1].

[0116] The information density refers to the content of effective information in a unit page, and too sparse or redundant will lead to ambiguity of the focus and affect the audience to grasp. The higher the information density score is, the greater the information density is. The weight values used for weighting are defined by the user or the developer. For example, if six weights are set as 0.3, 0.3, 0.3, 0.3, 0.1 and 0.1, then the information density score = 0.3×(A1+A2+A3+A4)-0.1×(A5+A6).

[0117] S2-3, calculating a logical coherence score of the non-title text in the slide; determining whether the logical coherence according to the logical coherence score; evaluating whether the text is logically smooth by specifying a score range and combining the logical coherence score. Specifically, the multiple non-title text vectors of a slide are processed to obtain multiple vectors and form a similarity matrix; for any two adjacent vectors in the similarity matrix, the similarity values are calculated respectively, and the logical coherence score is calculated according to the multiple similarity values;

[0118] The multiple vectors are represented as ; wherein n represents the number of non-title texts under a slide; represents the i-th vector;

[0119] The similarity value of two adjacent vectors is represented as ;

[0120] The logical coherence score is represented as .

[0121] If the logical coherence score is not in the specified score range, i.e., it is determined that the logic is not coherent, the false positive program is started, i.e.:

[0122] For each non-title text, use regular expression matching to find out whether there is any format of number sequence at the beginning of the non-title text; if there is, it is a false positive, the logical coherence score is invalid, i.e., it is determined that the logic is coherent; otherwise, the logical coherence score is valid, i.e., it is determined that the logic is not coherent.

[0123] S2-4, calculate the topic cutting degree of non-title text and title text in the slide; determine whether the text is topic cutting according to the topic cutting degree; the topic cutting degree refers to the degree of fit between the text content and the topic.

[0124] Specifically, the non-title text is denoted as C1, and the title text is denoted as C2; the text vector similarity between C1 and C2 is calculated, i.e., the topic cutting degree.

[0125] Among them, the calculation of the text vector similarity is: C1 and C2 are converted into vectors C5 and C6 respectively, and then the similarity value of C5 and C6 is calculated.

[0126] Determine whether the topic cutting degree is within the specified numerical range; if it is within the range, it means that the topic cutting degree is high; otherwise, it means that the topic cutting degree is low.

[0127] S2-5, calculate the readability score of the text in the slide; in this embodiment, the calculation of the readability score is:

[0128] For a slide, the non-title text is subjected to word segmentation processing and sentence segmentation processing, the words and sentences are extracted, and the words are subjected to part-of-speech tagging to obtain a tagged type; the tagged type includes adverbs, conjunctions and others;

[0129] Match the words with the preset general Chinese word library to determine whether the words are rare words;

[0130] Determine whether the words are complex words according to the number of characters in the words; a complex word is a word with more than 3 characters;

[0131] Determine whether the words are function words according to the tagged type; function words include adverbs and conjunctions;

[0132] Calculate the proportion of the number of words of rare words, complex words and function words in the non-title text, denoted as D1, D2 and D3 respectively; calculate the number of words, and then divide by the preset maximum number of words, denoted as D4;

[0133] D1, D2, D3 and D4 are converted into percentages, and then multiplied by 100; D4 is truncated to the upper limit, so that the maximum value of D4 does not exceed the preset truncation value; specifically, if the truncation value is 30 and D4 is 31, D4 is changed to 30; if D4 is 29, D4 remains unchanged.

[0134] D1, D2, D3 and D4 are weighted and added to obtain a readability penalty score.

[0135] The readability score = 100 - readability penalty score.

[0136] The higher the readability score, the higher the readability, and vice versa.

[0137] S2-6, check if there are any forbidden words in the preset forbidden word list in the text of the slide, and count the degree of use of the forbidden words;

[0138] S2-7, identify whether the number of text words in the slide meets the preset standard; specifically, calculate the ratio of the number of text words in the slide to the number of text words in the slide document, denoted as B1; the preset standard is that B1 is within a preset range, for example, for a set of preset themes, the theme contents include "background problem", "method and result" and "core data display and argumentation process", "conclusion", and the assigned ranges are 10%-15%, 60%-75% and 15%-20% respectively.

[0139] The proportion of the text of a theme in the total text word count affects the speed of the defense speech and the proportion of the speech content. Appropriate proportion can help the speaker control the speech speed and interaction time of different blocks.

[0140] S2-8, identify whether the slide contains page number annotations;

[0141] S2-9, identify whether the text contains non-standard use of punctuation marks; non-standard use of punctuation marks includes continuous multiple same punctuation marks (such as continuous multiple commas), mixed use of Chinese and English punctuation marks, incomplete paired symbols (such as parentheses and quotation marks), and extra spaces. Punctuation standard check can be realized by regular expression.

[0142] S2-10, identify whether the text contains errors, and output the errors and the corrected words corresponding to the errors.

[0143] S3, generate a check report to show the check results and calculation results, and generate optimization suggestions such as correction of slide title and shortening of sentence length according to the check results and calculation results.

[0144] The inspection report is whether the presentation slide is associated with a preset theme, a presentation information density score, a logical coherence score, and whether the text is logically coherent, on topic, and readable.

[0145] The embodiment also provides a high school graduation defense slide content intelligent inspection system, which applies the above inspection method; the system comprises a plurality of modules.

[0146] A slide information extraction module is configured to extract text of each slide in a slide document, wherein the text comprises title text and non-title text.

[0147] A theme coverage inspection module is configured to inspect whether the content of each slide covers a preset theme; the preset theme comprises a plurality of themes; the slide and the preset theme are one-to-one mapped and associated; the preset theme is, for example, "catalog", "background", "meaning", "content scheme", "technical route", "algorithm design", and the like, which is preset in the system by a user or a developer; when detecting, it is judged whether all the titles and main contents of the entire slide document contain the set theme; the slide associated with the preset theme is mapped and processed and recorded.

[0148] A text information density calculation module is configured to calculate an information density score of the non-title text of the slide.

[0149] A text logical relationship inspection module is configured to calculate a logical coherence score of the non-title text in the slide; whether the text is logically coherent is judged according to the logical coherence score.

[0150] A text on-topic degree inspection module is configured to calculate an on-topic degree of the non-title text and the title text in the slide; whether the text is on topic is judged according to the on-topic degree.

[0151] A text readability inspection module is configured to evaluate the readability of the text in the slide to obtain a readability score.

[0152] A forbidden word inspection module is configured to inspect whether there is a forbidden word in a preset forbidden word table in the text, and to count the use degree of the forbidden word.

[0153] A text length structure inspection module is configured to identify whether the number of words of the text of the slide meets a preset standard.

[0154] A page number annotation inspection module is configured to identify whether the slide contains a page number annotation; specifically, a regular expression is used to detect whether the slide contains a page number annotation.

[0155] A punctuation specification checking module is configured to identify whether the text contains non-standard punctuation. Non-standard punctuation scenarios include multiple identical punctuation marks in succession (e.g., multiple commas in succession), mixed use of Chinese and English punctuation, incomplete paired symbols (e.g., brackets, quotation marks), and extra spaces.

[0156] A misspelling checking module includes a misspelling identification model and a whitelist. The misspelling identification model is configured to identify whether the text contains misspelled words. The whitelist stores custom words / phrases to avoid false positives from the misspelling identification model. For example, if a user studies a new field that generates new terms, the new terms may be incorrectly identified as misspelled words by the misspelling identification model. The whitelist can be used to avoid such false positives. Further, the whitelist interacts with the banned word checking module, and the banned word checking module ignores words / phrases in the whitelist.

[0157] The checking report also shows whether the word count meets the requirements, whether the page number is correctly labeled, whether the punctuation is used correctly, and whether the text contains misspelled words.

[0158] A report generation module is configured to generate a checking report based on the checking results and calculation results of the modules.

[0159] The checking report shows whether the slides are associated with the preset theme, the information density score, the logical coherence score, and whether the text is logically coherent, on topic, and readable. The user can adjust the defects in the slide document based on the evaluation results or scores and optimize the improvements.

[0160] The modules are decoupled from each other, and the user can configure and select any number of modules to run during detection.

[0161] Further, the report generation module generates rebuttal questions based on the text. The report generation module also includes a general question template library. The checking report also displays rebuttal questions and general questions.

[0162] The pre-prepared rebuttal questions and general questions can help students rehearse the key points of the rebuttal and guide the reviewers to focus on the preset topic, thereby improving the efficiency of the rebuttal in both directions.

[0163] Based on the disclosure and teachings of the above specification, those skilled in the art of the present application can also make changes and modifications to the above embodiments. Therefore, the present application is not limited to the specific embodiments disclosed and described above. Some modifications and changes to the present application should fall within the scope of the claims of the present application. In addition, although some specific terms are used in the specification, these terms are only for convenience of explanation and do not constitute any limitation on the present application.

Claims

1. A method for intelligent checking of contents of graduation defense slide show, characterized in that, The method comprises the following steps: S1, structuring a slide document; the slide document comprises a plurality of slides; for the text in each slide, the text is divided into title text and non-title text respectively; S2, performing the following checks to obtain the check results or calculation results corresponding to each check: S2-1, checking whether the content of each slide covers a preset theme; the preset theme comprises a plurality of themes; the slide and the preset theme are one-to-one mapped and associated; displaying the preset theme that is not covered; one of the preset themes corresponds to one keyword dictionary; the keyword dictionary stores one or more theme keywords; performing theme matching between the theme keywords and the title text; if any of the theme keywords matches the title text successfully, the preset theme is mapped and associated with the title text; S2-2, calculating the information density score of the non-title text of the slide; S2-3, calculating the logical coherence score of the non-title text in the slide; whether the logical coherence score is logical coherence is determined according to the logical coherence score; S2-4, calculating the text vector similarity between the non-title text and the title text in the slide as the on-topic degree; whether the text is on-topic is determined according to the on-topic degree; S2-5, calculating the readability score of the text in the slide; the calculation of the readability score is: For one of the slides, the non-title text is subjected to word segmentation processing and sentence segmentation processing, the words and sentences are extracted, and the words are subjected to part-of-speech tagging to obtain a tagged type; the tagged type includes adverbs, conjunctions, and others; matching the words with a preset general Chinese word library to determine whether the words are rare words; determining whether the words are complex words according to the number of characters of the words; the complex words are words with more than 3 characters; determining whether the words are function words according to the tagged type; the function words include adverbs and conjunctions; calculating the proportion of the rare words, the complex words, and the function words in the non-title text, respectively denoted as D1, D2, and D3; calculating the number of words and dividing by a preset maximum word value, denoted as D4; converting D1, D2, D3, and D4 into percentages and multiplying by 100; performing upper limit truncation processing on D4 to ensure that the maximum value of D4 does not exceed a preset truncation value; weighting and adding D1, D2, D3, and D4 to obtain a readability penalty score; The readability score = 100 - readability penalty score; S2-6, checking whether there are any forbidden words in a preset forbidden word list in the text of the slide and counting the usage degree of the forbidden words; S3, generating a check report to display the check results and the calculation results.

2. The method for intelligent checking of the content of a graduation thesis presentation slide according to claim 1, characterized in that, The check items further comprise: S2-7, identifying whether the number of characters in the text of the slide meets a preset standard; S2-8, identifying whether the slide contains a page number annotation; S2-9, identifying whether the text contains non-standard punctuation marks; S2-10, identifying whether there is a wrong word in the text, and outputting the wrong word and the correction word corresponding to the wrong word.

3. The method of claim 1, wherein the method further comprises: The calculation of the information density score is: For one of the slides, the non-title text is processed by word segmentation and sentence segmentation, and the vocabulary and sentence are extracted; Compare the vocabulary with the preset professional field library to determine whether the vocabulary is a professional vocabulary, and calculate the proportion of the professional vocabulary in the non-title text, denoted as A1; Determine whether the vocabulary appears only in the non-title text; the unique vocabulary is denoted as unique vocabulary; Calculate the proportion of the unique vocabulary in the non-title text, denoted as A2; Extract the number of mathematical symbols in the non-title text, calculate the proportion of the number of mathematical symbols in the non-title text, denoted as A3; Calculate the number of words, and divide by the preset maximum number of words to get A4; Determine whether the vocabulary is a stop word, and calculate the proportion of the stop word in the non-title text, denoted as A5; Determine whether the vocabulary is repeated; the repeated vocabulary is denoted as repeated word; Calculate the proportion of the repeated word in the non-title text, denoted as A6; A1, A2, A3, A4, A5, A6 are weighted and added to obtain the information density score; wherein A1, A2, A3, A4 are positively weighted, and A5, A6 are negatively weighted, and the value range of the information density score is [0, 1].

4. The method of claim 1, wherein the method further comprises: One of the slides, the title text includes 1, and the non-title text includes multiple; Vectorize the multiple non-title text of one of the slides to obtain multiple vectors and form a similarity matrix; in the similarity matrix, the similarity value of any two adjacent vectors is calculated, which is the logical coherence score.

5. The method of claim 2, wherein the method further comprises: In S2-7, the ratio of the number of words in the slide to the number of words in the slide document is calculated, denoted as B1; the preset standard is that B1 is in a preset range.

6. A system for intelligent checking of contents of graduation thesis slides, characterized by, A kind of college graduation defense slide content intelligent inspection method is applied as claimed in any one of claims 1 to 5;The system includes multiple modules: Slide information extraction module, for extracting the text of each slide in the slide document, the text includes title text and non-title text; Theme coverage checking module, for checking whether the content of each of the slides covers the preset theme;The preset theme includes multiple;The slide and the preset theme are one-to-one mapped and associated; Text information density calculation module, for calculating the information density score of the non-title text of the slide; Text logical relationship checking module, for calculating the logical coherence score of the non-title text in the slide;According to the logical coherence score, it is judged whether it is logically coherent; Text relevance checking module, for calculating the relevance of the non-title text and the title text in the slide;According to the relevance, it is judged whether the text is relevant; Text readability checking module, for evaluating the readability of the text in the slide to obtain a readability score. The disabling word checking module is configured to check whether a disabling word in a preset disabling word list exists in the text and to count a usage degree of the disabling word. The report generating module is configured to generate a checking report according to checking results and calculation results of the modules.

7. The intelligent inspection system for the content of college graduation defense slides according to claim 6, characterized in that, Further comprising: The text length structure checking module is configured to identify whether a text word number of the slide conforms to a preset standard. The page number marking checking module is configured to identify whether the slide contains a page number marking. The punctuation standard checking module is configured to identify whether the text contains a non-standardly used punctuation symbol. The wrong word checking module comprises a wrong word identifying model and a white list. The wrong word identifying model is configured to identify whether the text contains a wrong word. The white list stores self-defined words / characters, and is configured to avoid false identification by the wrong word identifying model.

Citation Information

Patent Citations

  • Powerpoint quality detection method and device, electronic equipment and storage medium

    CN120046599A