Method and apparatus for evaluating and processing generated presentation documents.
By using multimodal AI technology to analyze PPT documents and perform multi-dimensional evaluation, the problems of incomplete PPT information extraction and context dependence in existing technologies are solved, achieving efficient PPT document optimization and knowledge utilization of the RAG system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING QDING INTERCONNECTION TECHNOLOGY CO LTD
- Filing Date
- 2026-01-09
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies cannot effectively parse visual element information in PPTs when building RAG systems, resulting in difficulty in extracting core information, fragmented content with serious contextual dependencies, and neglect of speaker notes, leading to knowledge loss.
By using multimodal AI technology to analyze PPT documents, textual information, visual elements, and structural information are obtained. Multidimensional evaluation processing is then performed, including assessments of text-image separation, contextual integrity, structuring degree, and information density to noise ratio. A comprehensive evaluation result is generated, and optimization suggestions are provided.
It significantly improves the automation level and accuracy of PPT document quality assessment, provides accurate basis for optimization and improvement decisions, and enhances the knowledge utilization efficiency of the RAG system.
Smart Images

Figure CN122089133A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method and apparatus for evaluating and processing generated presentation documents. Background Technology
[0002] Presentation documents (PowerPoint, PPT) are one of the most essential mediums for sharing knowledge, conducting training, and reporting work within enterprises. A large number of plans, reports, and training materials are stored in PPT format, forming a valuable knowledge base. Effectively utilizing these PPT documents to enable RAG (Retrieval-Augmented Generation) question-answering systems to learn from and answer employee questions is a key step in enhancing the value of AI applications within enterprises.
[0003] When building a RAG system, related technologies use general document parsing tool libraries to extract text and images from PPTs, extracting all text box content and images from the PPT, and then performing chunking and vectorization. However, this method does not extract information completely and destroys the spatial layout relationship between elements on the page, making it impossible for the RAG system to accurately understand the context information. Summary of the Invention
[0004] The embodiments of this application aim to at least partially address one of the technical problems in the related art. To this end, the embodiments of this application propose a method, apparatus, device, and medium for retrieving generated presentation documents for evaluation processing, thereby improving the accuracy of presentation document evaluation. The embodiments of this application provide a method for evaluating and processing a retrieved presentation document. This method includes: parsing the presentation document to obtain its text information, visual elements, and structural information, wherein the text information includes presentation text and / or presentation notes; performing multi-dimensional evaluation processing based on the text information, visual elements, and structural information to obtain at least two of the following: a text-image separation evaluation result, a contextual integrity evaluation result, a structured degree evaluation result, and an information density to noise ratio evaluation result; and weighting at least two of the following: text-image separation evaluation result, contextual integrity evaluation result, structured degree evaluation result, and information density to noise ratio evaluation result, to obtain a comprehensive evaluation result for the presentation document.
[0005] In some implementations, a multi-dimensional evaluation process is performed based on at least one of textual information, visual elements, and structural information to obtain an image-text separation evaluation result. This includes: determining, based at least on textual information and visual elements, whether the main information carrier in the presentation document is text or an image, and determining whether the image has a corresponding textual explanation, to obtain a core information carrier evaluation result; determining, based at least on textual information and visual elements, whether the textual description in the presentation document is associated with the corresponding image, to obtain an image-text matching evaluation result; and obtaining an image-text separation evaluation result based on the core information carrier evaluation result and the image-text matching evaluation result.
[0006] In some implementations, a multi-dimensional evaluation process is performed based on at least one of textual information, visual elements, and structural information to obtain a context integrity evaluation result, including: determining the page proportion of presentation notes text in the presentation document based at least on textual information and structural information to obtain a presentation notes text usage rate evaluation result; determining the page content independence in the presentation document based at least on textual information, visual elements, and structural information to obtain a page independence evaluation result; and obtaining a context integrity evaluation result based on the presentation notes text usage rate evaluation result and the page independence evaluation result.
[0007] In some implementations, a multi-dimensional evaluation process is performed based on at least one of textual information, visual elements, and structural information to obtain a structuredness evaluation result, including: determining whether the layout of the presentation document is a standard layout based at least on structural information to obtain a layout evaluation result; determining the page order logic and chapter division logic in the presentation document based at least on textual information and structural information to obtain a logical clarity evaluation result; and obtaining a structuredness evaluation result based on the layout evaluation result and the logical clarity evaluation result.
[0008] In some implementations, a multi-dimensional evaluation process is performed based on at least one of textual information, visual elements, and structural information to obtain an information density and noise ratio evaluation result, including: determining whether the textual information in the presentation document is concise content based at least on textual information to obtain a text conciseness evaluation result; determining content-irrelevant elements in the presentation document based at least on visual elements to obtain a visual noise evaluation result; and obtaining an information density and noise ratio evaluation result based on the text conciseness evaluation result and the visual noise evaluation result. In some implementations, the method further includes: determining, based on the comprehensive evaluation results, the content to be optimized in the presentation document, and generating document optimization suggestions for the content to be optimized, wherein the content to be optimized includes images lacking text explanations and / or fragmented pages, and the document optimization suggestions include adding a first presentation annotation text to the images and / or adding a second presentation annotation text to the fragmented pages; and generating a document evaluation report based on the comprehensive evaluation results and the document optimization suggestions, wherein the document evaluation report includes the document optimization suggestions and location information of the content to be optimized in the presentation document. In some implementations, the method further includes: in response to receiving a translation instruction for a presentation document, generating a text document based on textual information, visual elements, and structural information.
[0009] In some implementations, the method further includes: generating questions and corresponding answers for any page or core knowledge content of a presentation document based on textual information, visual elements, and structural information; and performing retrieval and debugging based on the questions and answers.
[0010] In some implementations, the method further includes: determining the association information between pages and content in the presentation document based on textual information, visual elements, and structural information; and generating logical relationship data of the presentation document based on the association information.
[0011] The embodiments of this application provide an evaluation processing apparatus for retrieving generated presentation documents. The apparatus includes: a parsing module for parsing the presentation document to obtain text information, visual elements, and structural information, wherein the text information includes presentation text and / or presentation annotation text; an evaluation module for performing multi-dimensional evaluation processing based on the text information, visual elements, and structural information to obtain at least two of the following: a text-image separation evaluation result, a context integrity evaluation result, a structured degree evaluation result, and an information density to noise ratio evaluation result; and a processing module for weighting at least two of the text-image separation evaluation result, context integrity evaluation result, structured degree evaluation result, and information density to noise ratio evaluation result to obtain a comprehensive evaluation result for the presentation document.
[0012] An embodiment of this application provides an electronic device, which includes: a memory, and one or more processors communicatively connected to the memory; the memory stores instructions executable by one or more processors, which are executed by one or more processors to cause the one or more processors to implement the steps of the method of any of the above embodiments.
[0013] Embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method of any of the above embodiments.
[0014] In the above embodiments, the evaluation processing method for the retrieved presentation document includes: parsing the presentation document to obtain its text information, visual elements, and structural information, wherein the text information includes the presentation text and / or presentation notes; performing multi-dimensional evaluation processing based on the text information, visual elements, and structural information to obtain at least two of the following: a text-image separation evaluation result, a context integrity evaluation result, a structured degree evaluation result, and an information density to noise ratio evaluation result; and weighting at least two of the following: a text-image separation evaluation result, a context integrity evaluation result, a structured degree evaluation result, and an information density to noise ratio evaluation result, to obtain a comprehensive evaluation result for the presentation document. The above method can accurately extract and structure text information, visual elements, and document framework from the original presentation document, and then construct key evaluation dimensions such as text-image separation, contextual integrity, structuring degree, and information density and noise ratio. Through intelligent weighting and fusion of the above multi-dimensional evaluation results, the RAG system can automatically generate comprehensive and objective presentation document evaluation results, significantly improving the automation level and evaluation depth of presentation document quality evaluation, providing accurate decision-making basis for document optimization and improvement, and improving evaluation efficiency and accuracy. Attached Figure Description
[0015] Figure 1 A flowchart illustrating a method for evaluating generated presentation documents, provided for implementation of this application; Figure 2 A flowchart illustrating another method for evaluating generated demonstration documents provided in this application; Figure 3 A schematic diagram of an evaluation processing apparatus for retrieving generated presentation documents, provided as another embodiment of this application; Figure 4 A block diagram of an electronic device provided for another embodiment of this application. Detailed Implementation
[0016] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0017] Presentation documents (PowerPoint, PPT) are one of the most essential mediums for sharing knowledge, conducting training, and reporting work within enterprises. A large number of plans, reports, and training materials are stored in PPT format, forming a valuable knowledge base. Effectively utilizing these PPT documents, enabling RAG (Retrieval-Augmented Generation) question-and-answer systems to learn from them and answer employee questions (such as "Please tell me the core architecture of solution XX"), is a key step in enhancing the value of AI applications within enterprises.
[0018] When enterprises use AI question-answering systems based on Retrieval-Augmented Generation (RAG) technology, using a large number of presentation slides (PPT, .pptx) as knowledge sources presents the following technical problems: (1) High coupling between text and images, making it difficult to extract core information: The core feature of PPT is its combination of text and images, with a large amount of key information contained in pictures, diagrams (such as flowcharts and architecture diagrams) and tables, while the text is often just a highly summarized summary of key points. The RAG data processing flow has difficulty effectively parsing the information in these visual elements, resulting in serious knowledge loss. (2) Fragmented content and serious context dependence: The content of PPT is divided into independent slides, and the content of each slide is usually highly summarized and fragmented. Its complete meaning depends heavily on the speaker's oral explanation and the logical order between pages. When AI (Artificial Intelligence) retrieves a single page, it often cannot accurately understand its true meaning due to a lack of sufficient context. (3) The value of “Speaker Notes” information is ignored: The “Speaker Notes” area of PPT usually contains the most detailed and richest explanations and background information on the page content. This part of the content is crucial for AI to understand the slides, but the data loading process of the RAG system often ignores this valuable information.
[0019] When building a RAG system, related technologies use a general document parsing tool library to extract text and images from PPTs, extracting all text box content and images from the PPT, and then performing chunking and vectorization. However, this method has the following fundamental technical defects when dealing with the special format of PPT: (1) "Looking at the text but not the images", incomplete information extraction: Existing parsing tools are mainly text-based. They can extract the words in the text boxes, but they cannot understand the structured information and logical relationships carried by visual elements such as SmartArt graphics, flowcharts, and architecture diagrams. This results in the most core and information-dense parts of the PPT being completely ignored. (2) "Taking things out of context", unable to reconstruct the context: The tool extracts the text box content of each PPT page in isolation, completely destroying the spatial layout relationship between the elements on the page, as well as the logical relationship between pages. AI only gets a bunch of disordered and fragmented text fragments, and cannot restore its complete context. (3) Ignoring key “notes” results in serious information loss: Most general parsing tools do not extract the content of the “speaker notes” area by default, which leads to the waste of this “golden information” containing a lot of background knowledge and in-depth explanation.
[0020] Therefore, this application proposes a method for evaluating and processing retrieved presentation documents, which can deeply understand the multimodal and fragmented characteristics of PPT documents, comprehensively evaluate whether a PPT is "AI-friendly", and guide users on how to transform it into a high-quality knowledge asset that can be effectively utilized by the RAG system.
[0021] Figure 1 This is a flowchart illustrating a method for evaluating generated demonstration documents, provided as an embodiment of this application.
[0022] like Figure 1 As shown, the evaluation processing method 100 for retrieving the generated presentation document includes, for example, steps S110-S130.
[0023] Step S110: Parse the presentation document to obtain its text information, visual elements, and structural information. The text information includes the presentation text and / or presentation notes. For example, a presentation document may include text box content (such as text boxes on the presentation document page, notes on the presentation document page), visual content (such as images, charts, flowcharts, architecture diagrams, etc.), and structural content (such as the page order and title structure of the presentation document). Parsing is achieved through a multimodal large model or a specialized computer vision (CV) model. Through parsing, the text information, visual elements (such as text descriptions of images, chart types, data, titles and conclusions, arrows, shapes, text, etc. in flowcharts / architecture diagrams), and structural information of the presentation document can be obtained.
[0024] Step S120: Based on text information, visual elements and structural information, perform multi-dimensional evaluation processing to obtain at least two of the following evaluation results for the presentation document: text-image separation degree evaluation result, context integrity evaluation result, structure degree evaluation result, and information density to noise ratio evaluation result.
[0025] For example, the AI large model performs multi-dimensional evaluation processing based on textual information, visual elements, and structural information (including image-text separation evaluation, context integrity evaluation, structuring degree evaluation, and information density to noise ratio evaluation). Points can be deducted for those that do not meet the evaluation conditions of each dimension, so that each dimension will get an evaluation score as the evaluation result.
[0026] Step S130: Weight at least two of the following evaluation results: text-image separation degree evaluation result, context integrity evaluation result, structure degree evaluation result, and information density to noise ratio evaluation result, to obtain a comprehensive evaluation result of the presentation document. For example, the AI big model performs weighted processing based on at least two of the following: text-image separation evaluation results, context integrity evaluation results, structuredness evaluation results, and information density-to-noise ratio evaluation results. The weights of each part are set according to the requirements to obtain a comprehensive score as the comprehensive evaluation result of the demonstration document, and optimization suggestions are generated based on the comprehensive evaluation result.
[0027] As can be seen, the presentation document evaluation processing method of this application can accurately extract and structure text information, visual elements and document framework from the original presentation document, and then construct key evaluation dimensions such as text-image separation degree, contextual integrity, degree of structuring and information density and noise ratio. Through intelligent weighting and fusion of the above multi-dimensional evaluation results, the RAG system can automatically generate comprehensive and objective presentation document comprehensive evaluation results, which significantly improves the automation level and evaluation depth of presentation document quality evaluation, provides accurate decision-making basis for document optimization and improvement, and improves evaluation efficiency and accuracy.
[0028] In one example, a multi-dimensional evaluation process is performed based on at least one of textual information, visual elements, and structural information to obtain the text-image separation evaluation result. This includes: determining, based at least on textual information and visual elements, whether the main information carrier in the presentation document is text or an image, and determining whether the image contains a corresponding textual explanation, to obtain the core information carrier evaluation result; determining, based at least on textual information and visual elements, whether the textual description in the presentation document is associated with the corresponding image, to obtain the text-image matching evaluation result; and obtaining the text-image separation evaluation result based on the core information carrier evaluation result and the text-image matching evaluation result.
[0029] Specifically, the image-text separation assessment comprises two parts: a core information carrier assessment and an image-text matching assessment. The core information carrier assessment evaluates whether the document's core information is primarily carried by text (textual information) or heavily relies on complex images or screenshots (visual elements) that AI cannot directly understand. If an image lacks a corresponding text explanation, or if the text on the image cannot be recognized by OCR (Optical Character Recognition), a deduction is made, resulting in the core information carrier assessment result (assessment score). The image-text matching assessment evaluates the relevance of the text description on the page to the content of its accompanying image, resulting in the image-text matching score. The sum of the two assessment results is used as the image-text separation assessment result.
[0030] In the above embodiments, through multi-dimensional image-text separation evaluation, the core information carrier type of the presentation document can be automatically identified and the correlation between images and text can be analyzed, thereby significantly improving the accuracy of understanding the document information structure.
[0031] In one example, a multi-dimensional evaluation process is performed based on at least one of textual information, visual elements, and structural information to obtain a contextual integrity evaluation result, including: determining the page proportion of presentation notes text in the presentation document based on at least textual information and structural information to obtain a presentation notes text usage rate evaluation result; determining the page content independence in the presentation document based on at least textual information, visual elements, and structural information to obtain a page independence evaluation result; and obtaining a contextual integrity evaluation result based on the presentation notes text usage rate evaluation result and the page independence evaluation result.
[0032] Specifically, the contextual integrity assessment comprises two parts: a presentation note text usage rate assessment and a page independence assessment. The presenter note usage rate (presentation note text usage rate assessment) is one of the most critical metrics for PowerPoint presentations. The percentage of pages containing effective text content in the "presenter notes" is calculated; the richer and more detailed the notes, the higher the score. The page independence assessment evaluates whether the information contained on a single slide constitutes a relatively independent and complete knowledge point. Pages that are too fragmented or require multiple pages to understand will be penalized. The sum of the two assessment results is used as the final contextual integrity assessment score.
[0033] In the above embodiments, by conducting multi-dimensional contextual integrity assessment, the usage of presentation notes can be automatically quantified and the independence of page content can be analyzed, thereby significantly improving the accuracy of judging the internal logic and information completeness of the presentation document.
[0034] In one example, a multi-dimensional evaluation process is performed based on at least one of textual information, visual elements, and structural information to obtain a structuredness evaluation result, including: determining whether the layout of the presentation document is a standard layout based at least on structural information to obtain a layout evaluation result; determining the page order logic and chapter division logic in the presentation document based at least on textual information and structural information to obtain a logical clarity evaluation result; and obtaining a structuredness evaluation result based on the layout evaluation result and the logical clarity evaluation result.
[0035] Specifically, the structured presentation assessment includes layout evaluation and logical clarity evaluation. The layout evaluation checks whether the built-in standard layouts in PowerPoint (such as the "Titles and Contents" layout) are used, rather than being arbitrarily pieced together from multiple independent text boxes. The logical clarity evaluation assesses whether the page order and chapter divisions of the entire PowerPoint presentation have a clear logical thread. The score obtained by summing the two evaluation results is the final structured presentation assessment result.
[0036] In the above embodiments, by evaluating the degree of structure in a multi-dimensional manner, the format standardization of the presentation document can be automatically identified and the clarity of its internal logical structure can be analyzed, thereby significantly improving the efficiency and accuracy of judging the overall organizational quality of the document.
[0037] In one example, a multi-dimensional evaluation process is performed based on at least one of textual information, visual elements, and structural information to obtain an information density and noise ratio evaluation result, including: determining whether the textual information in the presentation document is concise based at least on textual information to obtain a text conciseness evaluation result; determining content-irrelevant elements in the presentation document based at least on visual elements to obtain a visual noise evaluation result; and obtaining an information density and noise ratio evaluation result based on the text conciseness evaluation result and the visual noise evaluation result. Specifically, the information density and noise ratio assessment includes text refinement assessment and visual noise assessment. Text refinement assessment evaluates whether the text on the page represents concise key points, rather than large blocks of unrefined text. Visual noise assessment evaluates whether the page contains too many purely decorative visual elements unrelated to the content, which may interfere with the AI's recognition of core information. The sum of the two assessment results is used as the final information density and noise ratio assessment result.
[0038] In the above embodiments, by evaluating information density and noise ratio in multiple dimensions, the conciseness of document text can be automatically identified and visual interference elements unrelated to the content can be effectively identified, thereby significantly improving the quantitative evaluation ability of information transmission efficiency and visual interference level of presentation documents.
[0039] In one example, the method for retrieving the generated presentation document evaluation further includes: based on the comprehensive evaluation results, identifying content to be optimized in the presentation document, and generating document optimization suggestions for the content to be optimized, wherein the content to be optimized includes images lacking text explanations and / or fragmented pages, and the document optimization suggestions include adding a first presentation annotation text to the images and / or adding a second presentation annotation text to the fragmented pages; and generating a document evaluation report based on the comprehensive evaluation results and the document optimization suggestions, wherein the document evaluation report includes document optimization suggestions and location information of the content to be optimized in the presentation document. For example, the comprehensive evaluation results include weighting at least two of the following: text-image separation evaluation results, context integrity evaluation results, structuredness evaluation results, and information density-to-noise ratio evaluation results, to obtain the final evaluation result. The weight of each part can be set according to actual needs. The content to be optimized includes the content corresponding to the deduction items in the text-image separation evaluation results, context integrity evaluation results, structuredness evaluation results, and information density-to-noise ratio evaluation results. Thus, document optimization suggestions are generated for the content to be optimized and presented in the form of a document evaluation report.
[0040] For example, the scores of each dimension are weighted and averaged to obtain a comprehensive score (comprehensive evaluation result). At the same time, for each deduction item (content to be optimized), PPT-specific optimization suggestions (document optimization suggestions) are generated. For example: (1) For the deduction item: "The system architecture diagram on page 15 is a pure image and lacks text explanation." The generated suggestion is: "Please describe in detail the components of this architecture diagram and their interrelationships in the "Speaker Notes" area on this page so that the AI can understand this diagram." (2) For the deduction item: "The content on pages 8-12 is highly fragmented and cannot be understood independently on a single page." The generated suggestion is: "Please add a summary introduction to the entire topic (pages 8-12) in the "Speaker Notes" on page 8 (the starting page of this topic) to provide the necessary context." The document evaluation report will highlight specific pages with problems in the PPT preview (location information) and provide optimization suggestions. For example, for pages lacking notes, it will explicitly prompt, "Please add detailed speaker notes to this page."
[0041] In the above embodiments, it is possible to automatically identify specific content to be optimized in the presentation document, such as images lacking text explanations and pages with fragmented content, and generate an evaluation report containing specific optimization suggestions and location hints. This significantly improves the pertinence and operability of document evaluation, effectively assists users in quickly locating problems and making precise content optimizations, and greatly improves the efficiency and effectiveness of document quality improvement.
[0042] In one example, the evaluation processing method for retrieving the generated presentation document further includes: in response to receiving a translation instruction for the presentation document, generating a text document based on textual information, visual elements, and structural information.
[0043] Specifically, for an optimization suggestion, the system can provide a "translate into AI-friendly format" function (translation instruction). The AI model will automatically and intelligently reorganize the content of the PPT, especially the key points of the pages and the speaker's notes, to generate a long document (text document) in Markdown or Word format with a clear structure and complete context, as a better source of RAG knowledge.
[0044] In one example, the evaluation processing method for retrieving the generated presentation document further includes: generating questions and corresponding answers for any page or core knowledge content of the presentation document based on textual information, visual elements, and structural information; and performing retrieval, generation, and debugging based on the questions and answers.
[0045] Specifically, after gaining a deep understanding of the PPT content, the AI model can automatically generate several "question-answer" pairs for each page or each core knowledge point, and directly inject these question-answer pairs into the RAG system's fine-tuning data or index to improve the accuracy of answers to specific questions.
[0046] In one example, the evaluation processing method for retrieving the generated presentation document further includes: determining the association information between pages and content in the presentation document based on textual information, visual elements, and structural information; and generating logical relationship data of the presentation document based on the association information.
[0047] Specifically, the system can analyze the page transitions and content relationships (related information) of the entire PPT, generate a visual "logic relationship diagram" (logic relationship data), help users understand the overall presentation, and check for problems such as logical breaks or circular arguments.
[0048] Figure 2 This is a flowchart illustrating another method for evaluating and retrieving generated demonstration documents, provided as an embodiment of this application.
[0049] like Figure 2 As shown, the evaluation processing method 200 for retrieving the generated presentation document includes, for example, S201-S212.
[0050] S201, User uploads PPT document (.pptx). S202, Document Deep Analysis Module. For example, after a user uploads a PPT document, the system performs multimodal and multi-level in-depth analysis on it.
[0051] S203, Text and Notes Extraction. For example, not only should the textbox content (presentation text) on all slide pages be extracted, but more importantly, all the text in the corresponding "Speaker Notes" section (presentation notes text) on each page should also be extracted.
[0052] S204, Visual Element Analysis. For example, multimodal large models or specialized computer vision (CV) models are used to parse non-text elements on a page. For images, image captioning is performed to generate textual descriptions of the image content. For charts, the chart type is identified, and attempts are made to extract its title, data, and conclusions. For flowcharts / architecture diagrams, shapes, arrows, and text are identified, and attempts are made to understand the logical relationships or hierarchical structure they represent.
[0053] S205, Structure and Sequence Analysis. For example, record the page order and heading structure of the document (if heading placeholders are used in the layout).
[0054] S206, AI-powered friendliness rating engine. For example, the AI scoring engine evaluates and scores the parsed information, and its evaluation dimensions are specifically tailored to the characteristics of PPT.
[0055] S206A, Image-text separation evaluation. For example, this includes assessment of core information carriers and assessment of text-image matching.
[0056] S206B, Context Integrity Assessment. Examples include evaluation of presentation note text usage and evaluation of page independence.
[0057] S206C, Assessment of the degree of structuring. Examples include layout evaluation results and logical clarity evaluation.
[0058] S206D, Information Density to Noise Ratio Assessment. Examples include text refinement assessment and visual noise assessment.
[0059] S207, the weighted calculation yields the comprehensive score. For example, the scores of each dimension are weighted and averaged to obtain a comprehensive score (comprehensive evaluation result).
[0060] S208, Intelligent Optimization Suggestion Generation Module. S209 generates targeted optimization suggestions. For example, for each deduction item, generate PPT-specific optimization suggestions (document optimization suggestions).
[0061] S210, front-end visual report generation. S211 presents the analysis results (document evaluation report) in the form of radar charts, scorecards, and highlighted question pages. For example, the analysis report will highlight the pages in the PPT preview that have specific problems and provide optimization suggestions.
[0062] S212, users can view reports and optimize documents. The evaluation and processing method for the generated presentation documents proposed in this application achieves the following: (1) A pioneering AI-oriented PPT document quality evaluation model: For the first time, a set of quantitative standards specifically for evaluating the "AI friendliness" of PPT documents is defined and implemented, especially the two core evaluation dimensions of "text-image separation degree" and "contextual integrity" (with particular emphasis on speaker notes). (2) Comprehensive understanding and evaluation of multimodal content: The core technology of this application goes beyond pure text and can use multimodal AI technology to understand the visual elements (figures, tables, flowcharts) in PPT at the content level and perform correlation analysis with text information, thereby achieving a comprehensive evaluation of the integrity of PPT information. (3) Mining and revitalizing the implicit knowledge value of "speaker notes": For the first time, this application takes "speaker notes" as the core element of evaluation and optimization, and proposes a method to transform this part of implicit knowledge, which has been neglected for a long time and has a huge amount of information, into explicit knowledge assets that can be efficiently utilized by the RAG system.
[0063] Figure 3 A schematic diagram of an evaluation processing apparatus for retrieving generated demonstration documents, provided as another embodiment of this application.
[0064] like Figure 3 As shown, the evaluation processing device 300 for retrieving generated presentation documents includes: a parsing module 310, an evaluation module 320, and a processing module 330.
[0065] The parsing module 310 is used to parse the presentation document to obtain the text information, visual elements and structural information of the presentation document, wherein the text information includes presentation text and / or presentation notes text; The evaluation module 320 is used to perform multi-dimensional evaluation processing based on text information, visual elements and structural information to obtain at least two of the following evaluation results for the presentation document: text-image separation degree evaluation result, context integrity evaluation result, structure degree evaluation result and information density to noise ratio evaluation result. The processing module 330 is used to perform weighted processing on at least two of the following: the text-image separation degree evaluation result, the context integrity evaluation result, the structure degree evaluation result, and the information density to noise ratio evaluation result, to obtain a comprehensive evaluation result of the presentation document.
[0066] For example, the evaluation module 320 is further configured to determine, based at least on text information and visual elements, whether the main information carrier in the presentation document is text or an image, and to determine whether the image has a corresponding text explanation, thereby obtaining a core information carrier evaluation result; to determine, based at least on text information and visual elements, whether the text description in the presentation document is associated with the corresponding image, thereby obtaining a text-image matching degree evaluation result; and to obtain a text-image separation degree evaluation result based on the core information carrier evaluation result and the text-image matching degree evaluation result.
[0067] For example, the evaluation module 320 is also used to determine the page proportion of the presentation notes text in the presentation document based at least on text information and structural information, and obtain the presentation notes text usage rate evaluation result; to determine the page content independence in the presentation document based at least on text information, visual elements and structural information, and obtain the page independence evaluation result; and to obtain the context integrity evaluation result based on the presentation notes text usage rate evaluation result and the page independence evaluation result.
[0068] For example, the evaluation module 320 is further configured to determine whether the layout of the presentation document is a standard layout based at least on structural information, and obtain a layout evaluation result; determine the page order logic and chapter division logic in the presentation document based at least on text information and structural information, and obtain a logical clarity evaluation result; and obtain a structuredness evaluation result based on the layout evaluation result and the logical clarity evaluation result.
[0069] For example, the evaluation module 320 is further configured to determine whether the text information in the presentation document is concise content based at least on the text information, and obtain a text conciseness evaluation result; determine content-irrelevant elements in the presentation document based at least on visual elements, and obtain a visual noise evaluation result; and obtain an information density to noise ratio evaluation result based on the text conciseness evaluation result and the visual noise evaluation result. For example, the evaluation processing apparatus 300 for retrieving generated presentation documents further includes: a first generation module, configured to determine, based on a comprehensive evaluation result, content to be optimized in the presentation document, and generate document optimization suggestions for the content to be optimized, wherein the content to be optimized includes images lacking text explanations and / or fragmented pages, and the document optimization suggestions include adding a first presentation annotation text to the images and / or adding a second presentation annotation text to the fragmented pages; and to generate a document evaluation report based on the comprehensive evaluation result and the document optimization suggestions, wherein the document evaluation report includes document optimization suggestions and location information of the content to be optimized in the presentation document. For example, the evaluation processing apparatus 300 for retrieving the generated presentation document further includes: a second generation module, configured to generate a text document based on text information, visual elements, and structural information in response to receiving a translation instruction for the presentation document.
[0070] For example, the evaluation processing device 300 for retrieving generated presentation documents further includes: a third generation module, used to generate questions and corresponding answers for any page or core knowledge content of the presentation document based on text information, visual elements and structural information; and to perform retrieval, generation and debugging based on the questions and answers.
[0071] For example, the evaluation processing apparatus 300 for retrieving the generated presentation document further includes: a fourth generation module, used to determine the association information between pages and content in the presentation document based on text information, visual elements and structural information; and to generate logical relationship data of the presentation document based on the association information.
[0072] It is understood that the specific functional implementation of the presentation document evaluation processing device 300 used for retrieval and generation can be referred to the presentation document evaluation processing method used for retrieval and generation above, and will not be repeated here.
[0073] Figure 4 A block diagram of an electronic device provided for another embodiment of this application.
[0074] An embodiment of this application provides an electronic device, which includes: a memory, and one or more processors communicatively connected to the memory; the memory stores instructions executable by one or more processors, which are executed by one or more processors to cause the one or more processors to implement the steps of the method of any of the above embodiments.
[0075] like Figure 4 As shown, for ease of understanding, embodiments of this application illustrate a specific electronic device 400.
[0076] Electronic device 400 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 400 may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0077] like Figure 4 As shown, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. The RAM 403 may also store various programs and data required for the operation of the electronic device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0078] Multiple components in electronic device 400 are connected to input / output (I / O) interface 405. These components include: input unit 406, such as a keyboard or mouse; output unit 407, such as various types of displays or speakers; storage unit 408, such as a hard disk or optical disk; and communication unit 409, such as a network interface card (NIC), modem, or wireless transceiver. Communication unit 409 allows electronic device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0079] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods described above. For example, in some embodiments, any one or more of the methods described above can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by the computing unit 401, one or more steps of any one or more of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to perform any one or more of the methods described above by any other suitable means (e.g., by means of firmware).
[0080] Embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method of any of the above embodiments.
[0081] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this application, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0082] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0083] In the description of this application, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this application, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0084] In the description of this application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc., indicating the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.
[0085] Furthermore, the terms "first," "second," etc., used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance, or implicitly specifying the number of technical features indicated in this embodiment. Therefore, features defined with terms such as "first" and "second" in the embodiments of this application can explicitly or implicitly indicate that the embodiment includes at least one of those features. In the description of this application, the word "multiple" means at least two or more, such as two, three, four, etc., unless otherwise explicitly and specifically defined in the embodiments.
[0086] In this application, unless otherwise explicitly specified or limited in the embodiments, the terms "installation," "connection," "joining," and "fixing" appearing in the embodiments should be interpreted broadly. For example, a connection can be a fixed connection, a detachable connection, or an integral part; it can also be a mechanical connection, an electrical connection, etc. Of course, it can also be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication between two components, or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific implementation.
[0087] In this application, unless otherwise expressly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "on top of," and "over" the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.
Claims
1. A method for evaluating generated presentation documents, characterized in that, The method includes: The presentation document is parsed to obtain its text information, visual elements, and structural information, wherein the text information includes presentation text and / or presentation notes text. Based on the text information, the visual elements, and the structural information, a multi-dimensional evaluation process is performed to obtain at least two of the following evaluation results: text-image separation degree evaluation result, context integrity evaluation result, structure degree evaluation result, and information density-to-noise ratio evaluation result of the presentation document. The comprehensive evaluation result of the presentation document is obtained by weighting at least two of the following: the text-image separation degree evaluation result, the context integrity evaluation result, the structuring degree evaluation result, and the information density to noise ratio evaluation result.
2. The method according to claim 1, characterized in that, Based on at least one of the text information, the visual elements, and the structural information, a multi-dimensional evaluation process is performed to obtain the image-text separation degree evaluation result, including: Based at least on the text information and the visual elements, determine whether the main information carrier in the presentation document is text or an image, and determine whether the image has a corresponding text explanation, to obtain the core information carrier evaluation result; Based at least on the text information and the visual elements, determine whether the text description in the presentation document is associated with the corresponding image, and obtain the image-text matching degree evaluation result; Based on the evaluation results of the core information carrier and the evaluation results of the image-text matching degree, the evaluation results of the image-text separation degree are obtained.
3. The method according to claim 1, characterized in that, Based on at least one of the textual information, the visual elements, and the structural information, a multi-dimensional evaluation process is performed to obtain the context integrity evaluation result, including: Based at least on the text information and the structural information, determine the page percentage of the presentation notes text in the presentation document to obtain the presentation notes text usage rate evaluation result; Based at least on the text information, the visual elements, and the structural information, the independence of page content in the presentation document is determined, and a page independence evaluation result is obtained; The context integrity assessment result is obtained based on the usage rate assessment result of the demonstration notes text and the page independence assessment result.
4. The method according to claim 1, characterized in that, Based on at least one of the textual information, the visual elements, and the structural information, a multi-dimensional evaluation process is performed to obtain the structuredness evaluation result, including: Based at least on the structural information, determine whether the layout of the presentation document is a standard layout, and obtain the layout evaluation result; Based at least on the text information and the structural information, determine the page order logic and chapter division logic in the presentation document, and obtain the logical clarity evaluation result; Based on the layout evaluation results and the logical clarity evaluation results, the structuredness evaluation results are obtained.
5. The method according to claim 1, characterized in that, Based on at least one of the text information, the visual elements, and the structural information, a multi-dimensional evaluation process is performed to obtain the information density and noise ratio evaluation result, including: Based at least on the text information, determine whether the text information in the presentation document is concise content, and obtain a text conciseness assessment result; Based at least on the visual elements, identify content-irrelevant elements in the presentation document to obtain visual noise evaluation results; Based on the text refinement evaluation results and the visual noise evaluation results, the information density and noise ratio evaluation results are obtained.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: Based on the comprehensive evaluation results, the content to be optimized in the presentation document is identified, and document optimization suggestions are generated for the content to be optimized. The content to be optimized includes images lacking text explanations and / or fragmented pages. The document optimization suggestions include adding a first presentation note text to the images and / or adding a second presentation note text to the fragmented pages. Based on the comprehensive evaluation results and the document optimization suggestions, a document evaluation report is generated, wherein the document evaluation report includes the document optimization suggestions and location information of the content to be optimized in the presentation document.
7. The method according to any one of claims 1-5, characterized in that, The method further includes: In response to receiving a translation instruction for the presentation document, a text document is generated based on the text information, the visual elements, and the structural information.
8. The method according to any one of claims 1-5, characterized in that, The method further includes: Based on the text information, the visual elements, and the structural information, questions and corresponding answers are generated for any page or any core knowledge content of the demonstration document. The search results are generated and debugged based on the question and the answer.
9. The method according to any one of claims 1-5, characterized in that, The method further includes: Based on the text information, the visual elements, and the structural information, determine the association information between pages and content in the presentation document; Based on the aforementioned association information, logical relationship data for the demonstration document is generated.
10. An apparatus for retrieving and evaluating generated presentation documents, characterized in that, The device includes: The parsing module is used to parse the presentation document to obtain the text information, visual elements and structural information of the presentation document, wherein the text information includes presentation text and / or presentation notes text; The evaluation module is used to perform multi-dimensional evaluation processing based on the text information, the visual elements, and the structural information to obtain at least two of the following evaluation results for the presentation document: text-image separation degree evaluation result, context integrity evaluation result, structure degree evaluation result, and information density to noise ratio evaluation result. The processing module is used to perform weighted processing on at least two of the following: the text-image separation degree evaluation result, the context integrity evaluation result, the structuring degree evaluation result, and the information density to noise ratio evaluation result, to obtain a comprehensive evaluation result of the presentation document.