Method, device and equipment for generating PowerPoint, medium and program product
By automatically generating precisely matched slides and lecture notes through a large language model and combining it with the principles of educational psychology, the problem of existing presentations requiring manual operation and updating is solved, an automated, multimodal presentation experience is achieved, and the engagement and efficiency of users and audiences are improved.
Patent Information
- Application Number
- CN202510906368.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-17
AI Technical Summary
Existing presentations require users to manually operate slides and provide oral explanations, which prevents the audience from actively participating. Moreover, when the content is updated, the entire presentation needs to be remade, which wastes time and is inefficient.
It automatically generates precisely matched slides and lecture notes through a large language model, combines the principles of educational psychology, and provides automated multimodal content presentation, including voice synchronization and interactive elements, to adapt to different device screen sizes.
It achieves precise matching and synchronous display of slides and lecture notes, improves user experience, saves time, supports active audience participation and facilitates content updates.
Smart Images

Figure CN120804346A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to the field of computers, and more specifically to a method, an apparatus, a device, a computer-readable storage medium, and a computer program product for generating a presentation. BACKGROUND
[0002] Machine learning techniques based on large language models have developed excellent language understanding and generation capabilities. These models not only can accurately parse complex natural language instructions of humans, but also can quickly generate various forms of content according to different needs.
[0003] In information dissemination, an effective way to deliver knowledge is to seamlessly combine visual information such as slides with real-time explanations by the speaker. Through the presentation of slides, the audience can quickly grasp the new context. Through the presentation of slides, abstract concepts are made concrete, complex mechanisms are analyzed step by step, and the audience's sense of immersion is enhanced to increase cognitive depth. At the same time, the courseware based on slides can be modified and used multiple times, reducing repetitive labor and facilitating focus on the content of the speech. SUMMARY
[0004] According to an example embodiment of the present disclosure, a method, an apparatus, a device, a computer storage medium, and a computer program product for generating a presentation are provided.
[0005] In a first aspect of the present disclosure, a method for generating a presentation is provided, the method comprising obtaining a question for generating a presentation. The method further comprises displaying a presentation for the question, the presentation being generated by a target model based on the question, and the presentation comprising at least a slide, a script corresponding to the slide, or multi-modal content.
[0006] In a second aspect of the present disclosure, an apparatus for generating a presentation is provided, the apparatus comprising an obtaining module configured to obtain a question for generating a presentation. The apparatus comprises a displaying module configured to display a presentation for the question, the presentation being generated by a target model based on the question, and the presentation comprising at least a slide, a script corresponding to the slide, or multi-modal content.
[0007] In a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processing unit; at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method described in the first aspect of the present disclosure.
[0008] In a fourth aspect of the disclosure, there is provided a computer readable storage medium having stored thereon machine executable instructions, which when executed by a device, cause the device to perform the method described according to the first aspect of the disclosure.
[0009] In a fifth aspect of the disclosure, there is provided a computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method described according to the first aspect of the disclosure.
[0010] The summary is provided to introduce a selection of concepts that are further described in the detailed description below. It is not intended to identify key or essential features of the disclosure or to delineate the scope of the disclosure. Other features of the disclosure will be apparent from review of the disclosure herein. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 A schematic diagram showing an example environment in which embodiments of the disclosure can be implemented;
[0012] Figure 2 A flow diagram showing a method for generating a presentation according to embodiments of the disclosure;
[0013] Figure 3 A further process flow diagram showing a method for generating a presentation according to certain embodiments of the disclosure;
[0014] Figure 4 A schematic diagram showing a process for generating a presentation according to embodiments of the disclosure;
[0015] Figure 5 A schematic block diagram showing an example apparatus according to some embodiments of the disclosure;
[0016] Figure 6 A block diagram showing an example device that can be used to implement embodiments of the disclosure.
[0017] In all of the drawings, like or similar reference numerals are used to refer to like or similar elements throughout different views. DETAILED DESCRIPTION
[0018] The names of the messages or information exchanged between the plurality of apparatuses in the embodiments of the disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information. It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0019] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the use range, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0020] For example, when receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will need to obtain and use the personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium, etc. performing the operation of the technical solutions of the present disclosure according to the prompt information.
[0021] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be a pop-up window manner, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0022] It can be understood that the above notification and obtaining of the authorization of the user are only illustrative, and do not limit the implementation manners of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manners of the present disclosure.
[0023] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0024] In the description of the embodiments of the present disclosure, the term "comprising" and similar terms thereof should be understood as open-ended including, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. can refer to different or same objects unless explicitly stated. The following can also include other explicit and implicit definitions.
[0025] In the related art, the current presentation only has slides, and usually needs the user to manually operate the slides and cooperate with oral explanation. However, the user needs a lot of time to explain on site. Secondly, the current slides cannot actively involve the audience, and when the content to be demonstrated is updated, the entire presentation needs to be re-produced.
[0026] To at least address the aforementioned and other potential issues, embodiments of the present disclosure provide a method for generating presentations. The method implemented in the present disclosure accurately matches each slide in a presentation with the corresponding lecture notes, converts the lecture notes into speech, and presents them synchronously with the slides, thereby providing a complete automated presentation process, saving time and improving the user experience.
[0027] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Figure 1 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Figure 1 As shown, according to some embodiments of the present disclosure, a user 102 such as a student or a teacher can input the subject or course content that needs to be learned or taught into a target model 104 such as a large language model. As an example, the input content can be the name of a subject, or it can be a course chapter or unit, such as mathematics, physics, history, artificial intelligence and the like. In some embodiments, the user 102 can also input the learning goals or teaching goals that they hope to achieve through the generated courseware content or teaching content 106. It should be understood that although the teaching scenario is used as an example, the method implemented in the present disclosure can also be applied to other scenarios, such as product display, conference discussion, speech, project management and the like, and the present disclosure does not impose any restrictions on this.
[0028] According to some embodiments of the present disclosure, the user 102 can input prompt words such as specific information or topics that the user desires to master, such as generating content about a presentation. In some embodiments of the present disclosure, the user can also specify the difficulty level of the generated content, such as elementary, intermediate, advanced, etc. Additionally or alternatively, in some embodiments, the user 102 can specify the teaching style of the courseware content 106 to be generated for different types of students to be taught. For example, in some embodiments, depending on whether the students to be taught are middle school students or college students, the user can choose whether the voice content is more formal or colloquial, and whether the generated teaching content is more vivid and interesting or more concise and highlights the key points.
[0029] According to some embodiments of the present disclosure, upon receiving a request input by a user 102, a target model 104, such as a large language model, can automatically generate courseware content 106 that conforms to the principles of educational psychology, including slides and corresponding lecture text. According to embodiments of the present disclosure, the target model 104 can be pre-trained by reading a large amount of educational psychology literature or relevant knowledge of the principles. For example, the target model 104 can be pre-trained to determine the presentation method of teaching content to avoid excessive information that interferes with student learning, and to promote student learning outcomes through multimodal learning methods such as visual and auditory learning.
[0030] According to some embodiments of the present disclosure, the target model 104 can autonomously generate the course content 106 by analyzing the user 102 input. For example, in some embodiments, the course content 106 includes one or more slides or PPTs 108, 112 and corresponding scripts 110, 114 that are associated with each other. In some embodiments of the present disclosure, the target model 104 can distribute the course content 106 into the slides 108-112, such as a slide 108 for formula derivation and a slide 112 for example explanation, respectively.
[0031] In some embodiments of the present disclosure, each of the one or more slides 108-112 generated by the target model 104 can include a respective title, key content and explanation, and corresponding images. In some embodiments, each of the one or more slides 108-112 intelligently generated by the target model 104 based on the teaching content can include corresponding interactive elements, such as operable graphs, adjustable parameter demonstrations, etc., to help students understand complex concepts through interaction with these elements.
[0032] Additionally or alternatively, in some embodiments of the present disclosure, the target model 104 can automatically optimize the visual presentation of the slides according to educational design principles, such as optimizing the layout of each slide, including responsive layout, position of title, body, images, font system, and color palette, etc., to ensure clarity and consistency of visual effects and improve user experience.
[0033] According to some embodiments of the present disclosure, the target model 104 can also simultaneously generate the scripts 110 and 114 corresponding to the one or more slides 108-112. In some embodiments, the scripts 110 and 114 are auxiliary materials for teachers when explaining the slides, and can include more detailed explanations and supplementary extended content. For example, in some embodiments of the present disclosure, the scripts 110 and 114 generated by the target model 104 can provide in-depth explanations of each knowledge point. In some embodiments of the present disclosure, the target model 104 can also adjust the language style of the scripts 110 and 114 based on the preferences of the teachers or the needs of the students to generate more formal academic language scripts or more interactive scripts. In some embodiments, the scripts 110 and 114 can also include prompts for interaction with students, such as questions, case analysis, etc., to promote student interaction with the course content 106 and help students focus their attention.
[0034] In some embodiments of the present disclosure, for example, target model 104 may organize slides 108 through 112 and corresponding lecture notes 110 and 114 using a "one slide, one lecture" markup language format to form a standardized output format. In other words, the slides generated by target model 104 can correspond one-to-one with the lecture notes. The methods implemented in the present disclosure enable precise matching of the content of the slides generated by target model 104 with the lecture notes.
[0035] According to the methods implemented in the present disclosure, the target model 104 can generate corresponding slides and corresponding audio notes in a streaming manner. In some embodiments, for example, the target model 104 can use asynchronous loading and lazy loading to enable parallel loading and rendering of slide text, images, animations, and other content. This allows the content of each slide to be displayed immediately after generation, without having to wait for all slides to load. Users can interact with the slides in real time, such as by clicking buttons and dragging sliders.
[0036] In some examples, the backend can automatically call a voice API to convert the speech into speech, which is then displayed synchronously with the slides, providing a complete presentation experience. This allows users to more deeply understand the content they need to learn by combining multimodal content such as images, videos, and interactive elements. In some embodiments, the target model can also directly generate a video for the presentation based on questions posed by the user.
[0037] As understood by those skilled in the art, the server instance where the target model 104 resides can be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The servers can be connected directly or indirectly via wired or wireless communication, which is not limited in this application.
[0038] The user device used by user 102 can be any type of mobile computing device, including a mobile computer (e.g., a personal digital assistant, laptop computer, notebook computer, tablet computer, netbook computer, etc.), a mobile phone (e.g., a cellular phone, a smartphone, etc.), a wearable computing device (e.g., a smartwatch, a head-mounted device, including smart glasses, etc.), or other types of mobile devices. In some embodiments, the user device can also be a stationary computing device, such as a desktop computer, a game console, a smart TV, etc. It should be understood that if the user device has sufficient computing power, the user device can replace the server to complete the above operations, or the user device and the server can jointly complete the above operations.
[0039] It should be understood that the architecture and functionality of the example environment 100 is described for illustrative purposes only and is not meant to limit the scope of the present disclosure. Embodiments of the present disclosure can be applied to other environments with different structures and / or functionality, without departing from the scope of the present disclosure.
[0040] The processes according to embodiments of the present disclosure will be described below in detail with reference to other drawings. For ease of understanding, the specific data mentioned in the following description are all exemplary and are not used to limit the protection scope of the present disclosure. It can be understood that the embodiments described below can also include additional actions not shown and / or can omit the actions shown, and the scope of the present disclosure is not limited in this respect.
[0041] Figure 2 A flowchart of a method 200 for generating a presentation according to certain embodiments of the present disclosure is shown. In this embodiment, the present method can be performed by an application program of a user device of a user 102. In block 202, a question for generating a presentation is obtained. The input provided by the user can include learning objectives, question descriptions or specific requirements, how to make product presentation introductions, how to make conference presentations, etc. For example, the user can raise the question of "help me explain the Pythagorean theorem". In some embodiments, the presentation can include multiple slides and corresponding scripts.
[0042] The target model can determine the presentation based on these questions. In some embodiments, the presentation can include a structured text format of the slides and the script, including which knowledge points, product information pictures, charts or examples need to be shown, and how to explain these contents in detail through the explanation text. In some embodiments, each page of the slides in the slides is interleaved with the corresponding script, and the slides are in a structured text format.
[0043] In some embodiments, the target model can adjust the slides and the script. For example, the target model can adjust the target style of the slides, such as the format, content, size, style, etc. of the images, text, charts, animations, audio, video of the multi-modal content in the slides, so that the adjusted slides are more beautiful, suitable for the user to read and maintain attention.
[0044] In some embodiments, the script can also be converted into voice content and matched with the slides one by one, so that the user can listen to the corresponding teaching content while watching. In some embodiments, the layout of the slides can also be adjusted based on the size of the display screen of different devices to adapt to different display devices.
[0045] At block 204, a presentation is displayed for the question. In some embodiments, the presentation is generated by the target model based on the question, and the presentation includes at least slides, a script corresponding to the slides, or multi-modal content. In some embodiments, the content of the slides and the script can be interleaved. In some embodiments, one or more of the multi-modal content can be outputted in sequence based on an output priority of the one or more of the multi-modal content. For example, the text content can be inputted first, and the corresponding images and videos can be outputted later, so that the user can see the script without waiting, improving the user experience. In some embodiments, the multi-modal content includes text, images, charts, animations, audio, video, interactive components, and the like.
[0046] According to the method implemented by the present disclosure, the precise matching of each slide in the slides with the corresponding script, the synchronization with the slides, and the diversified content display can be achieved, thereby providing a complete automated presentation process and improving the user experience.
[0047] Figure 3 Another process flow diagram of the method 300 for generating a presentation according to some embodiments of the present disclosure is shown. As shown at block 302, a user can input the content desired to be learned. In embodiments of the present disclosure, the input of the user can include a course topic outline, or more detailed content requirements, such as specific subject knowledge points, required teaching material types, and interactive elements, and the like. Figure 3
[0048] At block 304, according to some embodiments of the present disclosure, after receiving the requirements inputted by the user, the content generation module of the target model can generate courseware content conforming to the principles of educational psychology based on the inputted content. As an example, the content generation module can be a large language model based on natural language processing or deep learning models, which is not limited by the present disclosure. In some embodiments, the courseware content can include slides and corresponding scripts.
[0049] In some embodiments, for example, the target model can extract the intention in the user input through morphological analysis, syntactic analysis, and semantic understanding, and determine the user requirements in combination with understanding the background information and context of the user. In some embodiments, the target model can retrieve and integrate information based on a pre-trained related knowledge base to determine the script. The target model can subdivide the generated script into multiple easy-to-understand explanation points. The target model can determine the characteristics of these explanation points, such as data characteristics, difficulty, complexity, and interactivity of the explanation points.
[0050] The target model can generate corresponding suitable multi-modal content for these knowledge points, for example, chart content can be suitable for explaining complex data, text can be used to explain key concepts, and interactive components can be used for testing or experiments.
[0051] As an example, when a user inputs a desire to learn about the rotation and revolution of the earth, the content generation module can automatically generate images based on this input, such as dynamic images of the rotation and revolution of the earth, trajectory graphs, etc., and insert them into the slides. In some examples, the content generation module can also generate scripts corresponding to these images to facilitate understanding, such as the script "The rotation of the earth refers to the rotation of the earth around the rotation axis from west to east, which appears to rotate counterclockwise from above the north pole and clockwise from above the south pole."
[0052] According to some embodiments of the present disclosure, at block 306, the content interleaving output module can use a specific markup language to interleavingly output the generated slide content and scripts in a manner corresponding to the slides and scripts, for example, in a manner of one page of slides corresponding to one page of scripts or one page of slides corresponding to several pages of scripts, so as to form standardized output content. For example, as an example, the output content can include: <script>Let's learn about the rotation of the earth today...< / script>; <ppt>Slide content of the Earth's rotation...< / ppt> .
[0053] At block 308, the interactive element generation module can also analyze the content characteristics based on the generated content and add interactive functions. For example, in some embodiments, the interactive element generation module can determine which types of interactive elements by analyzing the content characteristics. For example, in some embodiments, the interactive element generation module can generate an interactive earth rotation animation, in which students can control the rotation speed, direction and angle of the earth, so as to observe the day and night change phenomenon of the earth under different rotation conditions. Students can control the rotation speed by dragging or clicking the controller to speed up or slow down, or adjust the direction of rotation.
[0054] In some embodiments, students can also select different geographical positions such as the equator, the north pole or the south pole to observe the length change of the local day and night. Additionally or alternatively, in some embodiments, in yet another embodiment, when the user inputs a mathematical graph or formula, the interactive element generation module can also generate an interactive chart or function image. In some embodiments, when the user inputs a demonstration of a physical experiment, the interactive element generation module can generate a corresponding adjustable simulation experiment and associated parameters. According to some embodiments of the present disclosure, the interactive element generation module can support the implementation of interactive functions by automatically referencing necessary Web libraries.
[0055] In some embodiments, for example, a user can change the style of an existing element in a slide by clicking on the corresponding interactive element. For example, when a test question in a slide is answered correctly, the text can turn green, while when a test question in a slide is answered incorrectly, the text can turn red. In this way, the content, color, font, size, playing status, etc. of multi-modal content such as text, image, chart, audio, video, animation can be changed.
[0056] According to some embodiments of the present disclosure, at 310, the style and layout optimization module can automatically optimize the visual presentation or target style of the generated content according to the educational design principles, for example, including responsive layout, color matching, font system, and visual hierarchy, so as to ensure that the generated content can have a good visual experience on different devices.
[0057] For example, in some embodiments, the style and layout optimization module can automatically adjust the layout based on the screen size for different devices such as mobile phones, tablets, and desktop computers. For devices with smaller screens such as smart phones, the style and layout optimization module can automatically adjust the content to a portrait layout and optimize the size of the text and pictures to make them more suitable for touch operation. For devices with larger screens of desktop computers, the style and layout optimization module can provide a more complex layout such as multi-column display, so as to facilitate students to view the lecture notes and interactive content at the same time. In some embodiments, the style and layout optimization module can automatically adjust the color and font so that students can clearly read and maintain attention.
[0058] At block 312, the generated content can be distributed based on the target model. For example, at block 314, the generated slide content can be distributed to a front-end rendering engine. The front-end rendering engine can structure and organize the slide content into different parts such as images, text, audio, video, etc. The front-end rendering engine can ensure that the content can be quickly loaded and smoothly displayed on different devices. Then the slides are gradually presented based on user interaction or time control. For example, in some embodiments, the next or previous slide can be presented based on clicking the "next page" or "previous page" button. In some embodiments, the slides can be automatically switched after a certain time interval by setting a timer.
[0059] According to some embodiments of the present disclosure, at block 316, the speech can also be distributed to the speech conversion module. For example, the speech conversion module can invoke a speech API (e.g., based on speech synthesis TTS) to convert the speech into fluent speech. At block 318, the speech synchronization module can coordinate the speech playback to be synchronized with the slide transition. For example, during the speech playback, the speech synchronization control module can implement the content displayed in the slide to be synchronized with the speech content, so that the speech playback and the slide presentation remain in pace. For example, the speech synchronization control module can automatically highlight the relevant diagram part when the speech plays to “earth rotation”, so that the student can understand the diagram content in synchronization when hearing the relevant explanation. Additionally or alternatively, the presentation including the slides and the speech can also be converted into video content by adjusting and rendering the aligned timeline, etc., so that the user can obtain the corresponding information by watching the video.
[0060] According to some embodiments of the present disclosure, at block 320, the student can interact with the generated slides on the user interface by clicking, sliding, zooming, etc. For example, in an embodiment, when the student selects a time point in the interactive chart, the data of the time point can be updated and displayed in real time, and the corresponding explanation can be popped up, so as to facilitate the student to understand the meaning behind the data. In some embodiments, the user can provide feedback based on the generated content, and the feedback can be used as input to generate the presentation again.
[0061] For example, in some embodiments, the type of feedback can be determined, and according to some embodiments of the present disclosure, the type of feedback can include feedback on the content of the slides, feedback on the content of the speech, feedback on the interaction, etc. For example, the feedback on the content of the slides can include feedback on the integrity, accuracy, visual design (font, color, layout), chart, image and animation effect of the content in the slides. The feedback on the content of the speech can include feedback on the logical structure of the explanation, the clarity of the expression, the depth, the language style, the adaptability to the target audience, etc. In some embodiments, the feedback on the interaction can include feedback on the user's behavior data, such as playback time, pause frequency, number of times of playing back a specific slide, turning to a certain part of the content, etc.
[0062] By analyzing these feedbacks, the content in the presentation can be further modified. For example, in some embodiments, the content of the text in the slides, the color of the text, the font of the text, the size, etc. can be modified so that the information is easy to read. In some embodiments, the color scheme can also be adjusted, the color scheme can be reselected, the charts and images can be added or optimized, and the complexity of the animation can be adjusted or unnecessary animation effects can be removed, etc. In some embodiments, the script can also be optimized, for example, the tone, formality, professionalism, etc. of the explanation content can be optimized, personalized adjustment can be made based on the personal preferences and styles of the user, and adaptation can be made for specific audiences.
[0063] By the method implemented by the present disclosure, the target model such as a large language model can autonomously generate and simultaneously present multi-modal outputs such as voice, text, web pages, visualizations, pictures, etc. to provide an automatic and personalized content generation and presentation process for education, product introduction, conference seminars, etc. so that the user can listen and watch at the same time to achieve an immersive experience.
[0064] Figure 4 A schematic diagram of a process 400 for generating a presentation according to an embodiment of the present disclosure is shown. As shown, a user can input a requirement in an application 402 implemented according to the method of the present disclosure, such as explaining artificial intelligence, and click a send button to put forward their learning requirement to the target model. Figure 4
[0065] After receiving this request, the target model can automatically determine the knowledge and information associated with the question and include these contents in the presentation of the corresponding slide(s) 404 and the corresponding script(s) 414. In some embodiments, a curve or chart can be included in the slide 404 for displaying the corresponding data information to facilitate the user to more intuitively understand complex data, such as the historical evolution process of reinforcement learning. In some embodiments, the script 414 can include a more detailed description corresponding to the slide 404 for introducing the relevant concepts in the slide.
[0066] In some embodiments, the slide 404 can further include corresponding textual content 408 to facilitate the user to more clearly and intuitively understand some concepts or definitions associated with the content. In some embodiments, the slide 404 can further include corresponding animations 410 to visually demonstrate the corresponding content of the defined concepts. Additionally or alternatively, in some embodiments, the slide 404 can further include video, audio, or other elements to facilitate the user to visually understand the content. In some embodiments, the slide 404 can further include interactive components 412 such as sliders, which can be dragged by the user to change the slide to display different content. Through the method implemented by the present disclosure, the user can be helped to more intuitively understand the content to be learned, and through the combination of text and images and interactive display, the complex concepts can be made more easily understood.
[0067] Figure 5 A schematic block diagram of an apparatus 500 for generating a presentation is shown according to some embodiments of the present disclosure. The apparatus 500 can be implemented by software, hardware, or a combination of both. As shown, the apparatus 500 includes an obtaining module 510 and a displaying module 530. Figure 5
[0068] In some embodiments, the obtaining module 510 is configured to obtain a question for generating a presentation. In some embodiments, the displaying module 530 is configured to display the presentation for the question, the presentation being generated by a target model based on the question, and the presentation including at least a slide, a script corresponding to the slide, or multi-modal content.
[0069] In some embodiments, the apparatus 500 further includes an adjusting module configured to adjust the script, including converting the script into speech content, respectively; and matching the speech content with the slide.
[0070] In some embodiments, the apparatus 500 further includes a generating module configured to generate a video for the presentation.
[0071] In some embodiments, the apparatus 500 further includes an outputting module configured to display one or more of the slide, the script corresponding to the slide, or the multi-modal content in a streaming output. In some embodiments, the outputting module is configured to sequentially output the one or more of the multi-modal content based on an output priority of the one or more of the multi-modal content.
[0072] In some embodiments, the multi-modal content includes one or more of text, image, chart, animation, audio, video, and interactive component. In some embodiments, the apparatus 500 further includes a feedback module configured to receive feedback of the presentation from a user and regenerate the presentation based on the feedback.
[0073] In some embodiments, the feedback module is configured to determine a type of the feedback, wherein the type of the feedback comprises feedback on content of the slide, feedback on content of the script, feedback on the interaction. In some embodiments, the apparatus 500 further comprises a modification module configured to modify the presentation based on the type of the feedback, including modifying the multi-modal content of the slide and modifying the language style, tone, voice of the script.
[0074] In some embodiments, the modification module is configured to modify the slide including modifying content, color, font, size of text of the slide, modifying color, size of image of the slide, regenerating another image on the slide, modifying content, color, font, size of chart of the slide, regenerating another chart on the slide, modifying content of animation, audio, video of the slide, or regenerating another animation, audio, video on the slide.
[0075] In some embodiments, the apparatus 500 further comprises a changing module configured to change the multi-modal content in the slide in response to the user interaction with the interactive component in the slide. In some embodiments, the presentation is generated by the target model based on the question, including the user needs in the question are analyzed by the target model; the script is generated by the target model based on the user needs; the script is decomposed into one or more talking points by the target model; the multi-modal content associated with the one or more talking points is generated by the target model; and the slide is generated by the target model based on the multi-modal content.
[0076] In some embodiments, the multi-modal content associated with the one or more talking points is generated by the target model, including: features of the one or more talking points are determined by the target model; the multi-modal content suitable for presentation of the one or more talking points is determined by the target model based on the talking features of the one or more talking points and the user needs; and the multi-modal content is presented on the slide by the target model.
[0077] In some embodiments, the first slide and the second slide are adjusted by the target model, including the layout of the slide is adjusted by the target model based on the size of the display screen of the different devices.
[0078] Figure 6 A block diagram of an example device 600 that can be used to implement embodiments of the present disclosure is shown. It should be understood that Figure 6 The device 600 shown is merely an example and should not be construed as limiting the functionality and scope of the implementations described herein. For example, the device 600 can correspond to the user devices described herein in connection with Figure 1 the processes described above. For another example, the device 600 can correspond to the electronic device of the third aspect of the summary section. Figures 1 to 5 The device 600 shown is merely an example and should not be construed as limiting the functionality and scope of the implementations described herein. For example, the device 600 can correspond to the user devices described herein in connection with Figure 1 the processes described above. For another example, the device 600 can correspond to the electronic device of the third aspect of the summary section. Figures 1 to 5 The device 600 shown is merely an example and should not be construed as limiting the functionality and scope of the implementations described herein. For example, the device 600 can correspond to the user devices described herein in connection with Figure 1 the processes described above. For another example, the device 600 can correspond to the electronic device of the third aspect of the summary section.
[0079] like Figure 6 As shown, device 600 is in the form of a general-purpose computing device. Components of device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit 610 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of device 600.
[0080] Device 600 typically includes multiple computer storage media. Such media can be any available media accessible to device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory) or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks or any other media, which can be used to store information and / or data (e.g., training data for training) and can be accessed within device 600.
[0081] The device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 6 As shown in FIG, a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. Memory 620 may include a computer program product 625 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
[0082] The communication units 640 enable communications with other computing devices over a communication medium. Additionally, the functionality of the components of device 600 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating with each other through a communication connection. Thus, device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.
[0083] The input device 650 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 660 can be one or more output devices, such as a display, a speaker, a printer, etc. The device 600 can also communicate with one or more external devices (not shown), such as a storage device, a display device, etc., through the communication unit 640, as needed, communicate with one or more devices that enable a user to interact with the device 600, or any devices (e.g., a network card, a modem, etc.) that enable the device 600 to communicate with one or more other computing devices. Such communication can be carried out via an Input / Output (I / O) interface (not shown).
[0084] According to example implementations of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to example implementations of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above. According to example implementations of the present disclosure, a computer program product is provided having a computer program stored thereon, which when executed by a processor implements the method described above.
[0085] Various aspects of the disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0086] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0087] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0088] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0089] While various implementations of the present disclosure have been described above, the foregoing descriptions are illustrative and non-exhaustive, and are not intended to limit the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for generating a presentation, comprising: Get the questions used to generate the presentation; as well as The presentation for the question is displayed, where the presentation is generated by a target model based on the question, and the presentation includes at least slides, lecture notes corresponding to the slides, or multimodal content.
2. The method according to claim 1, further comprising adjusting the lecture script, comprising: Converting the lecture notes into voice content respectively; as well as Matching the voice content with the slides.
3. The method of claim 2, wherein displaying the presentation for the problem further comprises generating a video for presentation.
4. The method of claim 1 , wherein displaying the presentation for the question comprises: The slides, the lecture notes corresponding to the slides, or one or more of the multimodal contents are displayed in a streaming output.
5. The method of claim 4 , wherein displaying one or more of the multimodal contents in a streaming output comprises: The one or more contents in the multimodal contents are output in sequence based on an output priority of the one or more contents in the multimodal contents.
6. The method according to claim 1, wherein the multimodal content includes one or more of text, images, charts, animations, audio, video, and interactive components.
7. The method according to claim 6, further comprising: receiving user feedback on the presentation; as well as The presentation is regenerated based on the feedback.
8. The method of claim 7, wherein regenerating the presentation based on the feedback comprises: Determining the type of the feedback, wherein the type of feedback includes feedback on the content of a slide, feedback on the content of the lecture, and feedback on interaction; Modifying the presentation based on the type of feedback includes: modifying the multimodal content of the slideshow; as well as Modify the language style, voice, and intonation of the speech.
9. The method of claim 8, wherein modifying the slideshow comprises: Modify the content, color, font, and size of the text on the slide; Modify the color and size of the image of the slide; regenerating another image on the slide; Modify the content, color, font, and size of the chart on the slide; regenerating another chart on the slide; Modify the animation, audio, and video content of the slideshow; or Regenerate another animation, audio, video on the slide.
10. The method according to claim 9, further comprising: In response to the user interacting with an interactive component in the slideshow, the multimodal content in the slideshow is changed.
11. The method of claim 1 , wherein the presentation is generated by the goal model based on the question, comprising: The user needs in the problem are analyzed by the target model; The lecture notes are generated by the target model based on the user needs; The lecture script is decomposed into one or more explanation points by the target model; Multimodal content associated with the one or more teaching points is generated by the target model; and The slides are generated by the target model based on the multimodal content.
12. The method according to claim 11, wherein the multimodal content associated with the one or more teaching points is generated by the target model, comprising: The characteristics of one or more explanation points are determined by the target model; The multimodal content suitable for presentation at the one or more explanation points is determined by the target model based on the explanation features of the one or more explanation points and the user needs; and The multimodal content is presented on the slide by the target model.
13. The method according to claim 1, wherein adjusting the slide by the target model comprises: The layout of the slides is adjusted by the target model based on the size of the display screens of different devices.
14. A device for generating a presentation, comprising: an acquisition module configured to acquire questions for generating a presentation; as well as The display module is configured to display the presentation for the question, where the presentation is generated by the target model based on the question, and the presentation includes at least slides, lecture notes corresponding to the slides, or multimodal content.
15. An electronic device comprising: at least one processing unit; At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1 to 13.
16. A computer program product having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
PowerPoint generation method and device, equipment and storage medium
CN114021541A
Powerpoint generation method and device
CN116579308A
Lost property management system and program
JP7371843B1
Enhanced Knowledge Delivery and Attainment Using a Question Answering System
US20200097598A1
Communication system for updating lecture text using slide specific user data
WO2025026802A1