Slide interactive editing method and device based on artificial intelligence

Through artificial intelligence technology, a slide editing method based on human-computer interaction interface is provided. After the user selects the target object, the system generates candidate content, which solves the problem of cumbersome operation and inconsistent content of traditional PPT editing tools, and realizes efficient and accurate slide editing.

CN120578445APending Publication Date: 2025-09-02BEIJING SANSAN SMART EDUCATION TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510739193.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Traditional PPT editing tools are cumbersome and inefficient when inserting pictures. The content of the picture is difficult to match the document theme and style, and existing tools are difficult to deeply integrate.

Method used

Through artificial intelligence technology, a slide editing method based on human-computer interaction interface is provided. After the user selects the target object, enters the editing requirements through the editing component, the system generates candidate content and updates the target object, supporting the accurate analysis and generation of text and graphic requirements.

Benefits of technology

The editing process is simplified, editing efficiency is improved, and the generated content is accurately matched with user expectations, meets personalized needs, and optimizes the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578445A_ABST
    Figure CN120578445A_ABST
Patent Text Reader

Abstract

The invention provides a slide interactive editing method and device based on artificial intelligence, and the method comprises the steps: displaying a slide through a human-computer interaction interface, enabling a user to select a picture or a picture occupying frame as a to-be-edited object, triggering an editing assembly, entering an editing interface, and inputting an editing demand; the artificial intelligence generates at least one candidate content according to requirements, and after the user confirms the target content, the system applies the target content to the target object to complete updating; according to the invention, efficient, accurate and personalized slide editing can be realized by means of artificial intelligence, and meanwhile, the interaction experience is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based slide interactive editing method and device. Background Art

[0002] In modern office environments, PowerPoint (PowerPoint presentations) are a common presentation tool. However, traditional PowerPoint editing tools have the following issues: First, when inserting images, users need to use search engines or other methods to search for them, which is cumbersome and inefficient. Second, even if users find the appropriate image, the image content may not be consistent with the theme and style of the current document. Third, existing document image editing tools often exist as independent programs or windows, making them difficult to deeply integrate with PowerPoint editors.

[0003] Therefore, there is an urgent need for a more efficient and convenient slide interactive editing method to reduce the user's operation steps and time cost. Summary of the Invention

[0004] In view of this, the embodiments of the present application provide an artificial intelligence-based interactive slide editing method, device, electronic device and storage medium, which can use artificial intelligence to achieve efficient, accurate and personalized slide editing while optimizing the interactive experience.

[0005] The technical solution of the embodiment of the present application is implemented as follows: In a first aspect, an embodiment of the present application provides an artificial intelligence-based interactive slide editing method, wherein the slide is displayed through a human-computer interaction interface, and the method includes: In response to a selection operation on a target object in the slide, determining the selected target object as an object to be edited; wherein the target object includes a picture or a picture placeholder; In response to a triggering operation on an editing component, an editing interface of the editing component is displayed in the human-computer interaction interface; wherein the editing interface is used to edit the object to be edited; In response to an input operation on the editing interface, displaying an editing requirement in the editing interface; In response to a submission operation for the editing requirement, displaying at least one candidate content in the editing interface; wherein the at least one candidate content is generated by the artificial intelligence based on the editing requirement; In response to a confirmation operation on target content in the at least one candidate content, the target object is updated based on the target content.

[0006] In a second aspect, an embodiment of the present application further provides an artificial intelligence-based interactive slide editing device, which displays the slides through a human-computer interaction interface, and the device includes: A selection module, configured to respond to a selection operation on a target object in the slide and determine the selected target object as an object to be edited; wherein the target object includes a picture or a picture placeholder; A trigger module, configured to respond to a trigger operation on an editing component and display an editing interface of the editing component in the human-computer interaction interface; wherein the editing interface is used to edit the object to be edited; An input module, configured to respond to input operations on the editing interface and display editing requirements in the editing interface; a display module configured to display at least one candidate content in the editing interface in response to a submission operation for the editing requirement; wherein the at least one candidate content is generated by the artificial intelligence based on the editing requirement; An updating module is configured to respond to a confirmation operation on a target content in the at least one candidate content and update the target object based on the target content.

[0007] In a third aspect, an embodiment of the present application further provides an electronic device comprising: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to execute the artificial intelligence-based interactive slide editing method described in any one of the first aspects.

[0008] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the interactive slide editing method based on artificial intelligence as described in any one of the first aspects is executed.

[0009] The embodiments of the present application have the following beneficial effects: The embodiment of the present application greatly reduces tedious manual operations and significantly improves editing efficiency through a simple and smooth interactive process, from selecting the target object to confirming the update; at the same time, after the user clearly inputs the requirements, artificial intelligence generates candidate content to ensure that the editing results are accurately in line with expectations, reducing the probability of repeated modifications; and provides at least one candidate content, bringing users diverse choices to meet personalized needs; in addition, the reasonable human-computer interaction interface design provides timely operation feedback, optimizes the user experience, and gives full play to the advantages of artificial intelligence in providing creative inspiration and expanding editing possibilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0011] Figure 1 10 is a flow chart of steps S101-S105 provided in an embodiment of the present application; Figure 2 It is a flowchart of steps S201-S202 provided in an embodiment of the present application; Figure 3 Schematic diagram of the process of steps S301-S302 provided in the embodiment of the present application; Figure 4 4 is a flow chart of steps S401-S402 provided in an embodiment of the present application; Figure 5 This is one of the interactive slide editing interfaces based on artificial intelligence provided by the embodiment of the present application; Figure 6 This is the second diagram of the interactive editing interface of a slideshow based on artificial intelligence provided by an embodiment of the present application; Figure 7 Schematic diagram of the structure of an artificial intelligence-based interactive slide editing device provided in an embodiment of the present application; Figure 8 It is a schematic diagram of the composition structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0012] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can be implemented out of sequence, and steps without logical context can be reversed or implemented simultaneously. In addition, those skilled in the art, under the guidance of the contents of this application, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.

[0013] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0014] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.

[0015] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0016] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.

[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0018] See also Figure 1 , Figure 1 This is a flow chart of steps S101-S105 of the interactive slide editing method based on artificial intelligence provided by the embodiment of the present application, which will be combined with Figure 1 Steps S101-S105 are shown for explanation.

[0019] In step S101, in response to a selection operation on a target object in the slide, the selected target object is determined as an object to be edited; wherein the target object includes a picture or a picture placeholder.

[0020] Here, the user selects a target object on a slide through a human-computer interface (e.g., a computer software interface or mobile app interface). The target object here is limited to an image or image placeholder. This step is the starting point for editing, identifying the element the user wants to modify as the object to be edited, providing an object foundation for subsequent editing operations.

[0021] In step S102, in response to a triggering operation on an editing component, an editing interface of the editing component is displayed in the human-computer interaction interface; wherein the editing interface is used to edit the object to be edited.

[0022] Here, after the user triggers the editing component, the human-computer interaction interface will display the editing interface of the editing component. The editing interface is the core area for users to interact with the editing function, specifically for performing various editing operations on the edited object, providing users with a centralized operation space.

[0023] In step S103, in response to an input operation on the editing interface, the editing requirement is displayed in the editing interface.

[0024] Here, the user enters the editing requirements for the object to be edited in the editing interface. The editing interface will display these input editing requirements in real time, allowing the user to clearly see their editing intentions and facilitate modification or confirmation.

[0025] In step S104, in response to a submission operation for the editing requirement, at least one candidate content is displayed in the editing interface; wherein the at least one candidate content is generated by the artificial intelligence based on the editing requirement.

[0026] After a user submits their editing request, the AI ​​system generates at least one candidate content based on that request and displays it in the editing interface. Leveraging AI's powerful computing and generation capabilities, it can quickly provide users with a variety of possible editing outcomes, broadening their editing ideas and choices.

[0027] In step S105 , in response to a confirmation operation on a target content in the at least one candidate content, the target object is updated based on the target content.

[0028] Finally, after the user selects the target content from the displayed candidate content and confirms it, the system will update the previously selected target object based on the target content. After the entire editing process is completed, the user's editing intention is implemented in the slideshow through appropriate content generated by artificial intelligence.

[0029] The aforementioned approach uses AI to automatically generate candidate content, eliminating the need for users to manually perform complex editing operations and saving considerable time and effort. This approach is particularly effective when users lack design experience or inspiration, as it quickly provides multiple options for selection. AI can generate high-quality editing content based on extensive data and algorithms, providing users with professional editing suggestions to ensure that images or image placeholders in slides better align with the overall style and requirements. The intuitive human-computer interaction interface and clear editing process make it easy for users to get started and quickly complete editing tasks. Furthermore, the presentation of multiple candidate content increases the fun and exploratory nature of editing.

[0030] In some embodiments, the editing requirement is a text requirement or a graphic requirement; the method further includes: When the editing requirement is the text requirement, determining a first target feature based on the text requirement, so that the artificial intelligence generates the at least one candidate content based on the first target feature; When the editing requirement is the graphic and text requirement, a second target feature is determined based on the graphic and text requirement, so that the artificial intelligence generates the at least one candidate content based on the second target feature; wherein the graphic and text requirement includes a reference image and / or a requirement description for a reference image, and the reference image is a picture in the slide and / or a specific uploaded picture.

[0031] Here, editing requirements are divided into text requirements and graphic requirements. Different types of requirements correspond to different information characteristics. Classification processing helps to improve the relevance and accuracy of generated content.

[0032] When a text request is generated, the system performs semantic analysis and keyword extraction to extract feature information that accurately reflects the user's editing intent. For example, if a user enters "adjust the image to a retro style," the first target features might include "retro style" and "image processing." Based on these features, AI generates candidate content that matches the retro style.

[0033] When the editing requirement is a graphic and text requirement, the graphic and text requirement includes a reference image and / or a description of the requirements for the reference image. The reference image can be an existing image in the slideshow or a specific image uploaded by the user. The system will perform image feature analysis on the reference image, such as color distribution, texture features, object recognition, etc., and combine it with the description of the requirements to comprehensively determine the second target feature. For example, if a user uploads a landscape reference image and describes "I hope the generated image has a similar color combination and seasonal atmosphere", the second target feature will include the color features, seasonal features, etc. of the reference image, thereby generating similar candidate content.

[0034] In some embodiments, see Figure 2 , Figure 2This is a flow chart of steps S201-S202 provided in an embodiment of the present application. When the editing requirement is the text requirement, the artificial intelligence generates the at least one candidate content through steps S201-S202, which will be explained in combination with each step.

[0035] In step S201, the text requirement is received, and semantic analysis is performed on the text requirement to obtain keyword information; wherein, the keyword information includes positive prompt words and negative prompt words, the positive prompt words are used to guide the generation direction, and the negative prompt words are used to exclude negative features.

[0036] In step S202, the keyword information is used as input information, and the at least one candidate content is generated through a text graph model.

[0037] Here, the textual requirements are semantically analyzed to extract positive and negative prompt words. Positive prompt words clearly define the characteristics that users expect in generated content, such as "retro," "dreamlike," and "technical," providing the AI ​​with a general direction for generation. Negative prompt words are used to eliminate undesirable characteristics and help filter out candidate content that does not meet the requirements. This classification allows the AI ​​to more accurately understand user needs and improve the quality and relevance of generated content.

[0038] The semantic parsing process includes lexical analysis, syntactic analysis, semantic role labeling, etc. Through in-depth analysis of text requirements, complex natural language is converted into keyword information that can be understood and processed by computers, providing a structured data foundation for subsequent text-based graph model input.

[0039] The text-generated graph model is a generative model based on deep learning that generates images based on input text descriptions. Trained with extensive data, the model learns the complex mapping relationship between text and images, and can generate images with corresponding features based on keyword information. In this embodiment of the present application, the text-generated graph model is a Flux model.

[0040] In some embodiments, see Figure 3 , Figure 3 This is a flow chart of steps S301-S302 provided in an embodiment of the present application. When the editing requirement is the graphic and text requirement, the artificial intelligence generates the at least one candidate content through S301-S302, which will be explained in combination with each step.

[0041] In step S301, the image and text requirements are received, visual features are extracted from the reference image and text features are extracted from the requirement description, and generation conditions are determined based on the visual features and the text features.

[0042] In step S302, the generation condition is used as input information to generate the at least one candidate content through a graph-to-graph model.

[0043] First, visual features are extracted from the reference image, such as color distribution (including dominant tones and color contrast), texture features (such as smooth, rough, and regular textures), object shape (identifying specific objects and their outlines in the image), and spatial layout (the positional relationships of objects in the image). Multi-level, multi-dimensional feature analysis of the reference image is then performed to comprehensively capture the visual style and content of the reference image. Next, text features are extracted from the requirement description, and operations such as word segmentation, part-of-speech tagging, and named entity recognition are performed on the requirement description to extract key information, such as the user's requirements for image style (such as "retro," "modern," and "cartoon"), themes (such as "natural scenery," "urban architecture," and "portraits"), and specific elements (such as "adding flower elements" and "removing clutter from the background"). This helps the AI ​​understand the user's editing intent expressed through text.

[0044] The system matches and correlates visual and textual features. For example, if the reference image is a retro-style landscape and the requirement mentions "adding some warm tones," the generation criteria will include elements such as the retro style, the landscape theme, and the warm tones adjustment. This generation criteria ensures that the candidate content generated subsequently matches the visual style of the reference image and meets the specific requirements specified by the user through the text description.

[0045] The graph-based graph model generates candidate content using a determined generation condition as input information. The graph-based graph model is a deep learning model that can transform images and generate new images based on input conditions. It is trained based on a large amount of image data to learn the mapping relationship and generation rules between images. During the generation process, the model will modify, optimize, or recreate the reference image based on the generation conditions to generate multiple candidate contents with different details and styles. In the embodiment of the present application, the graph-based graph model is the OminiGen model.

[0046] In some embodiments, the human-computer interaction interface includes a first display area (eg Figure 5 or Figure 6 "PPT editor area") and the second display area (e.g. Figure 6 The slideshow is displayed in the first display area, and the editing interface is displayed in the second display area. The second display area is hidden when the editing function is not enabled; The text requirements are input in the following ways: In response to a triggering operation on the editing component, displaying the editing interface in the second display area; wherein the editing interface includes an input box component; In response to the input operation on the input box component, the input text requirement is displayed in the input box component.

[0047] Here, the human-computer interaction interface is divided into a first display area and a second display area, which helps users quickly distinguish between the slide display area and the editing operation area. The slides are displayed in the first display area, making it easier for users to focus on viewing and browsing the slide content; the editing interface is displayed in the second display area, and when editing operations are required, users can focus their attention on this area. The second display area is hidden when the editing function is not enabled. This design effectively reduces visual interference on the interface, making the interface more concise and allowing users to focus more on the slides themselves. The editing interface is only displayed through corresponding operations when the user has editing needs, which improves the usability of the interface and user experience.

[0048] The user triggers the editing component (such as clicking the edit button) to display the editing interface in the second display area. The editing interface includes an input box component. The user can display the entered text requirements in the input box component through input operations on the input box component (such as keyboard input, voice input, etc.).

[0049] In some embodiments, the editing interface further includes a reference image upload component, and the image and text requirements are input in the following manner: In response to a trigger operation on the reference image component, a reference image upload interface is displayed, the image uploaded through the reference image upload interface is used as the reference image, and a preview of the uploaded reference image is displayed in the input box component; and in response to an input operation on the input box component, the input requirement description is displayed in the input box component; Alternatively, it is determined based on the requirement description whether to use the selected picture in the target object as the reference picture.

[0050] Here, when the user triggers the reference image upload component (for example, by clicking the "Reference Image" button), the system responds by displaying the reference image upload interface. The user selects and uploads an image through the reference image upload interface. After the upload is successful, the system displays a preview of the uploaded reference image in the input box component.

[0051] Users can enter their requirements in the input box component while or after uploading a reference image. The input box component provides a centralized place for users to express their editing intentions, allowing them to describe their requirements for the image's style, elements, adjustment direction, and other aspects in detail.

[0052] The system also has the ability to determine whether to use an image in a selected target object as a reference image based on the requirement description. For example, in a slideshow editing scenario, if the user has already selected an image in the slideshow as the target object, the system can analyze keywords in the requirement description (such as "adjust the color tone based on this image") and automatically determine whether to use the selected image as a reference image, further simplifying the user's operation process.

[0053] In some embodiments, see Figure 4 , Figure 4 This is a flow chart of steps S401-S402 provided in an embodiment of the present application, which will be explained in conjunction with each step.

[0054] In step S401, a generation animation of the at least one candidate content being generated is displayed in the editing interface, and a polling request is periodically sent to a backend via a timer.

[0055] In step S402, when the backend generates any candidate content among the at least one candidate content, the corresponding generated animation is replaced with the generated candidate content based on the polling result; wherein the polling result is the image connection and size information returned by the backend.

[0056] Here, a generation animation is displayed in the editing interface. This animation indicates to users that the system is processing their request, preventing them from becoming confused or anxious due to an unresponsive interface while waiting. Generation animations can take the form of loading progress bars, spinning icons, dynamic renderings, and more, enhancing the user experience through visual feedback.

[0057] The frontend uses a timer to periodically send polling requests to the backend to check the status of candidate content generation. Timed polling is a simple and reliable way to query backend status, suitable for scenarios that require real-time information on generation progress but don't require complex communication mechanisms (such as WebSocket). The timer interval can be adjusted based on actual needs to balance real-time performance and server load.

[0058] Based on the polling results (i.e., the image link and size information returned by the backend), the frontend replaces the corresponding generated animation with the generated candidate content. This dynamic replacement mechanism ensures that users can see the generated candidate content in a timely manner without having to manually refresh the page or perform other operations. The image link is used to load the generated image, and the size information ensures that the image is displayed at the appropriate size in the editing interface.

[0059] The following is a complete description of the embodiments of the present application. Figure 5 and Figure 6 , Figure 5 This is one of the interactive slide editing interfaces based on artificial intelligence provided by the embodiment of the present application. Figure 6This is the second diagram of the interactive editing interface of the slide based on artificial intelligence provided by the embodiment of the present application, such as Figure 5 As shown, there are three pictures in the PPT editor area, labeled as Picture 1, Picture 2, and Picture 3. Users can select existing pictures or picture placeholders on the slide by clicking the mouse to prepare for subsequent editing operations.

[0060] like Figure 6 As shown, after selecting an image or an image placeholder, the user can activate the AI ​​assistant function. In this dialog box, the user has multiple input methods: either enter text requirements in the dialog box to clearly state the specific requirements for image editing; or click the reference image button to upload an image, and the source of the reference image is relatively flexible. It can be an externally uploaded image or a specified image in the PPT editor area, for example Figure 6 Image 1 in the image. After the user completes the input, the AI ​​Assistant dialog box displays an animation that says "Image Generating..." to inform the user that the system is processing. Eventually, the system generates a new image. The user clicks the generated image to replace the previously selected image (e.g., Image 1), completing the image editing operation.

[0061] In summary, the embodiments of the present application have the following beneficial effects: (1) The embodiment of the present application directly completes the full-link operation of "object selection → requirement input → AI generation → content replacement" through an interactive interface, solving the fragmentation problem of traditional PPT editing requiring cross-platform image search, improving image editing efficiency, and supporting intelligent rewriting of existing image placeholders, avoiding the tedious operation of users manually adjusting the layout.

[0062] (2) The embodiment of the present application automatically extracts positive / negative keywords through semantic analysis, thereby improving the accuracy of AI generation, reducing the generation of invalid content, and supporting the fusion analysis of reference graph feature extraction and demand description, thereby improving the consistency of generated content with the document theme.

[0063] (3) Text requirements are precisely controlled through the text-based image model. The negative prompt word mechanism can filter out content that does not match the scenario and supports the joint modeling of the visual features (color / texture / composition) of the reference image and the text description, thereby improving the consistency coefficient between the generated image and the overall visual style of the PPT.

[0064] (4) The separation design of the first display area (canvas area) and the second display area (editing area) prevents content preview and parameter adjustment from interfering with each other, improving editing efficiency. The hidden editing panel automatically folds when inactive, keeping the PPT editing interface tidy and reducing visual interference.

[0065] (5) The embodiment of the present application supports three reference image acquisition methods: automatically capturing the currently selected image, uploading a local image, and calling from the slide library, which can cover a wide range of usage scenarios. The intelligent parsing function of the requirement description can automatically associate the selected image, reducing the user's repeated operations.

[0066] (6) The embodiment of the present application greatly reduces the average perceived waiting time by generating animation and polling mechanism, and the incremental content loading strategy allows users to preview part of the generated results in advance, thereby improving the user experience.

[0067] (7) The embodiment of the present application adopts a short-cycle polling strategy, which triggers a local update immediately when the AI ​​generates the first candidate content, avoiding frequent requests from putting pressure on the server. The image size information carried in the polling response is automatically matched with the slide layout parameters to ensure that the aspect ratio and resolution of the inserted image match the placeholder frame, reducing the workload of typesetting adjustments.

[0068] Based on the same inventive concept, the embodiments of the present application also provide an artificial intelligence-based interactive slide editing device corresponding to the artificial intelligence-based interactive slide editing method in the first embodiment. Since the principle of solving the problem by the device in the embodiments of the present application is similar to the above-mentioned artificial intelligence-based interactive slide editing method, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0069] like Figure 7 As shown, Figure 7 : is a schematic diagram of the structure of an artificial intelligence-based interactive slide editing device 700 provided in an embodiment of the present application, wherein the slides are displayed through a human-computer interaction interface; the artificial intelligence-based interactive slide editing device 700 includes: A selection module 701 is configured to respond to a selection operation on a target object in the slide and determine the selected target object as an object to be edited; wherein the target object includes a picture or a picture placeholder; A trigger module 702 is configured to respond to a trigger operation on an editing component and display an editing interface of the editing component in the human-computer interaction interface; wherein the editing interface is used to edit the object to be edited; An input module 703 is configured to respond to input operations on the editing interface and display editing requirements in the editing interface; Display module 704, configured to display at least one candidate content in the editing interface in response to a submission operation for the editing requirement; wherein the at least one candidate content is generated by the artificial intelligence based on the editing requirement; The updating module 705 is configured to respond to a confirmation operation on a target content in the at least one candidate content and update the target object based on the target content.

[0070] It should be understood by those skilled in the art that Figure 7 The functions implemented by each unit in the illustrated artificial intelligence-based interactive slideshow editing device 700 can be understood by referring to the related description of the aforementioned artificial intelligence-based interactive slideshow editing method. Figure 7 The functions of the various units in the illustrated artificial intelligence-based interactive slideshow editing device 700 may be implemented by a program running on a processor, or by a specific logic circuit.

[0071] In a possible implementation, the editing requirement is a text requirement or a graphic requirement; the display module 704 further includes: When the editing requirement is the text requirement, determining a first target feature based on the text requirement, so that the artificial intelligence generates the at least one candidate content based on the first target feature; When the editing requirement is the graphic and text requirement, a second target feature is determined based on the graphic and text requirement, so that the artificial intelligence generates the at least one candidate content based on the second target feature; wherein the graphic and text requirement includes a reference image and / or a requirement description for a reference image, and the reference image is a picture in the slide and / or a specific uploaded picture.

[0072] In a possible implementation, when the editing requirement is a text requirement, the artificial intelligence generates the at least one candidate content in the following manner: Receive the text requirement, perform semantic analysis on the text requirement, and obtain keyword information; wherein the keyword information includes positive prompt words and negative prompt words, the positive prompt words are used to guide the generation direction, and the negative prompt words are used to exclude negative features; The keyword information is used as input information, and the at least one candidate content is generated through a text graph model.

[0073] In a possible implementation, when the editing requirement is a graphic and text requirement, the artificial intelligence generates the at least one candidate content in the following manner: receiving the image-text requirement, extracting visual features from the reference image and text features from the requirement description, and determining generation conditions based on the visual features and the text features; The generation condition is used as input information to generate the at least one candidate content through a graph-to-graph model.

[0074] In a possible implementation, the human-computer interaction interface includes a first display area and a second display area, the slide is displayed in the first display area, and the editing interface is displayed in the second display area, and the second display area is hidden when the editing function is not enabled; The text requirements are input in the following ways: In response to a triggering operation on the editing component, displaying the editing interface in the second display area; wherein the editing interface includes an input box component; In response to the input operation on the input box component, the input text requirement is displayed in the input box component.

[0075] In a possible implementation, the editing interface further includes a reference image upload component, and the image and text requirements are input in the following manner: In response to a trigger operation on the reference image component, a reference image upload interface is displayed, the image uploaded through the reference image upload interface is used as the reference image, and a preview of the uploaded reference image is displayed in the input box component; and in response to an input operation on the input box component, the input requirement description is displayed in the input box component; Alternatively, it is determined based on the requirement description whether to use the selected picture in the target object as the reference picture.

[0076] In a possible implementation, the display module 704 further includes: Displaying a generation animation of the at least one candidate content being generated in the editing interface, and periodically sending a polling request to a backend through a timer; When the backend generates any candidate content among the at least one candidate content, the corresponding generation animation is replaced with the generated candidate content based on the polling result; wherein the polling result is the image connection and size information returned by the backend.

[0077] The above-mentioned interactive slide editing device based on artificial intelligence has the following beneficial effects: (1) The embodiments of this application are based on artificial intelligence technology. Based on the text or image requirements input by the user, they accurately analyze and generate candidate content that meets the user's expectations. Whether it is a simple text description or a complex combination of text and image requirements, AI can deeply understand and output personalized editing results, providing users with a highly intelligent and customized slide editing experience, meeting the editing needs of different users in diverse scenarios.

[0078] (2) The embodiment of the present application is simple to operate, from selecting the object to be edited to inputting the editing requirements, and then generating and selecting candidate content to complete the update. This greatly shortens the editing time of the slides, significantly improves the production efficiency of PPT, and enables users to complete slide creation more quickly.

[0079] (3) For text requirements, AI extracts precise keywords through semantic analysis to ensure that the generated content meets the user's intention. For graphic and text requirements, it integrates the visual features of the reference image and the text features of the requirement description, and uses professional models to generate high-quality candidate content, effectively improving the quality of the generated content and ensuring the presentation effect of the final slide.

[0080] (4) The embodiments of this application achieve seamless integration of the editor and AI functions. Users can complete the entire process from requirement input to content generation in the editing interface without switching between different software or platforms. This integration method reduces the user's learning cost and provides a smoother and more natural interactive experience.

[0081] (5) The embodiments of the present application can be widely used in various fields such as office, education, and design. In office scenarios, it can help employees quickly create high-quality presentations; in the field of education, it can facilitate teachers to create vivid and interesting teaching materials; in the design industry, it can provide designers with creative inspiration and efficient editing tools. The embodiments of the present application have broad application prospects.

[0082] like Figure 8 As shown, Figure 8 This is a schematic diagram of the structure of an electronic device 800 provided in an embodiment of the present application. The electronic device 800 includes: A processor 801, a storage medium 802 and a bus 803, wherein the storage medium 802 stores machine-readable instructions executable by the processor 801. When the electronic device 800 is running, the processor 801 communicates with the storage medium 802 via the bus 803, and the processor 801 executes the machine-readable instructions to perform the steps of the artificial intelligence-based interactive slide editing method described in the embodiment of the present application.

[0083] In actual application, the various components in the electronic device 800 are coupled together via a bus 803. It is understood that the bus 803 is used to achieve connection and communication between these components. In addition to the data bus, the bus 803 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 8 Various buses are labeled as bus 803.

[0084] The electronic device has the following beneficial effects: (1) The embodiments of this application are based on artificial intelligence technology. Based on the text or image requirements input by the user, they accurately analyze and generate candidate content that meets the user's expectations. Whether it is a simple text description or a complex combination of text and image requirements, AI can deeply understand and output personalized editing results, providing users with a highly intelligent and customized slide editing experience, meeting the editing needs of different users in diverse scenarios.

[0085] (2) The embodiment of the present application is simple to operate, from selecting the object to be edited to inputting the editing requirements, and then generating and selecting candidate content to complete the update. This greatly shortens the editing time of the slides, significantly improves the production efficiency of PPT, and enables users to complete slide creation more quickly.

[0086] (3) For text requirements, AI extracts precise keywords through semantic analysis to ensure that the generated content meets the user's intention. For graphic and text requirements, it integrates the visual features of the reference image and the text features of the requirement description, and uses professional models to generate high-quality candidate content, effectively improving the quality of the generated content and ensuring the presentation effect of the final slide.

[0087] (4) The embodiments of this application achieve seamless integration of the editor and AI functions. Users can complete the entire process from requirement input to content generation in the editing interface without switching between different software or platforms. This integration method reduces the user's learning cost and provides a smoother and more natural interactive experience.

[0088] (5) The embodiments of the present application can be widely used in various fields such as office, education, and design. In office scenarios, it can help employees quickly create high-quality presentations; in the field of education, it can facilitate teachers to create vivid and interesting teaching materials; in the design industry, it can provide designers with creative inspiration and efficient editing tools. The embodiments of the present application have broad application prospects.

[0089] The embodiment of the present application also provides a computer-readable storage medium, which stores executable instructions. When the executable instructions are executed by at least one processor 801, the artificial intelligence-based interactive slide editing method described in the embodiment of the present application is implemented.

[0090] In some embodiments, the storage medium can be a magnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface storage, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various devices including one or any combination of the above memories.

[0091] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0092] As an example, executable instructions may, but do not necessarily, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (for example, files storing one or more modules, subroutines, or code portions).

[0093] By way of example, executable instructions may be deployed to be executed on one computing device, or on multiple computing devices at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.

[0094] The computer-readable storage medium has the following advantages: (1) The embodiments of this application are based on artificial intelligence technology. Based on the text or image requirements input by the user, they accurately analyze and generate candidate content that meets the user's expectations. Whether it is a simple text description or a complex combination of text and image requirements, AI can deeply understand and output personalized editing results, providing users with a highly intelligent and customized slide editing experience, meeting the editing needs of different users in diverse scenarios.

[0095] (2) The embodiment of the present application is simple to operate, from selecting the object to be edited to inputting the editing requirements, and then generating and selecting candidate content to complete the update. This greatly shortens the editing time of the slides, significantly improves the production efficiency of PPT, and enables users to complete slide creation more quickly.

[0096] (3) For text requirements, AI extracts precise keywords through semantic analysis to ensure that the generated content meets the user's intention. For graphic and text requirements, it integrates the visual features of the reference image and the text features of the requirement description, and uses professional models to generate high-quality candidate content, effectively improving the quality of the generated content and ensuring the presentation effect of the final slide.

[0097] (4) The embodiments of this application achieve seamless integration of the editor and AI functions. Users can complete the entire process from requirement input to content generation in the editing interface without switching between different software or platforms. This integration method reduces the user's learning cost and provides a smoother and more natural interactive experience.

[0098] (5) The embodiments of the present application can be widely used in various fields such as office, education, and design. In office scenarios, it can help employees quickly create high-quality presentations; in the field of education, it can facilitate teachers to create vivid and interesting teaching materials; in the design industry, it can provide designers with creative inspiration and efficient editing tools. The embodiments of the present application have broad application prospects.

[0099] In the several embodiments provided in this application, it should be understood that the disclosed methods and electronic devices can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0100] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0101] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0102] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, platform server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage media include various media that can store program code, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks, or optical disks.

[0103] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. An artificial intelligence-based interactive slide editing method, characterized in that: The slideshow is displayed via a human-computer interaction interface, and the method includes: In response to a selection operation on a target object in the slide, determining the selected target object as an object to be edited; wherein the target object includes a picture or a picture placeholder; In response to a triggering operation on an editing component, an editing interface of the editing component is displayed in the human-computer interaction interface; wherein the editing interface is used to edit the object to be edited; In response to an input operation on the editing interface, displaying an editing requirement in the editing interface; In response to a submission operation for the editing requirement, displaying at least one candidate content in the editing interface; wherein the at least one candidate content is generated by the artificial intelligence based on the editing requirement; In response to a confirmation operation on target content in the at least one candidate content, the target object is updated based on the target content.

2. The method according to claim 1, characterized in that The editing requirement is a text requirement or a graphic requirement; the method further includes: When the editing requirement is the text requirement, determining a first target feature based on the text requirement, so that the artificial intelligence generates the at least one candidate content based on the first target feature; When the editing requirement is the graphic and text requirement, a second target feature is determined based on the graphic and text requirement, so that the artificial intelligence generates the at least one candidate content based on the second target feature; wherein the graphic and text requirement includes a reference image and / or a requirement description for a reference image, and the reference image is a picture in the slide and / or a specific uploaded picture.

3. The method according to claim 2, characterized in that When the editing requirement is a text requirement, the artificial intelligence generates the at least one candidate content in the following manner: Receive the text requirement, perform semantic analysis on the text requirement, and obtain keyword information; wherein the keyword information includes positive prompt words and negative prompt words, the positive prompt words are used to guide the generation direction, and the negative prompt words are used to exclude negative features; The keyword information is used as input information, and the at least one candidate content is generated through a text graph model.

4. The method according to claim 2, characterized in that When the editing requirement is a graphic and text requirement, the artificial intelligence generates the at least one candidate content in the following manner: receiving the image-text requirement, extracting visual features from the reference image and text features from the requirement description, and determining generation conditions based on the visual features and the text features; The generation condition is used as input information to generate the at least one candidate content through a graph-to-graph model.

5. The method according to claim 2, characterized in that The human-computer interaction interface includes a first display area and a second display area, the slide is displayed in the first display area, the editing interface is displayed in the second display area, and the second display area is hidden when the editing function is not enabled; The text requirements are input in the following ways: In response to a triggering operation on the editing component, displaying the editing interface in the second display area; wherein the editing interface includes an input box component; In response to the input operation on the input box component, the input text requirement is displayed in the input box component.

6. The method according to claim 5, characterized in that The editing interface also includes a reference image upload component, and the image and text requirements are input in the following way: In response to a trigger operation on the reference image component, a reference image upload interface is displayed, the image uploaded through the reference image upload interface is used as the reference image, and a preview of the uploaded reference image is displayed in the input box component; and in response to an input operation on the input box component, the input requirement description is displayed in the input box component; Alternatively, it is determined based on the requirement description whether to use the selected picture in the target object as the reference picture.

7. The method according to claim 1, characterized in that The method further comprises: Displaying a generation animation of the at least one candidate content being generated in the editing interface, and periodically sending a polling request to a backend through a timer; When the backend generates any candidate content among the at least one candidate content, the corresponding generation animation is replaced with the generated candidate content based on the polling result; wherein the polling result is the image connection and size information returned by the backend.

8. An interactive slideshow editing device based on artificial intelligence, characterized in that: The slide is displayed via a human-computer interaction interface, and the device comprises: A selection module, configured to respond to a selection operation on a target object in the slide and determine the selected target object as an object to be edited; wherein the target object includes a picture or a picture placeholder; A trigger module, configured to respond to a trigger operation on an editing component and display an editing interface of the editing component in the human-computer interaction interface; wherein the editing interface is used to edit the object to be edited; An input module, configured to respond to input operations on the editing interface and display editing requirements in the editing interface; a display module configured to display at least one candidate content in the editing interface in response to a submission operation for the editing requirement; wherein the at least one candidate content is generated by the artificial intelligence based on the editing requirement; An updating module is configured to respond to a confirmation operation on a target content in the at least one candidate content and update the target object based on the target content.

9. An electronic device, characterized in that: include: A processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the storage medium communicate via the bus, and the processor executes the machine-readable instructions to perform the artificial intelligence-based interactive slide editing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for interactively editing slides based on artificial intelligence as claimed in any one of claims 1 to 7 is executed.

Citation Information

Cited By

  • Media content processing method and device, equipment, storage medium and program product

    CN121547648A