System and method for generating new content for presentation prepared in presentation application

By using server-side generative AI tools and prompt word generation engines, the challenges users face when creating demo content are solved, enabling the rapid generation of customized content, simplifying the process and improving efficiency.

CN120936997APending Publication Date: 2025-11-11MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480021128.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-11
Filing Date
2024-04-28
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

When creating demo content, users often find it difficult to quickly find or create content that meets their needs, resulting in a time-consuming process or exceeding their capabilities.

Method used

By using server-side generative AI tools and prompt word generation engines, the content of multimedia files is reconstructed to generate new content, which is then transmitted to the client-side demo application to enhance the demo effect.

Benefits of technology

It simplifies the process for users to create content that meets their needs in presentations, improving efficiency and flexibility, and allowing users to quickly generate customized content as needed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120936997A_ABST
    Figure CN120936997A_ABST
Patent Text Reader

Abstract

A data processing system includes: a server having a processor and a network interface; and a memory including programming instructions, the programming instructions including a cue generation engine. When executed by the processor, the instructions cause the server to implement a service to: receive a plurality of media files from a presentation application on a client device; reconstructing the content of the media file into a form compatible with a generative artificial intelligence (AI) tool; constructing, using the cue generation engine, a cue for the generative AI tool using content of the media file in a form compatible with the generative AI tool, the cue including instructions for generating new content by fusing content from the plurality of media files; receiving the new content from the generative AI tool; and transmitting the new content to the presentation application on the client device to enhance a presentation generated with the presentation application.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Presentation software applications are powerful tools for creating visual items to support presentations or convey information. Typically, presentation applications allow users to prepare a series of slides to be displayed or presented during a presentation. Each slide can include various media or content elements, such as text, images, audio files, and video files. Presentation applications allow users to add and arrange these elements to configure each slide as needed.

[0002] However, sometimes users may not have the need or ability to find exactly what they want to include in a presentation. This raises a technical problem: users then have to find or create content that better meets the presentation's needs. Determining what content is needed and what users must process can be frustratingly time-consuming or simply beyond their capabilities. Summary of the Invention

[0003] In one general aspect, this disclosure explains a data processing system comprising: a server having a processor, a network interface, and a memory, the memory including programming instructions including a prompt word generation engine. When executed by the processor alone or in combination with other processors, the instructions cause the server to perform the following services: receiving multiple media files from a demo application on a client device or from the web; reconstructing the content of the media files into a form compatible with a generative artificial intelligence (AI) tool; using the prompt word generation engine, constructing prompt words for the generative AI tool using the content of the media files in a form compatible with the generative AI tool, the prompt words including instructions for generating new content by fusing content from the multiple media files; receiving the new content from the generative AI tool; and transmitting the new content to the demo application on the client device to enhance a demo generated using the demo application. The multiple media files being fused may be of different types.

[0004] In another general aspect, this disclosure explains a demo application executed by a client device including a processor and memory, the demo application being stored in non-transitory memory of the client device and including executable instructions that, when executed by the processor, cause the processor, alone or in combination with other processors, to: present a user interface for generating a demo, the user interface being configured to receive a plurality of media files, each media file including content to be included in the demo; and, in response to a user instruction, invoke a service to reconstruct the content of the media files in a form compatible with a generative artificial intelligence (AI) tool; construct a prompt for the generative AI tool using the content of the media files in the form compatible with the generative AI tool, the prompt including instructions for generating new content by fusing content from the plurality of media files; and transmit the new content to the demo application on the client device. The application is configured to display the new content in the user interface, the user interface having controls for using the new content to enhance the demo generated using the demo application.

[0005] In another general aspect, this disclosure describes a method for providing a service to generate new content by combining two sets of input content, the service being supported by a server having a processor, a network interface, and memory, the memory including programming instructions including a prompt word generation engine, the programming instructions causing the server to implement the method when executed by the processor alone or in combination with other processors. The method includes: receiving a plurality of media files from a demo application on a client device; reconstructing the content of the media files in a form compatible with a generative artificial intelligence (AI) tool; using the prompt word generation engine, constructing prompt words for the generative AI tool using the content of the media files in the form compatible with the generative AI tool, the prompt words including instructions for generating new content by fusing content from the plurality of media files; receiving the new content from the generative AI tool; and transmitting the new content to the demo application on the client device to enhance the demo generated using the demo application.

[0006] The choice of concepts to introduce a simplified form is provided in this summary, which will be further described in the detailed description below. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to the implementation of any or all of the shortcomings mentioned in any part of this disclosure. Attached Figure Description

[0007] The accompanying drawings are for illustrative purposes only and not as limitations, depicting one or more implementations according to this teaching. In the drawings, similar reference numerals refer to the same or similar elements. Furthermore, it should be understood that the drawings are not necessarily drawn to scale.

[0008] Figure 1 An example system is described on which various aspects of this disclosure can be implemented to generate new image content.

[0009] Figure 2 Another example system is described on which aspects of this disclosure can be implemented to generate various types of new content.

[0010] Figure 3 Another example system is described on which various aspects of this disclosure can be implemented to leverage a voice assistant to generate a variety of new content.

[0011] Figure 4A Describing for, for example Figure 1-3 Examples of user interfaces for those systems.

[0012] Figure 4B Depicting the subsequent stages Figure 4A An example of a user interface.

[0013] Figure 5 Describing for, for example Figure 1-3 Another example of the user interface of those systems.

[0014] Figure 6 A flowchart depicts a method for generating new content in accordance with this disclosure.

[0015] Figure 7 This is a block diagram illustrating an example software architecture, whose various parts can be used in conjunction with the various hardware architectures described in this document.

[0016] Figure 8 This is a block diagram illustrating the components of an example machine configured to read instructions from a machine-readable medium and execute any of the features described herein. Detailed Implementation

[0017] As mentioned above, presentation software applications allow users to prepare a series of slides for use during a presentation, where each slide can include various media or content elements. However, sometimes users may not have the exact content they want to include in their presentation. This raises a technical problem: users must then find or create content that better meets the presentation's needs. Deciding what content is needed and what content users must process can be frustrating, time-consuming, or simply beyond their capabilities.

[0018] In order to provide a technical solution to this technical problem, Figure 1 An example system 100 is described, on which various aspects of this disclosure can be implemented. (As in...) Figure 1 As shown, demo application 106 is running on computer 102. Computer 102 can be any computerized device, including but not limited to: desktop computer, laptop computer, server, tablet computer, or even smartphone. In some examples, demo application 106 can be a dedicated application, or it can be a browser communicating with a server providing the demo application as a service. An example of a demo application is... of As described above, the presentation application allows users to prepare a series of slides to be displayed or presented during a presentation. Each slide can include various media or content elements, such as text, images, audio files, and video files. Presentation application 106 allows users to add and arrange these elements to configure each slide as needed.

[0019] Figure 1 This illustration depicts a specific example of an image a user of a demo application 106 wants to use in a demo that does not yet exist or that the user does not possess. The user has two different images, and corresponding image files 108 and 110 have been loaded into the demo application 106. However, in this scenario, each of the two different images includes elements that the user would prefer to combine in a single image. For example, the first image file 108 provides an image of a person, and the second image file 110 provides an image of a flower. The user actually wants a single image of a person holding a flower. Specifically, the user might want a single image of a person similar to the person in the first image file 108 holding a flower of the type found in the second image file 110.

[0020] To obtain the desired content and resolve the technical issue, demo application 106 is able to access the service. Figure 1 In the example, the service is supported by server 116, which communicates with computer 102 via network 101 (e.g., the Internet). Figure 1 As shown, the demo application 106 sends a request, including image files, to the server 116. This request will include a prompt entered by the user: using two image files 108 and 110 as input, generate a new image of a person holding flowers.

[0021] To generate the image desired by the user, server 116 can utilize image generation artificial intelligence (AI) tools. In recent years, AI has made significant progress in the field of image generation. Specifically, text-to-image refers to the process of generating images from text descriptions using machine learning algorithms and deep neural networks. The goal is to train AI models to interpret natural language descriptions and create images that accurately represent those descriptions. Text-to-image technology is based on a class of neural networks using generative adversarial networks (GANs), which include two main components: a generator and a discriminator. The generator creates new images based on given text input, while the discriminator evaluates the realism of the generated images and provides feedback to the generator. This process is repeated iteratively until the generator produces a realistic image that matches the text description. Current examples of image generation tools include OpenAI's DALL-E and Stability AI's Stable Diffusion.

[0022] Server 116 includes a prompt generation engine 118. This prompt generation engine 118 is software configured to construct input instructions or prompts for image generation AI tool 114. The current image generation AI tool 114 is trained to operate on text and therefore does not readily accept image data as input. The current image generation AI tool 114 also operates more efficiently with prompts in English. Therefore, the prompt generation engine 118 can have several tasks. If the instruction from the user of demo application 106 is a human language other than English, the prompt generation engine 118 can invoke language translation service 105 via network 101 to translate the user's instruction into English. The prompt generation engine 118 also converts image files 108 and 110 into a more compatible form for input into image generation AI tool 114.

[0023] Therefore, the prompt word generation engine 118 can send a request, including an image file, to the computer vision AI tool 112 via network 101. For example, in Figure 1As shown, this could be a Hypertext Transfer Protocol (HTTP) request. Computer vision is a field of AI that focuses on enabling machines to interpret and understand visual data from the world around us, including images and videos. Computer vision AI tool 112 converts the image data from image files 108 and 110 into a more textual format, such as JSON (JavaScript Object Notation), that can be ingested by image generation AI tool 114. JSON is a lightweight data-interchange format commonly used to transfer data between servers and web applications. JSON is not an image format, but rather a way of representing structured data (such as images) using text or numerical values. Therefore, the application programming interface (API) of computer vision AI tool 112 can return the results of image analysis in JSON format. For example, these APIs can analyze images and return JSON responses that include information about objects and features detected in the image, such as the location of faces, the presence of text, or the identification of specific objects or landmarks. Examples of such APIs include Microsoft's Computer Vision API and Google's Cloud Vision API. The prompt generation engine 118 receives the JSON data for image files 108 and 110 via HTTP responses.

[0024] Then, the prompt generation engine 118 constructs prompts for the image generation AI tool 114. The prompts or requests constructed by the prompt generation engine 118 include JSON data or other data in a format compatible with another AI tool, as well as text descriptions or instructions provided by the user of application 106 on how to blend or combine the images. In this example, the instructions are for an image of a person holding flowers. Therefore, all the data from the prompts in engine 118 is in a form easily ingested by the image generation AI tool 114.

[0025] There are many data formats that current AI tools or models can discover and readily utilize. These include, but are not limited to, text, comma-separated values ​​(CSV), JavaScript Object Notation (JSON), Joint Group of Image Experts (JPG), and Portable Web Graphics (PNG). In addition to natural language text, these formats allow graphics and photographs to be understandable to current AI models. As used herein and in the appended claims, the terms “format compatible with AI tools” or “AI tool-compatible format” and similar terms will refer to file or data formats that are easily ingested by AI tools. What one or more formats a particular AI model will readily ingest will vary depending on the configuration of that particular AI model. However, generally, formats used by applications from multiple developers (such as those examples listed above, e.g., CSV, JSON, text or rich text, JPG, PNG, etc.) will be compatible with different AI models.

[0026] As in Figure 1 As shown, the image generation AI tool 114 returns one or more images generated using data from two input image files 108 and 110 and a user description. The AI ​​tool can generate multiple alternative images that a user can choose from based on the same input. Because the input to the AI ​​tool 114 includes data from two input image files 108 and 110, the resulting images are more likely to represent people and flowers that are visually similar to those in the input images. This solves the technical problem of the AI ​​tool 114 randomly generating the appearance of people, flowers, or other requested image elements and allowing the user to customize some aspects of the resulting output image based on the input.

[0027] The service on server 116 will then return the merged or alternative image to the demo application 106, as shown in Figure 1 As shown in the image. The user can then choose from alternatives (if provided) and place the desired image into the presentation being prepared. Again, this solves the technical problem of a user wanting or needing a single image with specified content (where the user did not previously own that single image).

[0028] Figure 2 Another example system is described on which various aspects of this disclosure can be implemented to generate a variety of different types of new content. Typically, Figure 2 The example illustrations demonstrate that the concepts disclosed herein can be applied to a variety of different content types, not just combining the content of two images into a single fused image. The system 121 again includes a computer 102 with a processor and associated memory 103, which executes the demonstration application 106.

[0029] However, in this example, media or content files 120 and 122 can be of any type, including but not limited to: text, images, videos, audio, and code. Any combination of these content types can be merged or combined into a single file or object for inclusion in a presentation created using presentation application 106. For example, an image file combined with an audio file can result in the image of the image file being displayed as a still video object while the audio of the audio file is played. The resulting video object can then be incorporated into the presentation. Image and video files can be combined into a video object, where images are added to other images from a video file according to each user instruction. AI tools and other services can generate content based on different inputs as part of the merged content, such as those described herein, such as text to image (Dall-E / Stable Diffusion); text to video, text to audio, text to code (GitHub), audio to text (Audio LM), audio to audio (Audio LM), image to text (generative pre-trained transformer or GPT), and video to text.

[0030] Therefore, as in Figure 2 As shown, demonstration application 106 sends a request to service 126 on server 116 containing two or more content files 120 and 122. Service 126 includes a file type detection engine 128. Using metadata, file extensions, codecs, user input, or other means, file type detection engine 128 determines what type of content exists in content files 120 and 122, such as text, video, audio, images, etc. Using this determination, such as in... Figure 1 In this process, service 126 will invoke one or more format conversion tools 124 to render the contents of the file in a format compatible with the AI ​​tools, when needed. This may include invoking language translation service 105 as described above.

[0031] When the content data is in a format compatible with the AI ​​tool, as described above, the prompt generation engine 130 will construct prompts or requests for the generative AI tool 115. This generative AI tool may be, for example, Dall-E, StableDiffusion, or other image or video generation AI tools, or may include AI text generation tools such as some versions of GPT. The generative AI tool will also receive instructions from the user in the prompts from the prompt generation engine 130 and will merge or combine the content of media files 120 and 122 as instructed. The response with the merged content is returned to service 126, which will provide the merged content, including any alternatives, to the demo application 106 for the user to include in the presentation being prepared.

[0032] Figure 3 Another example system 123 is described, on which various aspects of this disclosure can be implemented to leverage voice assistants to generate various types of new content. Figure 3 The example in is similar to Figure 2 Examples, except in Figure 3 In the computer 102, there is an AI voice assistant.

[0033] AI voice assistants are computer programs that use artificial intelligence to process natural language commands and provide helpful responses to users. These assistants are designed to be hands-free and accessible, using speech recognition technology to understand and interpret verbal requests. It is by This is one example of an AI voice assistant developed for use on its Windows operating system. Other examples include Apple's Siri, Amazon's Alexa, and Google Assistant. These assistants are typically integrated into computers, mobile devices, smart speakers, and other smart home devices, allowing users to control their technology and access information solely using their voice. AI voice assistants use a variety of technologies to understand and respond to user requests, including natural language processing, machine learning, and speech recognition. They are capable of performing a wide range of tasks.

[0034] exist Figure 3 In the example, computer 102 includes a microphone 107 for translating commands spoken by the user and a speaker 109 for outputting audio (e.g., verbal) information to the user. These devices enable AI voice assistant 104 to receive commands or issue audio prompts to the user. For example, a user can use AI voice assistant 104 to input commands to merge media files, such as “combine the two images on slide 8 into an image of a person holding flowers.” AI voice assistant 104 then interacts with presentation application 106 to initiate actions as described above. Figure 2 The example shows the merging of the contents of two media files, 120 and 122.

[0035] The AI ​​voice assistant 104 can also make suggestions such as, "Would you like Cortana to suggest images for your slides?" or "Would you like to combine the content of two media objects on slide 8 into one media object?" Over time, the AI ​​voice assistant will also learn typical user behavior and be able to make suggestions based on previous usage patterns. For example, if a user typically creates a merged image from two previous images when preparing a presentation, but hasn't done so yet when creating the current presentation, the AI ​​voice assistant 104 can prompt the user with audible verbal output, such as, "Would you like to upload any images you want to merge, or see suggestions using the images currently in the presentation?". Generally, if the user uploads any audio / video / images or other media files, suggestions are given to the user to use features of the presentation application (e.g., stock images or audio) to increase feature discoverability. Over time, the algorithm learns what the user's style is and improves suggestions in the present application's designer or suggestion pane. For example, if the user always inserts images with a flower or floral theme, the designer / suggestion pane will focus on suggestions containing a flower or floral theme.

[0036] Figure 4A Describing for, for example Figure 1-3 Examples of user interfaces for those systems. For example, in... Figure 4A As shown, the presentation application 106 can have an interface 300 with the following elements: Various states and toolbars 302 provide user controls at the top of the interface 300. The currently focused slide 308 is depicted in the central pane 310 and can include various content elements, including the contents of the first document 120 and the second document 122 discussed above. The left pane 304 can sequentially include smaller versions or thumbnails 306 of all the slides in the presentation. This allows the user to navigate and focus on any slide in the presentation.

[0037] The right pane 312 can be a designer or suggestion pane, which application 106 uses to suggest alternative configurations or layouts 314 for slide 308. By selecting any of these options 314, the user can reconfigure the focused slide 308. The pane 312 may also include user interface elements to invoke the content combination techniques described above. For example, when two documents 120 and 122 are added to slide 308, application 106 can respond using a prompt 320 that asks the user whether to combine the contents of the two documents, as described herein. The user can deactivate the prompt using a negative response. Alternatively, the user can affirmatively indicate the combination of document contents.

[0038] Figure 4B An example of the user interface in Figure 4 is depicted in a subsequent stage. Figure 4B In this context, the user has already responded affirmatively that the contents of the two files will be merged. Therefore, the prompt 320 is replaced by a new user interface element 322. This element 322 includes two options. First, the user can select a button 324 with instructions such as "Give me a surprise." Under this option, the system is instructed to combine or merge the contents of the two files without any specific instructions on how. Thus, the prompt-generated image 130, as discussed above, will submit a request to the generative AI tool 114 to combine the content into a single object without further instructions on how. The generative AI tool 114 will then make its own determination regarding the combination of content and can generate several different alternatives 328 that can be presented to the user in the right pane 312.

[0039] Alternatively, user interface element 322 includes a field 326 that prompts the user to enter instructions on how to merge or combine the contents of files 120, 122. In this case, the user will enter instructions, such as the example of "the person holding the flower" given above, to specify the relationship between the elements of files 120, 122 to be used in the combined output. Field 326 will accept any instructions from the user on how to combine the contents of the two files 120, 122 in the merged output.

[0040] Figure 5 Describing for, for example Figure 1-3 Another example of a user interface for those systems. Figure 5In the demonstration application 106, the user interface 300 includes interaction with the AI ​​voice assistant 104 as described above. Specifically, the AI ​​voice assistant 104 includes a user behavior machine learning (ML) tool 320. This tool 320 is trained using the user's behavior when preparing a presentation using application 106. As in the previously mentioned example, if the user typically requests a combination or fusion of content from multiple files when preparing a presentation, the ML tool 320 will learn this preference over time. Therefore, as described above, the AI ​​voice assistant 104, or other user interfaces of application 106, can prompt the user with instructions regarding whether to merge the content of multiple media files or objects from one or more slides. The AI ​​voice assistant 104 can use the computer's speakers to announce questions and use the computer's microphone to receive input responses. In another example, if the user regularly uses a specific theme or element, such as flowers or floral themes, the user behavior ML tool 320 will learn this behavior. Therefore, the user behavior ML tool 320 will include instructions for the prompt generation engine 130 to instruct the generative AI tool 114 to prepare content using a preferred theme or repeating elements, possibly including content from the two objects being merged. As a result, the suggested slides 330 in the right pane 312 will be influenced by the user behavior ML tool 320. In this way, the user's preferred themes or elements, as demonstrated by user behavior over time, will be reflected in the suggestions 330 presented to the user.

[0041] Figure 6 A flowchart illustrating a method for generating novel content according to this disclosure is provided. This is an example method of operation for a server 116, which includes a processor, a network interface, and a memory including programming instructions, such as a prompt word generation engine. When executed by the processor alone or in combination with other processors, the programming instructions cause the server to perform the following services: receiving 602 multiple media files from a demo application on a client device; reconstructing 604 the content of the media files into a form compatible with a generative artificial intelligence (AI) tool; using the prompt word generation engine, constructing 606 prompt words for the generative AI tool using the content of the media files in the form compatible with the generative AI tool, the prompt words including instructions for generating new content by fusing content from the multiple media files; receiving 608 the new content from the generative AI tool; and transmitting 609 the new content to the demo application on the client device to enhance the demo generated using the demo application.

[0042] Figure 7This is a block diagram 700 illustrating an example software architecture 702, the various parts of which can be used in conjunction with various hardware architectures described herein, which can implement any of the features described above. This software architecture can be represented in... Figure 3 The software shown in B that invokes service 301 or other components. Figure 7 This is a non-limiting example of a software architecture, and it will be appreciated that many other architectures can be implemented to facilitate the functionality described herein. Software architecture 702 can be implemented in, for example... Figure 8 The hardware of the machine 800 is executed, and the machine 800 specifically includes a processor 810, a memory 830, and an input / output (I / O) component 850. A representative hardware layer 704 is illustrated and can represent, for example... Figure 8 The machine 800. A representative hardware layer 704 includes a processing unit 706 and associated executable instructions 708. The executable instructions 708 represent executable instructions of the software architecture 702, including implementations of the methods, modules, etc., described herein. Hardware layer 704 also includes a memory / storage device 710, which also includes the executable instructions 708 and accompanying data. Hardware layer 704 may also include other hardware modules 712. The instructions 708 held by the processing unit 706 may be a portion of the instructions 708 held by the memory / storage device 710.

[0043] The example software architecture 702 can be conceptualized as layers, each providing various functionalities. For example, software architecture 702 may include layers and components such as an operating system (OS) 714, libraries 716, frameworks 718, applications 720, and a presentation layer 744. Operationally, applications 720 and / or other components within a layer can invoke API calls 724 to other layers and receive corresponding results 726. The illustrated layers are representative in nature, and other software architectures may include additional or different layers. For example, some mobile or dedicated operating systems may not provide a framework / middleware 718.

[0044] OS 714 can manage hardware resources and provide public services. OS 714 may include, for example, a kernel 728, services 730, and drivers 732. The kernel 728 can act as an abstraction layer between the hardware layer 704 and other software layers. For example, the kernel 728 may be responsible for memory management, processor management (e.g., scheduling), component management, networking, security settings, etc. Services 730 can provide other public services to other software layers. Drivers 732 can be responsible for controlling the underlying hardware layer 704 or interfacing with the underlying hardware layer 704. For example, depending on the hardware and / or software configuration, drivers 732 may include display drivers, camera drivers, memory / storage device drivers, peripheral device drivers (e.g., via Universal Serial Bus (USB)), network and / or wireless communication drivers, audio drivers, etc.

[0045] Library 716 can provide common infrastructure that can be used by application 720 and / or other components and / or layers. Library 716 typically provides functionality used by other software modules to perform tasks, rather than interacting directly with OS 714. Library 716 may include system libraries 734 (e.g., the C standard library) that provide functions such as memory allocation, string manipulation, and file operations. Additionally, library 716 may include API libraries 736, such as media libraries (e.g., libraries supporting the rendering and manipulation of image, sound, and / or video data formats), graphics libraries (e.g., OpenGL libraries for rendering 2D and 3D graphics on a display), database libraries (e.g., SQLite or other relational database functionalities), and web libraries (e.g., WebKit that provides web browsing functionality). Library 716 may also include a wide variety of other libraries 738 to provide numerous functionalities for application 720 and other software modules.

[0046] Framework 718 (sometimes referred to as middleware) provides higher-level common infrastructure that can be used by Application 720 and / or other software modules. For example, Framework 718 can provide various graphical user interface (GUI) functions, advanced resource management, or advanced location services. Framework 718 can provide a wide range of other APIs for Application 720 and / or other software modules.

[0047] Application 720 includes built-in application 740 and / or third-party application 742. Examples of built-in application 740 may include, but are not limited to: contact application, browser application, location application, media application, messaging application, and / or game application. Third-party application 742 may include any application developed by an entity other than a platform-specific vendor. Application 720 may use the functionality available via OS 714, library 716, framework 718, and rendering layer 744 to create a user interface for user interaction.

[0048] Some software architectures use virtual machines, such as the one exemplified by virtual machine 748. Virtual machine 748 provides an execution environment in which applications / modules can function as if they were running on a hardware machine (e.g., a physical machine). Figure 8 The virtual machine 748 executes in the same manner as on a host OS (e.g., OS 714) or hypervisor, and may have a virtual machine monitor 746 that manages the operation of the virtual machine 748 and interoperates with the host OS. A software architecture different from the external software architecture 702 may execute within the virtual machine 748, such as an OS 750, libraries 752, frameworks 754, applications 756, and / or a presentation layer 758.

[0049] Figure 8This is a block diagram illustrating components of an example machine 800 configured to read instructions from a machine-readable medium (e.g., a machine-readable storage medium) and execute any of the features described herein. The example machine 800 is in the form of a computer system within which instructions 816 (e.g., in the form of a software component) can be executed to cause the machine 800 to perform any of the features described herein. This example machine can be represented in... Figure 3 The hardware shown is used to call up service 301 or other components.

[0050] Thus, instruction 816 can be used to implement the modules or components described herein. Instruction 816 causes an unprogrammed and / or unconfigured machine 800 to operate as a specific machine configured to perform the described features. Machine 800 can be configured to operate as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, machine 800 can operate as a server machine or client machine in a server-client network environment, or as a node in a peer-to-peer or distributed network environment. Machine 800 can be embodied as, for example, a server computer, client computer, personal computer (PC), tablet computer, laptop computer, netbook, set-top box (STB), gaming and / or entertainment system, smartphone, mobile device, wearable device (e.g., smartwatch), and Internet of Things (IoT) device. Furthermore, although only a single machine 800 is illustrated, the term "machine" includes a collection of machines that execute instruction 816 individually or jointly.

[0051] Machine 800 may include processor 810, memory 830, and I / O components 850, which may be communicatively coupled via, for example, bus 802. Bus 802 may include various buses that couple various components of machine 800 via various bus technologies and protocols. In examples, processor 810 (including, for example, a central processing unit (CPU), graphics processing unit (GPU), digital signal processor (DSP), ASIC, or suitable combinations thereof) may include one or more processors 812a to 812n capable of executing instructions 816 and processing data. In some examples, one or more processors 810 may execute instructions provided or recognized by one or more other processors 810. The term "processor" includes multi-core processors, which include cores capable of executing instructions simultaneously. Although Figure 8 Multiple processors are shown, but machine 800 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors each with a single core, multiple processors each with multiple cores, or any combination thereof. In some examples, machine 800 may include multiple processors distributed among multiple machines.

[0052] Memory / storage device 830 may include main memory 832, static memory 834, or other memory and storage unit 836, both of which can be accessed by processor 810, such as via bus 802. Storage unit 836 and memories 832, 834 store instructions 816 embodying any one or more of the functions described herein. Memory / storage device 830 may also store temporary, intermediate, and / or long-term data for processor 810. Instructions 816 may also reside wholly or partially in memories 832, 834, storage unit 836, at least one processor in processor 810 (e.g., in an command register or cache memory), in the memory of at least one I / O component in I / O component 850, or any suitable combination thereof. Thus, memories 832, 834, storage unit 836, the memory in processor 810, and the memory in I / O component 850 are examples of machine-readable media.

[0053] As used herein, “machine-readable medium” refers to a device capable of temporarily or permanently storing instructions and data that cause machine 800 to operate in a particular manner, and may include, but is not limited to: random access memory (RAM), read-only memory (ROM), cache memory, flash memory, optical storage media, magnetic storage media and devices, cache memory, network-accessible or cloud storage, other types of storage, and / or any suitable combination thereof. The term “machine-readable medium” applies to a single medium or combination of media used to store instructions (e.g., instruction 816) for execution by machine 800, such that when executed by one or more processors 810 of machine 800, the instructions cause machine 800 to perform one or more features described herein. Therefore, “machine-readable medium” can refer to a single storage device, as well as a “cloud-based” storage system or storage network comprising multiple storage devices or equipment. The term “machine-readable medium” excludes signals themselves.

[0054] I / O component 850 may include a wide variety of hardware components suitable for receiving input, providing output, generating output, transmitting information, exchanging information, capturing measurement results, etc. The specific I / O component 850 included in a particular machine will depend on the type and / or function of the machine. For example, mobile devices such as mobile phones may include touch input devices, while headless servers or IoT devices may not include such touch input devices. Figure 8The specific examples of I / O components illustrated herein are by no means limiting, and other types of components may be included in machine 800. The grouping of I / O components 850 is for the simplicity of the discussion only, and the grouping is by no means limiting. In various examples, I / O components 850 may include user output components 852 and user input components 854. User output components 852 may include, for example, display components (e.g., liquid crystal display (LCD) or projectors) for displaying information, acoustic components (e.g., speakers), haptic components (e.g., vibration motors or force feedback devices), and / or other signal generators. User input components 854 may include, for example, alphanumeric input components (e.g., keyboards or touchscreens), pointing components (e.g., mouse devices, touchpads, or other pointing instruments), and / or haptic input components (e.g., physical buttons or touchscreens that provide the position and / or force of touch or touch gestures), configured to receive various user inputs, such as user commands and / or selections.

[0055] In some examples, I / O component 850 may include biometric component 856, motion component 858, environmental component 860, and / or position component 862, as well as various other physical sensor components. Biometric component 856 may include components for detecting bodily expressions (e.g., facial expressions, vocal expressions, hand or body gestures, or eye tracking), measuring biosignals (e.g., heart rate or brain waves), and identifying people (e.g., via voice-based, retinal, fingerprint, and / or face-based recognition). Motion component 858 may include, for example, accelerometers (e.g., accelerometers) and rotation sensors (e.g., gyroscopes). Environmental component 860 may include, for example, lighting sensors, temperature sensors, humidity sensors, pressure sensors (e.g., barometers), acoustic sensors (e.g., microphones for detecting ambient noise), proximity sensors (e.g., infrared sensing of nearby objects), and / or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment. The location component 862 may include, for example, a location sensor (e.g., a Global Positioning System (GPS) receiver), an altitude sensor (e.g., a barometric pressure sensor from which altitude can be derived), and / or an orientation sensor (e.g., a magnetometer).

[0056] I / O component 850 may include communication component 864, which implements various technologies operable to couple machine 800 to one or more networks 870 and / or one or more devices 880 via corresponding communication couplers 872 and 882. Communication component 864 may include one or more network interface components or other suitable devices to interface with one or more networks 870. Communication component 864 may include components, for example, adapted to provide wired communication, wireless communication, cellular communication, near field communication (NFC), Bluetooth communication, Wi-Fi, and / or communication via other modes. One or more devices 880 may include other machines or various peripheral devices (e.g., via USB coupling).

[0057] In some examples, communication component 864 may detect identifiers or include components suitable for detecting identifiers. For example, communication component 864 may include a radio frequency identification (RFID) tag reader, an NFC detector, an optical sensor (e.g., a one-dimensional or multi-dimensional barcode or other optical code), and / or an acoustic detector (e.g., a microphone for identifying audio signals of the tag). In some examples, location information may be determined based on information from communication component 862, such as, but not limited to, geographic location via Internet Protocol (IP) address, geographic location via Wi-Fi, cellular, NFC, Bluetooth, or other wireless station identification and / or signal triangulation.

[0058] Although various embodiments have been described, the description is intended to be exemplary and not limiting, and it should be understood that many more embodiments and implementations are possible within the scope of the embodiments. While many possible combinations of features are shown in the drawings and discussed in this detailed description, many other combinations of the disclosed features are possible. Unless specifically limited, any feature of any embodiment may be used in combination with or substituted for any other feature or element in any other embodiment. Therefore, it should be understood that any feature shown and / or discussed in this disclosure can be implemented together in any suitable combination. Therefore, the embodiments are not limited except for the appended claims and their equivalents. Furthermore, various modifications and changes may be made within the scope of the appended claims.

[0059] Typically, the functions described in this article (e.g., Figure 1-6The features shown herein can be implemented using software, firmware, hardware (e.g., fixed logic, finite state machines, and / or other circuitry), or a combination of these implementations. In the case of a software implementation, the program code performs a specified task when executed on a processor (e.g., a CPU or CPU). The program code can be stored in one or more machine-readable storage devices. The techniques described herein are characterized by system independence, meaning that these techniques can be implemented on various computing systems with a variety of processors. For example, an implementation may include an entity (e.g., software) that causes hardware to perform operations (e.g., processor function blocks, etc.). For example, a hardware device may include a machine-readable medium that can be configured to maintain instructions that cause the hardware device (including an operating system running thereon and associated hardware) to perform operations. Thus, the instructions can be used to configure the operating system and associated hardware to perform operations, and thereby configure or otherwise adapt the hardware device to perform the functions described above. The instructions can be provided by the machine-readable medium to the hardware elements executing the instructions in various different configurations.

[0060] In the following sections, further features, characteristics, and advantages of the invention will be described by way of items:

[0061] Project 1: A data processing system, comprising:

[0062] A server, which has a processor and a network interface; and

[0063] The memory includes programming instructions, which include a prompt word generation engine. When executed by the processor alone or in combination with other processors, the programming instructions cause the server to perform the following services:

[0064] Receive multiple media files from the demo application on the client device;

[0065] The content of the media files is reconstructed into a form compatible with generative artificial intelligence (AI) tools;

[0066] Using the aforementioned prompt word generation engine, prompt words for the generative AI tool are constructed using the content of the media files in a form compatible with the generative AI tool, the prompt words including instructions for generating the new content by fusing content from the multiple media files;

[0067] Receive the new content from the generative AI tool; and

[0068] The new content is transmitted to the demo application on the client device to enhance the demo generated using the demo application.

[0069] Project 2, the data processing system according to Project 1, wherein the service further includes a file detection engine, the file detection engine being used to determine the type of media in each of the plurality of media files, and to identify one or more corresponding format conversion tools invoked to reconstruct the content of the media files into a form compatible with the generative AI tool.

[0070] Project 3. According to the data processing system described in Project 1, the prompt word generation engine further includes a function to call a language translation service to translate the content of the multiple media files into English.

[0071] Project 4. According to the data processing system described in Project 1, wherein the plurality of media files are accompanied by user instructions on how to combine the content when the content of the plurality of media files is being merged, and the prompt word generation engine constructs the prompt words for the generative AI tool to include corresponding instructions on how to combine the content when the generative AI tool is merging the content of the plurality of media files.

[0072] Project 5. The data processing system according to Project 1, wherein the demonstration application is integrated with an AI voice assistant, the AI ​​voice assistant being configured to receive verbal user input regarding the content of the fused multiple media files.

[0073] Project 6. The data processing system according to Project 1, wherein the service uses the generative AI tool to generate multiple alternative versions of the new content and provides the alternative versions to the demonstration application.

[0074] Project 7. A demo application for execution by a client device including a processor and memory, the demo application being stored in the memory of the client device and including executable instructions that, when executed by the processor, cause the processor, alone or in combination with other processors, to:

[0075] A user interface for generating a presentation is presented, the user interface being configured to receive multiple media files, each media file including content to be included in the presentation; and

[0076] In response to user commands, the service is invoked as follows:

[0077] The content of the media files is reconstructed into a form compatible with generative artificial intelligence (AI) tools;

[0078] The content of the media files in a form compatible with the generative AI tool is used to construct prompts for the generative AI tool, the prompts including instructions for generating new content by fusing content from the multiple media files; and

[0079] The new content is transmitted to the demo application on the client device;

[0080] The application displays the new content in the user interface, which has controls for using the new content to enhance the presentation generated using the demo application.

[0081] Project 8, based on the demonstration application described in Project 7, further includes an element that asks the user whether to merge the content of the received media file.

[0082] Project 9, the demonstration application as described in Project 8, wherein the element is displayed in response to the application receiving the plurality of media files.

[0083] Project 10, a demonstration application based on Project 8, wherein the element includes a field for accepting user instructions on how the content of the multiple media files is combined by the generative AI tool.

[0084] Project 11, a demonstration application based on Project 8, wherein the elements include user controls for invoking the service without user instructions on how the content of the multiple media files is combined by the generative AI tool.

[0085] Item 12, the demonstration application according to Item 7, the user interface further includes a pane in which design suggestions for slides for the demonstration are displayed.

[0086] Project 13, a demo application according to Project 12, wherein the service returns multiple different versions of the new content, and the demo application displays the multiple different versions of the new content in a design suggestion pane having user controls for selecting which of the different versions to include in the demo.

[0087] Project 14. The demonstration application according to Project 7, wherein the application includes an interface for communicating with an artificial intelligence (AI) voice assistant.

[0088] Project 15. A demonstration application based on Project 14, wherein the application is configured to receive converted user input, the converted user input indicating when or how the application should use the service to merge the content of the multiple media files.

[0089] Project 16, the demonstration application according to Project 14, wherein the application is configured to use the AI ​​voice assistant to output audible prompts to the user for integrating the content of the multiple media files.

[0090] Project 17. A demonstration application based on Project 14, wherein the application includes access to a user behavior machine learning tool to learn user behavior regarding the content of the multiple media files, and the application outputs recommendations based on the learned user behavior.

[0091] Project 18. A method for providing a service to generate new content by combining two sets of inputs, the service being supported by a server having a processor, a network interface, and memory, the memory including programming instructions including a prompt word generation engine, the programming instructions, when executed by the processor alone or in combination with other processors, causing the server to implement the method, the method comprising:

[0092] Receive multiple media files from the demo application on the client device;

[0093] The content of the media files is reconstructed into a form compatible with generative artificial intelligence (AI) tools;

[0094] Using the aforementioned prompt word generation engine, prompt words for the generative AI tool are constructed using the content of the media files in a form compatible with the generative AI tool, the prompt words including instructions for generating new content by fusing content from the multiple media files;

[0095] Receive the new content from the generative AI tool; and

[0096] The new content is transmitted to the demo application on the client device to enhance the demo generated using the demo application.

[0097] Item 19. The method according to Item 18, wherein the plurality of media files are accompanied by user instructions on how to combine the content when the content of the plurality of media files is being merged, the method further comprising: constructing the prompt words for the generative AI tool to include corresponding instructions on how to combine the content when the generative AI tool is merging the content of the plurality of media files.

[0098] Project 20, the method according to Project 18, wherein the demonstration application is integrated with an AI voice assistant, and the method further includes receiving verbal user input regarding the content of the fused multiple media files.

[0099] In the foregoing detailed description, numerous specific details have been illustrated with examples to provide a thorough understanding of the teachings. It will be apparent to those skilled in the art, upon reading this specification, that various aspects can be practiced without these details. In other instances, well-known methods, processes, components, and / or circuits have been described in relatively high-level but undetailed terms to avoid unnecessarily obscuring aspects of the teachings.

[0100] While what is considered the best pattern and / or other examples has been described above, it should be understood that various modifications can be made therein, and the subject matter disclosed herein can be implemented in various forms and examples, and the teachings can be applied to many applications, only some of which are described herein. The appended claims are intended to cover any and all applications, modifications, and variations that fall within the true scope of this teaching.

[0101] Unless otherwise stated, all measurements, values, ratings, locations, sizes, dimensions, and other specifications set forth in this specification (including those in the following claims) are approximate, not precise. They are intended to have a reasonable range consistent with the functions they address and the conventions of the field to which they belong.

[0102] The scope of protection is defined solely by the following claims. This scope is intended and should be construed as consistent with the ordinary meaning of the language used in the claims when interpreted in accordance with this specification and subsequent examination history, and covers all structural and functional equivalents. Nevertheless, none of the claims are intended to cover subject matter that fails to meet the requirements of sections 101, 102, or 103 of the Patent Law, nor should they be interpreted in this manner. Any unintended encirclement of such subject matter is hereby claimed.

[0103] Apart from what has just been stated above, nothing that has been stated or shown is intended or should be construed as giving rise to any component, step, feature, object, benefit, advantage, or equivalent to any other thing that is known to the public, whether or not it is stated in the claims.

[0104] It will be understood that the terms and expressions used herein have the general meaning consistent with those in the relevant queries and fields of study, unless otherwise stated herein.

[0105] Relational terms such as "first" and "second" may be used solely to distinguish one entity or action from another, without necessarily requiring or implying any actual such relationship or order between such entities or actions. The terms "comprising," "including," and any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article of manufacture, or apparatus that includes a list of elements includes not only those elements but may include other elements not expressly listed or inherent to such a process, method, article of manufacture, or apparatus. An element preceded by "a" or "an" does not, without further constraints, exclude the presence of additional identical elements in the process, method, article of manufacture, or apparatus that includes that element.

[0106] An abstract of this disclosure is provided to allow the reader to quickly identify the nature of the technical disclosure. It should be understood that it is not intended to interpret or limit the scope or meaning of the claims. Furthermore, as can be seen from the foregoing detailed description, various features have been grouped together in various examples for the purpose of simplifying this disclosure. This approach of the disclosure should not be construed as reflecting an intention in any claim to require more features than those expressly stated in the claims. Rather, as reflected in the following claims, the subject matter of the invention lies in all features less than those in a single disclosed example. Therefore, the following claims are incorporated into the detailed description, wherein each claim is, in itself, a separately claimed subject matter.

Claims

1. A data processing system, comprising: Server 116, which has processor 90 and network interface 92; as well as Memory 91 includes programming instructions, including a prompt word generation engine 118, which, when executed by processor 90 alone or in combination with other processors, cause server 116 to perform the following services: Receive multiple media files 120, 122 from application 106 on client device 102; The content of the media files 120 and 122 is reconstructed into a form compatible with the generative artificial intelligence (AI) tool 114; Using the prompt word generation engine 118, prompt words for the generative AI tool 114 are constructed using the content of the media files 120, 122 in a form compatible with the generative AI tool 114. The prompt words include instructions for generating new content by fusing content from the plurality of media files 120, 122. Receive the new content from the generative AI tool 114; as well as The new content is transmitted to the application 106 on the client device 102 to enhance the work generated by the application 106.

2. The data processing system according to claim 1, wherein, The service also includes a file type detection engine 130, which is used to: determine the type of media in each of the plurality of media files 120, 122, and identify one or more corresponding format conversion tools 124 that are invoked to reconstruct the content of the media files 120, 122 into a form compatible with the generative AI tool 114.

3. The data processing system according to claim 1 or claim 2, wherein, The prompt word generation engine 118 also includes a function to call the language translation service 105 to translate the content of the plurality of media files 120, 122 into English.

4. The data processing system according to any one of claims 1-3, wherein, The plurality of media files 120, 122 are accompanied by user instructions on how to combine the content when the content of the plurality of media files 120, 122 is being merged, and the prompt word generation engine 118 is used to construct the prompt words for the generative AI tool 114 to include corresponding instructions on how to combine the content when the generative AI tool 114 is merging the content of the plurality of media files 120, 122.

5. The data processing system according to any one of claims 1-4, wherein, The application 106 is integrated with an AI voice assistant 104, which is configured to receive verbal user input regarding the content of the plurality of media files 120, 122.

6. The data processing system according to any one of claims 1-5, wherein, The service uses the generative AI tool 114 to generate multiple alternative versions of the new content and provides the alternative versions to the application 106.

7. A demonstration application 106 for execution by a client device 102 including a processor and memory 103, the demonstration application 106 being stored in the memory 103 of the client device 102 and including executable instructions that, when executed by the processor, cause the processor, alone or in combination with other processors, to: A user interface 300 for generating a presentation is presented, the user interface 300 being configured to receive a plurality of media files 120, 122, each media file including content to be included in the presentation; as well as In response to user commands, the service is invoked as follows: The content of the media files 120 and 122 is reconstructed 604 into a form compatible with the generative artificial intelligence (AI) tool 114; The content of the media files 120, 122 in a form compatible with the generative AI tool 114 is used to construct 606 prompts for the generative AI tool 114, the prompts including instructions for generating new content by fusing content from the plurality of media files 120, 122; as well as The new content is transmitted 609 to the demo application 106 on the client device 102; The application 106 displays the new content in the user interface 300, which has controls for using the new content to enhance the presentation generated by the presentation application 106.

8. The demonstration application 106 according to claim 7, wherein the user interface 300 further includes an element 320 that asks the user whether to merge the content of the received media files 120, 122.

9. The demonstration application 106 according to claim 8, wherein, The element 320 is displayed in response to the application 106 receiving the plurality of media files 120, 122.

10. The demonstrative application 106 according to claim 8 or claim 9, wherein, The element 322 includes a field 326 for receiving user instructions on how the content of the plurality of media files 120, 122 is combined by the generative AI tool 114.

11. The demonstration application 106 according to any one of claims 8-10, wherein, The element 322 includes a user control 324 for invoking the service without user instructions on how the content of the plurality of media files 120, 122 is combined by the generative AI tool 114.

12. The presentation application 106 according to any one of claims 7-11, wherein the user interface 300 further includes a pane 312 in which design suggestions for slides for the presentation are displayed.

13. The demonstration application 106 according to claim 12, wherein, The service returns multiple different versions 314 of the new content, and the demo application 106 displays the multiple different versions of the new content in the design suggestion pane 312, which has user controls for selecting which version of the different versions 314 to include in the demo.

14. The demonstration application 106 according to any one of claims 7-13, wherein, The application 106 includes an interface for communicating with an artificial intelligence (AI) voice assistant 104.

15. The demonstration application 106 according to claim 14, wherein, The application 106 is configured to receive converted user input instructing the application 106 when or how to use the service to merge the contents of the plurality of media files 120, 122.

16. The demonstrative application 106 according to claim 14 or claim 15, wherein, The application 106 is configured to use the AI ​​voice assistant 104 to output audible prompts to the user for integrating the content of the multiple media files 120, 122.

17. The demonstration application 106 according to any one of claims 14-16, wherein, The application 106 includes access to a user behavior machine learning tool 320 to learn user behavior regarding the content of the combined multiple media files 120, 122, and the application 106 outputs recommendations based on the learned user behavior.

18. A method for providing a service to generate new content by combining two sets of inputs, the service being supported by a server 116 having a processor 90, a network interface 92, and a memory 91, the memory 91 including programming instructions including a prompt word generation engine 118, the programming instructions, when executed by the processor 90 alone or in combination with other processors 90, causing the server 116 to implement the method, the method comprising: Receive multiple media files 120, 122 from the demo application 106 on the client device 102; The content of the media files 120 and 122 is reconstructed into a form compatible with generative artificial intelligence (AI) tools; Using the prompt word generation engine 118, prompt words for the generative AI tool 114 are constructed using the content of the media files 120, 122 in a form compatible with the generative AI tool 114. The prompt words include instructions for generating new content by fusing content from the plurality of media files 120, 122. Receive the new content from the generative AI tool 114; as well as The new content is transmitted to the demo application 106 on the client device 102 to enhance the demo generated using the demo application 106.

19. The method according to claim 18, wherein, The plurality of media files 120, 122 are accompanied by user instructions on how to combine the content when the content of the plurality of media files 120, 122 is being merged. The method also includes constructing the prompts for the generative AI tool 114 to include corresponding instructions on how to combine the content when the generative AI tool 114 is merging the content of the plurality of media files 120, 122.

20. The method according to claim 18 or claim 19, wherein, The demonstration application 106 is integrated with the AI ​​voice assistant 104, and the method also includes receiving verbal user input regarding the content of the multiple media files 120, 122.