Multi-user cues for generative artificial intelligence systems
By introducing multi-party interfaces and composite prompting mechanisms into CAD applications, the problem that AI models cannot accurately reflect collective intent in multi-user collaborative design is solved, thereby improving design quality and exploration efficiency.
Patent Information
- Application Number
- CN202480064041.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-05
- Filing Date
- 2024-09-20
- Publication Date
- 2026-05-01
AI Technical Summary
Existing computer-aided design (CAD) applications fail to effectively facilitate communication among multiple users in multi-user collaborative design processes. This results in 3D objects generated by AI models failing to accurately reflect the ideas and intentions of the collective user group, thus reducing design quality and exploration efficiency.
By generating a multi-party interface that communicates with at least a trained machine learning (ML) model, a first client device, and a second client device, inputs from different users are combined to generate a compound prompt, which is then transmitted to the ML model to generate digital content items that are displayed in the multi-party interface.
This enables AI systems to more accurately understand and reflect the collective thoughts and intentions of user groups without requiring a separate communication channel, thereby generating digital content that better meets actual needs.
Smart Images

Figure CN121970055A_ABST
Abstract
Description
Multi-user prompts for generative artificial intelligence systems
[0001] Cross-Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 588,657, filed October 6, 2023, entitled “TECHNIQUES FORENABLING MULTIPLE USERS TO CONTRIBUTE TO ONE OR MORE GENERATIVE ARTIFICIAL INTELLIGENCE SYSTEMS,” and claims priority to U.S. Patent Application No. 18 / 765,122, filed July 5, 2024, entitled “MultI-User Prompts for Generative Artificial Intelligence Systems.” The subject matter of these related applications is hereby incorporated by reference. Technical Field
[0002] The various implementation schemes involve computer-aided design and artificial intelligence as a whole, and more specifically, multi-user prompts for generative artificial intelligence systems. Background Technology
[0003] Designers interact with various types of digital content creation (DCC) applications to create diverse content, including but not limited to text, images, videos, and designs such as three-dimensional (3D) objects. During digital content creation, users typically experience a rapid generation and exploration process, where they explore various avenues for creating content that meets diverse design goals. For example, when drafting an initial version of a poem, a poet might list various keywords, synonyms, antonyms, and rhyming terms. Similarly, before choosing a specific art style to complete an illustration, an artist might apply different sketching techniques to draft illustrations in different art styles. During the design exploration phase for 3D objects, designers typically generate and evaluate various design alternatives for one or more 3D objects within a larger 3D design project. As is well known, even for relatively simple 3D objects, manually generating multiple designs can be extremely labor-intensive and time-consuming.
[0004] To address the aforementioned issues associated with the design exploration phase for 3D objects, various conventional computer-aided design (CAD) applications have been developed that implement artificial intelligence (AI) models, such as generative machine learning (ML) models, to automatically synthesize 3D objects in response to prompts provided by the user. In operation, the AI model responds to given prompts by executing various optimization algorithms to generate 3D object designs that satisfy one or more design characteristics specified by the user in the prompts. In some cases, the AI model generates a single 3D object design, which the user then incorporates into a larger 3D design project. In other cases, the AI model generates numerous alternative 3D object designs and presents these alternatives to the user for evaluation and selection.
[0005] One drawback of the above approach is that conventional CAD applications implementing AI models do not effectively facilitate communication between the AI model and a group of multiple users. In this regard, conventional CAD applications typically focus on one-to-one interaction between the AI model and a single user. In doing so, conventional CAD applications usually provide an interface that allows a single user to input prompts for the AI model, and the AI model generates output in response to the prompts. However, it is worth noting that these types of conventional interfaces do not provide a practical way for multiple users to communicate with both the AI model and other users. Therefore, when multiple users work together on a given 3D design project, different users are forced to interact with each other using separate communication channels (such as multi-user chat services or email threads) in order to generate prompts for the AI model.
[0006] Typically, during a one-on-one session, one user must refine different goals and input these cues into the AI model. This process often results in cues that do not capture and include all the information that the different users in the group want to provide to the AI model. Specifically, the user responsible for generating cues may not properly incorporate the goals and constraints expressed by the collective user group into the different cues. As a result, the cues may not reflect the collective user group's ideas and intentions regarding the design of various 3D objects. Without accurate cues, the AI model cannot generate 3D objects that accurately reflect the collective user group's ideas and intentions, which can significantly reduce overall design quality and inhibit exploration of the overall design space.
[0007] As shown above, what is needed in this field is a more efficient technology that uses artificial intelligence models to automatically generate designs. Summary of the Invention
[0008] In various implementations, a computer-implemented method for generating digital content includes: generating a multi-party interface communicating with at least a trained machine learning (ML) model, a first client device, and a second client device; combining a first input from the first client device and a second input from the second client device to generate a composite prompt; transmitting the composite prompt to the trained ML model for execution; receiving a digital content item generated in response to the composite prompt from the trained ML model; and displaying the digital content item in the multi-party interface.
[0009] At least one technical advantage of the disclosed technology over existing technologies lies in its ability to enable CAD applications to collect and aggregate input from multiple different users within a user group when generating prompts. This allows AI systems to more accurately understand the collective thoughts and intentions of the user group and generate digital content that more accurately reflects these collective thoughts and intentions. In this regard, the disclosed technology provides an automated process for collecting multiple prompts generated from multiple users within a shared interface and weighting these prompts before transmitting them to an AI model for execution. Collecting multiple prompts as input to the AI model and weighting these prompts allows group members to clarify the group's collective thoughts and intentions by emphasizing specific ideas and goals. Therefore, the disclosed technology enables AI models to better infer the thoughts and intentions of the user group and generate digital content that more accurately reflects these thoughts and intentions. Thus, the disclosed technology allows a group of users to use a CAD application to generate digital content that better matches the group's actual thoughts and intentions without requiring coordination among group members using separate communication channels. These technical advantages provide one or more technological advancements superior to existing methods. Attached Figure Description
[0010] To gain a more detailed understanding of the features described above in the various embodiments, reference can be made to the various embodiments for a more specific description of the inventive concept briefly outlined above, some of which are already illustrated in the accompanying drawings. However, it should be noted that the drawings illustrate only typical embodiments of the inventive concept and should therefore not be construed as limiting the scope in any way, and that other equivalent embodiments exist.
[0011] Figure 1 is a conceptual illustration of a system configured to implement one or more aspects of various implementation schemes; Figure 2 is a more detailed illustration of a design exploration application of Figure 1 according to various implementation schemes; Figure 3 is an exemplary illustration of a multi-user system including the intent management application of Figure 2 according to various implementation schemes; Figure 4 is an exemplary illustration of a multi-party interface of Figure 3 according to various implementation schemes; Figure 5 is an exemplary illustration of another multi-party interface of Figure 3 according to various other implementation schemes; Figure 6A is an exemplary illustration of a weighted interface implemented by the intent management application of Figure 3 according to various implementation schemes; Figure 6B is an exemplary illustration of another weighted interface implemented by the intent management application of Figure 3 according to various other implementation schemes; Figure 7 illustrates a flowchart of method steps for generating digital content items according to various implementation schemes; and Figure 8 depicts an architecture of a system in which implementation schemes of the present disclosure may be implemented. Detailed Implementation
[0012] In the following description, numerous specific details are set forth to provide a more thorough understanding of various embodiments. However, it will be apparent to those skilled in the art that the inventive concept can be practiced without one or more of these specific details. For purposes of explanation, multiple instances of similar objects are symbolized where necessary using reference numerals identifying the objects and brackets identifying the instances.
[0013] System Overview Figure 1 is a conceptual illustration of a system 100 configured to implement one or more aspects of various implementation schemes. As shown, in some implementations, system 100 includes, but is not limited to, a client device 110, a server device 160, and one or more remote machine learning (ML) models 190. Client device 110 includes, but is not limited to, a processor 112, one or more input / output (I / O) devices 114, and memory 116. Memory 116 includes, but is not limited to, a graphical user interface (GUI) 120, a design exploration application 130, and local data storage 140. Local data storage 140 includes, but is not limited to, one or more data files 142 and one or more design objects 144. Server device 160 includes, but is not limited to, a processor 162, one or more I / O devices 164, and memory 166. Memory 166 includes, but is not limited to, an intent management application 170, one or more trained ML models 180, and a design history 182. In some other implementations, system 100 may include any number and / or type of other client devices, server devices, remote ML models, or any combination thereof.
[0014] Any number of components of System 100 can be distributed across multiple geographical locations, or in any combination within one or more cloud computing environments. For exampleThis is implemented within encapsulated shared resources, software, and data. In some implementations, client device 110 and / or zero or more other client devices (not shown) may be implemented as one or more computing instances in a cloud computing environment, as part of any other distributed computing environment, or implemented independently. In various implementations, client device 110 may be used with any number and / or type of other devices (… For example (One or more other computing instances and / or display devices) are integrated into the user device. Some examples of user devices include, but are not limited to, desktop computers, laptop computers, smartphones, and tablet computers.
[0015] Generally, client device 110 is configured to implement one or more software applications. For illustrative purposes only, each software application is described as residing in memory 116 of client device 110 and executing on processor 112 of client device 110. In some embodiments, any number of instances of any number of software applications may reside in memory 116 and any number of other memories associated with any number of other computing instances, and execute in any combination on processor 112 of client device 110 and any number of other processors associated with any number of other computing instances. In the same or other embodiments, the functionality of any number of software applications may be distributed across any number of other software applications residing in memory 116 and any number of other memories associated with any number of other computing instances, and execute in any combination on processor 112 and any number of other processors associated with any number of other computing instances. Furthermore, subsets of the functionality of multiple software applications may be combined into a single software application.
[0016] Specifically, client device 110 is configured to implement design exploration application 130 to generate designs for one or more 3D objects. In operation, design exploration application 130 causes one or more ML models 180, 190 to synthesize designs for 3D objects based on any number of objectives and constraints. Design exploration application 130 then presents the designs as one or more design objects 144 to the user within the context of a design space. In some embodiments, the user can explore and modify one or more design objects via GUI 120. Additionally or alternatively, the user may also include at least one of the design objects 144 for additional design and / or manufacturing activities.
[0017] In various embodiments, processor 112 can be any instruction execution system, device, or apparatus capable of executing instructions. For example, processor 112 may include a central processing unit (CPU), digital signal processing unit (DSP), microprocessor, application-specific integrated circuit (ASIC), neural processing unit (NPU), graphics processing unit (GPU), field-programmable gate array (FPGA), controller, microcontroller, state machine, or any combination thereof. In some embodiments, processor 112 is a programmable processor that executes program instructions to manipulate input data. In some embodiments, processor 112 may include any number of processing cores, memory, and other modules for facilitating program execution.
[0018] Input / output (I / O) device 114 includes devices configured to receive input, such as a keyboard, mouse, etc. In some embodiments, I / O device 114 also includes devices configured to provide output, such as a display device, speaker, etc. Additionally or alternatively, I / O device 114 may also include devices configured to receive and provide input and output, such as a touchscreen, Universal Serial Bus (USB) port, etc.
[0019] Memory 116 includes storage modules or a collection of storage modules. In some embodiments, memory 116 may include various computer-readable media selected for their size, relative performance, or other capabilities: volatile and / or non-volatile media, removable and / or non-removable media, etc. Memory 116 may include cache, random access memory (RAM), storage devices, etc. Memory 116 may include one or more discrete memory modules, such as dynamic RAM (DRAM) dual in-line memory modules (DIMMs). Of course, various memory chips, bandwidths, and form factors can be alternatively selected. Memory 116 stores content (such as software applications and data) for use by processor 112. In some embodiments, a storage device (not shown) supplements or replaces memory 116. The storage device may include any number and type of external memory accessible to the processor 112 of client device 110. For example, but not limited to, the storage device may include a secure digital card (SD card), external flash memory, portable optical disc read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0020] The non-volatile memory included in memory 116 typically stores one or more application programs (including design exploration application 130) and data for processing by processor 112. For exampleThe data files 142 and / or design objects stored in local data storage 140. In various embodiments, memory 116 may include non-volatile memory, such as an optical disc drive, magnetic drive, flash memory drive, or other storage device. In some embodiments, separate data storage, such as one or more external data storage devices (“cloud storage”) connected via network 150, may supplement memory 116. In various embodiments, design exploration application 130 within memory 116 may be executed by processor 112 to implement the overall functionality of client device 110 to coordinate the operation of system 100 as a whole.
[0021] In various embodiments, memory 116 may include one or more modules for performing the various functions or techniques described herein. In some embodiments, one or more of the modules and / or applications included in memory 116 may be implemented locally on client device 110 and / or via a cloud-based architecture. For example, any of the modules and / or applications included in memory 116 may be implemented on a remote device communicating with client device 110 via a network interface or I / O device interface. For example It is executed on (such as smartphones, server systems, cloud computing platforms, etc.).
[0022] Design exploration application 130 resides in memory 116 and executes on processor 112 of client device 110. Design exploration application 130 interacts with the user via GUI 120. In some embodiments, design exploration application 130 and one or more separate applications (not shown) interact with the same user via GUI 120. In various embodiments, design exploration application 130 operates as a 3D design application to generate and modify an overall 3D design including one or more design objects 144. Design exploration application 130 interacts with the user via GUI 120 to allow direct user input (… For example One or more tools used to generate 3D objects, wireframe geometry, meshes, etc., or via a separate device. For example One or more design objects 144 are generated by a trained ML model 180, a remote ML model 190, or a standalone 3D design application. When one or more design objects 144 are generated via a standalone device, the design exploration application 130 uses one or more modalities ( For exampleThe design exploration application 130 generates prompts (text, voice, images, etc.) that effectively describe the design-related intent. Then, the design exploration application 130 causes one or more of the ML models 180, 190 to interact with the generated prompts to produce related design objects 144. The design exploration application 130 receives design objects 144 from one or more ML models 180, 190 and displays them within a GUI 120. Users can select design objects 144 for use via the GUI 120, such as incorporating them into larger 3D designs.
[0023] In some implementations, the design exploration application 130 may operate as another type of digital content creation (DCC) application. For example, the design exploration application 130 may operate as an image editor to generate and modify 2D or 3D images. In another example, the design exploration application 130 may operate as a video editor application to generate and modify audiovisual content. When the design exploration application operates as a DCC application, the design exploration application 130 interacts with the user via the GUI 120 to generate one or more content items directly or via ML models 180, 190. When generating content items via ML models 180, 190, the design exploration application 130 generates prompts that effectively describe the design-related intent for the specific type of digital content item to be generated. For example (Describing aspects of a 2D image or sketch). Therefore, the design exploration application 130 can generate various types of digital content items, including but not limited to text, computer-aided design (CAD) objects, geometry, images, sketches, videos, executable code, or audio recordings.
[0024] GUI 120 can be any type of user interface that allows a user to interact with one or more software applications via any number and / or type of GUI elements. GUI 120 can be displayed in any technically feasible manner on any number and / or type of stand-alone display devices, any number and / or type of displays integrated into any number and / or type of user devices, or any combination thereof. Design exploration application 130 can perform any number and / or type of operations to directly and / or indirectly display and monitor any number and / or type of interactive GUI elements and / or any number and / or type of non-interactive GUI elements within GUI 120. In some embodiments, each interactive GUI element implements one or more types of user interactions that automatically trigger corresponding user events. Some examples of interactive GUI element types include, but are not limited to, scrollbars, buttons, text input boxes, drop-down lists, and sliders. In some embodiments, design exploration application 130 organizes GUI elements into one or more container GUI elements (...). For example (panes and / or panes).
[0025] In some implementations, GUI 120 includes one or more communication channels with one or more other devices and / or entities. For example, design exploration application 130 may include a communication channel with a trained ML model 180 via intent management application 170. As will be discussed in further detail below, this can be achieved on two or more client devices 110 ( For example , 110(1), ..., 110(X)) and one or more ML models 180 ( For example , 180(1), ..., 180(Y)) and / or one or more remote ML models 190 ( For example A communication channel is established between 190(1), ..., 190(Z).
[0026] Local data storage 140 is part of the storage device in client device 110, storing one or more design objects 144 included in the overall 3D design and / or one or more data files 142 associated with the 3D design. For example, the overall 3D design of a building may include multiple stored design objects 144, including design objects 144 representing doors, windows, fixtures, walls, appliances, etc. Local data storage 140 may also include data files 142 associated with the generated overall 3D design. For example (Component files, metadata, etc.). Additionally or alternatively, local data storage 140 includes data files 142 associated with generating prompts for transfer to one or more ML models 180, 190. For example, local data storage 140 may store sketches, geometry, etc. For example (wireframes, meshes, etc.), images, videos, application status ( For example The local data storage 142 stores one or more data files, including camera angles used within the design space, tools selected by the user, and audio recordings. In some implementations, the local data storage 140 stores one or more digital content items created via the design exploration application 130. For example, the local data storage 140 may store 2D images created directly by the user or generated from one or more ML models 180, 190.
[0027] Design object 144 includes geometry, texture, images, and / or other components used by design exploration application 130 to generate an overall 3D design. In various embodiments, the geometry of a given design object refers to any multidimensional model of the physical structure, including CAD models, meshes, and point clouds, as well as circuit layouts, piping diagrams, freeform diagrams, and so on. In some embodiments, design exploration application 130 stores multiple design objects 144 for a given 3D design and stores multiple iterations of a given target object for ML modeling 180, 190. For example, a user can use design exploration application 130 to form an initial prompt and receive a first generated design object 144(1) from a trained ML model 180(1), then refine the prompt and receive a second generated design object 144(2) from the trained ML model 180(1).
[0028] Network 150 can be any technically feasible set of interconnected communication links, including a local area network (LAN), a wide area network (WAN), the World Wide Web, or the Internet. Network 150 enables communication between client device 110 and other devices in network 150 via wired and / or wireless communication protocols, including Bluetooth, Bluetooth Low Energy (BLE), Wi-Fi, cellular protocols, satellite networks, and / or Near Field Communication (NFC).
[0029] Server device 160 is configured to communicate with design exploration application 130 to generate one or more design objects 144. In operation, server device 160 executes intent management application 170 to process prompts generated by design exploration application 130, selects one or more ML models 180, 190 trained to generate design objects 144 in response to the content of the prompts, and inputs the prompts into the selected ML models 180, 190. Once the selected ML models 180, 190 generate design objects 144 in response to the prompts, server device 160 transmits the generated design objects to client device 110, whereby the generated design objects 144 can be used by design exploration application 130.
[0030] In various embodiments, processor 162 can be any instruction execution system, device, or apparatus capable of executing instructions. For example, processor 162 may include a central processing unit (CPU), digital signal processing unit (DSP), microprocessor, application-specific integrated circuit (ASIC), neural processing unit (NPU), graphics processing unit (GPU), field-programmable gate array (FPGA), controller, microcontroller, state machine, or any combination thereof. In some embodiments, processor 162 is a programmable processor that executes program instructions to manipulate input data. In some embodiments, processor 162 may include any number of processing cores, memory, and other modules for facilitating program execution.
[0031] Input / output (I / O) device 164 includes means configured to receive input, such as a keyboard, mouse, etc. In some embodiments, I / O device 164 also includes means configured to provide output, such as a display device, speaker, etc. Additionally or alternatively, I / O device 164 may further include means configured to receive and provide input and output respectively, such as a touch screen, universal serial bus (USB) port, etc.
[0032] Memory 166 includes storage modules or a collection of storage modules. In some embodiments, memory 166 may include various computer-readable media selected for their size, relative performance, or other capabilities: volatile and / or non-volatile media, removable and / or non-removable media, etc. Memory 166 may include cache, random access memory (RAM), storage devices, etc. Memory 166 may include one or more discrete memory modules, such as dynamic RAM (DRAM) dual in-line memory modules (DIMMs). Of course, various memory chips, bandwidths, and form factors can be alternatively selected. Memory 166 stores content (such as software applications and data) for use by processor 162. In some embodiments, storage devices (not shown) supplement or replace memory 166. The storage device may include any number and type of external memory accessible to the processor 162 of server device 160. For example, but not limited to, the storage device may include a secure digital card (SD card), external flash memory, portable optical disc read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0033] The non-volatile memory included in memory 166 typically stores one or more applications (including intent management application 170 and one or more trained ML models 180) and data for processing by processor 112. For example(Design History 182). In various embodiments, memory 166 may include non-volatile memory, such as an optical disc drive, magnetic drive, flash memory, or other memory. In some embodiments, separate data storage (such as one or more external data storage devices connected via network 150) may supplement memory 166. In various embodiments, intent management application 170 and / or one or more ML models 180 within memory 166 may be executed by processor 162 to enable the overall functionality of server device 160, thereby coordinating the operation of system 100 as a whole.
[0034] In various embodiments, memory 166 may include one or more modules for performing the various functions or techniques described herein. In some embodiments, one or more of the modules and / or applications included in memory 166 may be implemented locally on client device 110, server device 160, and / or via a cloud-based architecture. For example, any of the modules and / or applications included in memory 166 may be implemented on a remote device communicating with server device 160 via a network interface or I / O device interface. For example It can be executed on smartphones, server systems, cloud computing platforms, etc. Additionally or alternatively, the intent management application 170 can be executed on client device 110 and can communicate with a trained ML model 180 operating at server device 160.
[0035] In various implementations, the intent management application 170 receives prompts from the design exploration application 130 and inputs these prompts into the applicable ML models 180, 190. In some implementations, the intent management application 170 maintains a multi-party interface (not shown) that serves as a communication channel between one or more client devices 110 and the ML models 180, 190. In this case, the intent management application 170 and / or the ML models 180, 190 participating in the multi-party interface can process input provided by the client device 110 to determine whether the input includes prompts for the ML models 180, 190 to generate outputs. In some implementations, one or more of the ML models 180, 190 are trained to respond to specific types of input, such as being trained to respond to specific modalities (…). For example The intention management application 170 combines cues (text and images) to generate an ML model of the design object 144. In this case, the intention management application 170 processes the cues to determine the modality of the data included in the cues and identifies one or more ML models 180, 190 that have been trained to respond to such combinations of modalities. Upon identifying one or more applicable ML models, the intention management application 170 selects an ML model (text and images). For example The trained ML model 180(1) will input the prompt into the selected ML model 180(1).
[0036] The trained ML model 180 includes one or more generative ML models, which have been trained on a relatively large amount of existing data and any number of possible results. For example The training model 180 is trained on design objects 144 and user-provided evaluations to perform any number and / or type of prediction tasks based on patterns detected in existing data. In various embodiments, the remote ML model 190 is a trained ML model that communicates with server device 160 to receive cues via intent management application 170. In some embodiments, the trained ML model 180 is trained using various combinations of data from multiple modalities, such as text data, image data, sound data, etc. The trained ML model 180 and / or the remote ML model 190 trained using at least two data modalities are also referred to herein as multimodal ML models. For example, in some embodiments, one or more trained ML models 180 may include a third-generation generative pre-trained transformer (GPT-3) model, a specialized version of the GPT-3 model referred to as the “DALL-E2” model, a fourth-generation generative pre-trained transformer (GPT-4) model, etc. In various embodiments, the trained ML model 180 may be trained to generate design objects from various combinations of modalities. These combinations include text, CAD objects, geometry, images, sketches, videos, application status, audio recordings, etc.
[0037] Design history 182 includes data and metadata associated with one or more trained ML models 180 and / or one or more remote ML models 190 that generate design objects 144 in response to prompts provided by design exploration application 130. In some embodiments, design history 182 includes successive iterations of design objects 144 generated by a single ML model 180 in response to a series of prompts. Additionally or alternatively, design history 182 includes multiple design objects 144 generated by different ML models 180, 190 in response to the same prompt. In some embodiments, design history 182 includes design objects 144 generated by one or more users and / or one or more ML models 180, 190 (…). For example The server device 160 can use the design history 182 as training data to further train one or more ML models 180. Additionally or alternatively, the design exploration application 130 can retrieve the contents of the design history 182 and display the retrieved content to the user via a GUI 120.
[0038] Figure 2 is a more detailed illustration of the design exploration application 130 of Figure 1 according to various implementation schemes. As shown, in some implementation schemes, system 200 includes, but is not limited to, a GUI 120, a design exploration application 130, a local data storage 140, one or more data files 142, a server device 160, a remote ML model 190, and a multimodal cue 260. The GUI 120 includes, but is not limited to, a cue space 220 including one or more cue volumes 222 and a design space 230. The design exploration application 130 includes, but is not limited to, an intent manager 240 including one or more keyword datasets 242, one or more design objects 144, and a visualization module 250. The server device 160 includes, but is not limited to, an intent management application 170, one or more trained models 180, a design history 182, and one or more generated design objects 270. The multimodal cue 260 includes, but is not limited to, design intent text 262, one or more design files 264, and one or more design space references 266.
[0039] For illustrative purposes only, the functionality of the design exploration application 130 is described herein within the context of an exemplary interactive and linear workflow used to generate the generated design object 270 based on user-based design-related intent expressed during the workflow. The generated design object 270 includes, but is not limited to, one or more images, wireframe models, geometries and / or meshes for 3D design, and any amount (including none) and / or type of associated metadata.
[0040] As those skilled in the art will recognize, the techniques described herein are illustrative and not limiting, and can be modified and applied in other contexts without departing from the broader spirit and scope of the inventive concepts described herein. For example, during the entire process of generating and evaluating the design of the target 3D object, the techniques described herein can be modified and applied to generate any number of generated design objects 270 associated with any target content item in a linear, nonlinear, iterative, non-iterative, recursive, non-recursive manner, or any combination thereof. The target 3D object may include any number (including one) and / or type of target content items and / or target content item components.
[0041] For example, in some implementations, a generated design object 270 may be generated and displayed within GUI 120 during the first iteration, any portion (including all) of the design object 270 may be selected via GUI 120, and a first prompt multimodal prompt 260 may be set to equal the selected portion of the generated design object 270 to recursively generate a second generated design object 270 during the second iteration. In the same or other implementations, when generating each newly generated design object 270, the design exploration application 130 may display and / or re-display any number of GUI elements, generate and / or regenerate any number of data, or any combination thereof, any number and / or in any order.
[0042] In operation, the visualization module 250 of the design exploration application 130 provides a prompting space 220 and a design space 230 via a GUI 120. The user provides content for a multimodal prompt 260 via the prompting space 220. The design exploration application 130 processes this content to generate the multimodal prompt 260 and transmits it to the server device 160. The intent management application 170 identifies the modalities included in the data in the multimodal prompt 260 and identifies one or more trained ML models 180 and / or remote ML models 190 that have been trained to process the identified modalities. The intent management application 170 inputs the multimodal prompt into one or more of the identified ML models 180, 190. The ML models 180, 190 respond to the multimodal prompt 260 by generating one or more design objects 270. The visualization module 250 receives one or more generated design objects 270 and displays them in the prompting space 220 and / or the design space 230.
[0043] In various implementation schemes, design space 230 is a virtual workspace that includes the overall design that forms the content items. For example One or more renderings of the design object (overall 3D design) For example (The geometry of design object 144 and / or generated design object 270). In some implementations, design space 230 includes multiple design alternatives for the overall design. For example, design space 230 may graphically organize multiple designs including different combinations of design objects 144, 270. In this case, the user interacts with GUI 120 to navigate between design alternatives, thereby quickly analyzing trade-offs between different design options, observing trends in design options, constraining design space 230, selecting specific design options, and so on.
[0044] The prompt space 220 is a panel or volume in which a user can generate prompts, such as a multimodal prompt 260 and / or one or more prompt volumes 222. In some embodiments, the prompt space 220 is a panel, such as a window separate from the design space. For example, the prompt space 220 may include a multi-party interface (not shown) for communicating with a trained ML model 180 and / or one or more other client devices 110. Alternatively, in some embodiments, the prompt space 220 is a volume that overlays at least a portion of the design space. In this case, the user can invoke input areas of the prompt volume 222 and / or the multimodal prompt 260 at various locations within the design space 230.
[0045] The prompt volume 222 is a form of prompt indicating that an operation will be performed within the boundaries of a volume. The prompt volume 222 is a volume within the design space defined by a corresponding prompt definition, which specifies how an object appears and / or behaves within the boundaries of the prompt volume 222. The prompt volume 222 imposes a "range of influence" within the defined boundaries. For example Based on the influence volume of the boundary, modifications to the associated cue definition result in changes to the design objects within the boundary. For example, the cue definition allows the user to specify design intent text and / or non-text input for objects that at least partially overlap with cue volume 222. Cue volume 222 includes a set of features, including spatial location ( For example (position and orientation), boundaries (defined via text or via user input within prompt space 220), and shape ( For example (Spheres, cuboids, pyramids, irregular 3D shapes, etc.). In some implementations, the cue volume 222 includes a weighted area, a weighted gradient, and / or a linked cue volume (...). For example , hint volumes 222(1) to 222(x)). In this case, the linked hint volumes include other overlapping hint volumes and / or other hint volumes linked in the hierarchy.
[0046] In various implementations, when a user modifies the cue volume 222, the cue volume performs this by updating one or more design objects 144, 270 within the influence of the cue volume 222. For example, upon detecting a change to the cue definition, the cue volume 222 may receive the newly generated design object 270 and replace the existing design object 144 within the cue volume 222. Additionally or alternatively, in some implementations, the cue volume 222 applies weighted values corresponding to the weighted area and / or weighted gradient of the cue volume. When performing an update, the cue volume 222 may cause the design exploration application 130 to generate a message indicating the change and “transmit” the message to other linked cue volumes 222. In this case, the cue volume 222 propagates the change between linked cue volumes 222, allowing the user to modify multiple volumes within the design space 230 without applying a global change to the overall design space 230.
[0047] In various implementations, intent manager 240 determines the intent of input provided by the user. For example, intent manager 240 may include a natural language (NL) processor that parses text provided by the user. Additionally or alternatively, intent manager 240 may include an audio processor that processes audio data to identify words included in the audio data and parses the identified words. In some implementations, intent manager 240 is included in intent management application 170 and / or trained ML model 180. In this case, intent management application 170 and / or trained ML model 180 can determine the intent of input provided by the user via a multi-party interface.
[0048] In various implementations, intent manager 240 identifies one or more keywords in text data. In some implementations, intent manager 240 includes one or more keyword datasets 242, which intent manager 240 references when identifying one or more keywords included in text data. For example, keyword dataset 242 may include, but is not limited to, a 3D keyword dataset including any number and / or type of 3D keywords, a custom keyword dataset including any number and / or type of custom keywords, and / or user keywords including any number and / or type of user keywords. For example A user-specified dataset of words and / or phrases. Keywords may include specific words or phrases related to the design of 3D objects. For example(Indicative pronouns, technical terms, reference terms, etc.). For example, a user can enter a regular sentence ("I want a hinge to connect here") in the input area within cue space 220. The intent manager identifies "hinge," "connect," and "here" as words related to the ML models 180, 190 that generate design object 270. In this case, intent manager 240 can update cue space 220 by highlighting keywords, allowing the user to provide additional details (indicative pronouns, technical terms, reference terms, etc.). For example (Non-text data) to be included in multimodal prompts 260.
[0049] In various implementations, visualization module 250 displays design space 230 and / or prompt space 220 via GUI 120. In some implementations, visualization module 250 updates prompt space 220 and / or design space 230 based on user input and / or data received from server device 160. For example, visualization module 250 may initially respond to a user's call to a prompt via a hotkey or tag menu within prompt space 220 to receive data to be included in multimodal prompt 260 by displaying an input area. When the user initially enters a text phrase, visualization module 250 may respond to intent manager 240 by updating the input area to highlight keywords and / or displaying a contextual input area adjacent to at least one keyword to identify one or more keywords. In this way, design exploration application 130 iteratively receives multimodal input data to be included in multimodal prompt 260.
[0050] In various implementations, the design explorer application 130 receives text and / or non-text data via an input area included in the prompt space 220 for inclusion in the multimodal prompt 260. When non-text data is provided, the user can retrieve stored data from local data storage 140, such as one or more stored data files 142. For example (This includes stored geometry, stored CAD files, audio recordings, stored sketches, etc.). Additionally or alternatively, the user can retrieve content from the design history 182 and add content to the input area. In this case, the content from the design history 182 is stored in one or more data files 142, which the user retrieves from the local data storage 140.
[0051] Multimodal cue 260 is a data modality that includes two or more data modalities that specify the user's design intent. For example The design exploration application 130 receives multiple types of data and constructs a multimodal cue 260 to include each of the multiple data types. For example, the user may initially write design intent text 262 involving a sketch. Then, the design exploration application 130 receives the sketch (text data, image data, audio data, etc.). For example (Stored sketches or sketches entered by the user into the input design area). Upon receiving a sketch, the design exploration application 130 can then generate a multimodal cue 260 to include both design intent text 262 and the sketch. In some embodiments, the multimodal cue 260 may include multiple data inputs of the same modality. For example, the multimodal cue 260 may include multiple design intent texts 262 (…). For example , 262(1), 262(2), etc.) and / or multiple design files 264 (e.g., 264(1), 264(2), etc.).
[0052] Additionally or alternatively, in some implementations, the intent management application 170 may generate a multimodal cue 260. For example, a first user may initially write design intent text 262 into a multi-party interface. A second user may then provide a sketch to the multi-party interface. Upon receiving the sketch, the intent management application 170 may then generate a multimodal cue 260 to include both the design intent text 262 and the sketch. As will be discussed in further detail below, the intent management application 170 may also weight each component of the multimodal cue 260. For example, the intent management application 170 may weight each component of the multimodal cue 260 according to the user (…). For example Each input provided by the first user is weighted by 0.8, and each input provided by the second user is weighted by 0.2. Alternatively, the intent management application 170 may individually weight each component of the multimodal cue 260. For example (Weigh 0.4 for the design intent text 262 and 0.6 for the sketch).
[0053] Design intent text 262 includes textual data describing the user's intent. For example, design intent text may include descriptions of the target 3D design object ( For example The description of the characteristics of the handle (made of titanium). In some implementations, the design exploration application 130 generates design intent text 262 from different types of data input. For example, the intent manager 240 can perform NL processing to identify words included in an audio recording. In this case, the design exploration application 130 generates design intent text 262 including the identified words.
[0054] Design file 264 includes one or more files that the user adds to be included in multimodal prompt 260. For example (CAD files, stored text, audio recordings, stored geometry, etc.). In some implementations, design file 264 may include text data ( For example(Text description, physical dimensions, etc.). In various implementations, the user can add multiple design files 264 to be included in the multimodal prompt 260. In some implementations, the design exploration application 130 converts various types of data into design files 264. For example, the user can record audio via an input area. In this case, the design exploration application 130 can store the audio recording as a design file 264. The design file 264 may include one or more modalities (text description, physical dimensions, etc.). For example (Text data, video data, audio data, image data, etc.)
[0055] In some implementations, design space reference 266 may include one or more references to prompt space 220 and / or design space 230. For example, a user may enter text referencing a specific application state ( For example Examples include prompts like "Make the currently selected item brighter" or "Generate seats for the car in this view." In this case, the design exploration application 130 determines the application state that the user is referencing. The design exploration application 130 can then include that reference as a design space reference 266 in the multimodal prompt 260.
[0056] In various implementations, the intent management application 170 receives and processes the multimodal cue 260 to identify the modality of the content of the multimodal cue 260. For example, the intent management application 170 identifies the modality of design intent text 262, one or more design files 264, and / or one or more design space references 266 included in the multimodal cue 260. For example, the intent management application 170 may identify a combination of text, image, and video modalities included in the multimodal cue. The intent management application 170 identifies at least one ML model 180, 190 trained with the combination of modalities and selects one of the identified ML models 180, 190. The intent management application 170 executes the selected ML model by inputting the multimodal cue 260 into the selected ML model. The selected ML model generates a design object 270 in response to the multimodal cue 260. In some implementations, the server device includes the generated design object 270 in a design history 182. In this case, the generated design object 270 is part of the design history 182 and is used as training data to train one or more trained ML models 180. For example (Further training the selected ML model, training other ML models, etc.).
[0057] Figure 3 is an exemplary illustration of a multi-user system 300 including the intent management application 170 of Figure 2, according to various embodiments. As shown, in some embodiments, the multi-user system 300 includes, but is not limited to, multiple client devices 110 ( For example110(1), ..., 110(X)), server device 160 and one or more remote ML models 190 ( For example , 190(1), ..., 190(Z)). Multiple client devices 110 include, but are not limited to, design exploration applications 130 ( For example Examples of 130(1), ..., 130(X)) are explored in this design, each including a cue space of 220 ( For example , 220(1), ..., 220(X)). Server device 160 includes, but is not limited to, intent management application 170 and multiple trained ML models 180 ( example like , 180(1), ..., 180(Y)). The intent management application 170 includes, but is not limited to, a multi-party interface 310. In some other embodiments, system 100 may include any number and / or type of other client devices, server devices, remote ML models, or any combination thereof.
[0058] In various implementations, the intent management application 170 manages the multi-party interface 310. The multi-party interface 310 is a communication channel that includes one or more client devices 110, one or more trained ML models 180, and / or one or more remote ML models 190 as participants. Client devices 110(1)-110(X) transmit input to the multi-party interface 310 via corresponding cue spaces 220. The intent management application 170 processes the input received from client devices 110(1)-110(X) to identify cue for at least one ML model 180, 190. The intent management application 170 transmits the identified cue to the applicable ML model 180, 190, wherein the applicable ML model 180, 190 responds to the identified cue by generating digital content items as output.
[0059] In various implementations, the multi-party interface 310 is a GUI that displays input provided by participants in a shared communication channel. For example, the multi-party interface 310 may display text, images, video, audio, and other input provided by client devices 110(1)-110(X) and / or ML models 180(1)-180(Y), 190(1)-190(Z). The input can be directed to other users ( For example "Henry, when is the meeting?" or directed to at least one of the ML models 180 and 190 that generate digital content. For example"I need to design a logo for this sailboat." In this case, the intent management application 170 can process each input and transmit one or more inputs to at least one of the ML models 180 and 190 to generate an output.
[0060] In various implementations, intent management application 170 determines the likelihood that an input is directed to at least one ML model 180, 190. For example, intent management application 170 may include intent manager 240 (not shown) which determines the likelihood that a particular input transmitted from client device 110(X) to multi-party interface 310 is directed to ML model 180(1). In this case, intent manager 240 may generate a confidence score for a first input and compare the confidence score to a predetermined threshold. When intent manager 240 determines that the confidence score exceeds the predetermined threshold, intent management application 170 determines that the input includes a prompt for ML model 180(1) and transmits the prompt to ML model 180(1).
[0061] Additionally or alternatively, in various embodiments, the intent management application 170 determines the likelihood that an output generated by a given ML model 180, 190 is responsive to one or more inputs provided by client devices 110(1)-110(X). In this case, the intent management application 170 may generate a confidence score for the output, wherein the confidence score for the output indicates whether the output is responsive to the input. In some embodiments, the intent management application 170 collects multiple inputs and transmits the multiple inputs to multiple ML models 180, 190. In this case, the intent management application 170 may calculate a confidence score for the output provided by each of the ML models 180, 190.
[0062] In various implementations, the intent management application 170 combines and / or aggregates inputs from two or more client devices 110(1)-110(X) to generate a composite prompt for ML models 180, 190. For example, the intent management application 170 may combine text input transmitted from client device 110(1) and text input transmitted from client device 110(X0) to generate a composite prompt. In some implementations, the composite prompt may include a multimodal prompt that includes inputs from two or more modalities. These inputs may be from the same client device 110(1)-110(X0). For exampleThe input can be received from multiple client devices 110 (1) or from 110. For example, a first user may provide input including a text description, and a second user may provide an image. In this case, the intent management application 170 may generate a multimodal prompt that includes both inputs.
[0063] In some implementations, ML models 180 and 190 can generate outputs that can be transmitted to different ML models 180 and 190. For example, multiple users can provide outputs for a first ML model 180 and 190 ( For example The input of the remote ML model 190(Z) is used to generate prompts for the second ML model ( For example The trained ML model 180(Y) is received as input. In another example, multiple users can provide input to a second ML model 180(Y) to generate digital content items. For example The generated design object 270 is output, and then provided for the third ML model ( For example The input of 190(1) is used to evaluate the generated design object 270 generated by the second ML model 180(Y).
[0064] In various implementations, the intent management application 170 displays inputs provided by client devices 110(1)-110(X), ML models 180(1)-180(Y), and / or remote ML models 190(1)-190(Z) in a multi-party interface. Additionally or alternatively, the inputs provided to the multi-party interface 310 are selectable and can be passed to the design exploration application 130. For example, a first user can select a generated design object 270 and move the generated design object into the design space 230.
[0065] Figure 4 is an exemplary illustration of the multi-party interface 310 of Figure 3 according to various embodiments. As shown, the multi-party interface 400 includes, but is not limited to, a first user input 402 including a first text prompt 422, a second user input 404 including a second text prompt 442, and an AI response 406 including a text response 462 and a generated digital content item 464.
[0066] In various implementations, the intent management application 170 generates a multi-party interface 310 that communicates with multiple client devices 110 and one or more ML models 180, 190. As shown, the multi-party interface 400 displays information generated by user participants (…). For example (402, 404) and non-user participants ( For exampleThe GUI provides input (AI response 406). In some implementations, the intent management application 170 processes each of the inputs 402, 404 provided by the user participant to determine whether a portion of the inputs 402, 404 is directed to a non-user participant. For example, the intent management application 170 may include an intent manager 240 that determines the likelihood that user inputs 402, 404 are directed to a non-user participant. In this case, the intent manager 240 may generate a confidence score for the corresponding inputs 402, 404 and compare the corresponding confidence scores to a predetermined threshold. When the intent manager 240 determines that the confidence score exceeds the predetermined threshold, the intent manager 240 determines that the inputs 402, 404 include inputs for non-user participants (AI response 406). For example The prompt is given to ML model 180(1), and the prompt is transmitted to ML model 180(1). Alternatively, in some implementations, a non-user participant determines whether to respond to inputs 402 and 404.
[0067] In some implementations, the intent management application 170 may extract cue 422 included in input 402 and / or cue 442 included in input 404. In this case, the intent management application 170 may combine and / or aggregate cue 422, 442 to generate a composite cue. The intent management application 170 may then transmit the composite cue to a trained ML model 180(1) to generate a response. For example, the intent management application 170 may extract cue 422 (the type of requested text output) from input 402 and cue 442 (the topic of requested text output) from input 404. In this case, the intent management application 170 may generate a single composite cue containing both cue 422 and 442, which specify different requirements for output to the ML model 180(1). Alternatively, in some implementations, the intent management application 170 may transmit each cue 422, 442 separately. In this case, the ML model 180(1) may be trained to receive multiple inputs and generate output in response to each input.
[0068] The trained ML model 180(1) responds to the received compound prompts by generating a digital content item 464 as output. The generated digital content item 464 has properties that respond to the ideas and constraints specified by the combination of prompts 422, 442. While generating the output, the trained ML model 180(1), acting as a non-user participant, provides an AI response 406 to the multi-party interface 400. The AI response 406 includes text input 462. The text input 462 indicates that the non-user participant is responding to prompts 422, 442. The AI response 406 also includes the digital content item 464. Once the digital content item 464 is published to the multi-party interface 400, each participant can select the digital content item 464 for local use. For example, a first user can select the digital content item 464 and store it as a local copy in a local data storage 140(1). The local copy can then be added to the design space 230 included in a design exploration application 130(1) executed locally on the client device 110(1).
[0069] Figure 5 is an exemplary illustration of another multi-party interface 310 of Figure 3 according to various other embodiments. As shown, the multi-party interface 500 includes, but is not limited to, a first user input 502 including a first text prompt 522, a second user input 504 including a second text prompt 542, a first AI response 506 including a text response 562 and a generative prompt 564, and a second AI response 508 including a text response 582 and a digital content item 584.
[0070] Multi-party interface 500 is similar to multi-party interface 500. As shown in the figure, the multi-party interface includes two inputs 502 and 504 from the user participant and input from a non-user participant (…). For example Two inputs 506 and 508 to the remote ML model 190 (1) and the trained ML model 180 (Y). In various implementations, the intent management application 170 and / or each non-user participant can determine whether either of the inputs 502 and 504 provided by the user participant contains a prompt for the ML models 180 and 190 to generate an output. For example, the intent management application 170 can execute the intent manager 240 to parse the text of input 502 to determine whether the user is referring to a target for the output. For example (Image). Intent Manager 240 can also identify other keywords for modifying the target ("I don't know the words to describe it"), and identify the different types of output that ML models 180 and 190 will produce ( For example(a cue for generating an image). In this case, the intent management application 170 may transmit the cue 522 included in the input 502 to the remote ML model 190 (1) to generate a text cue, instead of transmitting the cue 522 to the trained ML model 180 (Y) to generate an image.
[0071] Alternatively, in some implementations, a remote ML model 190(1) and / or a trained ML model 180(Y) may process input 502 to determine whether to respond to input 502 by generating output. In this case, each of the remote ML model 190(1) and / or the trained ML model 180(Y) may generate a confidence score for input 502. Each of the remote ML model 190(1) and / or the trained ML model 180(Y) may then compare the confidence score to a predetermined threshold. If the confidence score exceeds the predetermined threshold, the remote ML model 190(1) and / or the trained ML model 180(Y) may do so. For example, the remote ML model 190(1) may calculate a first confidence score for input 502 that exceeds a first predetermined threshold. In this case, the remote ML model 190(1) may detect a cue 522 included in input 502 and respond by generating text output ( For example Generative prompt 564) responds to prompt 522. The trained ML model 180(Y) can compute a second confidence score for input 502 that does not exceed a second predetermined threshold. In this case, the trained ML model 180(Y) does not respond to input 502.
[0072] In various implementations, the ML model 190(1) responds to multiple inputs. For example, the intent management application 170 may determine that each input 502, 504 includes a cue for the remote ML model 190(1). The intent management application 170 may then generate a composite cue containing cue 522, 542 extracted from the inputs 502, 504. The intent management application 170 may then transmit the composite cue to the ML model 190(1), whereby the remote ML model 190(1) receives the composite cue as input. Alternatively, in some implementations, the intent management application 170 transmits each cue 522, 542 as a separate input. In this case, the remote ML model 190(1) responds to the multiple cue 522, 542 by generating an output.
[0073] In various implementations, ML model 190(1) can generate outputs that are transmitted to different ML models 180, 190. For example, multiple users can provide inputs 502, 504 to a remote ML model 190(1) to generate prompts for a trained ML model 180(Y) to receive as input. The remote ML model 190(1) can then produce a generative prompt 564 as output. In some implementations, the remote ML model 190(1) transmits the generative prompt 564 to a multi-party interface. For example, the remote ML model 190(1) issues an AI response 506, which includes text input 562 responding to inputs 502, 504 and a generative prompt 564. In this case, intent management application 170 can transmit the generative prompt to a trained ML model 180(Y) so that the trained ML model 180(Y) generates digital content items in response to the generative prompt 564. For example (Image). Alternatively, in some implementations, a remote ML model 190(1) generates output and transmits that output directly to a trained ML model 180(Y). In this case, the remote ML model 190(1) can provide an AI response 506 indicating that the transmission of output has occurred without the output being published to the multi-party interface 500.
[0074] In various implementations, a trained ML model 180(Y) generates digital content items in response to a generative prompt 564 generated by a remote ML model 190(1). In this case, the trained ML model 180(Y) responds to the generative prompt 564 by generating a digital content item 584 that conforms to the goals, constraints, and requirements included in the generative prompt 564. Upon generating the digital content item 584, the trained ML model 180(Y) can then transmit input 508, which includes a textual response 582 to the generative prompt 564 and the digital content item 584. In this case, the digital content item 584 can be copied from the multi-party interface 500 for local use by each participant of the multi-party interface 500.
[0075] Figure 6A is an exemplary illustration of a weighted interface 600 implemented by the intent management application 170 of Figure 3 according to various embodiments. As shown, the weighted interface 600 includes, but is not limited to, a first user 602 including a first prompt 622, a second user 604 including a second prompt 624, an ML model 608, a first weighted length 612, a second weighted length 614, and a digital content item 630.
[0076] In operation, the intent management application 170 can generate a weighted interface 600 to enable one or more user participants of the multi-party interface 310 to interact with the input ( For exampleThe input is weighted (as prompted by prompts 622 and 624). In some implementations, the weights can be specified based on the user. In this case, the intent management application 170 can apply user-specific weight values to each input provided by the user. Alternatively, weights can be specified for each individual input. In this case, the intent management application 170 can apply weight values for each individual input. The intent management application 170 can then transmit the weighted input or include the weighted input in the composite prompts used for the ML models 180 and 190.
[0077] In some implementations, the weighted interface 600 indicates relative weights based on the input to the weighted interface 600. For example, the weighted interface 600 may include a graphical interface that indicates the relative weighting of the input to the ML model 608 based on the inverse of the distance to the user participant. As shown, the weighting length 614 is shorter than the weighting length 612, thereby indicating the weight value relative to the weight used for the first cue 622. For example The first weight value is 0.4, which is larger than the weight value used for the second hint 624. For example The second weight value is 0.7). The ML model 608 responds to a weighted combination of prompts by generating digital content items 630. For example (For example, the combination of prompt 622 and the first weight value, and the combination of prompt 624 and the second weight value), the numeric content item is more responsive to prompt 624 than to prompt 622. For example (This generates a horse with some frog-like characteristics).
[0078] Figure 6B is an exemplary illustration of another weighted interface 650 implemented by the intent management application 170 of Figure 3 according to various other embodiments. As shown, the weighted interface 650 includes, but is not limited to, a first user 602 including a first prompt 622, a second user 604 including a second prompt 624, an ML model 608, a third weighted length 652, a fourth weighted length 654, and a second digital content item 670.
[0079] As shown in the figure, the relative weights indicated by weighted interface 650 differ from those indicated in weighted interface 600. Weighted interface 650 indicates the relative weights of the inputs to ML model 608 based on the reciprocal of the distance to the user participant. For example, the third weighting length 652 is shorter than the fourth weighting length 654, thus indicating the weight value relative to the weight used for the first cue 622. For example The fourth weight value (0.4) is larger than the weight value used for the second hint 624. For example The third weight is 0.8). The ML model 608 responds to a weighted combination of prompts by generating digital content items 670. For example(For example, the combination of prompt 622 and the first weight value, and the combination of prompt 624 and the second weight value), the numeric content item is more responsive to prompt 622 than to prompt 624. For example (This generates a frog with some horse-like features).
[0080] Figure 7 illustrates flowcharts of method steps for generating digital content items according to various implementation schemes. Although the method steps are described with reference to the systems in Figures 1 through 6B, those skilled in the art will understand that any system configured to implement the method steps in any order falls within the scope of the implementation schemes.
[0081] As shown in the figure, method 700 begins at step 702, where server device 160 receives one or more inputs from a first user. In various embodiments, intent management application 170 generates a multi-party interface 310 for a communication channel communicating with multiple user participants. In some embodiments, the multi-party interface 310 includes one or more ML models 180, 190 (…). For example The trained ML model 180(1) acts as a non-user participant. In this case, the intent management application 170 can receive various inputs from the participants. For example, the intent management application 170 can receive one or more first inputs from a first user via a first client device 110(1). In some embodiments, the intent management application 170 receives the first inputs continuously. Alternatively, in some embodiments, the intent management application 170 receives the first inputs scattered among inputs received from other participants. In various embodiments, the intent management application 170 displays the received first inputs in a graphical user interface representing the multi-party interface 310.
[0082] At step 704, server device 160 determines that one or more inputs from the first user include prompts for the ML model. In various embodiments, intent management application 170 parses the received first input to determine whether the first input includes at least one prompt for the ML model. For example, intent management application 170 may include intent manager 240, which parses text input to determine whether the text input includes prompts for the trained ML model 180(1). In some embodiments, intent manager 240 may generate a confidence score indicating that the first input includes prompts for the trained ML model 180(1). In this case, intent management application 170 determines that the first input includes prompts for the trained ML model 180(1) when the confidence score exceeds a predetermined threshold associated with the trained ML model 180(1).
[0083] At step 706, server device 160 receives one or more inputs from additional users. In various embodiments, intent management application 170 receives one or more second inputs from a second user via second client device 110 (2). In some embodiments, intent management application 170 receives second inputs continuously. Alternatively, in some embodiments, intent management application 170 receives second inputs scattered among inputs received from other participants. In various embodiments, intent management application 170 displays the received second inputs in a graphical user interface representing multi-party interface 310.
[0084] At step 708, server device 160 determines that one or more inputs from an additional user include a prompt for the ML model. In various embodiments, intent management application 170 parses the received second input to determine whether the second input includes at least one prompt for the ML model. For example, intent manager 240 may parse text input to determine whether the text input includes a prompt for the trained ML model 180(1). In some embodiments, intent manager 240 may generate a confidence score indicating that the second input includes a prompt for the trained ML model 180(1). In this case, intent management application 170 determines that the second input includes a prompt for the trained ML model 180(1) when the confidence score exceeds a predetermined threshold associated with the trained ML model 180(1).
[0085] At step 710, server device 160 determines whether input has been received from an additional user. In various embodiments, intent management application 170 determines whether input has been received from an additional user ( For example The intent management application 170 receives input from an additional user (or a third user). When the intent management application 170 determines that it will receive input from the additional user, it returns to step 706 to receive input from the additional user. Otherwise, the intent management application 170 determines that it will not receive input from the additional user, and proceeds to step 712.
[0086] At step 712, server device 160 determines whether to apply a weight to a prompt that is determined to have been included in the received input. In various embodiments, intent management application 170 determines whether to apply a weight value to a prompt identified in the first and second inputs. When intent management application 170 determines to apply a weight value, intent management application 170 proceeds to step 714. Otherwise, intent management application 170 determines not to apply a weight value. For example Generate a compound suggestion (where each suggestion in the suggestion is equally weighted), and proceed to step 716.
[0087] At step 714, server device 160 applies weights to the determined prompts. In various embodiments, intent management application 170 applies weight values to each prompt determined by intent management application 170 to be included in the first or second input. For example, intent management application 170 may receive weight values assigned to a specific user. For example (e.g., a first weight value assigned to a first user, a second weight value assigned to a second user, etc.). In this case, the intent management application 170 applies the weight value assigned to the user to each prompt generated by the user. Alternatively, in some embodiments, the intent management application 170 may receive a specific weight value for a particular prompt. For example, the intent management application 170 may receive the weight value via input to weighted interfaces 600, 650. In this case, the intent management application 170 applies the weight value assigned to the prompt.
[0088] At step 716, server device 160 transmits a prompt to trained ML model 180. In various embodiments, intent management application 170 transmits each of the determined prompts to trained ML model 180 (1). In some embodiments, intent management application 170 generates a composite prompt including each of the determined prompts and then transmits the composite prompt to trained ML model 180 (1). Alternatively, intent management application 170 may send each of the determined prompts individually to trained ML model 180 (1).
[0089] At step 718, server device 160 determines whether to select one or more additional ML models. In various embodiments, intent management application 170 determines whether to select one or more additional models to generate digital content items as output. For example, intent management application 170 may select multiple ML models 180, 190, each of which responds to the same compound cue. In another example, intent management application 170 may determine that trained ML model 180(1) generates an output to be used as input for one or more additional ML models 180, 190. For example (Generative prompts, first digital content item). When the intent management application 170 determines that at least one additional ML model 180, 190 is selected, the intent management application 170 proceeds to step 720. Otherwise, the intent management application 170 determines that additional ML models 180, 190 are not selected, and ends method 700.
[0090] At step 720, server device 160 selects a cue for one or more selected ML models. In various embodiments, intent management application 170 selects a cue for one or more selected ML models 180, 190. In some embodiments, trained ML model 180(1) generates generative cue, or user participants provide additional cue in response to digital content items generated by trained ML model 180(1). For example (Including tips for digital content items). In this case, the intent management application 170 can determine one or more additional ML models 180, 190 applicable to generative tips or additional tips. For example (This generates an ML model of a digital content item of a type specified in the generative prompt or additional prompt). Then, the intent management application 170 may optionally receive one or more additional ML models 180, 190 that are generative prompts or additional prompts.
[0091] At step 722, the server device transmits the selected prompt to one or more selected ML models. In various embodiments, the intent management application 170 transmits generative or supplementary prompts to the selected ML models 180, 190. The selected ML models 180, 190 generate digital content items in response to the generative or supplementary prompts. While generating the digital content items, the selected supplementary ML models 180, 190 transmit the digital content items to the multi-party interface 310. In this case, the digital content items can be copied from the multi-party interface 310 for local use by each participant of the multi-party interface 310.
[0092] Figure 8 illustrates an architecture of a system 800 in which embodiments of the present disclosure may be implemented. This figure is in no way limiting or intended to limit the scope of the present disclosure. In various implementations, system 800 may be an augmented reality, virtual reality, or mixed reality system or device, a personal computer, a video game console, a personal digital assistant, a mobile phone, a mobile device, or any other device suitable for practicing one or more embodiments of the present disclosure. Furthermore, in various embodiments, any combination of two or more systems 800 may be coupled together to practice one or more aspects of the present disclosure.
[0093] As shown in the figure, system 800 includes a central processing unit (CPU) 802 and a system memory 804 that communicate via a bus path, which may include a memory bridge 805. The CPU 802 includes one or more processing cores and, in operation, is the main processor of system 800 that controls and coordinates the operation of other system components. The system memory 804 stores software applications and data used by the CPU 802. The CPU 802 runs software applications and optionally an operating system. For example The Northbridge chip's memory bridge 805 is connected via a bus or other communication path ( For example The hyperlink is connected to the I / O (input / output) bridge 807. It can be... For example The southbridge chip's I / O bridge 807 connects to one or more user input devices 808. For example The system receives user input via a keyboard, mouse, joystick, digitizer, touchpad, touchscreen, still or video camera, motion sensor, and / or microphone, and forwards the input to the CPU 802 via a memory bridge 805.
[0094] The display processor 812 is connected via a bus or other communication path ( For example The display processor 812 is a graphics subsystem including at least one graphics processing unit (GPU) and graphics memory. The graphics memory includes display memory for storing pixel data for each pixel of the output image. (PCI Express, accelerated graphics port, or HyperTransport link) For example (Frame buffer). The graphics memory can be integrated into the same device as the GPU, connected to the GPU as a separate device, and / or implemented within system memory 804.
[0095] The display processor 812 periodically delivers pixels to the display device 810. For example (Screen or conventional CRT, plasma, OLED, SED, or LCD-based monitors or televisions). Additionally, the display processor 812 can output pixels to a film recorder suitable for reproducing computer-generated images on photographic film. The display processor 812 can provide analog or digital signals to the display device 810. In various embodiments, one or more of the various graphical user interfaces described in Appendix AJ appended herein are displayed to one or more users via the display device 810, and one or more users can input data to and receive visual output from those various graphical user interfaces.
[0096] System disk 814 is also connected to I / O bridge 807 and can be configured to store content, applications, and data for use by CPU 802 and display processor 812. System disk 814 provides non-volatile storage for applications and data and may include fixed or removable hard disk drives, flash memory devices, and CD-ROMs, DVD-ROMs, Blu-ray, HD-DVDs, or other magnetic storage devices, optical storage devices, or solid-state storage devices.
[0097] The switch 816 provides connectivity between the I / O bridge 807 and other components such as the network adapter 818 and various plug-in cards 820 and 821. The network adapter 818 allows the system 800 to communicate with other systems via electronic communication networks and may include wired or wireless communication via local area networks and wide area networks such as the Internet.
[0098] Other components (not shown), including USB or other port connections, film recording devices, etc., may also be connected to I / O bridge 807. For example, an audio processor may be used to generate analog or digital audio outputs based on instructions and / or data provided by CPU 802, system memory 804, or system disk 814. The communication paths interconnecting the various components in FIG1 may be implemented using any suitable protocol such as PCI (Peripheral Component Interconnect), PCI Express (PCI-E), AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol, and connections between different devices may use different protocols, as known in the art.
[0099] In one embodiment, the display processor 812 incorporates circuitry optimized for graphics and video processing, including, for example, video output circuitry, and constitutes a graphics processing unit (GPU). In another embodiment, the display processor 812 incorporates circuitry optimized for general-purpose processing. In yet another embodiment, the display processor 812 may be integrated with one or more other system elements, such as a memory bridge 805, a CPU 802, and an I / O bridge 807, to form a system-on-a-chip (SoC). In a further embodiment, the display processor 812 is omitted, and the functions of the display processor 812 are performed by software executed by the CPU 802.
[0100] Pixel data can be provided directly from CPU 802 to display processor 812. In some embodiments of this disclosure, instructions and / or data representing a scene are provided to a rendering field or a set of server computers, each similar to system 800, via network adapter 818 or system disk 814. The rendering field uses the provided instructions and / or data to generate one or more rendered images of the scene. These rendered images can be stored in a digital format on a computer-readable medium and optionally returned to system 800 for display. Similarly, stereoscopic image pairs processed by display processor 812 can be output to other systems for display, stored on system disk 814, or stored in a digital format on a computer-readable medium.
[0101] Alternatively, CPU 802 provides display processor 812 with data and / or instructions defining the desired output image, and display processor 812 generates pixel data for one or more output images based on the data and / or instructions, including characterizing and / or adjusting the offset between stereoscopic image pairs. The data and / or instructions defining the desired output image may be stored in system memory 804 or graphics memory within display processor 812. In embodiments, display processor 812 includes 3D rendering capabilities for generating pixel data for the output image based on instructions and data defining the geometry, lighting and shadows, texturing, motion, and / or camera parameters of a scene. Display processor 812 may further include one or more programmable execution units capable of executing shader programs, tone mapping programs, etc.
[0102] Furthermore, in other embodiments, the CPU 802 or display processor 812 may be replaced or supplemented by any technically feasible form of processing means configured to process data and execute program code. Such processing means may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc. In various embodiments, any of the operations and / or functions described herein may be performed by the CPU 802, the display processor 812, or one or more other processing means, or any combination of these different processors.
[0103] CPU 802, rendering field and / or display processor 812 may employ any surface or volume rendering technique known in the art to create one or more rendered images based on provided data and instructions, including rasterization, scanline rendering, REYES or micropolygon rendering, ray casting, ray tracing, image-based rendering techniques and / or combinations of these techniques and any other rendering or image processing techniques known in the art.
[0104] In other envisioned embodiments, system 800 may be a robot or robotic device and may include CPU 802 and / or other processing units or devices, as well as system memory 804. In such embodiments, system 800 may or may not include the other elements shown in FIG1. System memory 804 and / or other memory units or devices in system 800 may include instructions that, when executed, cause the robot or robotic device represented by system 800 to perform one or more operations, steps, tasks, etc.
[0105] It should be understood that the systems illustrated herein are illustrative, and variations and modifications are possible. The connection topology, including the number and arrangement of bridges, can be modified as needed. For example, in some embodiments, system memory 804 is connected directly to CPU 802 instead of via a bridge, and other devices communicate with system memory 804 via memory bridge 805 and CPU 802. In other alternative topologies, processor 812 is shown connected to I / O bridge 807 or directly to CPU 802, instead of to memory bridge 805. In yet another embodiment, I / O bridge 807 and memory bridge 805 may be integrated into a single chip. Certain components illustrated herein are optional; for example, any number of expansion cards or peripheral devices may be supported. In some embodiments, switch 816 is eliminated, and network adapter 818 and expansion cards 1120, 1121 are directly connected to I / O bridge 807.
[0106] In summary, the disclosed techniques can be used to generate digital content items, including 3D object designs, based on design intents expressed by inputs provided by two or more users via a GUI. In various embodiments, an intent management application generates a multi-party interface that communicates with at least one or more AI models and two or more users. Users communicate with the multi-party interface via instances of design exploration applications executed on different client devices. Each design exploration application displays a cue space for the user to view a representation of the multi-party interface. In some embodiments, the cue space overlaps with the design space, where the user can invoke the multi-party interface from anywhere within the design space. Alternatively, in some embodiments, the cue space and design space are separate. The multi-party interface is a graphical interface that can be used to perform one or more iterations of: entering input and submitting cuees to an AI model, and receiving responses from the AI model. Multiple users enter input and submit cuees to an AI model via the multi-party interface.
[0107] In various implementations, the intent management application analyzes input from one or more users onto a multi-party interface. The intent management application determines whether the input is a prompt to be submitted to an AI model. For example, the intent management application parses and analyzes the content of text input to determine if the input includes a text prompt for an AI model, and if so, determines that the text prompt is for a specific type of AI model. The intent management application collects and aggregates multiple prompts generated by multiple users. In some implementations, the intent management application applies weights to individual prompts, such as by applying user weight values already assigned to individual users. Alternatively, the intent management application may apply weight values to specific prompts. The intent management application generates composite prompts based on multiple prompts and weight values.
[0108] The intent management application identifies one or more AI models trained to process compound cues. The intent management application inputs the compound cues into the identified AI model. The AI model may reside locally or remotely on a server device. The AI model, trained using the history of cues, generated digital content items, and evaluations of the generated digital content items, generates digital content items in response to the compound cues. In some implementations, the AI model generates digital content items that can be used in a design project. Alternatively, in some implementations, a generative AI model generates multiple digital content items. Each of the generated digital content items conforms to the characteristics specified by the compound cues. The intent management application displays one or more digital content items via a GUI in a multi-party interface, where the digital content items can be selected to move within the design space.
[0109] In some implementations, the intent management application determines that at least one additional AI model will generate a digital content item based on a composite prompt. In one example, the composite prompt is directed to multiple AI models of similar type to generate alternative digital content items. In this case, the intent management application feeds the composite prompt into each of the multiple AI models. In some implementations, the composite prompt causes the AI model to create a generative prompt that is directed to an additional AI model. In this case, the intent management application receives the generative prompt created by the AI model and feeds the generative prompt into the additional AI model. The additional AI model generates the digital content item based on the generative prompt.
[0110] At least one technical advantage of the disclosed technology over existing technologies lies in its ability to enable CAD applications to collect and aggregate input from multiple different users within a user group when generating prompts. This allows AI systems to more accurately understand the collective thoughts and intentions of the user group and generate digital content that more accurately reflects these collective thoughts and intentions. In this regard, the disclosed technology provides an automated process for collecting multiple prompts generated from multiple users within a shared interface and weighting these prompts before transmitting them to an AI model for execution. Collecting multiple prompts as input to the AI model and weighting these prompts allows group members to clarify the group's collective thoughts and intentions by emphasizing specific ideas and goals. Therefore, the disclosed technology enables AI models to better infer the thoughts and intentions of the user group and generate digital content that more accurately reflects these thoughts and intentions. Thus, the disclosed technology allows a group of users to use a CAD application to generate digital content that better matches the group's actual thoughts and intentions without requiring coordination among group members using separate communication channels. These technical advantages provide one or more technological advancements superior to existing methods.
[0111] 1. In various embodiments, a computer-implemented method for generating digital content, the method comprising: generating a multi-party interface communicating with at least a trained machine learning (ML) model, a first client device, and a second client device; combining at least a first input from the first client device and a second input from the second client device to generate a composite prompt; transmitting the composite prompt to the trained ML model for execution; receiving from the trained ML model a digital content item generated in response to the composite prompt; and displaying the digital content item in the multi-party interface.
[0112] 2. The computer-implemented method as described in Clause 1, further comprising determining that the first input includes at least one cue for the trained ML model.
[0113] 3. A computer-implemented method as described in Clause 1 or 2, wherein determining that the first input includes at least one cue for the trained ML model comprises: generating an input confidence score associated with the first input; and determining that the input confidence score exceeds a predetermined threshold.
[0114] 4. The computer-implemented method of any one of Clauses 1 to 3 further includes generating an output confidence score indicating whether the digital content item responds to at least one of the first input or the second input.
[0115] 5. A computer-implemented method as described in any one of Clauses 1 to 4, wherein combining at least the first input from the first client device and the second input from the second client device comprises: applying a first weight value to the first input; and applying a second weight value to the second input.
[0116] 6. A computer-implemented method as described in any one of Clauses 1 to 5, wherein the first weight value is assigned to a first user of the first client device, and the second weight value is assigned to a second user of the second client device.
[0117] 7. The computer-implemented method of any one of Clauses 1 to 6, further comprising: receiving the first weight value for the first input via a graphical user interface (GUI); and receiving the second weight value for the second input via the GUI.
[0118] 8. A computer-implemented method as described in any one of Clauses 1 to 7, wherein the digital content item includes a generative prompt, and the method further includes: performing a second trained ML model on the generative prompt to generate a second digital content item; and displaying the second digital content item in the multi-party interface.
[0119] 9. The computer-implemented method of any one of Clauses 1 to 8, wherein the composite prompt comprises at least intent text and non-text input, and wherein the non-text input comprises at least one of: computer-aided design (CAD) objects, geometry, images, sketches, videos, application status, or audio recordings.
[0120] 10. A computer-implemented method as described in any one of Clauses 1 to 9, wherein the intent text is included in the first input, and the non-text input is included in one or more non-text inputs constituting the second input.
[0121] 11. In various embodiments, one or more non-transitory computer-readable media include instructions that, when executed by one or more processors, cause the one or more processors to generate design content by performing the following steps: generating a multi-party interface communicating with at least a trained machine learning (ML) model, a first client device, and a second client device; combining at least a first input from the first client device and a second input from the second client device to generate a composite prompt; transmitting the composite prompt to the trained ML model for execution; receiving a digital content item generated in response to the composite prompt from the trained ML model; and displaying the digital content item in the multi-party interface.
[0122] 12. One or more non-transitory computer-readable media as described in Clause 11, wherein the server device generates the multi-party interface, and the non-transitory computer-readable medium further includes instructions that, when executed by the one or more processors, cause the one or more processors to perform the following steps: transmitting the composite prompt by the server device to a remote device executing the trained ML model.
[0123] 13. One or more non-transitory computer-readable media as described in Clause 11 or 12, wherein the digital content item includes one of the following: text, computer-aided design (CAD) objects, geometry, images, sketches, video, executable code, or audio recordings.
[0124] 14. One or more non-transitory computer-readable media as described in any one of clauses 11 to 13, wherein combining at least the first input from the first client device and the second input from the second client device comprises: applying a first weight value to the first input; and applying a second weight value to the second input.
[0125] 15. One or more non-transitory computer-readable media as described in any one of Clauses 11 to 14, further comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform the steps of: receiving a first weight value for the first input via a graphical user interface (GUI); and receiving a second weight value for the second input via the GUI.
[0126] 16. One or more non-transitory computer-readable media as described in Clauses 11 to 15, wherein the digital content item includes a generative prompt, and the non-transitory computer-readable medium further includes instructions that, when executed by the one or more processors, cause the one or more processors to perform the following steps: execute a second trained ML model on the generative prompt to generate a second digital content item; and display the second digital content item in the multi-party interface.
[0127] 17. One or more non-transitory computer-readable media as described in Clauses 11 to 16, wherein the composite prompt comprises at least intent text and non-text input, and wherein the non-text input comprises at least one of: computer-aided design (CAD) objects, geometry, images, sketches, videos, application status, or audio recordings.
[0128] 18. One or more non-transitory computer-readable media as described in Items 11 to 17, wherein the trained ML model is trained using at least a combination of a first modality associated with text and at least one other modality associated with non-text input.
[0129] 19. One or more non-transitory computer-readable media as described in Clauses 11 to 18, further comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform the following steps: execute a second trained ML model on the composite prompt to generate a second digital content item; and display the second digital content item in the multi-party interface.
[0130] 20. In various embodiments, a system includes: one or more memories storing instructions; and one or more processors coupled to the one or more memories, the one or more processors performing the following steps when executing the instructions: generating a multi-party interface communicating with at least a trained machine learning (ML) model, a first client device, and a second client device; combining at least a first input from the first client device and a second input from the second client device to generate a composite prompt; transmitting the composite prompt to the trained ML model for execution; receiving a digital content item generated in response to the composite prompt from the trained ML model; and displaying the digital content item in the multi-party interface.
[0131] Any and all combinations of any element of any claim and / or any element described in this application, in any manner, fall within the scope of this disclosure and protection.
[0132] Various embodiments have been described for illustrative purposes; however, these descriptions are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.
[0133] Various aspects of embodiments of this invention may be embodied as systems, methods, or computer program products. Therefore, aspects of this disclosure may take the form of entirely hardware implementations, entirely software implementations (including firmware, resident software, microcode, etc.), or implementations combining software and hardware aspects, all of which are generally referred to herein as “modules,” “systems,” or “computers.” Furthermore, any hardware and / or software technology, process, function, component, engine, module, or system described in this disclosure may be implemented as a circuit or a set of circuits. Additionally, aspects of this disclosure may take the form of computer program products embodied in one or more computer-readable media having computer-readable program code embodied thereon.
[0134] Any combination of one or more non-transitory computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination thereof. More specific examples (not an exhaustive list) of computer-readable storage media will include: electrical connections having one or more wires, portable computer floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the context of this document, a computer-readable storage medium can be any tangible medium that can contain or store programs for use with or in connection with an instruction execution system, device, or apparatus.
[0135] The foregoing description of aspects of this disclosure is based on flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks of the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine. When executed via a processor of a computer or other programmable data processing apparatus, the instructions implement the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such processors can be, but are not limited to, general-purpose processors, special-purpose processors, application-specific processors, or field-programmable gate arrays.
[0136] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each box in a flowchart or block diagram may represent a module, segment, or portion of code comprising one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions mentioned in the boxes may not appear in the order shown in the drawings. For example, two boxes shown consecutively may actually be executed substantially simultaneously, or these boxes may sometimes be executed in reverse order, depending on the functionality involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented by a system based on dedicated hardware or a combination of dedicated hardware and computer instructions that performs the specified function or action.
[0137] While the foregoing describes embodiments of this disclosure, other and additional embodiments of this disclosure may be designed without departing from the basic scope of this disclosure, the scope of which is defined by the appended claims.
Claims
1. A computer-implemented method for generating digital content, the method comprising: Generate a multi-party interface that communicates with at least a trained machine learning (ML) model, a first client device, and a second client device; Combine at least a first input from the first client device and a second input from the second client device to generate a composite prompt; The composite prompt is transmitted to the trained ML model for execution; Receive digital content items generated in response to the composite prompts from the trained ML model; And display the digital content items in the multi-party interface.
2. The computer-implemented method of claim 1, further comprising determining that the first input includes at least one cue for the trained ML model.
3. The computer-implemented method of claim 2, wherein determining that the first input includes at least one cue for the trained ML model comprises: Generate an input confidence score associated with the first input; And determine that the input confidence score exceeds a predetermined threshold.
4. The computer-implemented method of claim 1, further comprising generating an output confidence score indicating whether the digital content item responds to at least one of the first input or the second input.
5. The computer-implemented method of claim 1, wherein combining at least the first input from the first client device and the second input from the second client device comprises: Apply the first weight value to the first input; And apply the second weight value to the second input.
6. The computer-implemented method of claim 5, wherein the first weight value is assigned to a first user of the first client device, and the second weight value is assigned to a second user of the second client device.
7. The computer-implemented method as described in claim 5, further comprising: The first weight value for the first input is received via a graphical user interface (GUI); And receive the second weight value for the second input via the GUI.
8. The computer-implemented method of claim 1, wherein the digital content item includes generative prompts, and the method further comprises: A second trained ML model is executed on the generative prompt to generate a second digital content item; And display the second digital content item in the multi-party interface.
9. The computer-implemented method of claim 1, wherein the composite prompt comprises at least intent text and non-text input, and wherein the non-text input comprises at least one of the following: computer-aided design (CAD) objects, geometry, images, sketches, videos, application status, or audio recordings.
10. The computer-implemented method of claim 9, wherein the intent text is included in the first input, and the non-text input is included in one or more non-text inputs constituting the second input.
11. One or more non-transitory computer-readable media, the non-transitory computer-readable media comprising instructions that, when executed by one or more processors, cause the one or more processors to generate design content by performing the following steps: generating a multi-party interface for communicating with at least a trained machine learning (ML) model, a first client device, and a second client device; combining at least a first input from the first client device and a second input from the second client device to generate a compound prompt; The composite prompt is transmitted to the trained ML model for execution; Receive digital content items generated in response to the composite prompts from the trained ML model; And display the digital content items in the multi-party interface.
12. The one or more non-transitory computer-readable media of claim 11, wherein the server device generates the multi-party interface, and the non-transitory computer-readable medium further includes instructions that, when executed by the one or more processors, cause the one or more processors to perform the following steps: transmitting the composite prompt by the server device to a remote device executing the trained ML model.
13. One or more non-transitory computer-readable media as claimed in claim 11, wherein the digital content item includes one of the following: text, computer-aided design (CAD) objects, geometry, images, sketches, video, executable code, or audio recordings.
14. One or more non-transitory computer-readable media as claimed in claim 11, wherein combining at least the first input from the first client device and the second input from the second client device comprises: Apply the first weight value to the first input; And apply the second weight value to the second input.
15. The one or more non-transitory computer-readable media of claim 14, further comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform the steps of: receiving a first weight value for the first input via a graphical user interface (GUI); and receiving a second weight value for the second input via the GUI.
16. The one or more non-transitory computer-readable media of claim 11, wherein the digital content item includes a generative prompt, and the non-transitory computer-readable medium further includes instructions that, when executed by the one or more processors, cause the one or more processors to perform the following steps: executing a second trained ML model on the generative prompt to generate a second digital content item; and displaying the second digital content item in the multi-party interface.
17. The one or more non-transitory computer-readable media of claim 11, wherein the composite prompt comprises at least intent text and non-text input, and wherein the non-text input comprises at least one of: computer-aided design (CAD) objects, geometry, images, sketches, videos, application states, or audio recordings.
18. One or more non-transitory computer-readable media as claimed in claim 11, wherein the trained ML model is trained using at least a combination of a first modality associated with text and at least one other modality associated with non-text input.
19. The one or more non-transitory computer-readable media of claim 11, further comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform the following steps: executing a second trained ML model on the composite prompt to generate a second digital content item; and displaying the second digital content item in the multi-party interface.
20. A system comprising: One or more memories, wherein the one or more memories store instructions; and one or more processors coupled to one or more memories, the one or more processors performing the following steps when executing the instructions: generating a multi-party interface for communicating with at least a trained machine learning (ML) model, a first client device, and a second client device; combining at least a first input from the first client device and a second input from the second client device to generate a compound prompt; The composite prompt is transmitted to the trained ML model for execution; Receive digital content items generated in response to the composite prompts from the trained ML model; And display the digital content items in the multi-party interface.