Method, apparatus, device, storage medium and program product for generating content
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-08-11
AI Technical Summary
[0008]以此方式,能够在减少人工拆解、人工串联和工程调试的情况下,将用户以自然语言等形式表达的输入内容转化为描述生成方案的结构化信息,并据此通过多个任务的协同执行与评审生成可运行的第一交互式内容,从而提升多模态内容生成的一致性、可控性和可扩展性。
Smart Images

Figure CN122547253A_ABST
Abstract
Description
Technical Field
[0001] The examples in this article generally relate to the field of computers, and in particular to methods, apparatus, devices, computer storage media, and computer program products for generating content. Background Technology
[0002] With the development of artificial intelligence technology, virtual objects are gradually being introduced into human-computer interaction scenarios to accompany users and provide diverse services. For example, users can interact with virtual objects through dialogue. Therefore, how to enrich the ways to interact with virtual objects is worthy of attention. Summary of the Invention
[0003] In a first aspect of this document, a method for generating content is provided. The method includes: receiving input content on a first interface associated with a virtual object; and providing first interactive content generated based on the input content, wherein the first interactive content is generated based on the following process: generating structured information based on the input content, the structured information describing a scheme for generating the first interactive content; performing multiple tasks based on the structured information; and generating the first interactive content based on the results of the execution of the multiple tasks.
[0004] In a second aspect of this document, an apparatus for generating content is provided. The apparatus includes: a first receiving module configured to receive input content on a first interface associated with a virtual object; and a providing module configured to provide first interactive content generated based on the input content, wherein the first interactive content is generated based on the following process: generating structured information based on the input content, the structured information describing a scheme for generating the first interactive content; performing multiple tasks based on the structured information; and generating the first interactive content based on the results of the execution of the multiple tasks.
[0005] In a third aspect, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.
[0008] In this way, user input expressed in natural language or other forms can be transformed into structured information describing the generation scheme with reduced manual disassembly, manual connection, and engineering debugging. Based on this, the first interactive content that can be run can be generated through the collaborative execution and review of multiple tasks, thereby improving the consistency, controllability, and scalability of multimodal content generation.
[0009] It should be understood that the content described in this section is not intended to limit the key or important features of the examples in this article, nor is it intended to restrict the scope of the solution. Other features will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the various examples herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A schematic diagram of the example environment is shown; Figure 2 Flowcharts of example processes for multi-agent cooperative processing in some scenarios are shown; Figures 3A to 3I The diagrams show several example interfaces for content generation in various scenarios; Figure 4 Flowcharts illustrating example processes for generating content in several scenarios are shown; Figure 5 Schematic block diagrams of example devices for generating content in several scenarios are shown; and Figure 6 A block diagram of an electronic device capable of implementing multiple illustrative scenarios is shown. Detailed Implementation
[0011] The examples in this document will now be described in more detail with reference to the accompanying drawings. While some examples are shown in the drawings, it should be understood that solutions can be implemented in various forms and should not be construed as limited to the examples presented herein. Rather, these examples are provided to provide a more thorough and complete understanding of the solutions. It should be understood that the drawings and examples in this document are for illustrative purposes only and are not intended to limit the scope of protection of the solutions.
[0012] It should be noted that the headings of any section / subsection provided herein are not restrictive. Various examples are described throughout this document, and examples of any type may be included under any section / subsection. Furthermore, examples described in any section / subsection may be combined in any way with any other examples described in the same section / subsection and / or different sections / subsections.
[0013] In the description of the examples in this document, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an example" or "the example" should be understood as "at least one example". The term "some examples" should be understood as "at least some examples". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0014] The examples in this article may involve user data, data acquisition, and / or use. All of these aspects comply with relevant laws, regulations, and rules. In the examples, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, when implementing each example, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained through appropriate means, in accordance with relevant laws and regulations. The specific methods of notification and / or authorization can vary depending on the actual situation and application scenario; the scope of the solution is not limited in this regard.
[0015] In this manual and the sample solutions, any processing of personal information will be conducted only under legal grounds (such as obtaining the consent of the data subject or being necessary for the performance of a contract) and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.
[0016] As used in this document, the term "virtual object" refers to an object that can be interacted with by a user through an interface. This typically includes, but is not limited to, digital objects with visual appearance, audio attributes, and configuration information, and can be configured to respond to received requests based on that configuration information. Such virtual objects can be created by the user. Different users can correspond to different virtual avatars.
[0017] As used in this document, the term "interface" refers to the visual medium presented to and interacted with by an electronic device, which may include, but is not limited to, a session interface, an authoring input interface, a preview interface, and a publishing interface. The term "first interface" refers to an interface associated with a virtual object used to receive user input; the term "second interface" refers to an interface used to present at least one parameter used to generate the first interactive content and for the user to edit. The first interface and the second interface can be different display states of the same interface, or they can be different interfaces.
[0018] As used in this document, “input content” refers to information provided by the user via a first interface that drives content generation, and may include, but is not limited to, text (e.g., natural language descriptions, user messages), images, videos, audio, or any combination thereof.
[0019] As used herein, the term "interactive content" refers to digital content that can be interacted with by a user. This content typically includes, but is not limited to, at least one of the following: virtual avatars, virtual environments, virtual elements, motion animations, page layouts, interactive logic, and runtime logic code, and can provide corresponding feedback in response to user interactions (e.g., tapping, swiping, voice input, or text input). The term "first interactive content" refers to interactive content generated based on input content; the term "second interactive content" refers to interactive content generated during the generation or modification process for review or delivery. In some examples, interactive content may also be referred to as interactive works or interactive content.
[0020] The term "structured information" as used in this article refers to the engineering specifications that describe the scheme for generating the first interactive content. This typically includes, but is not limited to, at least one of the following: work type, character assets, scene assets, motion configuration, page structure, interaction rules, acceptance criteria, version information, and constraints. It serves as the unified basis for task breakdown, task execution, and result review. In some examples, structured information may also be referred to as engineering creation specifications (CreationSpec).
[0021] As used in this article, the term "task" refers to a separately executable processing unit for generating or modifying first interactive content, which may include, but is not limited to, at least one of image generation tasks, scene generation tasks, action generation tasks, page generation tasks, and code generation tasks.
[0022] The term "first intelligence module" as used in this article refers to the processing module used to transform input content into structured information, decompose the structured information into tasks, orchestrate functional units to execute tasks, and review the execution results. In some examples, the first intelligence module may also be called a creation agent.
[0023] As used in this article, the term "second intelligence module" refers to a processing module deployed on the first interface that receives input content, identifies requests, collects information associated with generating the first interactive content, and creates a generation task based on the input content. In some examples, the second intelligence module may also be referred to as the host agent.
[0024] As used in this document, the term "functional unit" refers to a capability unit that can be dynamically invoked by the first intelligent module to perform corresponding tasks, typically provided in the form of standardized interfaces, microservices, or plugins. The term "set of functional units" refers to a collection of multiple functional units that are available for selection and invocation. In some examples, functional units may also be referred to as skills, and the set of functional units may also be referred to as a skills layer. Functional units may include, but are not limited to, at least one of the following: information gathering functional units, content understanding functional units, solution generation functional units, image generation functional units, scene generation functional units, action generation functional units, and code generation and quality review functional units.
[0025] The term "virtual avatar resource" as used in this document refers to digital resources used to present virtual avatars, which may include, but are not limited to, models, textures, skeletons, clothing, etc., of virtual humans or virtual characters in two-dimensional (2D) or three-dimensional (3D) form. The term "primary avatar resource" refers to the virtual avatar resource corresponding to the virtual object, which can be created based on the user's associated information.
[0026] The term "virtual environment resources" as used in this article refers to digital resources used to construct virtual scenes or page environments, which may include, but are not limited to, three-dimensional virtual spaces, scenes, backgrounds, props, lighting, cameras, and page layouts.
[0027] The term "interaction logic" as used in this article refers to the logic that describes the interaction mode of at least one virtual element contained in the first interactive content, which may include, but is not limited to, triggering conditions, responses after triggering, state transitions, and corresponding running logic code.
[0028] The term "virtual element" as used in this article refers to an element that can be presented and interacted with in the first interactive content, which may include, but is not limited to, virtual avatars, props, scenes, and user interface (UI) controls.
[0029] The term "associated information" as used in this article refers to information associated with users and / or virtual objects, which may include, but is not limited to, the user's character settings, personality, background, historical interaction records, and interaction memories between the user and virtual objects.
[0030] The term "conversational interface" as used in this article refers to an implementation of a first interface used to receive and present messages between the user and virtual objects (or intelligent modules) in the form of a dialogue.
[0031] The term "evaluation information" as used in this article refers to the information obtained from performing consistency checks, dependency verifications, and operability reviews on the generated interactive content, used to determine whether the generated results meet the preset conditions.
[0032] As used in this document, the term "edit request" refers to a request to modify the generated first interactive content, which may be provided via a natural language message (e.g., a third message) or via an edit operation on parameters.
[0033] As used in this article, the term "input template" refers to a template used to guide users in providing input content, corresponding to a specific type of interactive content. The term "first type" refers to a type of interactive content; the term "first action" refers to the action performed by the user to create first type of interactive content (e.g., selecting a template tag).
[0034] The term "agent" as used in this article refers to a processing entity that can perceive input, reason and make decisions and perform corresponding actions. It can usually be implemented based on machine learning models and can invoke one or more functional units (skills) to complete tasks.
[0035] The term "multimodal" as used in this article refers to information or processing involving two or more modalities, such as text, images, video, audio, and 3D assets. Understanding multimodal authoring requirements refers to the process of jointly parsing input content from multiple modalities to determine the user's creative requirements.
[0036] The term "AgentLoop" as used in this article refers to an iterative processing loop consisting of task decomposition, task execution, result review, and loop correction, which is used to automatically return to task decomposition or task execution for iteration when the generated result does not meet the conditions.
[0037] The term "ContentAgent" as used in this article refers to a carrier that carries, runs, and provides the first interactive content to the outside world. It can load offline asset indexes, running configurations, prompts, interaction rules, and running logic code, and provide users with accessible and interactive content.
[0038] As used in this article, the term "offline asset" refers to resources (such as virtual avatars, props, scenes, motion animations, images, and videos) that are pre-generated and stored and can be invoked at runtime by the first interactive content; the term "offline asset index" refers to index information that organizes and references offline assets.
[0039] As mentioned above, with the development of artificial intelligence technology, virtual objects are gradually being introduced into human-computer interaction scenarios to accompany users and provide diverse services. For example, users can interact with virtual objects through dialogue. Therefore, how to enrich the ways to interact with virtual objects is worthy of attention.
[0040] This paper proposes a content generation scheme. According to this scheme, input content can be received on a first interface associated with a virtual object. Furthermore, first interactive content can be provided. The first interactive content is generated based on the input content. Specifically, the first interactive content is generated based on the following process: generating structured information based on the input content, the structured information describing the scheme for generating the first interactive content; performing multiple tasks based on the structured information; and generating the first interactive content based on the execution results of the multiple tasks.
[0041] In this way, users can drive content generation with input expressed in the form of natural language, etc. The system automatically completes creative understanding, structured specification generation, task decomposition, capability invocation and result review. This reduces manual decomposition and engineering debugging, and stably generates multimodal interactive content that is consistent in style, logically self-consistent and operable, while supporting subsequent controllable modification and capability expansion.
[0042] The following describes various examples of this scheme in further detail with reference to the accompanying drawings.
[0043] Example Environment Figure 1 A schematic diagram of example environment 100 is shown. (e.g.) Figure 1 As shown, environment 100 may include electronic device 110.
[0044] In this example environment 100, electronic device 110 can run an application 120 that supports user interface interaction. Application 120 can be any suitable type of application for generating content, including but not limited to: social applications, virtual object applications, or other suitable applications. User 140 can interact with application 120 via electronic device 110 and / or its attached devices.
[0045] exist Figure 1 In environment 100, if application 120 is active, electronic device 110 can use application 120 to present interface 150 for supporting content generation. In some examples, interface 150 may include a first interface, a second interface, a third interface, etc., as mentioned in the scheme.
[0046] In some cases, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some cases, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).
[0047] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 that support content generation in electronic devices 110.
[0048] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection can include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections. In some cases, server 130 and electronic device 110 can exchange signaling information through their communication connection.
[0049] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the scheme.
[0050] The following description of the example will continue with reference to the accompanying drawings.
[0051] Figure 2A flowchart of an example process 200 for multi-agent cooperative processing in some scenarios is shown. Process 200 can be implemented in... Figure 1 This occurs in an environment where the electronic device 110 or server 130 is used in conjunction with the environment described below. For ease of understanding, the following description of each stage of process 200 will be combined with... Figures 3A to 3I The example interface shown illustrates the interactive presentation of the corresponding steps, thereby combining the processing flow with the interaction chain.
[0052] like Figure 2 As shown, in box 202, a user can initiate a user request. For example, electronic device 110 can receive a user request on a first interface associated with a virtual object. For example, a user request may include input content entered by the user in natural speech.
[0053] like Figure 3A As shown, electronic device 110 can present interface 300A. For example, interface 300A can be a schematic example of a first interface associated with a virtual object. Electronic device 110 can present a visual representation 302 of the virtual object on interface 300A. As an example, the virtual object can be associated with a user.
[0054] In some cases, virtual objects are created based on user-associated information. For example, the visual appearance of a virtual object can be generated based on at least one piece of media content provided by the user. For instance, in addition to a visual appearance, a virtual object may possess other attributes different from those of the visual appearance. For example, the attributes of a virtual object may also include audio attributes. Audio attributes can be attributes that characterize sound features (e.g., timbre). Alternatively, the attributes of a virtual object may include virtual actions that the virtual object can perform, response rules for user requests, etc. For example, the attributes of a virtual object may be configured by the user. Alternatively, the configuration information of a virtual object may be configured by the user. Such configuration information can indicate the identity, name, personality, etc., of the virtual object.
[0055] In some cases, on the one hand, virtual objects can have a visual representation for interaction. On the other hand, virtual objects can also utilize appropriate models to drive their interactive behavior; such models can include any suitable machine learning model, such as a generative model. In some examples in this paper, the various interactive processes and / or generative operations performed by virtual objects can be actually performed using appropriate models corresponding to the virtual objects.
[0056] In some situations, electronic device 110 can display guidance information 304 on interface 300A (e.g., "You can tell me your needs, and you can also send me photos and videos if you want references~"). Guidance information 304 can be used to guide users in creating interactive content.
[0057] In some cases, electronic device 110 may present input component 305 on interface 300A. For example, electronic device 110 may obtain user input via input component 305. For example, a user may interact with a virtual object through input component 305.
[0058] As an example, electronic device 110 may receive user requests via input component 305. For example, such user requests may include input content entered by the user in natural language via input component 305. Such input content may include text content, image content (e.g., pictures or videos), and / or audio content. Alternatively, electronic device 110 may present a message corresponding to the input content on interface 300A. Such a message originates from the user. For example, such a message may be presented in the form of a bubble and may also be presented along with user identification information (e.g., avatar or name identifier) to indicate that the message originates from the user.
[0059] In some scenarios, electronic device 110 can display control 306 (e.g., "Post Work") on interface 300A. As an example, electronic device 110 can also receive user requests based on triggering operations on control 306. For instance, electronic device 110 can display, in response to triggering operations on control 306, a user request. Figure 3B The interface shown is 300B.
[0060] like Figure 3B As shown, electronic device 110 can present interface 300B. For example, interface 300B can be implemented as a schematic example of a first interface associated with a virtual object. For example, electronic device 110 can present input panel 308 on interface 300B. Input panel 308 can be used to receive user input content (e.g., text content or image content, etc.). For example, electronic device 110 can present prompt text on input panel 308 (e.g., "Enter your ideas or material references to instantly generate your own unique interactive artwork"). As an example, electronic device 110 can also provide an upload entry on input panel 308 (e.g., "Add Material"). Electronic device 110 can present multiple candidate media content in response to triggering the upload entry. For example, multiple candidate media content are obtained from a local media library (e.g., "Photo Album") or an online media library. Further, electronic device 110 can receive a selection of at least one media content from the multiple candidate media content to use at least one media content as user input content. Alternatively, the electronic device 110 may, based on a triggered operation on the upload entry, invoke a camera component (e.g., a camera) deployed on the electronic device 110 to capture at least one piece of media content as user input. As an example, such at least one piece of media content may include images and / or videos.
[0061] In some cases, electronic device 110 may present multiple template items on interface 300B. For example, the multiple template items may include template items 310-1 to 310-4 (e.g., corresponding to "#template1" to "#template4"). In some cases, electronic device 110 may receive a first operation. The first operation represents the creation of interactive content of a first type. As an example, electronic device 110 may present an input template corresponding to the first type on a first interface. As an example, electronic device 110 may present an input template corresponding to the first template item in input panel 308 in response to the selection of the first template item among multiple template items. For example, the first template item may be associated with the first type. As an example, the first operation may include a selection operation of the first template item. Further, electronic device 110 may receive user input content via an input template. For example, such an input template may contain an input structure combining preset text content with blank parts to be filled in by the user. For example, the preset text content is used to provide a fixed instruction framework or example sentence to guide the user to express themselves in a standardized format. For example, blank areas serve as editable spaces, allowing users to freely fill in or supplement specific descriptive information (such as game name, scene style, character settings, etc.) according to their needs. After completing the blank areas, users can submit the complete input to the virtual object to generate interactive content. This method, by guiding user input through input templates corresponding to the type of work, lowers the creative threshold for users and improves the completeness of the input content.
[0062] like Figure 3C As shown, electronic device 110 can present interface 300C. For example, interface 300C can be implemented as a schematic example of a first interface associated with a virtual object. For example, interface 300C is a conversational interface with a virtual object. As an example, electronic device 110 can present messages corresponding to input content on interface 300C, such as messages 312-1 to 312-3. For example, messages 312-1 and 312-2 correspond to image content (e.g., pictures or videos) input by the user. For example, message 312-3 can correspond to text content input by the user.
[0063] In some cases, electronic device 110 may present a fourth message (e.g., message 314) on a first interface (e.g., interface 300C). The fourth message describes how to create the first interactive content. For example, such a fourth message may be generated based on input content. For example, the fourth message may characterize the creation steps, required materials, interaction logic, and / or formatting specifications of the first interactive content. For example, the fourth message is not a fixed template but is dynamically generated based on the user's input content. In this way, targeted creation guidance can be provided based on the input content, thereby reducing the user's learning cost and improving the efficiency and accuracy of interactive content generation.
[0064] In some scenarios, electronic device 110 may provide first interactive content in response to receiving confirmation of a fourth message. For example, electronic device 110 may trigger the generation of first interactive content in response to receiving confirmation of a fourth message, and provide the first interactive content to the user after its generation. As an example, electronic device 110 may provide a confirmation control associated with the fourth message. Electronic device 110 may receive confirmation of the fourth message in response to a triggering operation of the confirmation control (e.g., a click operation). Alternatively, electronic device 110 may also receive confirmation of the fourth message in response to receiving a message from the user. Such a message may represent confirmation of the fourth message.
[0065] Alternatively, the electronic device 110 may, in response to receiving a user's adjustment request, present a fifth message on the first interface. The fifth message describes how to create the first interactive content, and the creation method corresponding to the fifth message differs from the creation method corresponding to the fourth message. For example, the electronic device 110 may provide modification controls (e.g., "Modify" or "Regenerate") associated with the fourth message. The electronic device 110 may receive the user's adjustment request in response to triggering the modification controls. Alternatively, the electronic device 110 may receive the user's adjustment request in response to receiving a message representing the adjustment request (e.g., "Don't want to do it this way," "Another way," etc.). For example, the fifth message may describe an alternative creation scheme to the user, such as adjusting the creation steps, changing the material type, altering the interaction logic, or simplifying the complexity. The electronic device 110 presents the fifth message and waits for the user's further confirmation or rejection, thus forming a flexible interaction loop until the user is satisfied and the first interactive content is finally generated.
[0066] Continue to refer to Figure 2In box 204, the main agent (also known as the HostAgent, corresponding to the second intelligent module) can handle user requests. For example, the main agent can receive user input. In some examples, virtual objects can be associated with the second intelligent module. The dialogue interaction between the virtual object and the user on the first interface can be implemented based on the second intelligent module.
[0067] In some scenarios, the second intelligent module is configured to create a generation task based on the input content. Further, the second intelligent module can send the generation task to the first intelligent module to generate first interactive content. In one scenario, the second intelligent module can parse the user request. In some scenarios, box 206 is executed in response to the user request representing a need to create interactive content, i.e., when identified as a creation requirement. In box 206, the second intelligent module can perform creation information collection. In some scenarios, the second intelligent module can create a generation task based on the input content in response to the user's input content meeting the creation conditions.
[0068] As an example, creation conditions can indicate one or more of the following dimensions: completeness of input content, clarity of input content, complexity of input content, and format conformity of input content. For example, the completeness of input content can indicate whether the user has provided enough information (such as theme, core elements, interaction goals, etc.) to support content generation without further prompting or supplementation. The clarity of input content can indicate whether the user's description contains obvious ambiguity or contradictions, and whether it can be accurately parsed and transformed by the intelligent module. The complexity of input content can indicate whether the requested interactive content is within the system's supported capabilities and whether it does not exceed the processing limits of the current model or resources. For example, the format conformity of input content can indicate whether the input content meets preset format requirements (e.g., using supported text structures, referencing valid material identifiers, etc.).
[0069] In some cases, in response to the input content not meeting the creation conditions, the second intelligent module can continue to collect user requests to receive supplementary content from the user.
[0070] like Figure 3DAs shown, electronic device 110 can present interface 300D. Interface 300D can be implemented as a schematic example of a first interface associated with a virtual object. As an example, interface 300D is a conversational interface with a virtual object. As an example, electronic device 110 can receive a first message from a user in the conversational interface (e.g., interface 300D). As an example, the first message can be text, text transcribed from speech, an image, or a file, etc. For example, the first message may include message 320-1 and message 320-2. Further, electronic device 110 can present a second message (e.g., message 322) in the conversational interface (e.g., interface 300D). The second message comes from a second intelligent module. As an example, electronic device 110 can present a second message in response to the first message representing a need to create first interactive content (e.g., a user inputs a clear statement such as "I want to create a popular science interactive Q&A" or "Help me generate a product experience mini-game").
[0071] In some scenarios, a second message can be used to retrieve information associated with the generation of the first interactive content. As an example, a second intelligent module can generate a second message in response to the first message indicating that the creation information does not meet the creation conditions, guiding the user to provide more detailed creation information. In one scenario, the second message can be presented in a question-and-answer format; for example, the second message may include text content such as "What theme do you want this interactive content to correspond to?" or "Do you have any existing images or video materials?". In another scenario, the second message may also include interactive components. Such interactive components may include multiple interactive items (e.g., multiple options) for the user to select, thereby reducing the cost of manual input. For example, message 322 is one form of presentation of the second message. For example, message 322 may include interactive components. For example, message 322 may include text content 324. For example, text content 324 may be a question description associated with an interactive item, used to clarify to the user the content dimension that needs to be selected. Message 322 may also include multiple interactive items, such as interactive items 326-1 and 326-2. These interactive items can be clickable buttons, labels, icons, or cards, each representing a specific solution or answer to be chosen. Electronic device 110 can utilize these multiple interactive items to obtain user input. For example, when a user clicks interactive item 326-1, electronic device 110 records this selection as the user's input for the current question, and can further trigger follow-up questions or directly enter the generation process. In one scenario, multiple interactive items can support single or multiple selections. In another scenario, interactive items can also include supplementary input boxes, allowing users to add personalized descriptions after selecting an option. Through this "question description + optional items" interaction mode, the second intelligent module can efficiently collect structured creative information within a limited number of dialogue rounds, reducing the user's text input burden and improving the accuracy of information parsing, making the entire creative guidance process smoother.
[0072] In some cases, message 322 may also include indicator element 328 (e.g., "1 / 3"). Indicator element 328 is used to visually show the user the position and progress of the currently presented interactive component in the entire information collection process. For example, indicator element 328 may indicate that the currently displayed interactive component is the first of multiple interactive components, and the denominator "3" indicates that the guide contains three consecutive interactive components. The second intelligent module can only obtain complete creation information after the user completes the operation of all components in sequence. In terms of visual presentation, indicator element 328 can take various forms such as numerical labels (e.g., "1 / 3", "2 / 3", "3 / 3"), dot progress bars, segmented progress bars, or step numbers with highlighted states, so that users can understand it at a glance. As the user completes the selection or input of interactive items one by one, the electronic device 110 will dynamically update the content of indicator element 328. For example, after the user confirms the current option, the indicator element will change to "2 / 3" in the next second message, and so on. This progress indicator mechanism effectively reduces users' waiting anxiety and uncertainty, helps them estimate the amount of subsequent operations, and enhances their sense of control over the entire creation process. At the same time, when the indicator element displays "3 / 3", users can realize in advance that this is the last question, thus focusing more on completing the final input. Once all interactive components are completed, the electronic device 110 can automatically switch to the stage of generating the first interactive content.
[0073] Alternatively, electronic device 110 may present control 330 (e.g., "Generate Directly") in message 322. Electronic device 110 may trigger the direct generation of first interactive content in response to a triggering action on control 330 (e.g., click, touch, or voice command), without requiring the user to select all interactive components one by one. In this case, the second intelligent module can use preset content to complete the missing creative information in the input content. For example, the preset content may be system-built-in default parameters or reasonable default values inferred from user history or contextual conversations (e.g., automatically filling in the theme direction based on historical creative parameters). In one scenario, after receiving the direct generation instruction, the second intelligent module first performs a completeness check on the currently collected input content, identifies all necessary fields that are not provided, then extracts the corresponding default values from the preset content library to fill in the missing fields, ultimately forming a complete generation task and sending it to the first intelligent module. In another scenario, electronic device 110 may display a lightweight confirmation prompt (e.g., "Will use default settings to complete the missing fields, continue?") before triggering direct generation to obtain the user's final consent and avoid expected deviations caused by auto-completion. This alternative mechanism satisfies the needs of users who want quick output while ensuring the robustness of the generation process, making the interaction more flexible and user-friendly.
[0074] In some scenarios, electronic device 110 can receive a third message from a user within a conversational interface. For example, the third message corresponds to specific input provided by the user in the current conversation turn, and can be free text, voice commands, or actions on preset options. For instance, the third message can correspond to input by the user via input components (e.g., text input boxes, voice input buttons, or file upload controls), such as a detailed description manually typed by the user, text generated through speech-to-text transcription, or uploaded attachments such as images or documents. Alternatively, the third message can also indicate the user's selection of an interactive item (e.g., interactive item 326-1 or interactive item 326-2) within an interactive component; that is, the user selects an item from multiple preset options through clicks, touches, or swipes, and this selection result is considered the content carrier of the third message.
[0075] As an example, the input content can include a first message and a third message. Such input content can be used by the second intelligent module to create a generation task. For instance, the first message corresponds to the creative requirements submitted by the user in the initial round, while the third message corresponds to supplementary information input by the user in the current or subsequent rounds via the input component, or specific parameters provided through the selection of interactive items in the interactive component. The two messages logically complement each other, together forming a relatively complete set of creative materials. During the creation process, the second intelligent module can merge and structure the coarse-grained information (such as theme, goal, and general form) in the first message with the fine-grained information (such as applicable audience, material type, and interaction method) in the third message, forming a generation task configuration containing complete fields. In one scenario, if the third message already covers all the necessary information required for the creative conditions, the second intelligent module can directly construct the generation task based on this and send it to the first intelligent module for execution. In another scenario, if the input content is still incomplete, the second intelligent module can also save it as an intermediate state and continue to supplement and follow up with subsequent messages until the information is complete before formally creating the generation task. By accumulating and integrating messages from multiple rounds of dialogue into input content, the second intelligent module is able to maintain creative coherence within the context, ensuring that the final generated interactive content fully reflects the user's real needs.
[0076] like Figure 3E As shown, electronic device 110 can display interface 300E. Electronic device 110 can respond to receiving a user's exit operation via the session interface by displaying, as shown... Figure 3E The interface shown is 300E. This exit action can be performed by the user clicking the back button (e.g., on...). Figure 3DThe actions triggered by control 325, such as pulling down to close the session, switching to other applications, or locking the screen, indicate that the user does not want to stay on the current screen for the time being. For example, electronic device 110 can display a prompt pop-up window 332 on interface 300E. The prompt pop-up window 332 is used to inform that the first interactive content is being generated. Its content may include progress prompt text (such as "Content is being generated, please wait..."), estimated remaining time, or current generation progress percentage, allowing the user to keep track of the task progress even after exiting the interface. The prompt pop-up window 332 can also display control 334-1 (e.g., "Cancel Generation") and control 334-2 (e.g., "Generate Directly"), allowing the user to cancel the current generation or confirm to continue generation. For example, when the user clicks control 334-1, electronic device 110 terminates the current generation task and closes the prompt pop-up window 332. Alternatively, when the user clicks control 334-2, electronic device 110 keeps the generation task running in the background while closing the pop-up window, allowing the user to operate the device freely without interference. Since the generation process may require accessing a large number of assets and involve a certain waiting time, if the user leaves the interface, an asynchronous notification can be sent to the user upon completion of the generation (see [link to documentation]). Figure 3F ).
[0077] Continue to refer to Figure 2 The second intelligent module can collect information such as work type, theme, character image, scene, materials, interaction method, and publishing format through multiple rounds of interaction, and can call the information collection function unit to standardize and complete missing fields. When the creative information collected through multiple rounds of interaction meets the preset conditions, the second intelligent module can create a generation task based on the input content, encapsulate the collected requirements, user-uploaded materials, context records, and initial constraints into a generation task, and send the generation task to the first intelligent module to generate the first interactive content.
[0078] Alternatively, in response to a user's request to engage in casual conversation, the second intelligent module can generate a reply corresponding to the user's request without triggering the content creation process. This approach, by only initiating the subsequent creation process when the user expresses a creative need, reduces ineffective interactions and resource consumption.
[0079] As an example, the first intelligent module (also known as a creative agent, such as CreationAgent) can generate first interactive content based on the input after receiving a generation task. As another example, the first intelligent module can execute multiple stages indicated by the AgentLoop. This AgentLoop can include stages such as task breakdown, task execution, result review, and loop correction, and revolves around structured information.
[0080] For example, in box 208, the first intelligent module can generate media content. For instance, the first intelligent module can invoke the content understanding function unit to perform semantic analysis and feature extraction on the text, images, videos, audio, and other materials in the input content. For example, for user-uploaded image or video materials, the first intelligent module can use them as a reference (e.g., referencing the style, clothing, or expression of a person in the image to generate a first image resource corresponding to the virtual object). Alternatively, the first intelligent module can use them as content to be displayed (e.g., as visual resources or interactive rewards provided during interaction), and the specific application method can be determined by the first intelligent module based on the analysis results.
[0081] Alternatively, the first intelligent module can invoke the scheme generation function unit to generate structured information (also known as engineering creation specifications, such as CreationSpec) based on the input content. This structured information describes the scheme for generating the first interactive content, and may include, but is not limited to, work type, character assets, scene assets, action configurations, page structure, interaction rules, acceptance criteria, version information, and constraints. In some examples, the structured information can constrain script logic (i.e., the interactive flow and responses of the first interactive content), user interaction methods, and the generation specifications and presentation of materials or assets, thus serving as a unified basis for subsequent task decomposition, task execution, and result review.
[0082] Furthermore, the first intelligent module executes multiple tasks based on structured information. In some cases, the first intelligent module can create multiple tasks based on structured information. Further, the first intelligent module can utilize multiple functional units to execute multiple tasks. For example, multiple functional units are selected from a set of functional units based on multiple tasks. For example, the first intelligent module can decompose structured information into sub-tasks such as image generation, scene generation, action generation, page generation, and code generation, and select sub-tasks from the set of functional units (e.g., ...) according to the task type. Figure 2 In the skill layer 210, the corresponding functional unit is dynamically selected and invoked. The set of functional units can be provided in the form of standardized interfaces or microservices and defined as pluggable, so that when a certain capability (such as the ability to adjust the camera) needs to be added or replaced, it can be achieved by accessing the corresponding functional unit without changing the overall main process.
[0083] In some cases, the selected multiple functional units may include at least one of a first functional unit, a second functional unit, and a third functional unit.
[0084] For example, the first functional unit is configured to acquire virtual avatar resources. As an example, virtual avatar resources may include a first avatar resource corresponding to a virtual object. Such a virtual object is created based on the user's associated information. For example, the first avatar resource may be associated with the visual appearance of the virtual object. Specifically, the first avatar resource may be associated with the visual appearance of the virtual object, including but not limited to the virtual object's hairstyle, face shape, skin tone, clothing, accessories, facial expressions, and overall style (e.g., cartoon, realistic, etc.). Alternatively, the virtual avatar resource may also include a new avatar resource dynamically generated by the first functional unit using a generative model (e.g., a generative model) based on structured information.
[0085] As an example, the second functional unit is configured to acquire virtual environment resources to provide spatial scene-level visual support for the first interactive content. Virtual environment resources can include various types, such as 3D virtual spaces, scenes, backgrounds, props, lighting, cameras, or page layouts. Specifically, a 3D virtual space can define the overall world framework in which the interactive content occurs, including basic elements such as terrain, skyboxes, and collision boundaries. For example, a scene can refer to a combination of environments in a specific context, such as an indoor room, an outdoor forest, or a futuristic city. As an example, a background can be a static image or a dynamic video used to highlight the main content. For example, props refer to virtual elements that can be manipulated or observed by the user during interaction, such as buttons, signs, pickable items, or decorative objects. For example, lighting can indicate the brightness and visual hierarchy of the virtual environment, including ambient light, directional light, spotlights, etc., and different lighting combinations can create different atmospheres such as warmth, suspense, or a technological feel. For example, a camera defines the user's observation perspective in the virtual environment and can be set to first-person, third-person, etc. For example, page layout is suitable for lightweight interactive content, encompassing page structure, scrolling logic, card arrangement, and button placement. During execution, the second functional unit can acquire virtual environment resources in various ways based on parameters in the structured information, such as environmental style, spatial complexity, and interaction scale. For instance, it can load matching scene templates from a preset resource library and dynamically construct a 3D environment using a procedural generation algorithm. In one scenario, the second functional unit can also adaptively adjust the acquired environmental resources, such as automatically matching the environmental color scheme to the main color of the first character resource, or adjusting the camera movement rhythm according to the theme of the interactive content. In another scenario, if environmental details are not explicitly specified in the structured information, the second functional unit can use default values (such as "general indoor scene," "neutral warm lighting," and "default overhead shot") to complete the configuration, ensuring a smooth transition in the environment generation process. Through the participation of the second functional unit, the first intelligent module can build an immersive and stylistically harmonious virtual environment for interactive content within a unified creative framework, thereby enhancing the overall expressiveness and user experience of the final generated content.
[0086] In some scenarios, the third functional unit is configured to generate the interaction logic associated with the first interactive content. Interaction logic is the core mechanism that determines "how the interactive content responds" and "how it functions," directly influencing the feedback effect and overall user experience. As an example, interaction logic may include trigger conditions, responses after triggering, and corresponding execution logic code. Specifically, trigger conditions can indicate what actions the user performs to activate the interaction, such as clicking a button, swiping the screen, using voice commands, hovering the mouse, entering specific text, or completing a prerequisite task. Responses after triggering can indicate the behavior the system should perform after the interaction is activated, such as navigating to a new page, playing animations or sound effects, displaying hidden information, recording scores, switching dialogue branches, updating data status, or calling external interfaces. For example, the corresponding execution logic code can be a concrete product of these conditions and responses implemented in an executable programming language, used to actually drive the operation of the interactive content. During the generation process, the third functional unit can automatically derive appropriate interaction rules based on factors such as the interaction type (e.g., multiple choice, drag-and-drop sorting, timed challenges, scenario simulations, or role-playing dialogues), user goals, content themes, and environmental resources in the structured information, and convert them into code or configuration data. In one scenario, the third functional unit can reuse existing interaction templates (e.g., "single choice interaction template" or "drag-and-drop matching template") and then instantiate and populate them according to specific parameters (e.g., number of questions, option content, correct feedback text), thereby quickly generating runnable interaction logic. In another scenario, if the structured information describes complex interaction processes (e.g., multi-branch narratives, conditional unlocking, scoring leaderboards), the third functional unit can also call the code generation model or rule engine to dynamically construct a new logic tree and output code files that conform to the specifications. In addition, the third functional unit can also perform preliminary verification of the generated interaction logic, such as checking whether the triggering conditions are reachable, whether the response will lead to an infinite loop, and whether the variable names are consistent, to ensure the executability of the logic. Ultimately, the interactive logic code generated by the third functional unit is packaged together with the image and environmental resources produced by other functional units to form a complete, deployable interactive content package, which is then delivered to the user by the first intelligent module. Through this clearly defined architecture, the generation of interactive logic is both efficient and flexible, adapting to a wide range of creative needs, from simple question-and-answer games to complex story-driven games.
[0087] In some examples, virtual avatar resources and virtual environment resources can be derived from a pre-defined offline 3D asset library, or they can be generated in real time by natural language-based asset generation capabilities.
[0088] Alternatively, the first intelligent module can perform result verification. It conducts consistency checks, dependency checks, and runnability reviews on the generated assets, pages, interaction logic, and runtime logic code. Consistency checks verify whether there are conflicts or inconsistencies between various outputs, such as whether the style of the virtual avatar matches the color scheme of the virtual environment, whether page elements referenced in the interaction logic actually exist, and whether animation triggering conditions match the physical rules of the scene. Dependency checks confirm that all resource references (such as image paths, sound effect files, code repository versions, third-party interfaces, etc.) are valid and complete, without broken links or undefined variables. Runnability reviews are performed through static analysis or a sandbox environment to detect issues such as syntax errors, infinite loops, out-of-bounds access, or performance bottlenecks in the code, ensuring that interactive content can start stably and run smoothly.
[0089] In some scenarios, generating the first interactive content based on the execution results of multiple tasks may include: generating a second interactive content based on the execution results of multiple tasks (e.g., image resources produced by a first functional unit, environmental resources produced by a second functional unit, and interactive logic code produced by a third functional unit). This second interactive content can be considered an intermediate version or a preliminary draft, used for internal evaluation and verification before formal delivery. Furthermore, in response to the evaluation information of the second interactive content not meeting certain conditions, the first interactive content can be regenerated based on structured information, thus forming a closed-loop iterative mechanism of "generation-evaluation-optimization". For example, the evaluation information for the second interactive content can indicate one or more of the following dimensions: content completeness, i.e., whether all preset pages, characters, props, and interactive nodes have been correctly generated without any omissions; visual style consistency, i.e., whether the art style of the image resources and environmental resources is coordinated and unified, and whether the color matching and lighting effects meet the expected settings; interactive logic fluency, i.e., whether each triggering condition can activate the response normally, whether the response behavior meets the user's operation expectations, and whether there are any situations of no response or erroneous response; code robustness, i.e., whether the running logic code has passed the syntax and static security checks, and whether it has the ability to capture exceptions and handle faults; and semantic alignment, i.e., whether the overall generated content accurately reflects the original creative requirements expressed by the user in the first and third messages. When the evaluation score of any one or more of the above dimensions is lower than the preset threshold, the first intelligent module can determine that the conditions are not met and automatically trigger the regeneration process. During the regeneration process, the first intelligent module can choose to retain the output that has passed the verification (such as image resources), and only perform targeted correction and regeneration on the substandard parts (such as interactive logic or page layout) based on the original structured information, thereby improving iteration efficiency. If the evaluation information still fails to meet the conditions after multiple regenerations, the first intelligent module can send a feedback signal to the second intelligent module. The second intelligent module then adjusts or refines the parameter settings in the structured information, triggering the entire process of generating interactive content again, until the generated second interactive content passes all review conditions and is finally delivered to the user as the official first interactive content. Through this rigorous review and iteration mechanism, the system can significantly reduce the error rate and rework costs of content generation, ensuring the reliability of the final delivery and user satisfaction.
[0090] As an example, after the evaluation information meets the conditions, the first intelligent module can output the first interactive content. For example... Figure 2 As shown, the initial version of the work may include offline assets 214 (such as offline asset index) and running configuration, prompts, interaction rules and running logic code carried by the work's intelligent agent 216 (also known as the third intelligent module).
[0091] In some cases, continue to refer to Figure 3C The electronic device 110 can provide the generated first interactive content via a first interface. (See reference) Figure 3C Electronic device 110 may present card 316 in a session interface (e.g., interface 300C). Card 316 may present a preview of the first interactive content (e.g., a thumbnail). Alternatively, electronic device 110 may present control 318 (e.g., "Go to Publish") in card 316. Electronic device 110 may present a preview of the first interactive content in response to a triggering operation on card 316 or control 318.
[0092] Alternatively, a completion notification can be sent to the user asynchronously after generation. For example... Figure 3F As shown, the electronic device 110 can display an interface 300F. The electronic device 110 can display a message 336 (e.g., "My interactive artwork for you has been generated! Click the [Play] button to see it!") on the interface 300F. It can also display a control 338 (e.g., "View Artwork"), through which the user can access a preview.
[0093] As an example, the preview interface can be accessed via control 318 in interface 300C or the "View Artwork" control 338 in interface 300F. Figure 3G As shown, the electronic device 110 can present an interface 300G. The first interactive content includes at least one virtual element. The interface 300G can present at least one virtual element contained in the first interactive content. For example, element 340 (a presentation of a first image resource corresponding to a virtual object) and element 342 (a presentation of a virtual environment resource). As an example, the first interactive content is associated with interaction logic that describes the interaction method of at least one virtual element. In the preview interface, this interaction logic is actually running, and users can interact with the virtual elements in real time through clicks, drags, voice input, swipes, or hovers, experiencing whether the triggering conditions are sensitive, the response feedback is accurate, and the page transitions are smooth. Simultaneously, the interface 300G can also provide auxiliary functions, such as interactive prompt overlays (used to highlight operable areas when the user first enters), progress indicators (used to display the current position in a multi-step interaction process), and operation result feedback (such as correct / incorrect prompts, score updates, etc.).
[0094] In some cases, electronic device 110 may receive an editing request associated with the first interactive content. (Continue to refer to...) Figure 2In box 212, a user's edit request (also known as a modification request) can be received. As an example, the second intelligent module can receive the user's edit request and obtain the user's editing requirements through multiple rounds of interaction. In some cases, the electronic device 110 can receive a third message from the user on the first interface. The third message can represent the edit request. (Continue to the previous section) Figure 3G The electronic device 110 can display control 350 (e.g., modify) on interface 300G. The electronic device 110 can display a session interface with virtual objects in response to a trigger operation on control 350.
[0095] like Figure 3H As shown, electronic device 110 can present interface 300H. For example, interface 300H can be implemented as a schematic example of a first interface associated with a virtual object. For example, electronic device 110 can present a conversation panel 352 in interface 300H. For example, interface 300H can also be referred to as a conversation interface with a virtual object. For example, electronic device 110 can present a message 354 from the virtual object (e.g., "Do you have any new ideas for the newly generated work?") in interface 300H to guide the user to provide information associated with the editing request. As an example, electronic device 110 can present an input component 356 in interface 300H. Electronic device 110 can receive user input via input component 356. Furthermore, electronic device 110 can receive a third message from the user. For example, the third message can correspond to the input received via input component 356.
[0096] In other scenarios, electronic device 110 may present a second interface. This second interface differs from the aforementioned session interface (such as interface 300C or 300D). It can be an independent parameter configuration panel, a floating sidebar, or a full-screen settings page, specifically designed to centrally display and adjust the key parameters required to generate the first interactive content. The second interface presents at least one parameter used to generate the first interactive content. These parameters may include, but are not limited to: the virtual object's appearance style (e.g., realistic, cartoon, traditional Chinese style), scene theme (e.g., indoor, outdoor, science fiction), interaction mode (e.g., multiple choice, drag-and-drop sorting, voice dialogue), content difficulty level, estimated duration, number of pages, background music type, color theme, and whether specific functions are enabled (e.g., scoring, timing, sharing). These parameters can originate from the automatic parsing and filling of user input by the second intelligent module, or they can be system default values or a comprehensive presentation of user history information. In terms of layout, the second interface typically uses grouping or pagination, grouping related parameters into the same area (e.g., "Appearance Settings," "Environment Settings," "Interaction Settings"), and providing clear labels and explanatory text to help users understand the function of each parameter.
[0097] Furthermore, the electronic device 110 can receive editing operations via a second interface. These editing operations are used to edit at least one of the aforementioned parameters, such as adjusting values via a slider, switching options via a drop-down menu, enabling or disabling functions via a switch, entering custom content via a text box, or selecting a main color using a color picker. Editing operations can take various interactive forms, such as single clicks, dragging, swiping, or keyboard input, with real-time feedback during the operation (e.g., synchronized updates to the preview image), allowing users to intuitively see the expected impact of parameter adjustments on the generated content. Such editing operations can correspond to editing requests associated with the first interactive content. When a user modifies any parameter on the second interface and confirms it, the operation is considered as initiating a modification request for the first interactive content to the system. Compared to submitting editing requests in natural language within a conversational interface, the second interface provides a structured and visual method for parameter adjustment, particularly suitable for users who prefer direct operation rather than textual expression, or for scenarios requiring batch fine-tuning of multiple parameters. After receiving an editing operation, the electronic device 110 can trigger a regeneration or incremental update process of the first intelligent module based on the modified parameter set, thereby quickly generating an updated version of the content reflecting the new parameter settings. In addition, the second interface can provide auxiliary controls such as "Restore Defaults," "Reset All," or "Compare to Original," helping users compare and make decisions between different parameter combinations. By introducing this parametric editing interface, the electronic device 110 retains the flexibility of natural language interaction while also meeting the needs of advanced users for precise control and efficient operation, further enriching the interactive dimensions of creation and editing. Continue to refer to Figure 2 After the first interactive content is generated, the user can submit an editing request, thus entering a modification iteration loop.
[0098] As an example, the first intelligent module can perform material analysis. Specifically, the first intelligent module will structurally decompose the generated first interactive content, extracting and labeling elements such as image resources, environmental resources, interactive logic code, page layout, and animation parameters to form a material map that can be identified and located in subsequent modification processes. For example, if a user requests to "change the color," the material analysis step needs to be able to accurately identify attribute fields related to visual presentation, rather than treating the entire content as an indivisible black box.
[0099] For example, the first intelligent module can perform multiple rounds of request confirmation. Since user editing requests are often vague or incomplete (e.g., "Make it look better" or "Adjust the atmosphere"), the first intelligent module can clarify specific modification goals through multiple rounds of dialogue with the user. This process may include asking follow-up questions about specific dimensions (e.g., "Do you want to adjust the scene lighting, color tone, or prop arrangement?"), providing candidate options for the user to choose from (e.g., "Here are three preset nighttime styles, which one do you prefer?"), or requesting the user to upload reference samples to clarify the direction. Through such request confirmation, vague natural language instructions can be transformed into actionable modification parameters.
[0100] Alternatively, the first intelligent module can generate structured information. Based on the modification requirements confirmed in box 236, the first intelligent module will convert them into a structured data format compatible with the initial creation process. This structured information may include the target identifier to be modified (such as which specific page or character), the modification type (such as replacement, addition, deletion, or adjustment), new parameter values (such as new color codes or new position coordinates), and modification constraints (such as maintaining style consistency with other elements), providing standardized input for subsequent task breakdown.
[0101] Alternatively, the first intelligent module can decompose the task. Based on the specific scope of modification described in the structured information, the first intelligent module breaks down the overall editing requirements into several independently executable sub-tasks. For example, modifying a character's appearance might be broken down into tasks such as "generating new character textures," "updating character skeletal bindings," and "replacing existing character references"; adding a scoring component might be broken down into sub-tasks such as "designing components," "writing scoring logic code," and "configuring scoring trigger conditions." These sub-tasks may have dependencies on each other and need to be arranged reasonably according to the execution order.
[0102] As an example, the first intelligent module can perform multiple tasks. Similar to the initial generation process, the first intelligent module utilizes multiple functional units (such as the first functional unit for updating image resources, the second functional unit for adjusting environmental resources, and the third functional unit for modifying interaction logic) to execute the aforementioned sub-tasks in parallel or sequentially. During this process, functional units and their outputs not involved in the editing request can be retained and reused, thereby avoiding unnecessary recalculation and improving modification efficiency.
[0103] As an example, the first intelligent module can perform result verification. This step is similar to the verification logic mentioned earlier, but with a different focus. At this stage, the verification needs to specifically check whether the modified content is consistent with the unmodified parts. For example, whether the newly generated character image matches the lighting and color tone of the original scene, whether the added interactive logic conflicts with the original logic tree, and whether the modified page navigation links are complete and reachable. Simultaneously, it also needs to verify whether the modifications truly meet the user's expressed requirements. If the verification fails, the material parsing can be re-executed or confirmation requested, forming a cyclical optimization mechanism embedded in the modification iteration loop. Through this complete modification iteration loop, the electronic device 110 can continuously respond to the user's refined needs after the initial generation, extending the one-time "content generation" capability into a sustainable "content co-creation" experience.
[0104] Furthermore, the first intelligent module can generate third interactive content. This third interactive content can be generated based on the edit request and the first interactive content. For example... Figure 2 As shown, the third interactive content may include offline assets (such as an offline asset index) and runtime configurations, prompts, interaction rules, and runtime logic code carried and run by the work's intelligent agent (also known as the third intelligent module).
[0105] In some cases, electronic device 110 can receive a user's publishing request to publish the first or third interactive content mentioned above. (Continue to refer to...) Figure 3GBefore publishing, users can configure or modify the title or music content of the first or third interactive content. For example, electronic device 110 can present entry 344 (e.g., "Select Music") in interface 300G. After clicking this entry, a music library panel or local file selector will pop up, from which background music or sound effects can be selected. The system supports listening, replacing or removing the current background music, and can also adjust the volume, loop mode, or trigger timing bound to specific interactive nodes. For example, electronic device 110 can present entry 346 (e.g., "Edit Title") in interface 300G. After clicking, users can modify the display name of the interactive content in a pop-up window or inline editing box. This name will be used for the title display of sharing cards, notification bars, and publishing pages, making it easy for others to identify and remember. For example, electronic device 110 can present publishing control 348 (e.g., "Publish Now") in interface 300G. After the user clicks this control, the system will execute the final publishing process, including resource packaging, uploading to the server, generating access links or QR codes, and can selectively synchronize to social platforms or share with designated users. In one scenario, before clicking the publish control 348, the electronic device 110 can perform pre-publishing checks, such as checking if the title is empty, if the music file is valid, and if the content is complete. If any problems are found, prompts will be provided to guide the user to make corrections. In another scenario, after publishing, the system can return a successful publishing notification and generate a new card in the session interface, displaying an access link, QR code, or sharing poster for the published content, facilitating user dissemination. Through this pre-publishing configuration mechanism, users can flexibly adjust the display details of the content in the final step before publishing, ensuring content personalization while reducing the cost of post-publishing modifications, making the entire creation-to-sharing process smoother and more controllable.
[0106] As an example, the first interactive content can be carried and run by a work-based intelligent agent, allowing users to interact with and consume it. For example... Figure 3I As shown, the electronic device 110 can present an interface 300I. For example, the interface 300I can be implemented as a viewing interface for an information stream. For example, the information stream can be associated with multiple media contents, and users can switch between different media contents in the information stream based on preset operations (e.g., up and down swiping operations) in the viewing interface. Alternatively, the information stream can also be associated with interactive content. For example, the electronic device 110 can present a published interactive work 358 (e.g., the first or third interactive content mentioned above) on the interface 300I, and present a control 360 (e.g., "Click to enter interaction"). After entering via the control 360, the user can interact with the virtual elements according to the interaction method described by the interaction logic.
[0107] In some examples, the interaction logic describes interaction methods that may include, but are not limited to: tapping virtual elements (e.g., touching the head of a virtual avatar to trigger a corresponding action or expression), tapping interface controls (e.g., selecting an option or triggering a special effect), swiping, voice input, or text input, etc., and feedback can be presented by the corresponding virtual elements, such as the virtual avatar making an action, changing its expression, or deforming; changes in the scene, props, or lighting; special effects presented in the front-end user interface; or dialogue responses generated by a machine learning model. In this way, interactive content can support diverse interaction methods, enhancing the playability of the content.
[0108] Based on the process described above, significant technical effects can be achieved in multi-intelligent module collaboration, information guidance, task execution, and iterative optimization: The second intelligent module parses user input and determines creation conditions, and actively guides information completion through question-and-answer interaction or optional components, supplemented by progress indicators and a "direct generation" shortcut, which reduces the user's input burden and improves the completeness and efficiency of information collection; With the help of a pluggable functional unit architecture and structured task decomposition, the first intelligent module can transform creative needs into parallel sub-tasks such as image, environment, and interaction logic and dynamically call corresponding capabilities. At the same time, through consistency checks, dependency verification, and operability reviews in the result review, the accuracy and reliability of the generated content are effectively guaranteed; After generation, it supports editing iteration based on natural language or parameter panels, and with the incremental update mechanism, it can flexibly respond to users' fine-grained modification needs; In addition, the design of preview interface, asynchronous notification, and pre-publish title and music configuration allows users to fully experience and improve the content before publishing, ultimately forming a complete closed loop from guidance, generation, review, modification to publishing, which significantly improves the overall efficiency, flexibility, and user satisfaction of interactive content creation.
[0109] Example process Figure 4 A flowchart of an example process 400 for generating content is shown, based on some scenarios. Process 400 can be implemented at electronic device 110. See below for reference. Figure 1 To describe process 400.
[0110] In box 410, electronic device 110 can receive input content on a first interface associated with the virtual object.
[0111] In box 420, electronic device 110 can provide first interactive content, which is generated based on input content, wherein the first interactive content is generated based on the following process: generating structured information based on the input content, the structured information describing a scheme for generating the first interactive content; performing multiple tasks based on the structured information; and generating the first interactive content based on the results of the execution of the multiple tasks.
[0112] In some cases, performing multiple tasks based on structured information includes: creating multiple tasks by a first intelligent module based on structured information; and performing multiple tasks by the first intelligent module using multiple functional units, wherein the multiple functional units are selected from a set of functional units based on the multiple tasks.
[0113] In this way, by using intelligent modules to parse structured information and dynamically schedule functional units, task creation and execution are decoupled, improving the system's scalability and execution efficiency.
[0114] In some cases, the multiple functional units include at least one of the following: a first functional unit configured to acquire virtual avatar resources; a second functional unit configured to acquire virtual environment resources; and a third functional unit configured to generate interactive logic associated with the first interactive content.
[0115] In this way, the functional units have a clear division of labor, each responsible for the image, environment, and interaction logic, which facilitates independent development and reuse.
[0116] In some cases, virtual avatar resources include primary avatar resources corresponding to virtual objects, which are created based on the user's associated information.
[0117] In this way, the generated interactive content can reuse the virtual object images that users have already created, enhancing personalization and consistency.
[0118] In some cases, the first interface is associated with a second intelligent module, which is configured to: create a generation task based on the input content; and send the generation task to the first intelligent module to generate the first interactive content.
[0119] In this way, the second intelligent module acts as the front-end interaction layer, responsible for understanding user requirements and creating tasks, while the first intelligent module acts as the back-end execution layer, responsible for generating content. This division of labor and collaboration improves processing capabilities.
[0120] In some cases, the first interface includes a session interface with a virtual object, and receiving input includes: receiving a first message from a user in the session interface; presenting a second message in the session interface, the second message originating from a second intelligent module, the second message being used to obtain information associated with the generation of the first interactive content; and receiving a third message from a user in the session interface, wherein the input includes the first message and the third message.
[0121] In this way, the required information is collected and generated through multiple rounds of dialogue, which reduces the burden of one-time input for users and ensures the integrity of the information.
[0122] In some cases, presenting a second message includes: presenting a second message in response to the first message representing the need to create the first interactive content.
[0123] In this way, the second intelligent module actively guides information collection after recognizing creative requirements, forming a natural dialogue process.
[0124] In some cases, generating first interactive content based on the execution results of multiple tasks includes: generating second interactive content based on the execution results of multiple tasks; and triggering the regeneration of first interactive content based on structured information in response to the failure of evaluation information of the second interactive content to meet certain conditions.
[0125] In this way, the quality of the generated content is verified through an automatic evaluation mechanism, and regeneration is triggered when the conditions are not met, thereby improving the reliability and quality of the final output.
[0126] In some cases, process 400 may further include: receiving an edit request associated with the first interactive content; and providing third interactive content, which is generated based on the edit request and the first interactive content.
[0127] This approach allows for incremental editing of already generated content without starting from scratch, thus improving iteration efficiency.
[0128] In some cases, receiving an edit request associated with the first interactive content includes: on the first interface, receiving a third message from the user, the third message representing the edit request.
[0129] In this way, users can directly submit modification requests using natural language, lowering the editing threshold.
[0130] In some cases, receiving an edit request associated with the first interactive content includes: presenting a second interface that presents at least one parameter for generating the first interactive content; and receiving an edit operation via the second interface that edits at least one parameter.
[0131] In this way, a visual editing interface for parameters is provided to meet users' needs for precise adjustments.
[0132] In some cases, the first interactive content includes at least one virtual element, and the first interactive content is associated with interaction logic that describes how the at least one virtual element is interacted with.
[0133] In this way, the generated interactive content not only includes static elements, but also has interactive behavioral logic, enhancing the richness of the content.
[0134] In some cases, providing the first interactive content includes: presenting a fourth message on a first interface, the fourth message describing how the first interactive content was created, the fourth message being generated based on input content; and providing the first interactive content in response to receiving confirmation of the fourth message.
[0135] This approach, which presents the creative plan to the user before its official generation and allows for user confirmation before execution, enhances transparency and controllability.
[0136] In some cases, receiving input content on a first interface associated with a virtual object includes: receiving a first operation, the first operation representing the creation of interactive content of a first type; presenting an input template corresponding to the first type on the first interface; and receiving input content based on the input template.
[0137] In this way, customized input templates are provided based on content type, guiding users to provide information in a structured manner, thereby improving the completeness and efficiency of information collection.
[0138] Example devices and equipment A corresponding apparatus for implementing the above methods or processes is also provided. Figure 5 A schematic structural block diagram of an example device 500 for generating content under certain circumstances is shown. Device 500 can be implemented as or included in electronic device 110. The various modules / components in device 500 can be implemented by hardware, software, firmware, or any combination thereof.
[0139] like Figure 5 As shown, the device 500 includes: a first receiving module 510 configured to receive input content on a first interface associated with a virtual object; and a providing module 520 configured to provide first interactive content, which is generated based on the input content, wherein the first interactive content is generated based on the following process: generating structured information based on the input content, the structured information describing a scheme for generating the first interactive content; performing multiple tasks based on the structured information; and generating the first interactive content based on the execution results of the multiple tasks.
[0140] In some cases, the device 500 also includes an execution module configured to: create multiple tasks based on structured information by a first intelligent module; and execute the multiple tasks by the first intelligent module using multiple functional units, the multiple functional units being selected from a set of functional units based on the multiple tasks.
[0141] In some cases, the multiple functional units include at least one of the following: a first functional unit configured to acquire virtual avatar resources; a second functional unit configured to acquire virtual environment resources; and a third functional unit configured to generate interactive logic associated with the first interactive content.
[0142] In some cases, virtual avatar resources include primary avatar resources corresponding to virtual objects, which are created based on the user's associated information.
[0143] In some cases, the first interface is associated with a second intelligent module, which is configured to: create a generation task based on the input content; and send the generation task to the first intelligent module to generate the first interactive content.
[0144] In some cases, the first interface includes a session interface with a virtual object, and the first receiving module is further configured to: receive a first message from a user in the session interface; present a second message in the session interface, the second message being from a second intelligent module, the second message being used to obtain information associated with the generation of the first interactive content; and receive a third message from a user in the session interface, wherein the input content includes the first message and the third message.
[0145] In some cases, the first receiving module is also configured to present a second message in response to a first message representing a need to create first interactive content.
[0146] In some cases, the providing module is also configured to: generate second interactive content based on the execution results of multiple tasks; and trigger the regeneration of the first interactive content based on structured information in response to the failure of the evaluation information of the second interactive content to meet the conditions.
[0147] In some cases, the device 500 also includes an editing module configured to: receive an editing request associated with the first interactive content; and provide third interactive content generated based on the editing request and the first interactive content.
[0148] In some cases, the editing module is also configured to receive a third message from the user on the first interface, the third message representing an editing request.
[0149] In some cases, the editing module is also configured to: present a second interface that presents at least one parameter for generating the first interactive content; and receive an editing operation via the second interface that is used to edit at least one parameter.
[0150] In some cases, the first interactive content includes at least one virtual element, and the first interactive content is associated with interaction logic that describes how the at least one virtual element is interacted with.
[0151] In some cases, the providing module is also configured to: present a fourth message on a first interface, the fourth message describing how the first interactive content was created, the fourth message being generated based on the input content; and provide the first interactive content in response to receiving confirmation of the fourth message.
[0152] In some cases, the first receiving module is also configured to: receive a first operation, the first operation representing the creation of interactive content of a first type; present an input template corresponding to the first type on a first interface; and receive input content based on the input template.
[0153] The modules included in device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some cases, one or more modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 500 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.
[0154] Figure 6 A block diagram of an electronic device 600 in which one or more examples may be implemented is shown. It should be understood that... Figure 6 The electronic device 600 shown is merely exemplary and should not be construed as limiting the functionality and scope of the examples described herein. Figure 6 The illustrated electronic device 600 can be used to implement the electronic device 110 discussed above.
[0155] like Figure 6 As shown, electronic device 600 is in the form of a general-purpose electronic device. Components of electronic device 600 may include, but are not limited to, one or more processing units or processors 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processor 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 600.
[0156] Electronic device 600 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof). Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 600.
[0157] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 6 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various examples.
[0158] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers, or another network node.
[0159] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0160] A computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. A computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0161] The flowcharts and / or block diagrams of the methods, apparatus, devices, and computer program products referred to herein describe various aspects. It should be understood that each block of the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0162] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0163] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0164] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0165] Various examples have been described above. The foregoing descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for generating content, comprising: Input content is received on the first interface associated with the virtual object; as well as Provide first interactive content, which is generated based on the input content, wherein the first interactive content is generated based on the following process: Based on the input content, structured information is generated, and the structured information describes the scheme for generating the first interactive content; Execute multiple tasks based on the structured information; as well as Based on the execution results of the multiple tasks, the first interactive content is generated.
2. The method according to claim 1, wherein performing multiple tasks based on the structured information comprises: Based on the structured information, the first intelligent module creates the multiple tasks; as well as The first intelligent module executes the multiple tasks using multiple functional units, which are selected from a set of functional units based on the multiple tasks.
3. The method of claim 2, wherein the plurality of functional units comprises at least one of the following: The first functional unit is configured to acquire virtual avatar resources; The second functional unit is configured to acquire virtual environment resources; The third functional unit is configured to generate interactive logic associated with the first interactive content.
4. The method according to claim 3, wherein the virtual avatar resource includes a first avatar resource corresponding to the virtual object, and the virtual object is created based on the user's association information.
5. The method according to claim 2, wherein the first interface is associated with a second intelligent module, and the second intelligent module is configured to: A generation task is created based on the input content; and The generation task is sent to the first intelligent module to generate the first interactive content.
6. The method of claim 5, wherein the first interface includes a session interface with the virtual object, and the receiving of input includes: The first message from the user is received in the chat interface. as well as In the conversation interface, a second message is presented. The second message comes from the second intelligent module and is used to obtain information associated with the generation of the first interactive content. as well as In the session interface, a third message is received from the user, wherein the input content includes the first message and the third message.
7. The method of claim 6, wherein presenting the second message comprises: In response to the first message representing the need to create the first interactive content, the second message is presented.
8. The method according to claim 1, wherein generating the first interactive content based on the execution results of the plurality of tasks comprises: Based on the execution results of the multiple tasks, a second interactive content is generated; as well as If the evaluation information of the second interactive content does not meet the conditions, the first interactive content is regenerated based on the structured information.
9. The method according to claim 1, further comprising: Receive an edit request associated with the first interactive content; as well as A third interactive content is provided, which is generated based on the edit request and the first interactive content.
10. The method of claim 9, wherein receiving the edit request associated with the first interactive content comprises: On the first interface, a third message is received from the user, the third message representing the edit request.
11. The method of claim 9, wherein receiving the edit request associated with the first interactive content comprises: A second interface is presented, which displays at least one parameter used to generate the first interactive content; as well as The editing operation is received via the second interface, and the editing operation is used to edit the at least one parameter.
12. The method of claim 1, wherein the first interactive content includes at least one virtual element, and the first interactive content is associated with interaction logic, the interaction logic describing the interaction mode of the at least one virtual element.
13. The method of claim 1, wherein providing the first interactive content comprises: On the first interface, a fourth message is presented, which describes how the first interactive content was created. The fourth message is generated based on the input content. as well as In response to receiving confirmation of the fourth message, the first interactive content is provided.
14. The method of claim 1, wherein receiving input content on the first interface associated with the virtual object includes: Receive a first operation, the first operation representing the creation of a first type of interactive content; On the first interface, an input template corresponding to the first type is presented; as well as Based on the input template, the input content is received.
15. An apparatus for generating content, comprising: The first receiving module is configured to receive input content on a first interface associated with the virtual object; as well as A module is configured to provide first interactive content, which is generated based on the input content, wherein the first interactive content is generated based on the following process: Based on the input content, structured information is generated, and the structured information describes the scheme for generating the first interactive content; Execute multiple tasks based on the structured information; as well as Based on the execution results of the multiple tasks, the first interactive content is generated.
16. An electronic device comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 14 when executed by the at least one processor.
17. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 14.
18. A computer program product, the computer program product being tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 14.