Method, apparatus, device and product for interface interaction
Patent Information
- Application Number
- CN202610957168.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-22
AI Technical Summary
[0007]应当理解,本内容部分中所描述的内容并非旨在限定本文中的示例的关键特征或重要特征,也不用于限制方案的范围。其它特征将通过以下的描述而变得容易理解。
Smart Images

Figure CN122802755A_ABST
Abstract
Description
Technical Field
[0001] The examples in this article generally relate to the field of computer science, and in particular to methods, apparatuses, devices, and products for user interface interaction. Background Technology
[0002] With the development of computer technology, more and more applications have integrated media content generation functions, allowing users to generate media content. Media content generation technology enables users to create multi-dimensional creative works based on relevant descriptions and other information, greatly lowering the barrier to entry for media content creation. Summary of the Invention
[0003] In a first aspect, a method for interface interaction is provided. The method includes: presenting a first interface displaying a conversation with an intelligent system; presenting a first component on the first interface, the first component being provided by the intelligent system and referring to a first media resource; in response to receiving a first operation on the first component, presenting a second interface displaying at least one attribute of the first media resource; and in response to receiving a second operation on the second interface, performing a first processing on the first media resource.
[0004] In a second aspect, an apparatus for interface interaction is provided. The apparatus includes: a first presentation module configured to present a first interface displaying a conversation with an intelligent system; a second presentation module configured to present a first media resource on the first interface, the first media resource being provided by the intelligent system; a third presentation module configured to present a second interface in response to receiving a first operation on the first media resource, the second interface presenting at least one attribute of the first media resource; and a processing module configured to perform a first processing on the first media resource in response to receiving a second operation on the second interface.
[0005] In a third aspect, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.
[0007] It should be understood that the content described in this section is not intended to limit the key or important features of the examples in this article, nor is it intended to restrict the scope of the solution. Other features will become readily apparent from the following description. Attached Figure Description
[0008] The above and other features, advantages, and aspects of the various examples herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A schematic diagram of the example environment is shown; Figures 2A to 2T Example interfaces for some scenarios are shown; Figure 3 The flowcharts show example processes of interface interactions in some scenarios; Figure 4 Schematic block diagrams of example devices for interface interaction in some scenarios are shown; and Figure 5 A block diagram of an electronic device capable of implementing multiple illustrative scenarios is shown. Detailed Implementation
[0009] The examples in the text will now be described in more detail with reference to the accompanying drawings. While some examples are shown in the drawings, it should be understood that solutions can be implemented in various forms and should not be construed as limited to the examples presented herein. Rather, these examples are provided to provide a more thorough and complete understanding of the solutions. It should be understood that the drawings and examples in this document are for illustrative purposes only and are not intended to limit the scope of protection of the solutions.
[0010] It should be noted that the headings of any section / subsection provided herein are not restrictive. Various examples are described throughout this document, and examples of any type may be included under any section / subsection. Furthermore, examples described in any section / subsection may be combined in any way with any other examples described in the same section / subsection and / or different sections / subsections.
[0011] In the description of the examples in this document, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an example" or "the example" should be understood as "at least one example". The term "some examples" should be understood as "at least some examples". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0012] The examples in this article may involve user data, data acquisition, and / or use. All of these aspects comply with relevant laws, regulations, and rules. In the examples, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, when implementing each example, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained through appropriate means, in accordance with relevant laws and regulations. The specific methods of notification and / or authorization can vary depending on the actual situation and application scenario; the scope of the solution is not limited in this regard.
[0013] In this manual and the sample solutions, any processing of personal information will be conducted only under legal grounds (such as obtaining the consent of the data subject or being necessary for the performance of a contract) and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.
[0014] As used in this document, the term "intelligent system" refers to a system capable of autonomous control based on machine learning models. An intelligent system is, for example, a virtual object or physical entity that can make decisions and autonomously execute actions based on machine learning models to achieve preset goals or complete preset tasks. An intelligent system can be an automated program that understands user intent and can utilize models or invoke tools to complete various types of tasks. In some contexts, examples of intelligent systems may include, but are not limited to: agents, bots, chatbots, digital avatars, intelligent customer service, digital assistants, etc. Alternatively, an intelligent system can also be an intelligent role implemented based on machine learning models. An "intelligent system" can process user requests based on generative models (e.g., language models, multimodal models) to perform specified types of tasks. In some cases, an intelligent system may also relate to virtual accounts, which may have corresponding avatars or nicknames.
[0015] The term “first interface” as used in this article refers to the interface used to display the conversation between the user and the intelligent system, which may include, but is not limited to, presenting historical conversations, reply messages from the intelligent system, and dialog panels of various components, which typically include conversation input boxes for receiving user input.
[0016] The term "second interface" as used in this document refers to a component used to present at least one attribute of a media resource and to support viewing, editing, adapting, regenerating and / or publishing the media resource. For example, it may include, but is not limited to, one or more workbench forms such as a media resource details panel, a media resource editor panel, or a lyrics editor panel.
[0017] The term "third interface" as used in this document refers to an interface used to present at least one session entry point, such as, but not limited to, a dialog history page that presents historical session entry points in an asset interface; in response to a triggering operation on a session entry point therein, the corresponding session can be entered and the first interface can be presented.
[0018] The term "fourth interface" as used in this article refers to an interface used to present multiple resources, such as, but not limited to, a page in the asset interface that presents generated song resources, lyrics resources, etc.; at least one of the multiple resources can be referenced or edited to initiate new creations.
[0019] As used in this document, the terms "first component" and "second component" refer to components presented in the first interface, provided by the intelligent system, and representing corresponding media resources. These may include, but are not limited to, components presented in card form during a conversation that represent song resources or lyrics resources. Specifically, the first component refers to the first media resource, and the second component refers to the second media resource; operations on the first or second component can be considered operations on the media resources they represent.
[0020] The term "third component" as used in this article refers to a component provided by the intelligent system in the first interface for obtaining one or more parameters, such as, but not limited to, one or more of the following: style selection component, timbre selection component, lyrics style selection component, continuation setting component, reference weight setting component, etc.
[0021] The term "media resources" as used in this article refers to resources associated with media content, typically including but not limited to song resources, lyrics resources, etc.; media resources can be generated by intelligent systems, or provided or uploaded by users.
[0022] The term "attribute" as used in this article refers to information that a media resource possesses that can be presented or edited, such as, but not limited to, one or more of the following: style description, lyrics, timbre, generation method, duration, cover art, and model identifier.
[0023] As used in this document, the term "action" refers to an interactive behavior performed by a user on a presented object through an interface, which may include, but is not limited to, one or more of the following: clicking, long pressing, swiping, dragging, text input, voice input, and gestures; the term "request" refers to an instruction triggered by the user's input or action to instruct the intelligent system to perform corresponding processing; and the term "message" refers to conversational content presented in the first interface that originates from the user or from the intelligent system.
[0024] The term “publishing” as used in this article refers to the process by which media resources are made available, distributed, or made public, including but not limited to submitting media resources to the distribution chain, making them publicly available, or sharing them.
[0025] The term "style parameter" as used in this article refers to parameters used to characterize the style of media content, such as including but not limited to pop, rock, folk, etc.; the term "timbre parameter" refers to parameters used to characterize the timbre of singing or playing, such as including but not limited to the timbre of different virtual singers or human voices; the term "generation method" refers to the method used to generate media resources, such as including but not limited to the generation model or generation mode adopted.
[0026] The term "lyrical style" as used in this article refers to information used to characterize the style of lyrics, which may include, but is not limited to, lyrics with different themes, sentence structures, or language styles.
[0027] As used herein, the term "first processing" refers to a subsequent processing of a first media resource in response to a second operation, based on the data, parameters, or selection results generated by the second operation. First processing may include, but is not limited to, at least one of the following: publishing the first media resource, editing the lyrics of the first media resource, editing the style description of the first media resource, editing the first media resource, selecting a reference segment from the first media resource, and adjusting the reference weight of the reference segment.
[0028] As mentioned above, with the development of computer technology, more and more applications have integrated media content generation functions, allowing users to generate media content. Media content generation technology enables users to generate multi-dimensional creative works based on relevant descriptions and other information, which greatly lowers the barrier to entry for media content creation.
[0029] A user interface interaction scheme is proposed. The scheme includes: presenting a first interface, the first interface displaying a conversation with an intelligent system; presenting a first component on the first interface, the first component being provided by the intelligent system and referring to a first media resource; in response to receiving a first operation on the first component, presenting a second interface, the second interface presenting at least one attribute of the first media resource; and in response to receiving a second operation on the second interface, performing a first processing on the first media resource.
[0030] In this way, users can obtain media resources provided by the intelligent system within the same session, and directly view the attributes and complete publishing, editing, and other processes via a second component that presents the media resource attributes. This consolidates the acquisition, attribute viewing, and publishing of media resources into a coherent interactive chain that coordinates the dialog panel and the workbench, reducing the number of jumps and operation steps for users between different functional pages and improving human-computer interaction efficiency. Furthermore, since the attribute viewing and publishing processes reuse the same session and workbench structure, there is no need to load and render independent pages for each process, which reduces the client's page rendering overhead and interface refresh frequency, reduces the client's cache usage, and reduces the number of page requests sent to the server, thereby reducing the server's query load and the bandwidth usage of end-to-cloud communication.
[0031] The following describes various examples of this scheme in further detail with reference to the accompanying drawings.
[0032] Example Environment Figure 1 A schematic diagram of example environment 100 is shown. (e.g.) Figure 1 As shown, example environment 100 may include electronic device 110.
[0033] In this example environment 100, electronic device 110 can run an application 120 that supports user interface interaction. Application 120 can be any suitable type of application for user interface interaction, including but not limited to: music creation applications or other suitable applications. User 140 can interact with application 120 via electronic device 110 and / or its attached devices.
[0034] exist Figure 1 In environment 100, if application 120 is active, electronic device 110 can use application 120 to present interface 150 for supporting interface interaction.
[0035] In some cases, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some cases, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).
[0036] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 that support user interface interaction in electronic devices 110.
[0037] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection can include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections. In some cases, server 130 and electronic device 110 can exchange signaling information through their communication connection.
[0038] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the scheme.
[0039] The following description of the example will continue with reference to the accompanying drawings.
[0040] Example Interaction The following combination Figures 2A to 2TThe following are examples of the aforementioned solutions based on a complete music creation interaction process: Users enter a conversation with the intelligent system through the conversation portal, express their creation requests in the conversation, and are guided by the intelligent system to select the creation method. Then, with the help of various components provided by the intelligent system, parameters such as style, timbre, and lyrics are gradually provided. After confirmation, media resources are generated and presented in the conversation. The attributes of the media resources are then viewed and published through the second component, which serves as a workbench. The media resources can also be adapted and continued in the second component, and the accumulated resources can be managed and reused in the asset interface.
[0041] In these examples, the entity performing each operation follows the same principle. Figure 1 The objects and their corresponding reference numerals are already defined in the document (e.g., electronic device 110, application 120, server 130, user 140). It is understood that the examples below can be implemented individually or sequentially in the order described above.
[0042] The basic idea behind the aforementioned solution is to consolidate the entire music creation process into a single conversation between the user and the intelligent system, enabling the user to gradually express and adjust their creative intentions through continuous dialogue, without having to repeatedly switch between disparate functional pages.
[0043] like Figure 2A As shown, interface 200A can be a new music creation interface, which includes a conversation input box 202. The conversation input box 202 may include mode selection (e.g., mode A), model selection (e.g., model A), upload controls, and send controls; the lower area of interface 200A may also display inspiration cards for providing creative inspiration. In response to receiving input from user 140 via the conversation input box 202, electronic device 110 initiates a conversation and presents a first interface showcasing the conversation with the smart system.
[0044] How users express their creative requests after entering a session and how the intelligent system guides them to choose a creative method will be referenced. Figure 2B Description. (In the user's...) Figure 2A Once in the conversation, users need to further express their specific creative requests. Considering that users often haven't yet formed a clear and complete creative decision at the beginning of the creative process, the intelligent system can respond to the user's creative request by presenting multiple entry points corresponding to different creative dimensions. This allows users to select a creative dimension through multiple entry points and present the components corresponding to the respective creative dimension, thus helping users create primary media resources.
[0045] In some cases, a second request is received on a first interface; in response to the second request, multiple entry points corresponding to different ways of creating media resources are presented; and in response to a triggering operation on the first entry point among the multiple entry points, a third component is presented, the first entry point corresponding to the first method, and the type of the third component being associated with the first method.
[0046] like Figure 2B As shown, the electronic device 110 can receive natural language input from the user 140 (e.g., message 204 "I want to create a song, with a unique and distinctive style") as a second request via a first interface. In response to the second request, the intelligent system presents multiple entry points corresponding to different creation methods on the first interface, such as a first entry point 206-1 corresponding to the method of selecting a song style first, an entry point 206-2 corresponding to the method of filling in lyrics first, and an entry point 206-3 corresponding to the method of directly generating a song. In response to a trigger operation on the first entry point 206-1, a third component of the type associated with that method will be presented, such as the following text. Figure 2C The style selection component shown.
[0047] In this way, users can quickly select the creation method through multiple entry points provided by the intelligent system, reducing the number of rounds of natural language back-and-forth dialogue to clarify creative intent; and since each round of natural language clarification usually requires triggering a model inference and an end-to-cloud interaction, replacing repeated natural language clarification with structured multiple entry points can reduce the number of model inferences and the number of rounds of end-to-cloud interactions, thereby reducing the computing power load on the server and the bandwidth usage of end-to-cloud communication.
[0048] Taking the method of first determining the song style as an example, in response to the triggering of the first entry point 206-1, a third component for obtaining style parameters will be presented, see [link to relevant documentation]. Figure 2C In some cases, a third component is provided by the intelligent system, whose component type is associated with the initiated request and characterizes the type of parameters it is used to acquire; and a fourth operation is received through the third component to acquire at least a portion of the parameters.
[0049] like Figure 2CAs shown, in response to the triggering of the first entry 206-1, the electronic device 110 can present message 208 "Select song style first" on the first interface. Further, the intelligent system can determine, based on message 208, that a style selection component needs to be presented. Further, the electronic device 110 can present a style selection component 210, the component type of which (for style selection) is associated with the first method corresponding to the first entry. The third component 210 presents multiple style tags recommended by the intelligent system (e.g., pop 210-1, rock 210-2, etc.) and supports multi-selection; in response to multi-selection and confirmation via the third component 210, the selected style is sent as at least a part of the parameters to the session, and a corresponding message (e.g., "Select pop and rock styles") is presented.
[0050] In some cases, the third component 210 may also display action recommendation entries to guide users to the next creation step (e.g., filling in lyrics 206-4, selecting a timbre 206-5, directly generating a song 206-6). It can be understood that the component type of the third component is associated with the initiated request. For example, when a user requests to fill in lyrics first, a third component for obtaining lyrics is presented accordingly, thus adapting the presented third component to the type of parameter to be obtained.
[0051] In this way, users can directly provide style parameters in a structured multi-selection format, reducing the number of dialogue rounds required to repeatedly describe and clarify the style through natural language. Furthermore, since a single structured confirmation replaces multiple rounds of natural language clarification, it reduces the number of intent parsing and model inference performed by the intelligent system and the number of rounds of edge-cloud interaction, thereby reducing the computing load on the server and the bandwidth usage of edge-cloud communication.
[0052] In addition to style parameters, timbre parameters can also be obtained through the corresponding type of third-party component, see [link to relevant documentation]. Figure 2D .
[0053] and Figure 2C Corresponding to the acquisition of stylistic parameters, timbre parameters can also be acquired via a third component of the appropriate type provided by the intelligent system, thereby enabling the parameters acquired through the third component to cover the timbre dimension. Therefore, in some cases, the parameters acquired through the third component may include timbre parameters.
[0054] like Figure 2DAs shown, the electronic device 110 can present multiple timbre selection components (including components 212-1 to 212-3), each corresponding to a different timbre. For example, component 212-1 corresponds to timbre A, component 212-2 corresponds to timbre B, and component 212-3 corresponds to timbre C. Users can select a timbre using these components to create a first media resource based on the selected timbre. For instance, when component 212-1 is selected, the intelligent system can generate a first media resource based on the selected timbre A.
[0055] In some scenarios, the timbre selection component can also include descriptive information about the corresponding timbre. For example, component 212-1 includes the name and description (e.g., warm, clean, etc.) of timbre A. In other scenarios, electronic device 110 can also present playback controls within the timbre selection component. For example, electronic device 110 can present playback controls 214 within component 212-1. When the playback controls are clicked, electronic device 110 can play the audio content corresponding to timbre A. In other scenarios, electronic device 110 can also present a "More Timbres" entry 216. Entry 216 can be used to present more timbre selection components. In some cases, when timbre parameters are lacking, the intelligent system can query and present the timbre selection component, and the presentation order can be arranged according to the conversation context to prioritize recommending timbres more suitable for the current song style.
[0056] In this way, users can directly select and listen to timbres in a structured manner, reducing the need for natural language descriptions and repeated confirmations for a specified timbre. Furthermore, since the timbre parameters are determined by a single structured selection, the number of related intent parsing and model inferences performed by the intelligent system is reduced, as are the number of rounds of edge-cloud interaction, thereby reducing the computing load on the server and the bandwidth usage of edge-cloud communication.
[0057] In addition to parameters such as style and timbre, lyrics can be obtained by first having the user select the lyrics style, see [link to relevant documentation]. Figure 2E After determining the style and timbre, the lyrics can be further determined. Considering that different lyric styles significantly affect the content of the generated lyrics, one concept of the aforementioned scheme is to first allow the user to choose from multiple candidate lyric styles, and then generate lyrics in the corresponding style, making the direction of lyric generation more focused. To this end, in some cases, a first interface presents multiple options corresponding to different lyric styles; a fifth operation is received to select the first option from the multiple options, the first option corresponding to the first lyric style; and based on this, a third component presents the third lyrics generated based on the first lyric style.
[0058] like Figure 2EAs shown, the first interface presents multiple options corresponding to different lyric styles, such as option 218-1 for lyric style A, option 218-2 for lyric style B, and option 218-3 for lyric style C. In response to the selection of one of these options (e.g., option 218-1 for lyric style A), a first lyric style is determined; accordingly, the intelligent system generates third lyrics based on the first lyric style and presents them in the third component (see [link to documentation]). Figure 2F In some cases, the order of multiple lyric style options can be adjusted based on the conversation context, placing the more suitable lyric style first.
[0059] In this way, users can directly select from multiple candidate lyric styles in a structured manner, reducing the need for natural language descriptions to express lyric style preferences. Furthermore, since a single structured selection can trigger the generation of lyrics for the corresponding lyric style, it can reduce the number of model inferences and edge-cloud interaction rounds used for intent clarification, thereby reducing the server's computing load and the bandwidth usage of edge-cloud communication.
[0060] The presentation and confirmation of the third lyrics generated based on the selected first lyric style, see [link to documentation]. Figure 2F . Undertake Figure 2E The selected first lyric style is used by the intelligent system to generate third lyrics, which are then presented in a third component for user confirmation or editing, thus avoiding the need for users to write lyrics from scratch. Therefore, in some cases, the third lyrics generated based on the first lyric style are presented in the third component; and a fourth operation is received through the third component, used to confirm or edit the third lyrics.
[0061] like Figure 2F As shown, based on Figure 2E The selected first lyric style (e.g., lyric style A) is used by the intelligent system to generate third lyrics through a lyric-writing model, and presented in two versions, such as card 220-1 for lyric A and card 220-2 for lyric B. In response to receiving confirmation or editing of the third lyrics via a third component (i.e., the fourth operation), the subsequent creation process begins. In some cases, action recommendation controls (e.g., "more rock" or "gentler lyrics") can be displayed below the third lyrics, prompting the intelligent system to adjust the lyrics accordingly.
[0062] In this way, users can directly confirm or edit multiple candidate lyrics versions, reducing the multiple rounds of back-and-forth interaction required to generate from scratch. Furthermore, since confirmation or partial editing is based on already generated candidate lyrics, the number of times the model is regenerated and the scale of input and output for model inference can be reduced, thereby reducing the computing load on the server and the bandwidth usage of edge-cloud communication.
[0063] Further editing of the lyrics can be done in the second interface, which serves as the workbench. See [link / reference] Figure 2G Whether Figure 2F Whether the lyrics generated in the first interface are those corresponding to the first media resource, or the lyrics corresponding to the already generated media resource, they can all be viewed and edited in the second interface, which serves as the workbench. One concept behind the aforementioned solution is to consolidate the presentation and editing of lyrics into the same second interface that presents the attributes of the media resource, thereby avoiding switching between different pages. Therefore, in some cases, at least one first parameter received via the second interface may include lyrics. The second interface presents the first lyrics corresponding to the first media resource and receives editing operations on the first lyrics to obtain the second lyrics, thereby generating the second media resource.
[0064] like Figure 2G As shown, when the corresponding lyrics are generated or clicked, the electronic device 110 can display a lyrics editor panel 224 in the interface 200G. The lyrics editor panel 224 can display an editing component 226. The editing component 226 displays the lyrics text of the third lyrics. The lyrics text 226 can correspond to the third lyrics to be confirmed or edited. If the first media resource is generated based on the current third lyrics, then the third lyrics can also correspond to the first lyrics. The electronic device 110 can display a confirmation control (a control displaying the text "Generate Song") in the lyrics editor panel 224. The electronic device 110 can receive confirmation of the third lyrics via the confirmation control. In some scenarios, users can edit the third lyrics via the lyrics editor panel 224. The specific editing method will be referred to... Figure 2H describe.
[0065] In this way, users can directly view and edit lyrics in the same workbench that displays song attributes, reducing the steps of switching between different pages to edit lyrics. Furthermore, since the presentation and editing of lyrics reuse the same workbench panel, there is no need to load a separate page for lyrics editing, which reduces the page rendering overhead and interface refresh frequency on the client side, and reduces the number of page requests sent to the server, thereby reducing the server's query load and the bandwidth usage of end-to-cloud communication.
[0066] exist Figure 2G Building upon the lyrics editing shown, to reduce the burden on users of rewriting entire lyrics, one concept of the aforementioned solution is to allow users to select only a portion of the lyrics and replace it locally, while leaving the rest unchanged. To this end, in some cases, in response to receiving a third operation for selecting a first portion of the first lyrics, a first control is presented; and in response to receiving first content via the first control, the first portion is replaced with a second portion, which is derived based on the first content.
[0067] like Figure 2H As shown, in the lyrics editor panel 224, in response to the selection of a single phrase in the lyrics (i.e., the third operation, selecting the first part 228), a first control 230 is presented. The first control 230 may include several configuration items and an input area for receiving the first content. In response to receiving the first content via the first control 230, the selected first part 228 is replaced with a second part obtained based on the first content. Here, the first content can be a prompt word for generating the second part for the user. In some cases, the second part can be generated by an intelligent system based on the first content and the context of the lyrics using a lyric-writing model. It is understood that, in addition to the above-mentioned method of editing single phrases, users can also manually modify entire sections of lyrics without restriction.
[0068] In this way, users can replace only the selected lyrics without rewriting the entire lyrics, reducing the number of editing steps for users. Furthermore, since generation is triggered only for the selected part, the length of the input and output of the model inference can be reduced, the corresponding amount of computation can be reduced, and the amount of data transmitted between the end and the cloud can be reduced, thereby reducing the computing load on the server and the bandwidth usage of the end-to-cloud communication.
[0069] After collecting parameters such as style, timbre, and lyrics, you can confirm these parameters before the final generation. See [link / reference]. Figure 2I In intelligent systems via Figures 2C to 2H After collecting parameters such as style, timbre, and lyrics, the interactive process proceeds to a confirmation stage before formal generation. One concept behind the aforementioned solution is that the intelligent system summarizes and displays the parameters collected in the preceding dialogue, allowing the user to confirm before generation, thereby avoiding invalid generation due to parameter deviations. To this end, in some cases, a second request is received on the first interface; a second message from the intelligent system, displaying at least one second parameter, is presented on the first interface; and in response to receiving confirmation of the second message, a first media resource generated by the intelligent system based on at least one second parameter is presented.
[0070] like Figure 2I As shown, via the aforementioned three components (see...) Figures 2C to 2H After collecting second parameters such as style, lyrics, and timbre, the intelligent system presents a second message on the first interface. This second message may include, for example, a confirmation card 232. The confirmation card 232 displays at least one second parameter (e.g., style "pop, rock," lyrics, and timbre A). In response to receiving confirmation of the second message (e.g., triggering the "Generate Song" control 234), the intelligent system generates a first media resource based on at least one second parameter and presents it on the first interface (see [link to documentation]). Figure 2JIt is understood that the second message 232 summarizes the various second parameters collected in the preceding dialogue, thus providing them for user confirmation before generation. In some scenarios, the electronic device 110 can also receive an editing operation on at least one of the second parameters via the confirmation control 232, allowing the second parameters to be edited before generating the first media resource.
[0071] In this way, users can centrally confirm the collected parameters before generation, reducing regeneration and repeated modifications caused by parameter deviations; and since generation is triggered only after confirmation, invalid generation requests and invalid model inferences can be reduced, thereby reducing the computing load on the server and the bandwidth usage of edge-cloud communication.
[0072] Accept Figure 2I The confirmation of parameters is handled by the intelligent system, which generates a first media resource based on the collected parameters and presents it directly within the session. One concept behind the aforementioned scheme is to embed the generated media resource within the same session as the intelligent system, facilitating viewing and subsequent processing by the user within the continuous session context, without requiring switching between generation and viewing. To this end, in some cases, a first component is presented, referring to the first media resource, and this first component is provided by the intelligent system.
[0073] In some cases, the primary media resource (e.g., the audio of a song) is generated by an intelligent system based on at least one secondary parameter using an appropriate model. For example... Figure 2J As shown, the intelligent system can provide a first media resource based on the aforementioned collected second parameters and trigger the electronic device 110 to present a first component on the first interface, such as in the form of a song card. In some cases, the intelligent system can generate multiple media resources based on at least one second parameter and present multiple media resources (including first component 234-1 and component 234-2) on the interface 200J through the electronic device 110. The multiple media resources include the first media resource (e.g., the media resource corresponding to the first component 234-1). The first component 234-1 presents the cover art, song title, duration, model identifier, collection control, and other operation controls, etc. After the first media resource is generated, the electronic device 110 can present a global playback control bar at the bottom of the interface 200J to support the user in playing the first media resource.
[0074] In this way, users can directly obtain the generated media resources in the same session with the intelligent system, reducing page jumps between generation and viewing; and since the media resources are presented directly in the session in the form of cards, there is no need to load a separate page for viewing results, which can reduce the page rendering overhead and interface refresh frequency on the client side, and reduce the number of page requests sent to the server, thereby reducing the server's query load and the bandwidth usage of end-to-cloud communication.
[0075] The process of users viewing, editing, and publishing attributes of primary media resources can be found in [link to documentation]. Figure 2K .exist Figure 2J After obtaining the first media resource, users typically need to further examine its attributes and decide whether to edit it further. A basic concept of the aforementioned solution is that, in response to an operation on the first media resource, a second interface, acting as a workbench, is deployed to present its attributes and supports first processing of the first media resource within this second interface, thereby integrating the viewing and processing of media resources into the same session context. First processing refers to the subsequent processing of the first media resource based on the data, parameters, or selection results generated by the second operation, in response to a second operation. First processing may include, but is not limited to, at least one of: publishing the first media resource, editing the lyrics of the first media resource, editing the style description of the first media resource, editing the first media resource, selecting a reference segment from the first media resource, and adjusting the reference weight of the reference segment. Therefore, in some cases, in response to receiving a first operation on the first media resource, a second interface is presented, displaying at least one attribute of the first media resource; and in response to receiving a second operation on the second interface, first processing is performed on the first media resource. It can be understood that the aforementioned... Figures 2A to 2J The described creation and presentation process, together with the process of viewing attributes and performing the first processing on the first media resource described in this diagram, constitute a complete interactive link from entering the session to performing the first processing on the media resource.
[0076] like Figure 2KAs shown, in response to a first operation on the first component 234-1 (representing the first media resource, such as song A) (e.g., clicking on the component), a second interface is presented in interface 200K. The second interface may be, for example, an editing panel 236, which displays at least one attribute of the first media resource (e.g., style description, lyrics, etc.) and may include a "Create Lyrics" entry. In response to receiving a second operation (e.g., triggering a "Publish Song" control) in the second interface (e.g., editing panel 236), a publishing interface for the first media resource is presented. The electronic device 110 can receive publishing parameters for the first media resource via the publishing interface. For example, it can receive publishing scope, title, work description, etc., via the publishing interface. After receiving confirmation of the publishing parameters, the electronic device 110 can publish the first media resource.
[0077] In this way, users can directly view and publish media resources within the same interface that displays attributes, reducing the number of jumps and operation steps required between different pages such as the details page and the publishing page for viewing attributes and publishing. Furthermore, since attribute viewing and publishing reuse the same workbench panel, without the need to load separate pages for viewing details and publishing, it can reduce the client's page rendering overhead and the number of interface refreshes, reduce the client's cache usage, and reduce the number of page requests sent to the server, thereby reducing the server's query load and the bandwidth usage of end-to-cloud communication.
[0078] exist Figure 2K Based on the viewing attributes shown, the second interface can be further adjusted to an editing state to support users in adapting media resources. One concept of the aforementioned solution is to receive parameter adjustment and adaptation requests for existing media resources via the second interface, which acts as a workbench, and to clearly specify the referenced media resource and request type in the form of a message on the first interface, thereby maintaining a clear conversational flow during the adaptation process. To this end, in some cases, the second interface receives a first request associated with a first media resource, the first request representing at least one first parameter received via the second interface; the first interface presents a first message corresponding to the first request, indicating a reference to the first media resource and representing the first request type; and the first interface presents a second media resource obtained based on the first request and the first media resource, the second media resource being generated by the intelligent system based on the first media resource and at least one first parameter.
[0079] like Figure 2LAs shown, in response to activating the edit state (e.g., triggering the "Modify" control 238), the second interface (e.g., edit panel 236) is adjusted to the edit state. When the second interface is in the edit state, the electronic device 110 can display multiple editing components in the interface 200L. The multiple editing components can include components 240 to 246-2. Different editing components are used to edit different parameters. For example, component 240 is used to edit the style description 240, component 242 can be used to edit lyrics, and component 244 can be used to adjust the generation method (e.g., model selection, shown as the "Model A" dropdown in the figure). Components 246-1 and 246-2 can correspond to different timbres to support the user in selecting the corresponding timbres. In some cases, component 242 can display the first lyrics corresponding to the first media resource, and its editing result can be used as the second lyrics. The song editor can also have a "Create Lyrics" entry. After the user edits at least one first parameter, the electronic device 110 can present a first message 249 (e.g., "Modify Song A") on a first interface. The first message 249 indicates a reference to the first media resource (Song A) and characterizes the type of the first request (e.g., modification or adaptation). In response to receiving the first request (e.g., a trigger for generation) via a second component, the intelligent system generates a second media resource based on the first media resource and at least one first parameter. It is understood that the at least one first parameter received via the second component may include at least one of a style parameter, a timbre parameter, and a generation method. In some scenarios, the editing panel 236 also presents multiple editing mode selection labels, such as mode A, mode B, mode C, etc. Different editing modes may correspond to different editing components. For example, when mode A is selected, the electronic device 110 can present a component for adjusting the relevance in the editing panel 236. Here, the relevance can represent the similarity between the second media resource to be generated and the first media resource, which can represent the reference weight of the first media resource when the intelligent system generates the second media resource. For example, when mode B is selected, the electronic device 110 can present a segment selection component in the editing panel 236, which is used to select a target segment of the first media resource as a reference for generating the second media resource.
[0080] In this way, users can directly adjust parameters and initiate regeneration within the same workbench that presents media resource attributes. The referenced media resources and request types are clearly identified through the first message in the first interface, reducing the number of steps involved in switching between different pages and repetitive input during the adaptation process. Furthermore, since the adaptation is based on existing first media resources and reuses the same workbench and session structure, it reduces the amount of data that needs to be re-provided and transmitted, reduces the processing scale of model regeneration, and reduces the number of page requests sent to the server, thereby reducing the server's computing load, query load, and bandwidth usage for end-to-cloud communication.
[0081] Figure 2L The song editor shown can include a "New Lyrics" control 241 to support creating and editing copies of lyrics from an existing song during the adaptation process. This example serves as... Figure 2G , Figure 2H The lyrics editing is a continuation of the adaptation scenario, so that the creation and editing of lyrics are also converged in the workbench.
[0082] like Figure 2M As shown, lyrics editing can be accessed through the "New Lyrics" function in the song editor. A copy of the lyrics 252 is created and displayed in the workbench, and the edited lyrics are then filled back into the lyrics area of the song editor via the "Complete Editing and Return" control 254 (see [link]). Figure 2L Lyrics area 242). When creating a new copy of lyrics, a corresponding session can also be sent to the intelligent system (e.g., "Create a copy of lyrics for song A"). Workbench tab 250 can be used to switch between song details, song arrangement, and lyrics editing modes. In some cases, as an edge case, if the song editor is closed or covered by another song editor, the "Complete editing and return" function will no longer be displayed, and the "Generate song" function at the bottom of the lyrics editor will be restored.
[0083] In this way, users can create copies of lyrics based on existing songs in the workbench and refill them after editing, reducing the steps of switching between lyrics and songs and repeatedly pasting. Furthermore, since the creation, editing and refilling of lyrics reuse the same workbench structure and there is no need to load a separate page for lyrics editing, it can reduce the page rendering overhead and interface refresh frequency on the client side, and reduce the number of page requests sent to the server, thereby reducing the server's query load and the bandwidth usage of end-to-cloud communication.
[0084] After adjusting parameters and initiating the adaptation via the second interface, the second media resource generated by the intelligent system can be found here. Figure 2N .
[0085] Accept Figure 2LThe adaptation request initiated via the second interface is used by the intelligent system to generate a second media resource based on the first media resource and the adjusted parameters, and then presented in the session for comparison with the original media resource. Therefore, in some cases, after receiving the first request via the second interface, the second media resource is presented on the first interface; the second media resource is derived from the first request and the first media resource.
[0086] like Figure 2N As shown, via the second interface (see...) Figure 2L Upon receiving the first request, the system presents a second media resource on the first interface, such as card 255-1 for song C and card 255-2 for song D. The second media resource is generated by the intelligent system based on the first media resource and at least one first parameter. It is understood that the second media resource and the first media resource are presented correspondingly in the same session, thus facilitating comparison between the media resources before and after the adaptation.
[0087] In this way, users can obtain and compare media resources regenerated from existing media resources within the same session, reducing the number of steps required to switch between the adapted results and the original media resources. Furthermore, since the regenerated media resources are directly presented in the session without the need to load a separate page for them, it reduces the client's page rendering overhead and the number of interface refreshes, as well as the number of page requests sent to the server, thereby reducing the server's query load and the bandwidth usage of end-to-cloud communication.
[0088] In addition to the aforementioned adaptations based on parameter adjustments, the proposed solution also supports the continuation of existing media resources; see [link to relevant documentation]. Figure 2O .and Figures 2L to 2N The aforementioned parameter-adjustment-based adaptation method also supports the continuation of existing media resources. One concept for continuation is that the user selects a segment of an existing media resource as the basis for continuation, and then the intelligent system generates follow-up content after that segment, thus reusing existing audio without requiring a complete rewrite. To this end, in some cases, a third media resource is presented on the first interface; a sixth operation is received to select a first segment of the third media resource; and a second request associated with the first segment is received. The resulting first media resource includes both the first and second segments, with the second segment derived from the first segment.
[0089] like Figure 2OAs shown, message 256 is presented in the first interface, and message 256 includes identifier 256-1. Identifier 256-1 corresponds to the third media resource. Further, the electronic device 110 can present the third media resource 258 in the interface 200O. The third media resource 258 can display original song information (e.g., song title, model), waveform, and confirmation control. In response to a sixth operation performed via the waveform or continuation start point selection control 260 (e.g., dragging a point on the waveform or manually entering the continuation start point), a first segment of the third media resource 258 is selected; and a second request associated with the first segment is received, such a second request may, for example, instruct the generation of the first media resource based on the selected first segment. The first media resource generated thereby includes the selected first segment and a second segment (i.e., the continuation part) obtained based on the first segment, the second segment being generated by the intelligent system based on the first segment via a generative model.
[0090] In this way, users can directly select segments and continue writing based on existing songs, reducing the steps required to recreate the song. Furthermore, since the continuation is based on the selected first segment and reuses existing audio content, it can reduce the amount of audio data that needs to be generated and transmitted, reduce the processing scale of the continuation model, and thus reduce the computing load on the server and the bandwidth usage of the end-to-cloud communication.
[0091] exist Figure 2O In scenarios such as song continuation, covers, and imitations with reference audio, it is often necessary to further control the degree of adherence to the reference audio. This example, as another implementation of parameter acquisition, uses an interactive component to set reference weights, presented by the intelligent system, to obtain the corresponding weight parameters.
[0092] like Figure 2P As shown, in scenarios with reference audio, such as cover songs, imitations, or sequels, the electronic device 110 can present a weight setting component 264. The weight setting component 264 can include multiple weights (such as weight A, weight B, and weight C, which correspond to style weight, reference audio weight, and degree of creative divergence, respectively), and can be adjusted by dragging a slider or inputting values, and sent to the session via "confirmation".
[0093] In this way, users can directly set parameters such as reference weights in a structured manner, reducing the need for repeated natural language descriptions to express similarity preferences. Furthermore, since the parameter can be provided with a single structured confirmation, the number of model inferences and edge-cloud interaction rounds used for parameter clarification can be reduced, thereby reducing the server's computing load and the bandwidth usage of edge-cloud communication.
[0094] The foregoing Figures 2A to 2P It primarily describes the process of creating, adapting, and continuing media resources within a single session; excluding through... Figure 2A In addition to entering a session through the conversation input box, you can also access historical sessions or reuse existing resources from the asset interface. See [link to relevant documentation]. Figures 2Q to 2T .
[0095] like Figure 2A As mentioned above, the session entry point to the first interface is not limited to the session input box 202 in the new music interface. One concept of the aforementioned solution is that historical session entry points can also be presented in the first interface that aggregates user assets, allowing users to resume and continue their previous creations. To this end, in some cases, at least one session entry point is presented in the third interface; a trigger operation is received on the first session entry point among the at least one session entry point; and the first interface is presented.
[0096] like Figure 2Q As shown, under the "Conversation History" tab 266 of the asset interface (an implementation of a third interface), one or more historical conversation entries (e.g., conversation cards 268) are presented. In response to a trigger operation on the first conversation entry, the corresponding conversation is entered, and the first interface displaying the conversation with the intelligent system is presented, allowing navigation to the latest message. It can be understood that the conversation entries described in this example are related to... Figure 2A The corresponding conversation input box 202 can all serve as the conversation entry point for entering the first interface.
[0097] In this way, users can directly access and continue their previous work from the asset interface, reducing the steps required to find and restore previous creations. Furthermore, since the entry point is anchored to the corresponding location and the existing session structure is reused, the amount of content that needs to be reloaded and rendered is reduced, as well as the number of requests sent to the server, thereby reducing the server's query load and the bandwidth usage of end-to-cloud communication.
[0098] In addition to historical sessions, the third interface can also aggregate generated media resources for reuse; see [link to relevant documentation]. Figure 2R .exist Figure 2Q Beyond the dialogue logs shown, the asset interface can also aggregate multiple generated media resources. One concept behind the aforementioned solution is that users can directly select resources in a fourth interface that aggregates multiple resources and initiate new creations accordingly, thus reusing existing work without starting from scratch. To this end, in some cases, multiple resources are presented in the fourth interface, including a first media resource; a third request associated with at least one of the multiple resources is received; and the fourth media resource obtained based on the at least one resource and the third request is presented in the first interface.
[0099] like Figure 2RAs shown, under the "Songs" tab 270 of the asset interface (an implementation of the fourth interface), multiple resources are presented, such as component 272 corresponding to song A and component 274 corresponding to song B. These multiple resources may include a first media resource, which may correspond to component 272, i.e., the first media resource may be song A. The user can initiate a third request associated with at least one resource via the "Edit" control 276. This third component may be, for example, an edit request or a generation request for at least one resource. Taking a media resource as an example, the third request may be an edit request for at least one attribute of the media resource. The edit request may include triggering component 272. When component 272 is triggered, the electronic device 110 can present panel 278 in interface 200R. Panel 278 is used to edit at least one attribute of the media resource. After at least one attribute is edited, the intelligent system can generate a fourth media resource based on the edited attribute and the referenced resource corresponding to the third request, and further present the fourth media resource in the first interface via the electronic device 110.
[0100] In this way, users can directly select resources and initiate adaptations from an interface that aggregates multiple resources, reducing the steps of searching and repetitive input to reuse existing resources. Furthermore, since the fourth media resources are directly presented in the session and based on existing resources, the amount of data that needs to be re-provided and transmitted can be reduced, as well as the number of page requests sent to the server, thereby reducing the server's query load and the bandwidth usage of end-to-cloud communication.
[0101] In addition to media resources, the aforementioned resources may also include lyrics resources, see [link to relevant documentation]. Figure 2S .and Figure 2R Corresponding to the aforementioned song resources, the multiple resources gathered by the fourth interface can also include lyrics resources, thus enabling lyrics to be stored and managed as reusable assets. Therefore, in some cases, the multiple resources also include lyrics resources, which are generated by the intelligent system through conversation.
[0102] like Figure 2S As shown, under the "Lyrics" tab of the asset interface, lyric resources, such as lyric card 280 and lyric card 282, are presented. These lyric resources are generated by the intelligent system through a session. In response to the reference control 284, the lyric resources can be referenced to the session or published; when the corresponding lyric card (e.g., lyric card 280) is triggered, the electronic device 110 can present a lyric editor panel 286 on the interface 200S. The lyric editor panel 286 displays details of the lyric resources and can trigger music generation via the "Generate Song" control. An example of referencing lyric resources to a session will be provided in [reference]. Figure 2T describe.
[0103] In this way, users can directly reuse the generated lyrics resources for subsequent creation, reducing the number of steps required to repeatedly create lyrics; and since existing lyrics resources are reused and there is no need to regenerate lyrics, the number of inferences in the lyric writing model and the amount of data exchanged between the cloud and the client can be reduced, thereby reducing the computing load on the server and the bandwidth usage of cloud-client communication.
[0104] Accept Figure 2S The lyrics resource allows users to reference it in the session and initiate new creations based on it. When the reference control 284 is triggered, as shown... Figure 2T As shown, the electronic device 110 can display an identifier 288 in the input box 202, indicating that 288 indicates that the lyrics resource is referenced. After referencing the session input box 202 of the new music interface, based on the at least one resource (lyrics resource) and the third request (indicating the generation of media resources based on the lyrics resource), a fourth media resource is presented in the first interface. In some scenarios, not only lyrics resources can be referenced, but media resources can also be referenced. For example, the electronic device 110 can display a reference control in component 272 to support the reference of the corresponding media resource. It can be understood that creation initiated by referencing lyrics resources, and Figure 2R The creation based on song resources corresponds to the implementation of presenting a fourth media resource on the first interface based on existing resources and a third request.
[0105] In this way, users can directly reference existing lyrics resources into the conversation and generate songs based on them, reducing the steps of repeatedly entering lyrics. Furthermore, since the song is generated based on existing lyrics resources and there is no need to repeatedly transmit lyrics, the amount of data in the end-to-cloud interaction and the processing scale of the lyrics writing process can be reduced, thereby reducing the computing load on the server and the bandwidth usage of end-to-cloud communication.
[0106] Example process Figure 3 A flowchart illustrating an example process 300 for interface interaction under certain conditions is shown. Process 300 can be implemented at electronic device 110. See below for reference. Figure 1 To describe process 300.
[0107] In box 310, the first interface is displayed, showing the conversation with the intelligent system. (As mentioned earlier...) Figures 2A to 2I As described, the first interface can be implemented as a conversation panel or dialogue panel that displays multiple rounds of dialogue with the intelligent system. Users can gradually express and adjust their creative intentions through natural language dialogue via the first interface.
[0108] In box 320, the first component is presented on the first interface. This first component is provided by the intelligent system and refers to the first media resource. (As mentioned above...) Figure 2J As described, after collecting the necessary parameters, the intelligent system generates a first media resource (e.g., song A) from the composition model and presents it in the first interface as a first component (e.g., song card) that refers to the first media resource, thereby directly embedding the creative product into the context of the conversation.
[0109] In box 330, in response to receiving a first operation on the first component, a second interface is presented, which displays at least one attribute of the first media resource. (As mentioned above...) Figure 2K As described, in response to a first operation on the first component (referring to the first media resource), a second interface is expanded on one side of the session as a workbench, presenting the style description, lyrics, timbre, and other attributes of the first media resource in read-only or edit mode.
[0110] In box 340, in response to receiving the second operation on the second interface, the first media resource is processed. (As mentioned above...) Figure 2K As described, in response to a second operation received via a publishing control on a second interface, the first media resource is published.
[0111] In this way, Process 300 integrates the gradual expression of creative intent, the generation of media resources, the viewing and adjustment of attributes, and the publication of resources into a unified interactive chain linked by a session. On the one hand, users do not need to repeatedly jump between multiple disparate functional pages to complete the closed loop from intent expression to resource publication within a continuous session context, thereby reducing the complexity of the interaction and the number of required operation steps. On the other hand, since each stage reuses the same session context, and the viewing and adjustment of attributes are centralized in the workbench, it can reduce the amount of interface content that the client needs to reload and render, and reduce the number of page requests sent to the server, thereby reducing the client's rendering overhead, the server's processing load, and the communication bandwidth usage between the client and the cloud.
[0112] In some cases, process 300 may also include one or more of the following. It is understood that the following examples may be implemented individually or in combination with each other.
[0113] In some cases, in response to receiving a second operation on the second interface, performing a first processing on the first media resource includes: receiving a first request via the second interface, the first request being associated with the first media resource; and presenting a second component on the first interface, the second media resource being obtained based on the first request and the first media resource (see...). Figure 2L , Figure 2N Thus, re-creative products based on existing media resources are directly presented in the conversation, which reduces both repetitive input and the amount of data that needs to be transmitted between the end and the cloud.
[0114] In some cases, process 300 further includes: on a first interface, presenting a first message corresponding to a first request, the first message indicating a reference to a first media resource (see...). Figure 2L ); and in some cases, the first message characterizes the type of the first request (see Figure 2L , Figure 2O Thus, the context of the conversation clearly preserves the object and intent upon which the re-creation is based, making it easier for users to recall and for intelligent systems to understand the context.
[0115] In some cases, receiving a first request via a second interface includes: receiving at least one first parameter via the second interface; and receiving a first request, wherein the first request represents the at least one first parameter, and the second media resource is generated by the intelligent system based on the first media resource and the at least one first parameter (see [link to documentation]). Figure 2L In some cases, the at least one first parameter includes at least one of a style parameter, a timbre parameter, and a generation method (see [reference]). Figure 2L This allows users to precisely control the re-creation process with structured parameters, enabling the intelligent system to generate second media resources that better meet expectations, thereby reducing repetitive generation caused by repeated trial and error.
[0116] In some cases, the at least one first parameter includes lyrics, and receiving the at least one first parameter via a second interface includes: presenting first lyrics in the second interface, the first lyrics corresponding to a first media resource; receiving an editing operation on the first lyrics to obtain second lyrics, wherein the second media resource is generated based on the second lyrics (see...). Figure 2G , Figure 2M Furthermore, in some cases, receiving an editing operation on the first lyrics includes: in response to receiving a third operation, presenting a first control, the third operation being used to select a first portion of the first lyrics; and in response to receiving first content via the first control, replacing the first portion with a second portion, the second portion being obtained based on the first content (see...). Figure 2H This allows users to fine-tune lyrics, either as a whole or line by line, reducing the amount of work involved in rewriting entire lyrics and also reducing the amount of content that the lyric-writing model needs to regenerate.
[0117] In some cases, presenting the first component on the first interface includes: receiving a second request on the first interface; presenting a second message on the first interface, the second message originating from the intelligent system, the second message displaying at least one second parameter; and, in response to receiving confirmation of the second message, presenting the first component, the first media resource being generated by the intelligent system based on the at least one second parameter (see [link to documentation]). Figure 2B , Figure 2I , Figure 2JTherefore, by having the intelligent system summarize and display the proposed parameters for user confirmation before the actual generation, invalid generation caused by parameter deviations can be reduced, thereby reducing unnecessary inference overhead of the composition model.
[0118] In some cases, process 300 further includes: presenting a third component on a first interface, the third component being provided by the intelligent system; and receiving a fourth operation via the third component to obtain at least a portion of the at least one second parameter (see...). Figures 2C to 2H In some cases, the component type of the third component is associated with the second request, and the component type characterizes the type of parameters that the third component uses to obtain (see [link to relevant documentation]). Figures 2C to 2F Therefore, the intelligent system provides a third component that matches the current intent on demand to collect parameters locally, which reduces the user's input cost and makes the collected parameters more standardized, making it easier for the server to generate subsequent data.
[0119] In some cases, presenting a third component on the first interface includes: in response to receiving a second request, presenting multiple entry points, each corresponding to a different method of creating a media resource; and in response to a triggering operation on a first entry point among the multiple entry points, presenting a third component, the first entry point corresponding to a first method, and the type of the third component being associated with the first method (see [link to documentation]). Figure 2B , Figure 2C This allows users to choose the starting point of their creation path based on their own preferences, making the presentation of the subsequent third component more aligned with the user's intentions.
[0120] In some cases, receiving a fourth operation via a third component includes: presenting third lyrics in the third component, the third lyrics being obtained based on a second request; and receiving a fourth operation via the third component, the fourth operation being used to confirm or edit the third lyrics (see [link to documentation]). Figure 2F , Figure 2G Furthermore, in some cases, presenting third lyrics in the third component includes: presenting multiple options on a first interface, the multiple options corresponding to different lyric styles; receiving a fifth operation for selecting a first option from the multiple options, the first option corresponding to a first lyric style; and in the third component, presenting third lyrics, the third lyrics being generated based on the first lyric style (see...). Figure 2E , Figure 2F Therefore, users can first define the style of the lyrics and then obtain the corresponding lyrics, making the generation direction of the lyric writing model more focused and reducing repeated generation caused by style mismatch.
[0121] In some cases, receiving a second request on the first interface includes: presenting a third media resource on the first interface; receiving a sixth operation for selecting a first segment of the third media resource; and receiving a second request associated with the first segment, wherein the first media resource includes a first segment and a second segment, and the second segment is derived based on the first segment (see [link to documentation]). Figure 2O This allows users to specify the starting point for the continuation, enabling the continuation model to generate a second segment that is more coherent in style and connection based on the first segment, reducing resource consumption caused by a complete rewrite.
[0122] In some cases, process 300 further includes: presenting at least one session entry on a third interface; receiving a trigger operation on a first session entry of the at least one session entry; and presenting a first interface (see...). Figure 2A , Figure 2Q This allows users to easily access or resume their creation session from either the new session entry or the history session entry, reducing the need to search for and reconstruct the creation context.
[0123] In some cases, process 300 further includes: presenting multiple resources on a fourth interface, the multiple resources including a first media resource; receiving a third request, the third request being associated with at least one of the multiple resources; and presenting a fourth media resource on a first interface, the fourth media resource being obtained based on the at least one resource and the third request (see [link to documentation]). Figure 2R , Figure 2T In some cases, multiple resources also include lyrics, which are generated by an intelligent system through conversation (see [link to resource]). Figure 2S As a result, users can directly reuse existing song or lyric resources to create new content, reducing both repetitive input and generation, as well as the amount of data that needs to be transferred between the client and the cloud.
[0124] Example devices and equipment A corresponding apparatus for implementing the above methods or processes is also provided. Figure 4 A schematic structural block diagram of an example device 400 for interface interaction is shown, according to some scenarios. Device 400 can be implemented as or included in electronic device 110. The various modules / components in device 400 can be implemented by hardware, software, firmware, or any combination thereof.
[0125] The first presentation module 410 is configured to present a first interface displaying a conversation with the intelligent system. The second presentation module 420 is configured to present a first media resource on the first interface, the first media resource being provided by the intelligent system. The third presentation module 430 is configured to present a second interface in response to receiving a first operation on the first media resource, the second interface displaying at least one attribute of the first media resource. The processing module 440 is configured to perform a first process on the first media resource in response to receiving a second operation on the second interface.
[0126] In some cases, the processing module 440 is also configured to: receive a first request via a second interface, the first request being associated with a first media resource; and present a second component on the first interface, the second media resource being obtained based on the first request and the first media resource.
[0127] In some cases, device 400 further includes a message presentation module configured to: present a first message on a first interface, the first message corresponding to a first request, the first message indicating a reference to a first media resource. In some cases, the first message characterizes the type of the first request.
[0128] In some cases, the processing module 440 is also configured to: receive at least one first parameter via a second interface; and receive a first request, the first request representing the at least one first parameter, wherein the second media resource is generated by the intelligent system based on the first media resource and the at least one first parameter.
[0129] In some cases, the at least one first parameter includes at least one of a style parameter, a timbre parameter, and a generation method.
[0130] In some cases, the at least one first parameter includes lyrics, and the first receiving module is further configured to: present the first lyrics in a second interface, the first lyrics corresponding to a first media resource; and receive editing operations on the first lyrics to obtain second lyrics, wherein the second media resource is generated based on the second lyrics.
[0131] In some cases, the processing module 440 is also configured to: in response to receiving a third operation, present a first control, the third operation being used to select a first part of the first lyrics; and in response to receiving first content via the first control, replace the first part with a second part, the second part being obtained based on the first content.
[0132] In some cases, the second presentation module 420 is further configured to: receive a second request on a first interface; present a second message on the first interface, the second message originating from the intelligent system, the second message displaying at least one second parameter; and, in response to receiving confirmation of the second message, present a first component, the first media resource being generated by the intelligent system based on the at least one second parameter.
[0133] In some cases, the device 400 further includes a first receiving module configured to: present a third component on a first interface, the third component being provided by the intelligent system; and receive a fourth operation via the third component to obtain at least a portion of the at least one second parameter.
[0134] In some cases, the component type of the third component is associated with the second request, and the component type characterizes the type of parameters that the third component uses to obtain.
[0135] In some cases, the first receiving module is also configured to: in response to receiving a second request, present multiple entry points, the multiple entry points corresponding to different ways of creating media resources; and in response to a triggering operation on the first entry point among the multiple entry points, present a third component, the first entry point corresponding to the first way, and the type of the third component being associated with the first way.
[0136] In some cases, the first receiving module is also configured to: present third lyrics in the third component, the third lyrics being obtained based on the second request; and receive a fourth operation through the third component, the fourth operation being used to confirm or edit the third lyrics.
[0137] In some cases, the first receiving module is also configured to: present multiple options on a first interface, the multiple options corresponding to different lyric styles; receive a fifth operation for selecting a first option among the multiple options, the first option corresponding to a first lyric style; and in a third component, present third lyrics, the third lyrics being generated based on the first lyric style.
[0138] In some cases, the second presentation module 420 is also configured to: present a third media resource on a first interface; receive a sixth operation for selecting a first segment of the third media resource; and receive a second request associated with the first segment, wherein the first media resource includes the first segment and a second segment, the second segment being derived from the first segment.
[0139] In some cases, the device 400 further includes a second receiving module configured to: present at least one session entry on a third interface; receive a trigger operation on a first session entry in the at least one session entry; and present a first interface.
[0140] In some cases, device 400 further includes a third receiving module configured to: present multiple resources, including a first media resource, on a fourth interface; receive a third request associated with at least one of the multiple resources; and present a fourth media resource on a first interface, the fourth media resource being obtained based on the at least one resource and the third request.
[0141] In some cases, multiple resources also include lyrics, which are generated by an intelligent system through conversation.
[0142] The modules included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some cases, one or more modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 400 can be implemented at least partially by one or more hardware logic components. By way of example, and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.
[0143] Figure 5 A block diagram of an electronic device 500 in which one or more examples may be implemented is shown. It should be understood that... Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the examples described herein. Figure 5 The electronic device 500 shown can be used to implement the electronic device 110 discussed above.
[0144] like Figure 5As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processing units or processors 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processor 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.
[0145] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.
[0146] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various examples.
[0147] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 500 can operate in a networked environment using logical connections to one or more other servers, networked personal computers, or another network node.
[0148] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0149] A computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. A computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0150] The flowcharts and / or block diagrams of the methods, apparatus, devices, and computer program products referred to herein describe various aspects. It should be understood that each block of the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0151] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0152] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0153] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0154] Various examples have been described above. The foregoing descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for interface interaction, comprising: The first interface is presented, which displays the conversation with the intelligent system; On the first interface, a first component is presented, which is provided by the intelligent system and refers to a first media resource; In response to receiving a first operation on the first component, a second interface is presented, the second interface presenting at least one attribute of the first media resource; as well as In response to receiving a second operation on the second interface, the first media resource is processed in a first manner.
2. The method according to claim 1, wherein the first processing of the first media resource in response to receiving the second operation on the second interface comprises: A first request is received via the second interface, and the first request is associated with the first media resource; as well as On the first interface, a second component is presented, which refers to a second media resource, which is obtained based on the first request and the first media resource.
3. The method according to claim 2, further comprising: On the first interface, a first message is presented, which corresponds to the first request and indicates a reference to the first media resource.
4. The method of claim 3, wherein the first message characterizes the type of the first request.
5. The method of claim 2, wherein receiving the first request via the second interface comprises: At least one first parameter is received via the second interface; as well as The system receives the first request, which represents the at least one first parameter, wherein the second media resource is generated by the intelligent system based on the first media resource and the at least one first parameter.
6. The method of claim 5, wherein the at least one first parameter includes lyrics, and receiving the at least one first parameter via the second interface includes: In the second interface, the first lyrics are presented, and the first lyrics correspond to the first media resource; The system receives editing operations on the first lyrics to obtain second lyrics, wherein the second media resource is generated based on the second lyrics.
7. The method of claim 6, wherein receiving the editing operation on the first lyrics comprises: In response to receiving a third operation, a first control is presented, the third operation being used to select a first part of the first lyrics; as well as In response to receiving first content via the first control, the first part is replaced with a second part, the second part being obtained based on the first content.
8. The method according to claim 1, wherein presenting the first media resource on the first interface includes: On the first interface, the second request is received; On the first interface, a second message is presented. The second message comes from the intelligent system and displays at least one second parameter. as well as In response to receiving confirmation of the second message, the first media resource is presented, the first media resource being generated by the intelligent system based on the at least one second parameter.
9. The method of claim 8, further comprising: On the first interface, a third component is presented, which is provided by the intelligent system; as well as The third component receives a fourth operation to obtain at least a portion of the at least one second parameter.
10. The method of claim 9, wherein the component type of the third component is associated with the second request, the component type characterizing the parameter type that the third component is used to acquire.
11. The method of claim 9, wherein presenting the interactive component on the first interface comprises: In response to receiving the second request, multiple entry points are presented, each corresponding to a different way of creating media resources; as well as In response to a triggering operation on a first of the plurality of entry points, the third component is presented, the first entry point corresponding to a first mode, and the type of the third component being associated with the first mode.
12. The method of claim 9, wherein receiving the fourth operation via the third component comprises: In the third component, third lyrics are presented, which are obtained based on the second request; as well as The third component receives the fourth operation, which is used to confirm the third lyrics or edit the third lyrics.
13. The method of claim 12, wherein presenting the third lyrics in the third component comprises: On the first interface, multiple options are presented, each corresponding to a different lyric style; Receive a fifth operation, the fifth operation being used to select a first option among the plurality of options, the first option corresponding to a first lyrics style; as well as In the third component, the third lyrics are presented, which are generated based on the style of the first lyrics.
14. The method of claim 8, wherein receiving the second request on the first interface comprises: The first interface presents third-party media resources; Receive a sixth operation, the sixth operation being used to select a first segment of the third media resource; as well as The second request is received, the second request being associated with the first segment, wherein the first media resource includes the first segment and a second segment, the second segment being derived from the first segment.
15. The method according to claim 1, further comprising: On the third interface, at least one session entry point is presented; Receive a trigger operation on the first session entry of the at least one session entry; as well as The first interface is displayed.
16. The method according to claim 1, further comprising: On the fourth interface, multiple resources are presented, including the first media resource; Receive a third request, the third request being associated with at least one of the plurality of resources; as well as On the first interface, a fourth media resource is presented, which is obtained based on the at least one resource and the third request.
17. The method of claim 16, wherein the plurality of resources further comprises lyrics resources, the lyrics being generated by the intelligent system through the session.
18. A device for interface interaction, comprising: The first presentation module is configured to present a first interface, which displays the conversation with the intelligent system. The second presentation module is configured to present a first component on the first interface, the first component being provided by the intelligent system, and the first component referring to a first media resource. A third presentation module is configured to present a second interface in response to receiving a first operation on the first media resource, the second interface presenting at least one attribute of the first media resource; as well as The publishing module is configured to perform a first processing on the first media resource in response to receiving a second operation on the second interface.
19. An electronic device comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 17 when executed by the at least one processor.
20. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 17.