Interface interaction method and device, equipment and storage medium

By acquiring audio information from a second user and generating associated media content, the intelligent system enriches the ways in which it interacts with users, solves the problem of limited interaction methods in existing technologies, and enhances the user experience.

CN122044734APending Publication Date: 2026-05-15BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-02-10
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, the interaction methods between intelligent systems and users are relatively simple and lack variety, making it difficult to meet the diverse interaction needs of users.

Method used

By receiving instructions from the first user, acquiring audio information from the second user, and using an intelligent system to generate media content associated with the audio information, the ways in which users interact are enriched.

Benefits of technology

This enables intelligent systems to utilize specific audio information to enhance the interactive experience between users when generating media content, thereby improving the richness and efficiency of user interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122044734A_ABST
    Figure CN122044734A_ABST
Patent Text Reader

Abstract

The invention provides an interface interaction method and device, equipment and a storage medium. The method comprises the steps that a first instruction is received, the first instruction is from a first user, and the first instruction is associated with the intelligent system; and outputting a first media content, the first media content being generated by the intelligent system based on first audio information of a second user, the first audio information being obtained based on a first operation of the second user on a first message, and the first message being sent to the second user based on a second instruction of the first user. In this way, the interaction mode between the users can be enriched.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The examples in this article generally relate to the field of computer science, and in particular to methods, apparatuses, devices, and computer-readable storage media for user interface interaction. Background Technology

[0002] With the rapid development of artificial intelligence technology, machine learning models are being applied to all aspects of people's lives. For example, users can trigger intelligent systems to perform various types of tasks by interacting with them. Taking chatbots as an example, people can engage in real-time conversations with chatbots in a chat interface, and the chatbots can use models to generate corresponding responses. Therefore, enriching the ways to interact with intelligent systems is a topic worthy of attention. Summary of the Invention

[0003] In a first aspect, a method for interface interaction is provided. The method includes: receiving a first instruction from a first user, the first instruction being associated with an intelligent system; and outputting first media content generated by the intelligent system based on first audio information from a second user, wherein the first audio information is obtained based on a first operation by the second user on a first message, and the first message is sent to the second user based on a second instruction from the first user.

[0004] In a second aspect, a method for interface interaction is provided. The method includes: presenting a first interface to a first user, the first interface being associated with an intelligent system; receiving a first instruction via the first interface; and, in response to the satisfaction of a preset condition, triggering the intelligent system to send first media content to a second user, wherein at least one audio attribute of the first media content is associated with the first user, and the preset condition is associated with the first instruction.

[0005] In a third aspect, an apparatus for interface interaction is provided. The apparatus includes: a first receiving module configured to receive a first instruction from a first user, the first instruction being associated with an intelligent system; and an output module configured to output first media content generated by the intelligent system based on first audio information from a second user, wherein the first audio information is obtained based on a first operation by the second user on a first message, and the first message is sent to the second user based on a second instruction from the first user.

[0006] In a fourth aspect, another device for interface interaction is provided. The device includes: a first presentation module configured to present a first interface to a first user, the first interface being associated with an intelligent system; a second receiving module configured to receive a first instruction via the first interface; and a first triggering module configured to trigger the intelligent system to send first media content to a second user in response to the fulfillment of a preset condition, wherein at least one audio attribute of the first media content is associated with the first user, and the preset condition is associated with the first instruction.

[0007] In a fifth aspect, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the methods of the first or second aspect.

[0008] In a sixth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the methods of the first or second aspect.

[0009] In a seventh aspect, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect or the second aspect.

[0010] In this way, when the intelligent system generates media content based on the instructions of the first user, it can use specific audio information (e.g., the first audio information of the second user) to generate first media content associated with the first audio information, thereby enriching the ways in which the intelligent system generates media content.

[0011] It should be understood that the content described in this section is not intended to limit the key or important features of the examples in this article, nor is it intended to restrict the scope of the solution. Other features will become readily apparent from the following description. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the various examples herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A schematic diagram of the example environment is shown; Figures 2A to 2F Example interfaces for some scenarios are shown; Figure 3A and Figure 3B Example interfaces for several other scenarios are shown; Figure 4A and Figure 4B Example interfaces for other scenarios are shown; Figure 5A and Figure 5B Example interfaces for some other scenarios are shown; Figure 6 The flowcharts show example processes of interface interactions in some scenarios; Figure 7 The flowcharts show example processes of interface interactions in several other scenarios; Figure 8 Schematic structural block diagrams of example devices for interface interaction in some scenarios are shown; Figure 9 Schematic block diagrams of example devices for interface interaction in several other scenarios are shown; and Figure 10 A block diagram of an electronic device capable of implementing multiple illustrative scenarios is shown. Detailed Implementation

[0013] The examples in the text will now be described in more detail with reference to the accompanying drawings. While some examples are shown in the drawings, it should be understood that solutions can be implemented in various forms and should not be construed as limited to the examples presented herein. Rather, these examples are provided to provide a more thorough and complete understanding of the solutions. It should be understood that the drawings and examples in this document are for illustrative purposes only and are not intended to limit the scope of protection of the solutions.

[0014] It should be noted that the headings of any section / subsection provided herein are not restrictive. Various examples are described throughout this document, and examples of any type may be included under any section / subsection. Furthermore, examples described in any section / subsection may be combined in any way with any other examples described in the same section / subsection and / or different sections / subsections.

[0015] In the description of the examples in this document, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an example" or "the example" should be understood as "at least one example". The term "some examples" should be understood as "at least some examples". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0016] The examples in this document may involve user data, data acquisition, and / or use. All of these aspects comply with relevant laws, regulations, and provisions. In the examples presented herein, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, when implementing each example, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained through appropriate means, in accordance with relevant laws and regulations. The specific methods of notification and / or authorization can vary depending on the actual situation and application scenario; the scope of the solution is not limited in this regard.

[0017] In this manual and the sample solutions, any processing of personal information will be conducted only under legal grounds (such as obtaining the consent of the data subject or being necessary for the performance of a contract) and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.

[0018] In this paper, an intelligent system refers to a system capable of autonomous control based on machine learning models. An intelligent system is, for example, a virtual object or physical entity capable of making decisions and autonomously executing actions based on machine learning models to achieve preset goals or complete preset tasks. An intelligent system can be an automated program that understands user intent and can utilize models or invoke tools to complete various types of tasks. In some contexts, examples of intelligent systems may include, but are not limited to: agents, bots, chatbots, digital avatars, intelligent customer service, digital assistants, etc. Alternatively, an intelligent system can also be an intelligent role implemented based on a machine learning model. An "intelligent system" can process instructions (or requests) from users based on generative models (e.g., language models, multimodal models) to perform specified types of tasks. In some cases, an intelligent system may also relate to virtual accounts, which may have corresponding avatars or nicknames.

[0019] In this document, "response content" can include various appropriate types of content provided to the user via a session interface, examples of which may include, but are not limited to, text, images, audio, video, code, documents, applications, etc. In some cases, "response content" may be generated by an intelligent system using a generative model. Alternatively, "response content" may also include content obtained through other means, such as through a search. "Response content" is provided to the user as a response to a user instruction (or request).

[0020] As mentioned above, with the rapid development of artificial intelligence technology, machine learning models are being applied to all aspects of people's lives. For example, users can trigger intelligent systems to perform various types of tasks by interacting with them. Taking chatbots as an example, people can engage in real-time conversations with chatbots in a chat interface, and the chatbots can use models to generate corresponding responses. Therefore, enriching the ways to interact with intelligent systems is a topic worthy of attention.

[0021] This paper proposes a user interface interaction scheme. The scheme includes: receiving a first instruction. The first instruction originates from a first user. The first instruction is associated with an intelligent system. Further, it can output first media content. The first media content is generated by the intelligent system based on first audio information from a second user. The first audio information is obtained based on a first operation by the second user on a first message. The first message is sent to the second user based on a second instruction from the first user.

[0022] In this way, a first user can obtain a second user's first audio information by sending a first message to the second user, thereby enriching the ways users interact. Furthermore, based on the first user's first instruction, the system can output first media content generated by the intelligent system based on the first audio information. Thus, when the intelligent system generates media content based on the first user's instruction, it can utilize specific audio information (e.g., the second user's first audio information) to generate first media content associated with the first audio information, enriching the ways the intelligent system generates media content.

[0023] This paper also proposes another interface interaction scheme. This scheme includes: presenting a first interface to a first user. The first interface is associated with an intelligent system. A first command is received via the first interface. Furthermore, in response to the fulfillment of preset conditions, the intelligent system can be triggered to send first media content to a second user. At least one audio attribute of the first media content is associated with the first user. The preset conditions are associated with the first command.

[0024] In this way, it is possible to receive a first command from a first user via a first interface associated with the intelligent system. Furthermore, in response to the fulfillment of preset conditions related to the first command, the intelligent system can be triggered to send first media content to a second user. This allows the first user to send first media content to the second user through the intelligent system, thereby enriching the interaction methods between the first and second users. In addition, at least one audio attribute of the first media content is associated with the first user, thus enabling the intelligent system to send media content associated with the first user's audio information to the second user, further enriching the interaction methods between the first and second users and improving the efficiency of user interaction.

[0025] The following describes various examples of this scheme in further detail with reference to the accompanying drawings.

[0026] Example Environment Figure 1 A schematic diagram of example environment 100 is shown. (e.g.) Figure 1 As shown, example environment 100 may include electronic device 110.

[0027] In this example environment 100, electronic device 110 can run an application 120 that supports user interface interaction. Application 120 can be any suitable type of application for user interface interaction, including but not limited to: artificial intelligence applications, media applications, social applications, or other suitable applications. User 140 can interact with application 120 via electronic device 110 and / or its attached devices.

[0028] exist Figure 1 In environment 100, if application 120 is active, electronic device 110 can use application 120 to present interface 150 for supporting interface interaction.

[0029] In some cases, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some cases, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).

[0030] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 that support user interface interaction in electronic devices 110.

[0031] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection can include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (Wi-Fi) connections. In some cases, server 130 and electronic device 110 can exchange signaling information through their communication connection.

[0032] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the scheme.

[0033] The following description of the example will continue with reference to the accompanying drawings.

[0034] Example Interaction Example 1: Figures 2A to 2F Example interfaces 200A to 200F are shown, illustrating interface interactions under various scenarios. Interfaces 200A to 200F can, for example, be... Figure 1 The electronic device 110 shown is provided.

[0035] In some cases, electronic device 110 can receive a first instruction. The first instruction originates from a first user. The first instruction is associated with an intelligent system. For example, electronic device 110 can receive a first instruction sent by the first user to the intelligent system. For example, electronic device 110 can receive a first instruction from the first user via a first interface.

[0036] In some cases, such as Figure 2A As shown, the electronic device 110 can present an interface 200A. Interface 200A can be implemented as a first interface. As an example, the first interface can be implemented as a conversational interface between a first user and an intelligent system.

[0037] In some contexts, an intelligent system refers to a system capable of autonomous control based on machine learning models. An intelligent system is, for example, a virtual object or physical entity that can make decisions and autonomously execute actions based on machine learning models to achieve preset goals or complete preset tasks. An intelligent system can be an automated program that understands user intent and can utilize models or invoke tools to accomplish various types of tasks. Examples of intelligent systems include, but are not limited to, agents, bots, chatbots, digital avatars, intelligent customer service representatives, and digital assistants. Alternatively, an intelligent system can also be an intelligent role implemented based on machine learning models. An "intelligent system" can process user instructions based on generative models (e.g., language models, multimodal models) to perform specified types of tasks. In some cases, an intelligent system can also refer to a virtual account, which may have a corresponding avatar or nickname.

[0038] In some cases, electronic device 110 may receive a first instruction in response to receiving an input message from a first user via a first interface. For example, electronic device 110 may receive the first user's input message via an input control in the first interface. For example, in some cases, electronic device 110 may present an input control 202 and a message display area 204 in the first interface (e.g., interface 200A). Electronic device 110 may receive the user's input message via the input control 202 and send the input message to the intelligent system. Accordingly, electronic device 110 may present the response content of the intelligent system in the message display area 204. In some cases, electronic device 110 may receive a first instruction in response to obtaining the first user's input message via the input control 202. For example, such an input message may instruct the intelligent system to generate media content. Such an input message may instruct a prompt for generating media content.

[0039] In some cases, electronic device 110 can receive a first instruction via a control. For example, electronic device 110 can present a control on a first interface (e.g., interface 200A). In some cases, electronic device 110 can present area 206 on the first interface (e.g., interface 200A). Area 206 can present at least one control (or tool component) associated with the intelligent system, such as control 208. In some scenarios, area 206 can also be referred to as a toolbar or action bar. Area 206 can, for example, be presented near input component 202 to facilitate easier use of tools related to the intelligent system by the user. For example, electronic device 110 can receive a first instruction via control 208.

[0040] Alternatively, the electronic device 110 may present at least one control, such as control 210-1 and control 210-2, in the message display area 204. As an example, the electronic device 110 may, in response to a trigger on control 210-1, send an input message corresponding to control 210-1 to the intelligent system, and present the input message in the message display interface 204 associated with the identification information (e.g., avatar) of the first user. For example, the electronic device 110 may receive a first instruction via control 210-1.

[0041] In some cases, electronic device 110 may display control 205 on a first interface. Electronic device 110 may receive a first instruction based on triggering control 205. For example, control 205 may be associated with a voice conversation mode or video conversation mode of a smart system. For example, electronic device 110 may display a voice conversation interface or video conversation interface with a smart system based on triggering control 205.

[0042] In some cases, electronic device 110 may, in response to receiving a first instruction, trigger the intelligent system to generate first media content.

[0043] In some cases, the controls mentioned above (e.g., control 208 or control 210-1, etc.) can be associated with a preset theme. The electronic device 110 can trigger a smart system to generate first media content based on the preset theme associated with a control (e.g., control 208 or control 210-1, etc.) upon triggering the control. For example, such a preset theme may include a holiday theme, an event theme, etc. Alternatively, the electronic device 110 can also present multiple candidate themes associated with the preset theme in response to triggering a control (e.g., control 208 or control 210-1, etc.). Figure 2B As shown, the electronic device 110 can present an interface 200B. Interface 200B can, for example, be implemented as a first interface. The electronic device 110 can present multiple candidate topics associated with a preset theme in the first interface (e.g., interface 200B), such as candidate topic 212-1, candidate topic 212-2, and candidate topic 212-3. For example, the preset theme is a holiday theme, and candidate topic 212-1 corresponds to text content associated with the holiday theme. Furthermore, in response to the selection of at least one of the multiple candidate topics, the electronic device 110 can trigger an intelligent system to generate first media content based on at least one topic.

[0044] In other examples, the electronic device 110 may also obtain the user's generation request through other suitable means and may generate first media content based on the user's generation request. In some cases, the generation request may indicate one or more parameters for controlling the generation of media content.

[0045] As an example, a user's request to generate content could be prompted by user input, such as, "Create a New Year's greeting video using audio information from friend A." As another example, a user's request could indicate reference media content, such as a reference image, that was captured or uploaded. The intelligent system could then generate a New Year's greeting video based on this reference image and audio information. In yet another example, a user's request could also indicate the selection of a media template. For instance, the intelligent system could generate initial media content based on the selected video template.

[0046] In some examples, the first media content may include any suitable type of media content with audio information. For example, the first media content may be audio content, which may include, for instance, song content, speech content, etc., generated based on the audio information of a second user. In another example, the first video content may be video content, for instance, in which at least a portion of the audio may be generated based on the audio information of a second user. In yet another example, the first media content may also include, for instance, graphic works, presentation documents, and other media content that has both visual and audio content.

[0047] In some cases, the first instruction may direct audio information used to generate the first media content. For example, the first media content is generated by an intelligent system based on audio information specified by the first instruction. In some examples, the audio information may direct feature representations or data characterizing one or more audio attributes. Such audio attributes may include, but are not limited to, timbre, rhythm, and other attributes related to the state of sound expression. Taking timbre as an example, the audio information may include, for example, feature representations or data characterizing a specified timbre. Such audio information may be provided to an audio generation model or an appropriate model with audio generation capabilities as a control condition to generate audio content that matches the specified audio attributes. For example, such feature representations or data may be provided to an audio generation model or a video model to guide the generation of media such that the audio of the media content matches the specified timbre.

[0048] In some examples, intelligent systems can utilize one or more appropriate machine learning models to generate media content. As an example, audio generation models such as the injection-diffusion model can be used to generate audio content based on audio information from a second user. Such audio information can, for example, be provided as control signals to the audio generation model.

[0049] In some examples, where timbre is used as an audio attribute, the audio generation model may also acquire other timbres to generate audio content related to multiple timbres. As an example, feature representations corresponding to multiple timbres can be provided to the audio generation model as control signals. Accordingly, the generated audio content may include, for example, dialogue content corresponding to different timbres.

[0050] Additionally, such audio content can be combined with preset image content or generated image content to generate video content. For example, an image generation model can be used to generate a corresponding set of video frames, and the final video content can be generated by temporally aligning the set of video frames and audio content.

[0051] In other scenarios, end-to-end generative models can be used to directly generate media content with both audio and visual elements. For example, multimodal models can be used to process input features from different modalities (e.g., audio information, cues, reference images, etc.) to directly generate video content that may include audio portions corresponding to a single or multiple timbres.

[0052] For example, electronic device 110 may present one or more audio options. Furthermore, electronic device 110 may determine audio information for generating the first media content based on the selection of at least one of the one or more audio options.

[0053] As an example, electronic device 110 may present configuration entry 214 on a first interface (e.g., interface 200B). Electronic device 110 may present an audio configuration interface associated with the intelligent system based on a trigger on configuration entry 214 (e.g., a click operation). For example, configuration entry 214 may be associated with the intelligent system's identification information (e.g., the name identifier "XXX"). Figure 2C As shown, electronic device 110 can present interface 200C. Interface 200C can be implemented as an audio configuration interface associated with a smart system. Electronic device 110 can present one or more audio options in the audio configuration interface. For example, electronic device 110 can present one or more audio options corresponding to label 216, such as audio option 218, in response to the selection of label 216 in the audio configuration interface. For example, audio option 218 indicates audio information of a first user (also referred to as third audio information). For example, the third audio information indicated by audio option 218 is determined based on audio content recorded by the first user. For example, electronic device 110 can obtain the audio content recorded by the first user via control 219 in the first interface (e.g., interface 200C) to determine the third audio information.

[0054] Alternatively, the electronic device 110 may, in response to a selection of label 220 in the audio configuration interface, present one or more audio options corresponding to label 220. For example... Figure 2D As shown, electronic device 110 can present interface 200D. Interface 200D can be implemented as an audio configuration interface. Electronic device 110 can have one or more audio options corresponding to label 220 in interface 200D, such as audio option 222-1, audio option 222-2, and audio option 222-3, etc. For example, the audio information indicated by one or more audio options corresponding to label 220 is associated with a user other than the first user. For example, audio option 222-1 can indicate the first audio information of a second user. For example, audio option 222-2 can indicate the second audio information of a third user. For example, the audio information corresponding to label 220 (e.g., the first audio information and the second audio information) is provided to the first user by the corresponding user.

[0055] Additionally or alternatively, the electronic device 110 may, in response to the selection of other labels (e.g., "Recommended," "Male Voice," "Female Voice") in the audio configuration interface, present at least one audio option corresponding to the other labels. As an example, the at least one audio option corresponding to the other labels includes system-preset audio options.

[0056] Additionally, the electronic device 110 can determine audio information for generating the first media content based on the selection of at least one audio option from one or more audio options presented in the audio configuration interface. For example, the electronic device 110 can determine audio information for generating the first media content based on the selection of at least one audio option from one or more audio options presented in the audio configuration interface. Figure 2D The selection of audio option 222-1 triggers the intelligent system to generate first media content based on the first audio information of the second user. Additionally or alternatively, the electronic device 110 can, based on the selection of audio option 222-1, trigger the intelligent system to generate first media content. Figure 2D The selection of audio option 222-2 triggers the intelligent system to generate first media content based on the second audio information from a third user. Additionally or alternatively, the electronic device 110 can also, based on... Figure 2C The selection of audio option 218 triggers the intelligent system to generate first media content based on the third audio information of the first user.

[0057] In some cases, the second user's first audio information is obtained based on the second user's first operation on the first message. The first message is sent to the second user based on the first user's instruction (also referred to as the second instruction). The electronic device 110 can send the first message in a session in response to the second instruction. An example process for obtaining the second user's first audio information will be described below.

[0058] In some examples, electronic device 110 can obtain a second instruction from the first user via an interactive interface related to the intelligent system. As an example, the second instruction could instruct the acquisition of first audio information from the second user. As another example, electronic device 110 could obtain the second instruction based on prompts input by the first user or the first user's triggering of preset interactive controls.

[0059] Furthermore, the intelligent system can send a first message to a session related to the second user based on a second instruction. As an example, the participants in this session can include at least the intelligent system and the second user. Additionally, the session can also be a group chat session, whose participants can include, for example, the intelligent system, the first user, the second user, or more users. In some examples, the first message may include a link message. For example, the second user can authorize the first user to access the second user's first audio information by clicking the link message; or, for example, the second user can also record audio by clicking the link message. An example of the audio recording process can be found below regarding... Figure 3B The example described.

[0060] In other examples, the intelligent system can also generate the content to be recorded first after receiving the second instruction. For example, electronic device 110 can present the message "I will ask the second user to record 'Happy New Year,' do you think it's appropriate?" Furthermore, the intelligent system can also obtain the first user's confirmation or modification instructions regarding the content to be recorded.

[0061] For example, if the first user deems the content to be recorded appropriate, the first user can, for instance, input preset content and click a preset control to instruct the intelligent system to send a first message, requesting the second user to record audio content corresponding to that content.

[0062] As another example, if the first user wishes to modify the content to be recorded, the first user can, for example, input an adjustment command. For instance, the first user could input "longer." Accordingly, the intelligent system can, for example, output new content to be recorded based on the adjustment command. Further, upon receiving confirmation of the new content to be recorded, the intelligent system sends a first message requesting the second user to record audio content corresponding to that content.

[0063] In some cases, the intelligent system can also provide the user with a link message corresponding to the acquisition of audio information. Further, the electronic device 110 can receive a forwarding or copying operation from the first user to send the link message to a session associated with the second user. As an example, such a session could be a one-on-one chat between the first and second users. Alternatively, such a session could be a group chat, with participants including the first user, the second user, and at least one other participant.

[0064] like Figure 2D As shown, electronic device 110 may provide control 224 in an audio configuration interface (e.g., interface 200D). For example, electronic device 110 may present a set of candidate sessions associated with a first user in response to a trigger on control 224.

[0065] Additionally or alternatively, the electronic device 110 may also display response content from the intelligent system in the message display area of ​​the first interface. Such response content may be provided based on a message from the first user. For example, the intelligent system may provide corresponding response content in response to receiving a message from the first user (e.g., "Get my friend's voice"). Such response content may include preset controls. These preset controls can be configured to trigger the presentation of the aforementioned set of candidate sessions.

[0066] Additionally or alternatively, the electronic device 110 may send a first message in at least one session based on the selection of at least one session from a set of candidate sessions. For example, the at least one session may include a session associated with a second user. For example, the participants in the session associated with the second user include at least the first user and the second user. For example, the session associated with the second user may be a one-way chat session including only the first user and the second user. Alternatively, the participants in the session associated with the second user may include at least three users. Figure 2E As shown, electronic device 110 can present interface 200E. For example, interface 200E can be implemented as a session interface associated with a second user. Such a session may include, for example, five users. Electronic device 110 can present a first message 226 from a first user (e.g., user 1) on the session interface (e.g., interface 200E). In some cases, the first message 226 may indicate a first reference to the second user. For example, the first message 226 may include text content 228 (e.g., "@user 2") associated with the identification information of the second user (e.g., the text identifier "user 2") to indicate a first reference to the second user.

[0067] In some cases, the first audio information is obtained based on the second user's first action on the first message. The following is based on... Figure 3A and Figure 3B The process of acquiring the first audio information will be described exemplarily.

[0068] Figure 3A and Figure 3B Example interfaces 300A and 300B are shown, illustrating interface interactions under certain circumstances. Interfaces 300A to 300B can, for example, be provided by... Figure 1 The electronic device 110 shown is provided.

[0069] In some cases, the first audio information corresponds to reference audio content from a second user. For example, electronic device 110 can be implemented as a client associated with the second user. Figure 3A As shown, electronic device 110 can present interface 300A to a second user. Interface 300A can be implemented as a session interface associated with a session of the second user. Electronic device 110 can present a first message 305 from a first user in the session. For example, electronic device 110 can provide an audio component to the second user in response to a first operation. For example, electronic device 110 can provide an audio component to the second user in response to a triggering of the first message 305 (e.g., a click operation). Alternatively, electronic device 110 can present control 310 associated with the first message 305. Electronic device 110 can provide an audio component to the second user in response to a triggering of control 310 (e.g., a click operation).

[0070] like Figure 3B As shown, electronic device 110 can present interface 300B to a second user. Interface 300B can be implemented as an interface associated with the second user. For example, electronic device 110 can provide audio component 315 to the second user through interface 300B. Further, electronic device 110 can acquire reference audio content from the second user through audio component 315. For example, electronic device 110 can invoke an audio acquisition unit (e.g., a recording device) deployed on electronic device 110 to acquire the reference audio content in response to the second user's pressing operation on the audio component. In some cases, electronic device 110 can record reference audio content through audio component 315. As an example, the reference audio content is associated with text content. For example, electronic device 110 can present text content 320 in interface 300. For example, the reference audio content is acquired by the second user reading text content 320 aloud.

[0071] In some examples, the text content 320 to be recorded may include preset text content. As an example, the text content 320 may be associated with a preset theme, such as "New Year's greetings." In some cases, such a theme may be determined based on a second instruction from a first user. For example, if the second user initiates the second instruction via an interactive control related to "New Year," the text content to be recorded may be related to "New Year's greetings."

[0072] In other examples, the text content 320 to be recorded may be related to a second instruction from a second user. For example, as described above, the intelligent system may generate the text content to be recorded based on the second instruction, and may also modify the text content to be recorded based on the adjustment operation of the first user.

[0073] In some cases, appropriate audio coding models can be used to process reference audio content to generate corresponding audio information. For example, Mel-spectral features of the reference audio content can be extracted, and features representing timbre can be extracted from these features as audio information. As another example, a pre-trained audio coding model can also be used to process the reference audio content to generate feature vectors that express timbre, which can then be used as audio information.

[0074] In some cases, text content 320 may be associated with a second instruction from a first user. For example, text content 320 may be configured based on the second instruction. For example, the second instruction may instruct the first user to select text content 320 from a plurality of candidate text contents. Alternatively, the second instruction may instruct the first user's input text to be used as text content 320.

[0075] In some scenarios, electronic device 110 or server 130 can utilize a pre-trained model to determine first audio information based on reference audio content. The pre-trained model can be implemented as a machine learning model capable of determining audio information based on audio content. This document does not intend to limit the specific implementation or training process of the machine learning model. For example, electronic device 110 can send the determined first audio information to a first user. Alternatively, electronic device 110 can also provide the acquired reference audio content to the first user, so that a client or server 130 associated with the first user can determine the first audio information based on the reference audio content.

[0076] In other cases, continue to refer to Figure 3A Upon receiving a trigger from the second user on the first message 305 or control 310, the electronic device 110 may, in response to the second user's association with the first audio information, provide a provision control (or authorization control) associated with the first audio information. The first audio information is determined based on the second user's historical configuration operations. Further, the electronic device 110 may, in response to a trigger on the provision control, provide the first audio information to the first user.

[0077] Continue to refer to Figure 2E The electronic device 110 can be implemented as a client associated with the first user. For example, the second instruction can also instruct a third user. Accordingly, the first message (not shown) can also instruct a second reference (e.g., "@user2") to the third user (e.g., "@user2").

[0078] In some cases, the first media content is also generated based on second audio information from a third user. This second audio information may be obtained based on a second operation performed by the third user on the first message. For example, the process of obtaining the second audio information can be referred to the exemplary description of the process of obtaining the first audio information above, and will not be repeated here.

[0079] In some scenarios, electronic device 110 can output first media content. For example, a smart system can output the first media content to a first user. As an example, electronic device 110 can respond to a first instruction by presenting a conversational interface with the smart system (e.g., a voice call interface or a video call interface). Figure 2F As shown, electronic device 110 can present interface 200F. Interface 200F can be implemented as a conversational interface with an intelligent system. Electronic device 110 can output first media content through interface 200F. As an example, the first media content can include at least one of audio content and image content (e.g., picture content or video content). Alternatively, the audio content corresponding to the first media content can also be singing content. For example, a first user can listen to the intelligent system sing or sing along with the intelligent system through interface 200F. For example, the first media content output by the intelligent system through interface 200F can be a singing voice. Such a singing voice can be associated with the first audio information of a second user. Such a singing voice can match one or more audio attributes of the second user. For example, such a singing voice can have a similar timbre to that of the second user. Alternatively, the first instruction can also instruct the output of the first media content to the first user at a preset time.

[0080] In some cases, electronic device 110 may also receive a third instruction from the first user. This third instruction is associated with the intelligent system. For example, electronic device 110 may receive the third instruction from the first user via a first interface (e.g., a session interface with the intelligent system). The first interface may be associated with the intelligent system. As an example, the first interface may be an appropriate interactive interface related to the intelligent system. For example, the first interface may be an interactive interface for a session. This session may be, for example, a one-on-one chat session with the intelligent system; or, the session may be a group chat session involving multiple participants, and the multiple participants may include at least the first user and the intelligent system. As yet another example, the first interface may also be other appropriate interfaces for initiating interaction requests to the intelligent system, such as a virtual interactive space, a real-time call interface, etc.

[0081] For example, electronic device 110 may provide controls on a first interface. Electronic device 110 may obtain third instructions via the controls. For example, the controls may be associated with a preset theme (e.g., a holiday theme, an event theme, etc.).

[0082] Alternatively, the third instruction can also indicate the selection of a fourth user. For example, electronic device 110 can present a set of candidate users in response to a triggering of a control. Further, electronic device 110 can receive the third instruction based on the selection of a fourth user from the set of candidate users. For example, the fourth user can be any user (e.g., a second user, a third user, or other users). As an example, electronic device 110 can trigger the intelligent system to output first media content to the fourth user in response to the satisfaction of a preset condition. As an example, the preset condition is related to the third instruction. For example, the preset condition can be determined based on a control used to obtain the third instruction. Alternatively, the third instruction is also associated with the configuration operation of the first user. For example, the preset condition can also be determined based on the configuration operation of the first user. For example, the preset condition can indicate the time for outputting the first media content. Alternatively, the preset condition can also indicate whether the fourth user allows the output of the first media content.

[0083] In some cases, the third instruction may specify a predetermined time. For example, the predetermined time may be determined based on a control associated with a preset theme. For instance, the predetermined time indicated by a control associated with a holiday theme may correspond to midnight on the holiday day. Alternatively, the predetermined time may also be determined based on a configuration operation associated with the third instruction. For instance, the predetermined time may be configured by a first user. For example, electronic device 110 or server 130 may trigger the intelligent system to output first media content to a fourth user in response to the arrival of the predetermined time.

[0084] Next based on Figure 4A and Figure 4B The process of an intelligent system outputting first media content to a fourth user is described exemplarily.

[0085] Figure 4A and Figure 4B Example interfaces 400A and 400B are shown for interface interaction according to other scenarios. Interfaces 400A to 400B can, for example, be provided by... Figure 1 The electronic device 110 shown is provided. The electronic device 110 can, for example, be implemented as a client associated with a fourth user.

[0086] In some situations, electronic device 110 may trigger the intelligent system to send a second message to a fourth user in response to the fulfillment of preset conditions. For example... Figure 4AAs shown, electronic device 110 can present interface 400A to a fourth user. Interface 400A can be implemented as a session interface. For example, electronic device 110 can present a second message 405 sent by the intelligent system in the session on interface 400A. For example, the participants in such a session include at least the intelligent system and the fourth user. As an example, the second message 405 can indicate a third reference to a first user (e.g., user 1) (e.g., "@user 1"). For example, the second message 405 can include first media content. For example, electronic device 110 can play the first media content or present a viewing interface of the first media content in response to the fourth user's triggering of the second message 405 (e.g., a click operation). For example, the first media content can be implemented as at least one of audio content and image content (e.g., picture content and / or video content).

[0087] In other situations, electronic device 110 may, in response to the fulfillment of preset conditions, trigger the intelligent system to initiate a real-time audio or video call with a fourth user to output first media content. For example... Figure 4B As shown, electronic device 110 can present interface 400B to a fourth user. Interface 400B can be implemented as an interface of a client associated with the fourth user (e.g., system desktop, application interface, etc.). Electronic device 110 can present reminder information 410 associated with the smart system. For example, reminder information 410 can be overlaid on other display content of interface 400B. For example, reminder information 410 can indicate a real-time audio call request or a real-time video call request from the smart system. Electronic device 110 can receive a real-time audio call request or a real-time video call request from the smart system in response to a confirmation operation on reminder information 410 (e.g., triggering control 415). Furthermore, electronic device 110 can output first media content to the fourth user. For example, in a real-time audio call with the smart system, electronic device 110 can play audio content corresponding to the first media content. For example, in a real-time video call with the smart system, electronic device 110 can play audio content corresponding to the first media content and present video content corresponding to the first media content.

[0088] Based on the process described above, it is possible for a first user to obtain a second user's first audio information by sending a first message to the second user, thereby enriching the interaction methods between users. Furthermore, based on the first user's first instruction, it is possible to output first media content generated by the intelligent system based on the first audio information. Thus, when the intelligent system generates media content based on the first user's instruction, it can utilize specific audio information (e.g., the second user's first audio information) to generate first media content associated with the first audio information, enriching the ways in which the intelligent system generates media content.

[0089] Example 2: Figure 5A and Figure 5B Example interfaces 500A and 500B are shown, illustrating interface interactions under certain circumstances. Interfaces 500A and 500B can, for example, be provided by... Figure 1 The electronic device 110 shown is provided.

[0090] In some cases, such as Figure 5A As shown, electronic device 110 can present interface 500A to a first user. Interface 500A can be implemented as a first interface. The first interface is associated with an intelligent system. For example, the first interface is a session interface associated with the intelligent system. Electronic device 110 can receive a first instruction via the first interface.

[0091] In some cases, electronic device 110 may receive a first instruction in response to receiving an input message from a first user via a first interface. For example, electronic device 110 may receive the first user's input message via an input control in the first interface. For example, in some cases, electronic device 110 may present an input control 502 and a message display area 504 in the first interface (e.g., interface 500A). Electronic device 110 may receive the user's input message via the input control 502 and send the input message to the intelligent system. Accordingly, electronic device 110 may present the response content of the intelligent system in the message display area 504. In some cases, electronic device 110 may receive a first instruction in response to obtaining the first user's input message via the input control 502. For example, such an input message may instruct the intelligent system to generate media content. Such an input message may instruct a prompt for generating media content.

[0092] In some cases, electronic device 110 can receive a first instruction via a control. For example, electronic device 110 can present a control on a first interface (e.g., interface 500A). In some cases, electronic device 110 can present area 506 on the first interface (e.g., interface 500A). Area 506 can present at least one control (or tool component) associated with the intelligent system, such as control 508. In some scenarios, area 506 can also be called a toolbar or action bar. Area 506 can, for example, be presented near input component 505 to facilitate easier use of tools related to the intelligent system by the user. For example, electronic device 110 can receive a first instruction via control 508.

[0093] Alternatively, the electronic device 110 may present at least one control, such as control 510-1 and control 510-2, in the message display area 504. As an example, the electronic device 110 may, in response to a trigger on control 510-1, send an input message corresponding to control 510-1 to the intelligent system, and present the input message in the message display interface 204 associated with the identification information (e.g., avatar) of the first user. For example, the electronic device 110 may receive a first instruction via control 510-1.

[0094] In some cases, electronic device 110 may trigger the intelligent system to send first media content to a second user in response to the fulfillment of preset conditions. For example, at least one audio attribute of the first media content is associated with the first user. In some cases, the first media content is generated by the intelligent system based on the first user's audio information (also referred to as third audio information). For example, the third audio information is determined based on audio content recorded by the first user. The process by which the first user selects the third audio information can be referred to the exemplary description above, and will not be repeated here. The third audio information may indicate at least one audio attribute.

[0095] In some cases, the audio information mentioned herein (including, but not limited to, first audio information, second audio information, and third audio information) may indicate feature representations or data characterizing one or more audio attributes. Such audio attributes may include, but are not limited to, timbre, rhythm, and other attributes related to the state of sound expression. Taking timbre as an example, the audio information may include, for example, feature representations or data characterizing a specified timbre. Such audio information may be provided to an audio generation model or a suitable model with audio generation capabilities as a control condition to generate audio content that matches the specified audio attribute. For example, such feature representations or data may be provided to an audio generation model or a video model to guide the generation of media such that the audio of the media content matches the specified timbre.

[0096] In some cases, the first media content includes at least first audio content and second audio content. For example, the first audio attribute of the first audio content is related to a first user. The second audio attribute of the second audio content is related to a second user. For example, the first audio attribute can be determined based on the third audio information of the first user. The second audio attribute can be determined based on the first audio information of the second user. For example, the first media content can be generated by an intelligent system based on the audio information of the second user (also referred to as the second audio information) and the audio information of the first user (also referred to as the third audio information). For example, electronic device 110 can obtain the selection of audio options corresponding to the first audio information and audio options corresponding to the third audio information through the audio configuration interface mentioned above, so as to trigger the intelligent system to generate the first media content based on the first audio information and the third audio information.

[0097] In some cases, the second audio attribute (e.g., corresponding to the first audio information) is obtained based on the second user's first operation on the first message. The first message is sent to the second user based on the first user's second instruction. The process of sending the first message to the second user based on the second instruction and the process of obtaining the first audio information based on the second user's first operation on the first message can be referred to the exemplary description above, and will not be repeated here.

[0098] As an example, preset conditions are associated with a first instruction. For example, preset conditions may be determined based on a control used to obtain the first instruction (e.g., control 508 or control 510-1). Alternatively, preset conditions may also be determined based on an input message corresponding to the first instruction. For example, the input message may indicate preset conditions for outputting the first media content. Alternatively, a third instruction may also be associated with a configuration operation of a first user. For example, preset conditions may also be determined based on a configuration operation of the first user. For example, preset conditions may indicate the time for outputting the first media content. Alternatively, preset conditions may also indicate whether a second user allows the output of the first media content.

[0099] In some cases, the first instruction may indicate a predetermined time. For example, the predetermined time may be determined based on a control associated with a preset theme (e.g., control 508 or control 510-1). For example, the predetermined time indicated by a control associated with a holiday theme may correspond to midnight on the holiday day. For example, the predetermined time may be determined based on an input message corresponding to the first instruction. For example, an input message may indicate a predetermined time. Alternatively, the predetermined time may also be determined based on a configuration operation associated with the first instruction. For example, the predetermined time may be configured by a first user. For example, electronic device 110 or server 130 may trigger the intelligent system to output first media content to a second user in response to the arrival of the predetermined time.

[0100] In some scenarios, electronic device 110 may trigger the intelligent system to send a second message to a second user during a session in response to the fulfillment of preset conditions. The second message includes the first media content. The participants in the session include at least the intelligent system and the second user. As an example, the process by which the intelligent system sends a second message to the second user during a session can be referenced above regarding... Figure 4A An exemplary description of the second message 405 in the document is not repeated here.

[0101] In other cases, the participants in the session may also include the first user. For example... Figure 5B As shown, electronic device 110 can present interface 500B. Interface 500B can be implemented as a session interface. The participants in the session include a first user, a second user, and an intelligent system. For example, electronic device 110 can present a second message 515 from the intelligent system in the session interface 500B. The second message 515 can indicate a reference to the second user (e.g., user 2) (e.g., "@user 2") and a reference to the first user (e.g., user 1) (e.g., "@user 1") to indicate that the second message 515 was sent by the intelligent system based on an instruction (e.g., a first instruction) from the first user. For example, the second message 515 is configured to trigger the playback of first media content or the presentation of a viewing interface for the first media content. For example, the first media content can be implemented as at least one of audio content and image content (e.g., picture content and / or video content).

[0102] In some cases, electronic device 110 may also trigger the intelligent system to initiate a real-time audio or video call with a second user in response to the fulfillment of preset conditions, so as to output the first media content. The real-time audio or video call can be referred to the exemplary description above, and will not be repeated here.

[0103] In some cases, the theme of the first media content is related to a predetermined time. For example, the predetermined time may correspond to a holiday. Accordingly, the theme of the first media content may correspond to a holiday theme associated with the predetermined time. For example, the text content corresponding to the first media content may be associated with a holiday theme. Alternatively, the visual attributes of the first media content may be related to the predetermined time. As an example, the visual content of the first media content may correspond to a holiday theme associated with the predetermined time. For example, if the holiday theme corresponds to New Year's Day, then the visual content of the first media content may include visual elements associated with New Year's Day.

[0104] In some cases, the first instruction also directs a third user. The electronic device 110 can also trigger the intelligent system to send second media content to the third user in response to the fulfillment of preset conditions. At least one audio attribute of the second media content is related to the first user. For example, the second media content may be generated by the intelligent system and third audio information from the first user. The third audio information may indicate at least one audio attribute of the second media content.

[0105] In some cases, the content of the first media differs from the content of the second media. For example, the content of the first media may also be generated based on the first reference information of the second user. The content of the second media may also be generated based on the second reference information of the third user. For example, the reference information may indicate the corresponding user's identification information (e.g., image identifiers and / or text identifiers), personal profile, historical interaction information (e.g., historical interactions with the first user), etc.

[0106] It should be noted that the user-related information mentioned in this document (including but not limited to the first and second reference information) was obtained and used with the knowledge and authorization of the relevant user.

[0107] Based on the process described above, it is possible to receive a first instruction from a first user via a first interface associated with the intelligent system. Furthermore, in response to the fulfillment of preset conditions associated with the first instruction, the intelligent system can be triggered to send first media content to a second user. This allows the first user to send first media content to the second user through the intelligent system, thereby enriching the interaction methods between the first and second users. In addition, at least one audio attribute of the first media content is associated with the first user, thus enabling the intelligent system to send media content associated with the first user's audio information to the second user, further enriching the interaction methods between the first and second users and improving the efficiency of interaction between users.

[0108] Example process Figure 6 A flowchart of an example process 600 for interface interaction under certain conditions is shown. Process 600 can be implemented at electronic device 110. See below for reference. Figure 1 To describe process 600.

[0109] like Figure 6 As shown, in box 610, electronic device 110 receives a first instruction from a first user, and the first instruction is associated with an intelligent system.

[0110] In frame 620, electronic device 110 outputs first media content, which is generated by intelligent system based on first audio information of second user. The first audio information is obtained based on first operation of second user on first message, and the first message is sent to second user based on second instruction of first user.

[0111] In some cases, process 600 further includes: in response to a second instruction, sending a first message in a session, wherein the participants in the session include at least a first user and a second user.

[0112] In this way, the first user can obtain the second user's first audio information by sending messages to the second user in the session, which helps to enrich the interaction between the first user and the second user.

[0113] In some cases, the participants in the session include at least three users, where the first message indicates a first reference to the second user.

[0114] In this way, the first user can send a first message in a multi-user session that indicates a reference to the second user, thereby enriching the ways in which the first user can obtain the first audio information of the second user.

[0115] In some cases, the second instruction also directs to a third user, and the first message also directs to a second reference to the third user.

[0116] In some cases, the first media content is also generated based on the second audio information of a third user, which is obtained based on the third user's second operation on the first message.

[0117] In this way, the first user can obtain the first audio information of the second user and the second audio information of the third user by sending the first message to the session, thereby improving the efficiency of the first user in obtaining the audio information of the relevant users.

[0118] In some cases, the primary media content is also generated based on the third-party audio information of the primary user.

[0119] In this way, it is possible to support the generation of first media content based on the third audio information of the first user and the first audio information of the second user, thereby enriching the generation methods of first media content.

[0120] In some cases, the first audio information corresponds to reference audio content from a second user, and the reference audio content is received based on the following process: providing an audio component to the second user in response to a first operation; and obtaining the reference audio content through the audio component.

[0121] In this way, a second user can obtain reference audio content through the audio component provided by the first message, thereby improving the efficiency of obtaining the second user's first audio information.

[0122] In some cases, obtaining reference audio content through the audio component includes: recording reference audio content through the audio component, the reference audio content being associated with text content, and the text content being associated with the second instruction.

[0123] In this way, the first user can obtain the corresponding reference audio content by specifying text content, thereby enriching the ways in which the first user obtains the first audio information from the second user.

[0124] In some cases, process 600 further includes: receiving a third instruction associated with the intelligent system; and triggering the intelligent system to output first media content to a fourth user in response to a preset condition being met, the preset condition being associated with the third instruction.

[0125] In some cases, the third instruction indicates a predetermined time, and in response to the satisfaction of preset conditions, triggers the intelligent system to output the first media content to the fourth user, including: in response to the arrival of the predetermined time, triggering the intelligent system to output the first media content to the fourth user.

[0126] In some cases, triggering the intelligent system to output the first media content to the fourth user includes: triggering the intelligent system to send a second message to the fourth user, the second message including the first media content; or triggering the intelligent system to initiate a real-time audio call or real-time video call with the fourth user to output the first media content.

[0127] In this way, the intelligent system can output the first media content to the fourth user through messages, real-time audio calls, or real-time video calls, thereby enriching the ways in which the intelligent system outputs media content to users and enriching the interaction between the first user and the fourth user.

[0128] In some cases, receiving a first instruction includes: presenting a control associated with a preset theme; and receiving a first instruction via the control, wherein first media content is associated with the preset theme.

[0129] In this way, it is possible to support the generation of primary media content based on a preset theme associated with the control.

[0130] Figure 7 A flowchart of an example process 700 for interface interaction under certain circumstances is shown. Process 700 can be implemented in electronic device 110. See below for reference. Figure 1 To describe process 700.

[0131] like Figure 7As shown in box 710, electronic device 110 presents a first interface to a first user, and the first interface is associated with an intelligent system.

[0132] In frame 720, electronic device 110 receives a first instruction via a first interface.

[0133] In frame 730, electronic device 110 responds to the fulfillment of a preset condition by triggering the intelligent system to send first media content to a second user, wherein at least one audio attribute of the first media content is related to the first user, and the preset condition is related to the first instruction.

[0134] In some cases, the first instruction indicates a predetermined time, and in response to the satisfaction of preset conditions, triggering the intelligent system to output the first media content to the second user includes: in response to the arrival of the predetermined time, triggering the intelligent system to output the first media content to the second user.

[0135] In this way, it is possible to support the intelligent system to output the first media content to the second user based on the predetermined time specified by the first instruction, thereby enriching the ways of outputting media content.

[0136] In some cases, the content theme or visual attributes of primary media content are related to the scheduled time.

[0137] In this way, the content theme or visual attributes of the first media content can be determined based on a predetermined time, thereby linking the generation process of the first media content with the predetermined time and enriching the ways in which media content is generated.

[0138] In some cases, the first instruction also instructs a third user, and process 700 further includes: in response to a preset condition being met, triggering the intelligent system to send second media content to the third user, wherein at least one audio attribute of the second media content is related to the first user.

[0139] In some cases, the first media content differs from the second media content. The first media content is also generated based on the first reference information of the second user, and the second media content is also generated based on the second reference information of the third user.

[0140] This method enables the sending of different media content to a second and third user based on a first instruction, thereby improving the efficiency of user interaction. Furthermore, it allows for the generation of corresponding media content based on the user's reference information, which enhances the relevance between the media content and the recipient, and further improves the efficiency of user interaction.

[0141] In some cases, in response to the satisfaction of preset conditions, triggering the intelligent system to send the first media content to the second user includes: in response to the satisfaction of preset conditions, triggering the intelligent system to send the first message to the second user in the session, the first message including the first media content, and the participants in the session include at least the intelligent system and the second user.

[0142] In some cases, the participants in the session also include the first user.

[0143] In some cases, the first media content includes at least a first audio content and a second audio content, where the first audio attribute of the first audio content is related to a first user, and the second audio attribute of the second audio content is related to a second user.

[0144] This method enables the generation of first media content, which includes first audio content and second audio content, thereby enriching the ways in which media content is generated and improving the efficiency of interaction between users.

[0145] In some cases, the second audio attribute is obtained based on the second user's first operation on the first message, which is sent to the second user based on the first user's second instruction.

[0146] Example devices and equipment A corresponding apparatus for implementing the above methods or processes is also provided.

[0147] Figure 8 A schematic structural block diagram of an example device 800 for interface interaction is shown, according to some scenarios. Device 800 can be implemented as or included in electronic device 110. The various modules / components in device 800 can be implemented by hardware, software, firmware, or any combination thereof.

[0148] like Figure 8 As shown, the device 800 includes: a first receiving module 810 configured to receive a first instruction from a first user, the first instruction being associated with an intelligent system; and an output module 820 configured to output first media content generated by the intelligent system based on first audio information from a second user, wherein the first audio information is obtained based on a first operation of the second user on a first message, and the first message is sent to the second user based on a second instruction from the first user.

[0149] In some cases, the device 800 also includes a sending module configured to: in response to a second instruction, send a first message in a session, wherein the participants in the session include at least a first user and a second user.

[0150] In some cases, the participants in the session include at least three users, where the first message indicates a first reference to the second user.

[0151] In some cases, the second instruction also directs to a third user, and the first message also directs to a second reference to the third user.

[0152] In some cases, the first media content is also generated based on the second audio information of a third user, which is obtained based on the third user's second operation on the first message.

[0153] In some cases, the primary media content is also generated based on the third-party audio information of the primary user.

[0154] In some cases, the first audio information corresponds to reference audio content from a second user, and the reference audio content is received based on the following process: providing an audio component to the second user in response to a first operation; and obtaining the reference audio content through the audio component.

[0155] In some cases, obtaining reference audio content through the audio component includes: recording reference audio content through the audio component, the reference audio content being associated with text content, and the text content being associated with the second instruction.

[0156] In some cases, the device 800 also includes a third receiving module configured to: receive a third instruction associated with the intelligent system; and, in response to a preset condition being met, trigger the intelligent system to output first media content to a fourth user, the preset condition being associated with the third instruction.

[0157] In some cases, the third instruction indicates a predetermined time, and the third receiving module is also configured to trigger the intelligent system to output the first media content to the fourth user in response to the arrival of the predetermined time.

[0158] In some cases, the third receiving module is also configured to: trigger the intelligent system to send a second message to the fourth user, the second message including the first media content; or trigger the intelligent system to initiate a real-time audio call or real-time video call with the fourth user to output the first media content.

[0159] In some cases, the first receiving module 810 is also configured to: present a control associated with a preset theme; and receive a first instruction via the control, wherein the first media content is associated with the preset theme.

[0160] Figure 9 A schematic structural block diagram of an example device 900 for interface interaction is shown, according to some scenarios. Device 900 can be implemented as or included in electronic device 110. The various modules / components in device 900 can be implemented by hardware, software, firmware, or any combination thereof.

[0161] like Figure 9As shown, the device 900 includes: a first presentation module 910 configured to present a first interface to a first user, the first interface being associated with an intelligent system; a second receiving module 920 configured to receive a first instruction via the first interface; and a first trigger module 930 configured to trigger the intelligent system to send first media content to the second user in response to a preset condition being met, wherein at least one audio attribute of the first media content is associated with the first user, and the preset condition is associated with the first instruction.

[0162] In some cases, the first instruction indicates a predetermined time, and the first trigger module 930 is also configured to trigger the intelligent system to output the first media content to the second user in response to the arrival of the predetermined time.

[0163] In some cases, the content theme or visual attributes of primary media content are related to the scheduled time.

[0164] In some cases, the first instruction also instructs a third user, and the device 900 further includes a second trigger module configured to: in response to a preset condition being met, trigger the intelligent system to send second media content to the third user, wherein at least one audio attribute of the second media content is related to the first user.

[0165] In some cases, the first media content differs from the second media content. The first media content is also generated based on the first reference information of the second user, and the second media content is also generated based on the second reference information of the third user.

[0166] In some cases, the first trigger module 930 is also configured to: in response to the satisfaction of a preset condition, trigger the intelligent system to send a first message to the second user in the session, the first message including first media content, and the participants in the session include at least the intelligent system and the second user.

[0167] In some cases, the participants in the session also include the first user.

[0168] In some cases, the first media content includes at least a first audio content and a second audio content, where the first audio attribute of the first audio content is related to a first user, and the second audio attribute of the second audio content is related to a second user.

[0169] In some cases, the second audio attribute is obtained based on the second user's first operation on the first message, which is sent to the second user based on the first user's second instruction.

[0170] The modules included in device 800 or 900 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some cases, one or more modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 800 or 900 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0171] Figure 10 A block diagram of an electronic device 1000 in which one or more examples can be implemented is shown. It should be understood that... Figure 10 The electronic device 1000 shown is merely exemplary and should not be construed as limiting the functionality and scope of the examples described herein. Figure 10 The electronic device 1000 shown can be used to implement the electronic device 110 discussed above.

[0172] like Figure 10 As shown, electronic device 1000 is in the form of a general-purpose electronic device. Components of electronic device 1000 may include, but are not limited to, one or more processing units or processors 1010, memory 1020, storage device 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. Processor 1010 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1020. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 1000.

[0173] Electronic device 1000 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1020 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1030 can be removable or non-removable media and may include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 1000.

[0174] Electronic device 1000 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 10 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 1020 may include computer program product 1025 having one or more program modules configured to perform various methods or actions of various examples.

[0175] The communication unit 1040 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 1000 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 1000 can operate in a networked environment using logical connections to one or more other servers, networked personal computers, or another network node.

[0176] Input device 1050 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1060 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 1000 can also communicate with one or more external devices (not shown) via communication unit 1040 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 1000, or with any device that enables electronic device 1000 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0177] A computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. A computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0178] The flowcharts and / or block diagrams of the methods, apparatus, devices, and computer program products referred to herein describe various aspects. It should be understood that each block of the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0179] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0180] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0181] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0182] Various examples have been described above. The foregoing descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for user interface interaction, comprising: Receive a first instruction, which comes from a first user and is associated with an intelligent system; as well as The system outputs first media content, which is generated by the intelligent system based on first audio information of the second user. The first audio information is obtained based on the second user's first operation on a first message, and the first message is sent to the second user based on the first user's second instruction.

2. The method according to claim 1, further comprising: In response to the second instruction, the first message is sent in the session, wherein the participants in the session include at least the first user and the second user.

3. The method of claim 2, wherein the participants in the session include at least three users, and wherein the first message indicates a first reference to the second user.

4. The method of claim 2, wherein the second instruction further instructs a third user, and the first message further instructs a second reference to the third user.

5. The method according to claim 4, wherein the first media content is further generated based on the second audio information of the third user, the second audio information being obtained based on the second operation of the third user on the first message.

6. The method according to claim 1, wherein the first media content is further generated based on the third audio information of the first user.

7. The method of claim 1, wherein the first audio information corresponds to reference audio content from the second user, and the reference audio content is received based on the following process: In response to the first operation, an audio component is provided to the second user; and The reference audio content is obtained through the audio component.

8. The method of claim 7, wherein obtaining the reference audio content through the audio component comprises: The reference audio content is recorded by the audio component, and the reference audio content is associated with the text content, which is related to the second instruction.

9. The method according to claim 1, further comprising: Receive a third instruction, which is associated with the intelligent system; as well as In response to the fulfillment of a preset condition, the intelligent system is triggered to output the first media content to the fourth user, and the preset condition is related to the third instruction.

10. The method of claim 9, wherein the third instruction indicates a predetermined time, and triggering the intelligent system to output the first media content to the fourth user in response to a preset condition being met comprises: In response to the arrival of the predetermined time, the intelligent system is triggered to output the first media content to the fourth user.

11. The method of claim 9, wherein triggering the intelligent system to output the first media content to the fourth user comprises: The intelligent system is triggered to send a second message to the fourth user, the second message including the first media content; or The intelligent system is triggered to initiate a real-time audio or video call with the fourth user to output the first media content.

12. The method of claim 1, wherein receiving the first instruction comprises: Present controls that are associated with a preset theme; as well as The first instruction is received via the control, and the first media content is related to the preset theme.

13. A method for user interface interaction, comprising: A first interface is presented to the first user, and this first interface is associated with the intelligent system; Receive the first instruction via the first interface; as well as In response to the fulfillment of a preset condition, the intelligent system is triggered to send first media content to the second user, wherein at least one audio attribute of the first media content is related to the first user, and the preset condition is related to the first instruction.

14. The method of claim 13, wherein the first instruction indicates a predetermined time, and triggering the intelligent system to output the first media content to the second user in response to a preset condition being met comprises: In response to the arrival of the predetermined time, the intelligent system is triggered to output the first media content to the second user.

15. The method of claim 13, wherein the first instruction further instructs a third user, the method further comprising: In response to the preset condition being met, the intelligent system is triggered to send second media content to a third user, wherein at least one audio attribute of the second media content is related to the first user.

16. The method of claim 15, wherein the first media content is different from the second media content, the first media content is further generated based on the first reference information of the second user, and the second media content is further generated based on the second reference information of the third user.

17. The method of claim 13, wherein the first media content includes at least first audio content and second audio content, a first audio attribute of the first audio content is related to the first user, and a second audio attribute of the second audio content is related to the second user.

18. A device for user interface interaction, comprising: The first receiving module is configured to receive a first instruction, which comes from a first user and is associated with an intelligent system. as well as The output module is configured to output first media content, which is generated by the intelligent system based on first audio information of a second user, wherein the first audio information is obtained based on a first operation of the second user on a first message, and the first message is sent to the second user based on a second instruction of the first user.

19. An electronic device comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 12 or 13 to 17 when executed by the at least one processor.

20. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 12 or 13 to 17.