Multimedia content generation method, electronic device, storage medium, and product

By generating multimedia content for digital clones during user-agent dialogues and pushing related object information, the problem of digital clone applications being limited to physical appearance is solved, achieving a more efficient information acquisition and interactive experience.

WO2026066123A1PCT designated stage Publication Date: 2026-04-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

In existing technologies, the application of digital avatars is mainly limited to providing an image, lacking further multimedia content generation and information recommendation functions, resulting in low user interaction efficiency.

Method used

By generating multimedia content for a digital avatar based on messages sent by the user during dialogue with the intelligent agent, and pushing information about related objects, the system utilizes generative models and natural language processing techniques to achieve multimedia content generation and information recommendation.

Benefits of technology

It improves the efficiency and interactive experience of users in obtaining information, enabling users to intuitively browse information related to content of interest, and enhances the application value of digital clones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025094503_02042026_PF_FP_ABST
    Figure CN2025094503_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of multimedia, and relates to a multimedia content generation method, an electronic device, a storage medium, and a product. The multimedia content generation method comprises: displaying a digital avatar of a user on a dialog interface, wherein the digital avatar is generated on the basis of an image from the user; on the dialog interface, displaying a message sent by the user; on the basis of the digital avatar and the message, generating multimedia content of the digital avatar; and on the dialog interface, displaying the multimedia content sent to the user, and information of an object associated with the multimedia content.
Need to check novelty before this filing date? Find Prior Art

Description

Method for generating multimedia content, electronic device, storage medium and product

[0001] Cross-reference to Related Applications

[0002] This application is based on and claims priority to the application with the Chinese application number 202411376097.6, the filing date of September 29, 2024, the disclosure of which is hereby incorporated by reference in its entirety into this application. TECHNICAL FIELD

[0003] The present disclosure relates to the field of multimedia technology, and in particular, to a method for generating multimedia content, an electronic device, a storage medium, and a product. BACKGROUND

[0004] In related technologies, a user can generate a digital avatar of the user based on an image provided by the user. For example, the user can select a preset template. The template includes some prompt information. Based on the prompt information, a digital avatar matching the template can be generated. For example, the template is a travel photo at a certain scenic spot, and the digital avatar can be a photo of the user traveling at the scenic spot. SUMMARY

[0005] This summary is provided to introduce a selection of concepts, which are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in limiting the scope of the claimed subject matter.

[0006] According to some embodiments of the present disclosure, a method for generating multimedia content is provided, including: displaying a digital avatar of a user in a conversation interface, the digital avatar being generated based on an image from the user; displaying a message sent by the user in the conversation interface; generating multimedia content of the digital avatar based on the digital avatar and the message; and displaying the multimedia content sent to the user and information of an object associated with the multimedia content in the conversation interface.

[0007] According to some embodiments of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, the processor being configured to execute a method for generating multimedia content according to any of the embodiments of the present disclosure based on instructions stored in the memory.

[0008] According to some embodiments of the present disclosure, a computer-readable storage medium is provided, having a computer program stored thereon, the program being executed by a processor to perform a method for generating multimedia content according to any of the embodiments of the present disclosure.

[0009] According to some embodiments of the present disclosure, a computer program product is provided, which, when running on a computer, causes the computer to implement the method for generating multimedia content of any of the embodiments of the present disclosure.

[0010] According to some embodiments of the present disclosure, a computer program is provided, comprising instructions which, when executed by a processor, cause the processor to perform the method for generating multimedia content of any of the embodiments described in the present disclosure.

[0011] Other features, aspects, and advantages of the present disclosure will become apparent from the following detailed description of the exemplary embodiments of the present disclosure with reference made to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0012] The preferred embodiments of the present disclosure will be described herein below with reference to the accompanying drawings. The accompanying drawings are used in the provision of further understanding of the present disclosure, and together with the specific description of the preferred embodiments below, form a part of the detailed description of the present disclosure. It should be understood that the accompanying drawings are only some embodiments of the present disclosure and do not constitute a limitation on the present disclosure. In the drawings:

[0013] FIG. 1 shows a flowchart of a method for generating multimedia content according to some embodiments of the present disclosure.

[0014] FIG. 2 shows a schematic diagram of a dialogue interface according to some embodiments of the present disclosure.

[0015] FIG. 3 shows a flowchart of a method for determining an associated object according to some embodiments of the present disclosure.

[0016] FIG. 4 shows a flowchart of a method for determining an associated object according to some other embodiments of the present disclosure.

[0017] FIG. 5 shows a schematic diagram of a dialogue interface according to some other embodiments of the present disclosure.

[0018] FIG. 6 shows a flowchart of a method for generating multimedia content of a digital avatar according to some embodiments of the present disclosure.

[0019] FIG. 7 shows a structural diagram of a device for generating multimedia content according to some embodiments of the present disclosure.

[0020] FIG. 8 shows a structural diagram of an electronic device according to some embodiments of the present disclosure.

[0021] FIG. 9 shows a structural diagram of a computer system according to some embodiments of the present disclosure.

[0022] It should be understood that the dimensions of the various parts shown in the drawings are not necessarily to scale. Identical or similar components are identified throughout the various figures with identical or similar reference numerals. Therefore, when a component is identified in one figure, it can not be further discussed in subsequent figures. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, but not all the embodiments. The description of the embodiments below is actually only illustrative, and should not be construed as any limitation on the present disclosure and its application or use. It should be understood that the present disclosure can be implemented in various forms, and should not be construed as being limited to the embodiments set forth herein.

[0024] It should be understood that the various steps recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this respect. Unless specifically stated otherwise, the relative arrangement of components and steps, numerical expressions, and numerical values set forth in these embodiments should be interpreted as merely exemplary, not limiting the scope of the present disclosure.

[0025] The term "comprise" and variations thereof used in the present disclosure means an open term that includes at least the recited elements / features, but does not exclude other elements / features. In addition, the term "include" and variations thereof used in the present disclosure means an open term that includes at least the recited elements / features, but does not exclude other elements / features. Therefore, include and comprise are synonymous. The term "based on" means "at least partially based on".

[0026] Throughout the specification, the term "one embodiment", "some embodiments", or "embodiments" means that the specific feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. For example, the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; and the term "some embodiments" means "at least some embodiments". Moreover, the appearance of the phrase "in one embodiment", "in some embodiments", or "in embodiments" at various places in the specification does not necessarily all refer to the same embodiment, but can refer to different embodiments.

[0027] It should be noted that the terms "first", "second", and the like in the present disclosure are merely intended to distinguish different devices, modules, or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules, or units. Unless otherwise specified, the terms "first", "second", and the like are not intended to imply a given order or any other manner of given order in time, space, ranking, or any other manner.

[0028] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that unless otherwise explicitly specified in the context, it should be understood as "one or more".

[0029] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0030] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described in detail in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics can be combined by any suitable means from the present disclosure that is clear to those skilled in the art.

[0031] It should be understood that the present disclosure does not limit how to obtain the image to be applied / processed. In one embodiment of the present disclosure, the image can be obtained from a storage device, such as an internal memory or an external storage device, and in another embodiment of the present disclosure, a photographing component can be activated to take a picture. It should be noted that the obtained image can be a captured image or a frame image from a captured video, and is not particularly limited thereto.

[0032] In the context of the present disclosure, an image can refer to any of a variety of images, such as a color image, a grayscale image, etc. It should be noted that the type of image is not specifically limited in the context of the present specification. In addition, the image can be any suitable image, such as a raw image obtained by a camera, or an image that has been subjected to certain processing of the raw image, such as preliminary filtering, de-aliasing, color adjustment, contrast adjustment, normalization, etc. It should be noted that the pre-processing operation can also include other types of pre-processing operations known in the art, which will not be described in detail here.

[0033] In the related art, the role of digital avatars often remains at the stage of providing various images for users, i.e., mainly meeting the user's demand for "aesthetics". There are few further applications of digital avatars.

[0034] The application provides a method for generating multimedia content. In the method, a digital avatar's multimedia content is generated based on a message sent by a user in a conversation between the user and an intelligent agent, and information about an object associated with the multimedia content is sent to the user. The digital avatar's multimedia content can be used to respond to the message sent by the user.

[0035] FIG. 1 shows a flowchart of a method for generating multimedia content according to some embodiments of the present disclosure. As shown in FIG. 1, the method according to the embodiments includes steps S102-S108.

[0036] In step S102, a digital avatar of the user is displayed on a conversation interface, and the digital avatar is generated based on an image from the user.

[0037] The conversation interface refers to a conversation interface between the user and the intelligent agent, such as a chat interface. In the conversation interface, the user can send messages of various types, such as text, voice, multimedia content (e.g., images, videos, etc.), files, etc., to the intelligent agent. The intelligent agent will respond to the user based on the content in the message sent by the user, such as replying to the user with one or more messages. The messages replied by the intelligent agent can also be of various types. In some embodiments, the conversation interface can also involve more than two objects participating in the conversation, such as a group chat interface.

[0038] The intelligent agent can be implemented in software, hardware, or a combination of software and hardware. The intelligent agent can be implemented based on a machine learning model, such as a large language model (LLM) or a foundation model. The machine learning model can be a generative model that outputs target content based on input information. The input information may, for example, include a prompt. The prompt may, for example, include a conversation between the user and the intelligent agent, or content extracted from the conversation, or other information.

[0039] The generative model may, for example, include a model that generates based on text or a model that generates based on images. The output of the generative model may, for example, include text, images, or a combination of both. Of course, the input or output of the generative model can also be other modalities of data, such as audio, video, or a combination of multiple types of data. The generative model may, for example, be a single-modal model, such as a model that generates text based on text (referred to as a "text-to-text model"), a model that generates images based on images (referred to as an "image-to-image model"), or a cross-modal model, i.e., a model whose input and output belong to different modalities, such as a model that generates images based on text (referred to as a "text-to-image model"). Alternatively, the input of the generative model can include multiple modalities, and the output can also include multiple modalities.

[0040] A digital avatar of a user refers to a virtual object created for the user with the user's authorization. The virtual object can be a two-dimensional or three-dimensional figure, whose figure is generated according to an image from the user. The image is an image authorized by the user, including an image uploaded or selected by the user, etc. For example, in the case that the image includes the user, the appearance of the generated digital avatar can be close to the user.

[0041] In generating the digital avatar, in addition to referring to the image from the user, description information sent by the user can also be referred to. Thus, the appearance of the generated digital avatar can be close to the image from the user and match the description information. For example, the user provides a photo of himself and adds a description "make the eyes bigger", and a digital avatar is generated which is close to the user in appearance but has bigger eyes.

[0042] The user can create a digital avatar through a creation interface of the digital avatar, or through a conversation interface. Taking the latter as an example, the user can send a message including an image to the intelligent agent, and send a message including an instruction to create a digital avatar (for example, "generate a digital avatar for me according to this photo"). In addition, the user can also send a message including some description information to the intelligent agent, such as "add natural makeup and whitening effect".

[0043] FIG. 2 shows a schematic diagram of a conversation interface according to some embodiments of the present disclosure. As shown in FIG. 2, the conversation interface 2 displays a conversation between the user and the intelligent agent A. The content of the message 21 sent by the user is "help me generate an avatar photo: princess, long skirt, morning light, natural makeup, whitening effect", and the intelligent agent A can generate a digital avatar based on the message and an image from the user. After generating the digital avatar, the digital avatar can be displayed through the message 22.

[0044] The digital avatar displayed in the conversation interface can be a three-dimensional virtual object. Alternatively, it can also be one or more pictures, for example, a screenshot of one or more angles of the three-dimensional virtual object. Of course, it can also be displayed in other forms, such as a video, etc., which will not be described here.

[0045] In step S104, the message sent by the user is displayed in the conversation interface.

[0046] The message sent by the user can be of various types. By performing semantic understanding on the message, it can be determined whether the message includes an instruction to generate multimedia content, or information related to the multimedia content to be generated, such as the theme of the multimedia content to be generated, contained in the message. The semantic understanding can be based on a natural language processing model.

[0047] In some embodiments, after displaying the digital avatar, the user sends the message to the intelligent agent so as to send an instruction to further generate multimedia content based on the digital avatar.

[0048] In step S106, multimedia content of the digital avatar is generated based on the digital avatar and the message sent by the user. The generated multimedia content can be an image, a video, or the like including the digital avatar, and the multimedia content is associated with the message. At least one of the action, the clothing, the scene, or the like of the digital avatar in the generated multimedia content can be different from the originally displayed digital avatar. That is, the generated multimedia content can be a result of the digital avatar being presented in another style, and the style is determined based on the message sent by the user.

[0049] For example, if the message sent by the user includes or is related to a specified theme, the generated multimedia content is also related to the theme.

[0050] In step S108, the multimedia content sent to the user and the information of the object associated with the multimedia content are displayed on the conversation interface. That is, the intelligent agent can send one or more messages to the user, and the messages include the generated multimedia content and the information of the object associated with the multimedia content.

[0051] The object associated with the multimedia content can directly appear in the multimedia content, such as an item held or worn by the digital avatar. Alternatively, an element in the multimedia content can be related to the object, for example, the multimedia content is an image of the digital avatar at a scenic spot, and the object associated with the multimedia content can be a ticket to the scenic spot, or a hotel near the scenic spot, or the like.

[0052] The information of the object can be description information of the object, such as a description text, a description image, or a description video. Alternatively, the information of the object can also be a jump control, and in response to a triggering operation on the jump control of the object, a display interface or a transaction interface of the object is displayed. Thus, the conversation interface can be used to recommend the item to the user.

[0053] Through the above embodiments, after creating the digital avatar of the user, the multimedia content of the digital avatar can be generated based on the message sent by the user, and the information of the object associated with the multimedia content can be pushed to the user at the same time. Thus, in the process of the conversation between the intelligent agent and the user, the multimedia content of the digital avatar of the user can be sent based on the content of the conversation, that is, the user is responded to with the multimedia content of the digital avatar and the associated object information. In this way, the information that the user wants to obtain can be associated with the digital avatar of the user, and the recommended object is also associated with the digital avatar. Therefore, the user can intuitively browse the information related to the content that the user is interested in or the content that the user can be interested in, and the efficiency of obtaining information by the user is improved.

[0054] Since the multimedia content generated by the digital avatar is only a part of the scenario of the user's conversation with the digital avatar, many conversations between the user and the digital avatar can not involve the above scenario. Therefore, in order to reasonably respond to the user, the message can be subjected to intent recognition, and in response to the message sent by the user having an intent of creating multimedia content based on the digital avatar, the multimedia content generated by the digital avatar is generated. The intent recognition of the message can be determined by keyword recognition, processing by an intent analysis model, or by a natural language model (such as a large language model). If the message sent by the user does not have an intent of creating multimedia content based on the digital avatar, the multimedia content generated by the digital avatar is not generated.

[0055] The generated multimedia content can have a certain theme. For example, the theme can be determined based on the message sent by the user, and the multimedia content can be generated according to the specified theme, or the theme can be extracted after the multimedia content is directly generated based on the message sent by the user. Then, based on the theme, an object associated with the multimedia content can be determined. The two methods of determining the associated object are described below.

[0056] FIG. 3 shows a flowchart of a method for determining an associated object according to some embodiments of the present disclosure. As shown in FIG. 3, the determining method of this embodiment includes steps S302 to S306.

[0057] In step S302, a theme of the multimedia content to be generated is determined based on the message sent by the user.

[0058] The theme can be directly extracted from the message, or indirectly determined based on the message. In some embodiments, reference information can be generated based on the message, which can be generated by expanding or responding to the message, and then the theme is extracted from the reference information.

[0059] The operation of extracting the theme can be implemented using a language model, a theme analysis model, or other natural language processing models. The operation of generating the reference information can be implemented by a generative model.

[0060] For example, the message sent by the user is "generate a birthday greeting video", and the theme "birthday greeting" can be directly extracted from the message. For another example, the message sent by the user is "what color of clothes do you think I am suitable for?", and after semantic understanding and analysis of the shape and attributes of the digital avatar, it is determined that the user is suitable for blue clothes, then the theme can be determined as "blue clothes".

[0061] In step S304, image materials are searched based on the theme.

[0062] For example, the topic can be directly searched as a search word for image material, or image material can be searched based on the topic and synonyms, near-synonyms, etc. of the topic. The image material can be in an image format, such as an image of certain articles or places; or the image material can also be description information or processing information of the image, such as color, filter, special effect, etc.

[0063] In some embodiments, an object associated with the topic is searched, and image material associated with the object is determined. Some topics determined based on the message of the user can be general, but cannot be corresponded to a specific product. For example, the topic is “blue clothes”, and existing blue clothes products, such as a product of a certain brand and a certain style, can be searched as an object associated with the topic. Then, image material associated with the product, such as product pictures, product elements, etc. are taken as image material. Thus, the multimedia content of the generated digital avatar can reflect the relationship between the digital avatar and the actual product, such as reflecting the effect of the user wearing the product. Therefore, the generated multimedia content has more reference value for the user, and the relationship with the recommended object is closer, and the user can efficiently obtain valuable information.

[0064] The image material associated with the object includes at least one of image material of the object, image material of an article provided by the object, and image material of a use effect of the object. Taking the object as a lipstick as an example, the image material of the object refers to the image material of the lipstick itself, and the image material of the use effect of the object refers to the image material of the color of the lips after using the lipstick. If the lipstick is an article provided by an object, the object can be a store selling the lipstick.

[0065] In step S306, the multimedia content of the digital avatar is generated based on the digital avatar and the image material. For example, the digital avatar and the image material can be fused or processed.

[0066] In some embodiments, the superimposition manner of the image material and the image of the digital avatar, such as angle, position, etc. can be determined based on the message of the user and the image material, and then image fusion processing is performed. In some embodiments, the digital avatar and the image material can also be fused based on an image generation model, such as a “graph-to-graph model”.

[0067] Since the image material is searched based on the topic, the generated multimedia content can also reflect the topic. Further, since the topic is determined based on the message sent by the user, the generated multimedia content can effectively respond to the message sent by the user.

[0068] The above embodiment determines the theme based on the message sent by the user, searches for image materials based on the theme, and generates the multimedia content of the digital avatar. Thus, the generated multimedia content can correspond to a specific object, so that when the information of the object is sent to the user for recommendation, the user can intuitively show the user the effect associated with the object through the multimedia content of the digital avatar. If the user is satisfied with the display effect, the user can further understand the recommended object according to the information of the object sent.

[0069] FIG. 4 shows a flowchart of a method for determining an associated object according to some embodiments of the present disclosure. As shown in FIG. 4, the generation method of this embodiment includes steps S402-S406.

[0070] In step S402, the multimedia content of the digital avatar is generated based on the digital avatar and the message.

[0071] In this embodiment, for example, the content in the message or the content after the message is expanded or replied can be directly used as the prompt information to generate the multimedia content of the digital avatar. Thus, the generated multimedia content can match the message sent by the user.

[0072] In step S404, the theme of the multimedia content of the digital avatar is identified.

[0073] The multimedia content of the digital avatar can be processed by using a processing model of the multimedia content to obtain the theme therein. For example, a description text of the multimedia content can be generated based on the multimedia content, and the theme can be extracted from the description text.

[0074] In step S406, the object associated with the multimedia content is determined based on the theme.

[0075] For example, the theme can be directly used as a search term to search for the object, or the object can be searched based on the theme and its synonyms, near-synonyms, etc.

[0076] Through the above embodiment, the multimedia content can be generated first, and then the theme is extracted therefrom to determine the associated object. Thus, the generated multimedia content can be more diverse. Although the generated multimedia content is generated based on the message of the user, the agent can refer to other information, such as the knowledge base of the agent, in the process of generating the multimedia content. In this way, the generated multimedia content can contain more elements to facilitate the response to the user's demand while recommending potential objects that the user can be interested in.

[0077] In some cases, the message sent by the user can explicitly include the intention of generating the digital avatar multimedia content, or can not explicitly include the intention. Since one of the main functions of the intelligent agent is to interact with the user, i.e., to answer the message sent by the user. Therefore, in some embodiments, it can also be judged whether to generate the digital avatar multimedia content based on the answer of the intelligent agent after the intelligent agent answers the user. The message answered by the intelligent agent can be text, or a combination of text and other types of messages, or other types of messages other than text.

[0078] In some embodiments, a response message of the intelligent agent to the message sent by the user is first generated; and in response to the specified object included in the response message, the digital avatar multimedia content is generated based on the digital avatar, the message and the specified object. Before the intelligent agent generates the response message, the intelligent agent can directly generate a response according to the message sent by the user, without considering whether the digital avatar multimedia content will be generated subsequently. For example, before generating the response by using the generative model, the prompt information is generated based on the message sent by the user, but the prompt information is not generated based on the indication of generating the digital avatar multimedia content. After the response is generated, it is determined whether the specified object is included in the response message, and whether the digital avatar multimedia content is generated based on the determination. In the case where the specified object is included in the response message, it is determined that the multimedia content is generated; otherwise, the multimedia content is not generated. The specified object can be an object suitable for generating the multimedia content, or a pre-set object.

[0079] For example, the message sent by the user to the intelligent agent indicates that the intelligent agent recommends the recently popular lipsticks, clothes, etc., and the intelligent agent can send a response to the request for these recommendations. And if a certain specified product (such as a specified category, brand product) is included in the response, the digital avatar multimedia content is generated based on the product.

[0080] Through this embodiment, the digital avatar multimedia content can be generated based on the specified object in the response of the intelligent agent. Therefore, the user can be responded more comprehensively through the sending of text, multimedia content, information of objects, etc., and the information acquisition efficiency and interaction experience of the user are improved.

[0081] The message sent by the user can explicitly describe the requirement, such as the element to be embodied in the generated multimedia content. However, in some cases, the requirement of the user is not very clear. The user can indicate the agent to generate the multimedia content of the digital avatar in an interrogative tone. For example, the user can send the message "What hairstyle do I suit?" In the message, the user does not explicitly indicate the generated hairstyle. The user can also give a relatively vague indication, such as "generate a travel video", but the user does not explicitly indicate where to travel. Therefore, in some embodiments, a response message can be sent to the user first, and then the multimedia content of the digital avatar is generated based on the response message. For example, in response to the indication of generating the multimedia content of the digital avatar in the message sent by the user, one or more prompt messages are generated based on the message; a message sent by the agent to the user is generated based on the one or more prompt messages; and the multimedia content of the digital avatar is generated based on the digital avatar and the one or more prompt messages. That is, the prompt information (prompt) of the response of the agent to the user is determined based on the message sent by the user, and then the response is generated based on the prompt information. For example, the message sent by the agent to the user is generated based on the prompt information and the generative model. Then, the prompt information and the digital avatar are processed by using the generative model for generating the multimedia content, to generate the multimedia content of the digital avatar. In this way, the user can more clearly determine the generation logic or key information of the multimedia content, so that the generated multimedia content is more referential, and the efficiency of the user to obtain information is improved.

[0082] FIG. 5 shows a schematic diagram of a conversation interface according to some other embodiments of the present disclosure. As shown in FIG. 5, the conversation interface 5 of this embodiment is the conversation interface of the user and the agent A. In this interface, the messages sent by the user are exemplarily aligned and displayed on the right side, and the messages sent by the agent A are exemplarily aligned and displayed on the left side.

[0083] In the process of the conversation, the user sends the message 51 "What hairstyle do I suit?". Since the message sent by the user is related to the appearance of the user, in some embodiments, the generation process of the multimedia content of the digital avatar can be triggered. Alternatively, the agent can first respond without considering whether to generate the multimedia content, and then generate the multimedia content based on the response.

[0084] The multimedia content to be generated can be determined to be related to the hairstyle based on the message sent by the user, and the multimedia content of the digital avatar with the new hairstyle is generated. Alternatively, as shown in FIG. 5, the agent first responds to the message 51 through the message 52, i.e., “I think the curly hairstyle is quite suitable for your face shape, and I have generated a digital avatar with the curly hairstyle and in a skirt, and you can take a look at it”, and then generates the image (or video, etc.) of the digital avatar based on the specific hairstyle and clothing, etc. in the response, and sends it to the user through the message 53.

[0085] In addition, the agent also sends the object associated with the generated multimedia content of the digital avatar to the user. In the example of FIG. 5, the associated object is “XX hair salon”, and the information of the associated object is sent to the user through the message 54. The message can be a text description, or a control. Therefore, if the user is satisfied with the effect of the digital avatar, the user can further understand the information of the related object through the message 54.

[0086] Of course, according to the needs, the agent can not send the response message to the user, but only send the information of the multimedia content and the object. Those skilled in the art can select as needed.

[0087] As described above, the digital avatar can be a three-dimensional avatar, so as to generate the multimedia content of each angle of the digital avatar. FIG. 6 shows a flowchart of a method for generating the multimedia content of the digital avatar according to some embodiments of the present disclosure. As shown in FIG. 6, the generating method of this embodiment includes steps S602 to S606.

[0088] In step S602, the theme of the multimedia content to be generated is determined based on the message sent by the user. The specific implementation of this step can refer to the foregoing embodiments, which will not be described here.

[0089] In step S604, the target part of the three-dimensional digital avatar is determined based on the theme.

[0090] For example, if the theme includes a description of a part of the body, the part is determined as the target part; if the theme is associated with a part of the body, the associated part is determined as the target part.

[0091] In some embodiments, the target part can be determined based on the image material associated with the theme. First, the image material associated with the theme is determined; the action part of the image material in the three-dimensional digital avatar is determined; and then the target part of the three-dimensional digital avatar is determined according to the action part. The image material can be pre-labeled to specify its action part. For example, if the material is a hat, the action part is the head; if the material is shoes, the action part is the feet. Of course, the action part of the image material can also be indicated by its name, classification, etc. The action part of the image material can be equivalent to the target part of the digital avatar, or the target part of the digital avatar can include the action part. For example, when the action part is the feet, the target part can be the user's legs and feet, or the user's whole body.

[0092] In step S606, the multimedia content of the digital avatar is generated based on the target part, the digital avatar and the message sent by the user. The generated multimedia content includes the target part of the digital avatar, and the appearance of the digital avatar presented can match or be associated with the message.

[0093] The above embodiments generate multimedia content based on three-dimensional digital avatars, and determine the part to be displayed based on the theme. Thus, the display mode of the digital avatar can be automatically determined, so that the generated multimedia content can match the message sent by the user.

[0094] The method of generating multimedia content of the embodiments of the present disclosure can be executed in whole or in part on the user device. For example, the steps related to display and interaction with the user can be executed on the user device. For the steps of processing using the generative model, they can be executed on the user device, or on the server, or partially on the user device and partially on the server.

[0095] The methods of the embodiments of the present disclosure are introduced above. The apparatuses for implementing the methods are further described below.

[0096] FIG. 7 shows a structural schematic diagram of a multimedia content generation apparatus according to some embodiments of the present disclosure. As shown in FIG. 7, the generation apparatus 70 of this embodiment includes: a first display module 701 configured to display the digital avatar of the user in the conversation interface, the digital avatar being generated based on the image from the user; a first display module 702 configured to display the message sent by the user in the conversation interface; a generation module 703 configured to generate the multimedia content of the digital avatar based on the digital avatar and the message; and a third display module 704 configured to display the multimedia content sent to the user and the information of the object associated with the multimedia content in the conversation interface.

[0097] In some embodiments, the generating module 703 is further configured to determine a theme of the multimedia content to be generated according to the message; search image materials based on the theme; and generate the multimedia content of the digital avatar based on the digital avatar and the image materials.

[0098] In some embodiments, the generating module 703 is further configured to search for an object associated with the theme; and determine image materials associated with the object.

[0099] In some embodiments, the image materials associated with the object include at least one of image materials of the object, image materials of an article provided by the object, and image materials of a use effect of the object.

[0100] In some embodiments, the generating apparatus 70 includes a determining module 705 configured to identify a theme of the multimedia content of the digital avatar; and determine an object associated with the multimedia content according to the theme.

[0101] In some embodiments, the generating module 703 is further configured to generate a response message of the agent to the message; and in response to the response message including a specified object, generate the multimedia content of the digital avatar based on the digital avatar, the message, and the specified object.

[0102] In some embodiments, the generating module 703 is further configured to, in response to the message including an indication of generating the multimedia content, generate one or more prompt messages based on the message; generate a message sent by the agent to the user based on the one or more prompt messages; and generate the multimedia content of the digital avatar based on the digital avatar and the one or more prompt messages.

[0103] In some embodiments, the generating module 703 is further configured to determine a theme of the multimedia content to be generated based on the message; determine a target part of the three-dimensional digital avatar based on the theme; and generate the multimedia content of the digital avatar based on the target part, the digital avatar, and the message.

[0104] In some embodiments, the generating module 703 is further configured to determine image materials associated with the theme; determine an action part of the image materials on the three-dimensional digital avatar; and determine a target part of the three-dimensional digital avatar according to the action part.

[0105] In some embodiments, the generating module 703 is further configured to perform intent recognition on the message; and in response to the message having an intent of creating the multimedia content based on the digital avatar, generate the multimedia content of the digital avatar.

[0106] In some embodiments, the information of the object is a jump control, and the generating apparatus 70 further includes a fourth display module 706 configured to display a display interface or a transaction interface of the object in response to a triggering operation on the jump control of the object.

[0107] It should be noted that each unit described above is a logical division according to the specific function implemented thereby, and is not intended to limit the specific implementation manner, for example, each unit can be implemented in software, hardware or a combination of software and hardware. In actual implementation, each unit described above can be implemented as an independent physical entity, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). In addition, each unit described above is indicated by a dashed line in the drawings, indicating that these units can not actually exist, and the operations / functions implemented thereby can be implemented by the processing circuit itself.

[0108] In addition, although not shown, the device can also include a memory, which can store various information generated by the device, each unit included in the device in operation, programs and data for operation, data to be transmitted by the communication unit, etc. The memory can be a volatile memory and / or a non-volatile memory. For example, the memory can include, but is not limited to, a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a read-only memory (ROM), a flash memory. Of course, the memory can also be located outside the device. Alternatively, although not shown, the device can also include a communication unit, which can be used for communication with other devices. In one example, the communication unit can be implemented in a suitable manner known in the art, for example, including communication components such as an antenna array and / or a radio frequency link, various types of interfaces, communication units, etc. Here will not be described in detail. In addition, the device can also include other components not shown, such as a radio frequency link, a baseband processing unit, a network interface, a processor, a controller, etc. Here will not be described in detail.

[0109] Some embodiments of the present disclosure also provide an electronic device. FIG. 8 shows a structural schematic diagram of an electronic device according to some embodiments of the present disclosure. For example, in some embodiments, the electronic device 8 can be various types of devices, for example, can include but is not limited to various types of mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. For example, the electronic device 8 can include a display panel for displaying data and / or execution results utilized in the scheme according to the present disclosure. For example, the display panel can be various shapes, such as a rectangular panel, an oval panel or a polygonal panel, etc. In addition, the display panel can not only be a flat panel, but also a curved panel, or even a spherical panel.

[0110] As shown in FIG. 8, the electronic device 8 of this embodiment includes a memory 81 and a processor 82 coupled to the memory 81. It should be noted that the components of the electronic device 8 shown in FIG. 8 are only exemplary and non-limiting, and the electronic device 8 can also have other components according to actual application needs. The processor 82 can control other components in the electronic device 8 to perform desired functions.

[0111] In some embodiments, the memory 81 is configured to store one or more computer readable instructions. When the processor 82 executes the computer readable instructions, the computer readable instructions are executed by the processor 82 to implement the method according to any of the above embodiments. For specific implementation of each step of the method and related explanations, please refer to the above embodiments, and the repeated parts will not be described here.

[0112] For example, the processor 82 and the memory 81 can directly or indirectly communicate with each other. For example, the processor 82 and the memory 81 can communicate through a network. The network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 82 and the memory 81 can also communicate with each other through a system bus, and the present disclosure does not limit the processor 82 and the memory 81.

[0113] For example, the processor 82 can be embodied as various appropriate processors, processing devices, etc., such as a central processing unit (CPU), a graphics processing unit (GPU), a network processing unit (NP), etc.; and can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic component, a discrete hardware component. The central processing unit (CPU) can be X86 or ARM architecture, etc. For example, the memory 81 can include any combination of various forms of computer readable storage media, such as volatile memory and / or non-volatile memory. The memory 81 may, for example, include a system memory, which stores, for example, an operating system, application programs, a boot loader, a database, and other programs, etc. Various application programs and various data, etc. can also be stored in the storage medium.

[0114] In addition, according to some embodiments of the present disclosure, various operations / processes according to the present disclosure, when implemented by software and / or firmware, can be installed from a storage medium or a network to a computer system with a dedicated hardware structure, such as the computer system 90 shown in FIG. 9, which, when various programs are installed, can perform various functions, including functions such as those described above, etc. FIG. 9 shows a structural schematic diagram of a computer system according to some embodiments of the present disclosure.

[0115] In FIG. 9, a central processing unit (CPU) 901 performs various processing in accordance with a program stored in a read only memory (ROM) 902 or a program loaded from a storage section 908 to a random access memory (RAM) 903. In the RAM 903, data required when the CPU 901 performs various processing and the like is also stored as necessary. The central processing unit is merely exemplary, and can be other types of processors, such as the various processors described above. The ROM 902, the RAM 903, and the storage section 908 can be various forms of computer readable storage media, as described below. Note that, although the ROM 902, the RAM 903, and the storage 908 are shown separately in FIG. 9, one or more of them can be combined or located in the same or different memory or storage modules.

[0116] The CPU 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output interface 905 is also connected to the bus 904.

[0117] The following components are connected to the input / output interface 905: an input section 909 including a touch panel, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, and the like; an output section 907 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage section 908 including a hard disk, a magnetic tape, and the like; and a communication section 909 including a network interface card such as a LAN card, a modem, and the like. The communication section 909 allows communication processing to be performed via a network such as the Internet. It is easily understood that, although the various devices or modules in the computer system 90 are shown in FIG. 9 as communicating through the bus 904, they can also communicate through a network or other means, where the network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.

[0118] A drive 910 is also connected to the input / output interface 905 as necessary. A removable medium 911 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 910 as necessary, so that a computer program read therefrom is installed in the storage section 908 as necessary.

[0119] In the case where the above-described series of processing is implemented by software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 911.

[0120] According to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the CPU 901, the above-described functions defined in the methods of the embodiments of the present disclosure are executed.

[0121] It should be noted that, in the context of the present disclosure, a computer readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer readable medium can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as part of a carrier wave, in which the computer readable program code is carried. Such a propagated data signal can take a variety of forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), or the like, or any suitable combination of the above.

[0122] The above computer readable medium can be included in the above electronic device; or can exist separately, without being assembled into the electronic device.

[0123] In some embodiments, a computer program including instructions which, when executed by a processor, causes the processor to carry out the method of any of the above embodiments is also provided. For example, the instructions can be embodied in a computer program code.

[0124] In an embodiment of the disclosure, computer program code to carry out operations of the disclosure described can be written in any of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0125] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that noted in the figures. For example, two blocks depicted contiguously can actually be substantially concurrent, and in some cases, the blocks can be executed in reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware-based systems and computer instructions.

[0126] The modules, components or units described in the embodiments of the present disclosure can be implemented by software or by hardware. In some cases, the name of the module, component or unit does not constitute a limitation on the module, component or unit itself.

[0127] The functionality described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, non- transitory machine-readable media can include RAM, ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, compact disk read-only memory (CD-ROM), digital versatile disk (DVD), Blu-ray, or another non-transitory medium suitable for storing non-transitory program code, wherein the above aforementioned memory is on a machine readable medium.

[0128] The above description is only some embodiments of the present disclosure and an explanation of the principles of the technology used. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with each other to form technical solutions with similar functions disclosed in the present disclosure (but not limited to).

[0129] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the application can be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure the understanding of this description.

[0130] In addition, while operations are depicted in a particular order, this should not be understood as requiring these operations to be performed in the particular order shown or in sequential order, as some other operations can be performed in parallel or concurrently. Also, while the above discussion has included several specific examples, these should not be construed as limiting the scope of the disclosure, as other configurations can fall within the scope of the present disclosure. For example, the various features of the foregoing examples can be combined in any combination. Likewise, other components not specifically recited can be included in some embodiments. The scope of the disclosure should be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. In addition, while the above examples have been described with reference to particular examples, those skilled in the art will understand that the examples described are merely illustrative of the principles and applications of the present disclosure. Numerous modifications can be made to the above-described examples without departing from the scope of the disclosure. It is therefore not intended that the present disclosure be limited to the examples described above. It is intended that claims define the scope of the disclosure and that the means recited in the claims be construed as including any and all features which would normally fall within the meaning of those elements unless otherwise expressed.

[0131] While certain aspects of the present disclosure have been described with reference to one or more particular embodiments, those skilled in the art will understand that many alternative implementations and adaptations of the present disclosure are possible. Accordingly, implementations of the present disclosure in their broader aspect are not limited to specific disclosed implementations, but rather include any implementation that falls within the scope of the appended claims and their equivalents.

Claims

1. A method for generating multimedia content, comprising: displaying, in a conversation interface, a digital avatar of a user, the digital avatar being generated based on an image from the user; displaying, in the conversation interface, a message sent by the user; generating, based on the digital avatar and the message, a multimedia content of the digital avatar; displaying, in the conversation interface, the multimedia content sent to the user and information of an object associated with the multimedia content.

2. The generation method of claim 1, wherein, The generating, based on the digital avatar and the message, a multimedia content of the digital avatar comprises: determining a theme of the multimedia content to be generated according to the message; searching image materials based on the theme; generating, based on the digital avatar and the image materials, the multimedia content of the digital avatar.

3. The generation method of claim 2, wherein, The searching image materials based on the theme comprises: searching an object associated with the theme; determining image materials associated with the object.

4. The generation method of claim 3, wherein, The image materials associated with the object comprise at least one of image materials of the object, image materials of an article provided by the object, and image materials of a use effect of the object. 5.The method of claim 1 or 2, further comprising: identifying a theme of the multimedia content of the digital avatar; determining the object associated with the multimedia content according to the theme.

6. The generation method of any one of claims 1 to 5, wherein, The generating, based on the digital avatar and the message, a multimedia content of the digital avatar comprises: generating a response message of an agent to the message; in response to the response message including a specified object, generating, based on the digital avatar, the message and the specified object, the multimedia content of the digital avatar.

7. The generation method of any one of claims 1 to 6, wherein, The generating, based on the digital avatar and the message, a multimedia content of the digital avatar comprises: in response to the message including an indication of generating the multimedia content, generating one or more prompt messages based on the message; generating, based on the one or more prompt messages, a message sent by an agent to the user; generating, based on the digital avatar and the one or more prompt messages, the multimedia content of the digital avatar.

8. The generation method of any one of claims 1 to 7, wherein, The digital avatar is a three-dimensional digital avatar, and the generating, based on the digital avatar and the message, a multimedia content of the digital avatar comprises: determining a theme of the multimedia content to be generated based on the message; determining a target part of the three-dimensional digital avatar based on the theme; generating, based on the target part, the digital avatar and the message, the multimedia content of the digital avatar.

9. The generation method of claim 8, wherein, The determining, based on the theme, a target part of the three-dimensional digital avatar comprises: determining image materials associated with the theme; determining an action part of the three-dimensional digital avatar for the image materials; determining the target part of the three-dimensional digital avatar according to the action part.

10. The generation method of any one of claims 1 to 9, wherein, The generating, based on the digital avatar and the message, a multimedia content of the digital avatar comprises: performing intent recognition on the message; in response to the message having an intent of creating a multimedia content based on the digital avatar, generating the multimedia content of the digital avatar.

11. The generation method of any one of claims 1 to 10, wherein, The information of the object is a jump control, and the generation method further includes: In response to a triggering operation on the jump control of the object, displaying a display interface or a transaction interface of the object. 12.An electronic device, comprising: a memory; and a processor coupled to the memory, the processor configured to execute, based on instructions stored in the memory, a generation method of multimedia content according to any one of claims 1 to 11. 13.A computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the generation method of multimedia content according to any one of claims 1 to 11. 14.A computer program product, which, when running on a computer, causes the computer to implement the generation method of multimedia content according to any one of claims 1 to 11. 15.A computer program, comprising: instructions that, when executed by a processor, cause the processor to perform the generation method of multimedia content according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Interaction method and device, equipment and storage medium

    CN117850937A

  • Image generation method and device, electronic equipment and storage medium

    CN117853600A

  • Message interaction method and device, equipment and storage medium

    CN118612520A

  • Multimedia content generation method, electronic equipment, storage medium and product

    CN118885628A

  • Systems and methods for animated clip generation

    US20140215360A1