Image display method and electronic equipment

By generating images that meet user needs, the problem of inconvenient image interaction in existing technologies has been solved, achieving more efficient and accurate image generation and interaction, and improving the user experience.

CN120950701APending Publication Date: 2025-11-14LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511072730.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

The existing images cannot meet the needs of users during interaction, making it difficult for users to find suitable images to interact with, resulting in a poor experience.

Method used

By determining the target user's input and emotional data, an image containing the target virtual avatar is generated. A generative large model or intelligent agent is used to generate an image that meets the user's needs and displays it in the chat interface.

Benefits of technology

It improves the convenience and accuracy of image interaction, enhances the user experience, meets users' personalized needs, and reduces the difficulty of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950701A_ABST
    Figure CN120950701A_ABST
Patent Text Reader

Abstract

The invention provides an image display method and electronic equipment. The method comprises the following steps: determining target input data input by a target user, and determining emotion data used for representing the emotion of the target user; determining first image data corresponding to the target virtual image; generating a target image based on the target input data, the emotion data and the first image data; wherein the target image comprises the target virtual image; image semantics corresponding to a target virtual image in the target image are used for representing the target input data and / or the emotion data; and displaying the target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image display technology, and more specifically, to an image display method and an electronic device. Background Technology

[0002] With the widespread adoption of smart devices, various applications offer a wide range of services. For example, users can interact with chat partners through applications such as chat apps, communication apps, and social apps. These chat partners can be other users, group chats, intelligent agents, or chatbots, etc.

[0003] During interaction, users can typically use existing images to interact, such as sending emojis. However, if the existing images cannot meet the current interaction needs, users will find it difficult to find suitable images to interact with, resulting in a poor user experience. Summary of the Invention

[0004] In view of this, the present disclosure provides an image display method and an electronic device that can improve the user experience when interacting with images.

[0005] One aspect of this disclosure provides an image display method, comprising:

[0006] Determine the target input data input by the target user, and determine the emotional data used to characterize the target user's emotions;

[0007] Determine the first image data corresponding to the target virtual image;

[0008] A target image is generated based on the target input data, the emotion data, and the first image data; wherein, the target image contains the target virtual image; the image semantics corresponding to the target virtual image in the target image are used to characterize the target input data and / or the emotion data;

[0009] Display the target image.

[0010] Optionally, the target input data includes data input by the target user into a chat interface dialog box; the method further includes: displaying an interactive control; the interactive control is used to send the target image in the chat interface.

[0011] Optionally, the target virtual avatar includes at least one of the following: the target user's preset virtual avatar, the target chat object's preset virtual avatar corresponding to the chat interface, and the target user's preset emoticon virtual avatar.

[0012] Optionally, determining the emotional data used to characterize the target user's emotions includes: determining emotional data used to characterize the target user's emotions based on the target input data and / or target interaction data; wherein the target interaction data includes interaction data between the target user and other users; the chat interface is used for interaction between user members included in the target group, and the target group includes the target user and the other users.

[0013] Optionally, the method for determining the target interaction data includes: determining historical interaction data between the target user and the interaction object; determining target interaction data that meets preset interaction conditions from the historical interaction data; and the target group includes the interaction object.

[0014] Optionally, the preset interaction conditions include at least one of the following: the interaction data targeted by the target input data; the interaction data within the target time period; the interaction data between the target user and a specified object in the interaction objects, wherein the target group includes the specified object.

[0015] Optionally, generating a target image based on the target input data, the emotion data, and the first image data includes: determining image performance description data based on the target input data and the emotion data; and generating a target image to characterize the image performance description data by using the first image data as a reference element.

[0016] Optionally, determining the emotional data used to characterize the target user's emotions includes: determining the emotional data used to characterize the target user's emotions based on at least one of the following: the target input data, the target interaction data, and the target user's expression style; wherein the target interaction data includes the interaction data between the target user and other users.

[0017] Optionally, the method for determining the expression style of the target user includes: determining the expression style of the target user based on the target user's historical interaction data.

[0018] Another aspect of this disclosure provides an image display device, comprising:

[0019] The first determining unit is used to determine the target input data input by the target user and to determine the emotional data used to characterize the target user's emotions.

[0020] The second determining unit is used to determine the first image data corresponding to the target virtual image;

[0021] The generation unit is configured to generate a target image based on the target input data, the emotion data, and the first image data; wherein the target image contains the target virtual image; and the image semantics corresponding to the target virtual image in the target image are used to characterize the target input data and / or the emotion data.

[0022] An image display unit is used to display the target image.

[0023] Optionally, the target input data includes data input by the target user into the chat interface dialog box; the image display unit is further configured to: display interactive controls; the interactive controls are used to send the target image in the chat interface.

[0024] Optionally, the target virtual avatar includes at least one of the following: the target user's preset virtual avatar, the target chat object's preset virtual avatar corresponding to the chat interface, and the target user's preset emoticon virtual avatar.

[0025] Optionally, the first determining unit is configured to: determine emotional data characterizing the emotions of the target user based on the target input data and / or target interaction data; wherein the target interaction data includes interaction data between the target user and other users; the chat interface is used for interaction between user members included in the target group, and the target group includes the target user and the other users.

[0026] Optionally, the method for determining the target interaction data includes: determining historical interaction data between the target user and the interaction object; determining target interaction data that meets preset interaction conditions from the historical interaction data; and the target group includes the interaction object.

[0027] Optionally, the preset interaction conditions include at least one of the following: the interaction data targeted by the target input data; the interaction data within the target time period; the interaction data between the target user and a specified object in the interaction objects, wherein the target group includes the specified object.

[0028] Optionally, the generation unit is configured to: determine image performance description data based on the target input data and the emotion data; and generate a target image to characterize the image performance description data, using the first image data as a reference element.

[0029] Optionally, the first determining unit is configured to: determine emotional data characterizing the target user's emotions based on at least one of the following: the target input data, the target interaction data, and the target user's expression style; wherein the target interaction data includes interaction data between the target user and other users.

[0030] Optionally, the method for determining the expression style of the target user includes: determining the expression style of the target user based on the target user's historical interaction data.

[0031] Another aspect of this disclosure provides an electronic device including one or more processors and one or more memories, wherein the memories are used to store executable instructions that, when executed by the processor, implement the method described above.

[0032] Another aspect of this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed, are used to implement the method described above.

[0033] Another aspect of this disclosure provides a computer program product including a computer program comprising computer executable instructions that, when executed, implement the method described above.

[0034] Another aspect of this disclosure provides an electronic device, including: a first application running on the electronic device, and a display unit;

[0035] The first application is configured to: parse at least one task based on the input, and invoke the target model to execute at least one task, for at least the following purposes:

[0036] Determine the target input data input by the target user, and determine the emotional data used to characterize the target user's emotions;

[0037] Determine the first image data corresponding to the target virtual image;

[0038] The target model is invoked, and the target input data, the emotion data, and the first image data are input into the target model to determine the target image generated by the target model; wherein, the target image contains the target virtual image; the image semantics corresponding to the target virtual image in the target image are used to characterize the target input data and / or the emotion data;

[0039] The display unit is invoked to display the target image.

[0040] According to some embodiments of this disclosure, by generating and displaying images based on user input data, the ease of image interaction for users can be improved, thereby enhancing the user experience.

[0041] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0042] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0043] Figure 1 This illustration schematically depicts an application scenario of an image display method according to an embodiment of the present disclosure;

[0044] Figure 2 A flowchart illustrating an image display method according to an embodiment of the present disclosure is shown schematically;

[0045] Figure 3 A schematic diagram of a chat interface according to an embodiment of the present disclosure is shown.

[0046] Figure 4 The schematic diagram illustrates a structural schematic of an image display device according to an embodiment of the present disclosure;

[0047] Figure 5 A block diagram schematically illustrates an electronic device suitable for implementing an image display method according to an embodiment of the present disclosure. Detailed Implementation

[0048] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0049] In the technical solution disclosed herein, the acquisition, storage, and application of user personal information are all information authorized by the user or fully authorized by all parties, comply with relevant laws and regulations, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation entry points are provided for users to choose to authorize or refuse. In the technical solution disclosed herein, the acquisition, collection, storage, use, processing, transmission, provision, disclosure, and application of data all comply with relevant laws and regulations, necessary confidentiality measures have been taken, and they do not violate public order and good morals.

[0050] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this disclosure all offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.

[0051] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0052] When using expressions such as "at least one of A, B, or C," it should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (e.g., "a system having at least one of A, B, or C" should include, but is not limited to, systems having A alone, having B alone, having C alone, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, features defined with "first" or "second" may explicitly or implicitly include one or more of the stated features.

[0053] With the widespread adoption of smart devices, various applications offer a wide range of services. For example, users can interact with chat partners through various applications; these partners can be other users, group chats, intelligent agents, or chatbots, etc. During this interaction, users typically use existing images, such as sending emojis. However, if the existing images cannot meet the current interaction needs, users may struggle to find suitable images. For instance, they may need to search for relevant images using search engines or other means, resulting in longer search times and a poor user experience.

[0054] This disclosure provides an image display method. In this method, a corresponding image can be generated in real time based on user input data, and the generated image can be displayed. This allows users to more quickly select a suitable image from the generated images for interaction, improving the convenience of image interaction and enhancing the user experience.

[0055] Specifically, generative large models or intelligent agents (such as text-to-graph models, graph-to-graph models, and image generation agents) can be used to generate corresponding images based on the user's actual needs. Accordingly, various types of user information can be obtained to help generate images that meet the user's requirements. For example, user input data during chats, user history, frequently used interactive images, chat partners, chat logs, and emotional information can be obtained to help generate images that meet the user's needs.

[0056] By combining user information to generate images, the efficiency and accuracy of image generation can be improved, the convenience of image interaction for users can be increased, and the user experience can be enhanced.

[0057] In a specific example, during user chat interactions, corresponding images, such as emoticons, can be generated based on the user's needs, allowing the user to easily select the appropriate emoticon for interaction. Specifically, emoticons can be generated based on the user's text or voice input in the chat interface, containing either the user's input text or voice message; alternatively, emoticons can be generated based on the user's preferred emoticon style or expression style, conforming to the user's usual emoticon style (cute, funny, or anime style, etc.) or expression style (simple, etc.). When engaging in image interaction, users can easily obtain images (emoticons) that meet their needs, thereby improving the convenience of image interaction and enhancing the user experience.

[0058] Furthermore, the method can also generate images based on virtual avatars. Specifically, the virtual avatar can be a virtual avatar set by the user, such as a virtual character, virtual animal, or virtual anime avatar set by the user; it can also be a virtual avatar set by the chat interaction object; it can also be a virtual avatar set by the chat interaction object for the user; it can also be a virtual avatar set by the user for the chat interaction object, and so on.

[0059] The generated images can include virtual avatars, specifically virtual avatars in different states. For example, virtual avatars with different expressions, actions, demeanors, colors, etc., can be generated according to actual needs, thus better meeting user requirements, reducing the difficulty of image generation, and improving the efficiency of image generation.

[0060] In a specific example, based on the user's input of "shake hands," an image of the two virtual avatars shaking hands can be generated, combining the user's chosen virtual avatar (which can be 2D or 3D) and the virtual avatar chosen by the chat partner. This allows the user to select and interact with the avatar. The generated image can also be a dynamic image.

[0061] Therefore, generating images containing virtual avatars can improve the efficiency and accuracy of image generation, enhance the convenience of image interaction for users, improve user experience, and also make it easier to meet user needs by using virtual avatars, thereby reducing the difficulty of image generation.

[0062] Figure 1 The illustration shows an application scenario of an image display method according to an embodiment of the present disclosure.

[0063] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0064] Users can use the first terminal device 101, the second terminal device 102, or the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0065] The first terminal device 101, the second terminal device 102, or the third terminal device 103 can be various electronic devices with a display screen and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0066] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, or the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0067] It should be noted that the image display method provided in this embodiment can generally be executed by a first terminal device 101, a second terminal device 102, a third terminal device 103, or a server 105. Correspondingly, the image display device provided in this embodiment can generally be located in the first terminal device 101, the second terminal device 102, the third terminal device 103, or the server 105. The image display method provided in this embodiment can also be executed by a server or server cluster that is different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the image display device provided in this embodiment can also be located in a server or server cluster that is different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0068] In an optional embodiment, the image display method provided in this disclosure can be executed by a first terminal device 101, a second terminal device 102, or a third terminal device 103. The device can communicate with a server 105 via a network 104, and the server 105 can execute the step of generating an image in the image display method and send the generated image to the terminal device for display.

[0069] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0070] Figure 2 A flowchart illustrating an image display method according to an embodiment of the present disclosure is shown schematically.

[0071] This method does not limit the specific entity that performs the execution. Optionally, this method can be executed through any electronic device or any software application, specifically a terminal device, server, cloud, client, etc.

[0072] like Figure 2 As shown, this method may include the following steps.

[0073] S201: Determine the target input data for the target user and determine the emotional data used to characterize the target user's emotions.

[0074] S202: Determine the first image data corresponding to the target virtual image.

[0075] S203: Generate a target image based on the target input data, emotion data, and first image data; wherein the target image contains a target virtual image; the image semantics corresponding to the target virtual image in the target image are used to characterize the target input data and / or emotion data.

[0076] S204: Display the target image.

[0077] This method can generate and display images based on user input data, which can improve the convenience of image interaction for users and enhance the user experience.

[0078] Furthermore, this method can also generate images based on input data, emotion data, and image data of virtual avatars, which can improve the efficiency and accuracy of image generation, enhance the convenience of image interaction for users, and improve the user experience.

[0079] This method can also generate images containing virtual avatars, which can improve the degree to which the generated images meet user needs, facilitate the fulfillment of user requirements, and reduce the difficulty of image generation.

[0080] In embodiments of this disclosure, user consent or authorization may be obtained before using or determining user information. For example, in the above method flow, a request may be sent to the user to determine target input data, emotion data, and first image data corresponding to the target virtual avatar. If the user consents or authorizes the determination of the target input data, emotion data, and first image data corresponding to the target virtual avatar, the above method flow may be executed. As another example, if the user consents or authorizes the generation of an image based on the target input data, emotion data, and first image data corresponding to the target virtual avatar, the above method flow may be executed.

[0081] In the embodiments of this disclosure, a corresponding operation entry point can be provided to the user, allowing the user to choose to agree or refuse to "determine user information and generate an image based on user information". Specifically, before "determining user information and generating an image based on user information", the user can provide an instruction to agree or refuse to "determine user information and generate an image based on user information" through the corresponding operation entry point. If the user agrees to "determine user information and generate an image based on user information", the above-described method flow can be executed. If the user refuses to "determine user information and generate an image based on user information", the expert decision-making process is initiated.

[0082] This disclosure does not limit the target user. The target user can be any user; for ease of description, in this disclosure, any user who needs to generate the image will be referred to as the target user.

[0083] This disclosure does not limit the target input data. Optionally, the target input data can be any data entered by the target user, or any data currently entered by the target user. Specifically, it can be any data entered by the target user in any interface, software, or program. For example, the target input data can be any data entered by the target user in image generation software or social applications; it can also be any data entered by the target user in an interactive interface or chat interface, thereby facilitating the determination of the user's interaction needs and generating corresponding images.

[0084] This disclosure does not limit the specific form of the target input data. The target input data can be in text, voice, or image form, or a mixed form, such as data containing both text and voice. Optionally, an image can be generated based on text-based target input data; alternatively, the target user's emotional data and voice text can be determined based on voice-based target input data to generate an image; or the target user's emotional data, image text, and image style information can be determined based on image-based target input data to generate an image. Specifically, the input image can be used as a reference image for image generation. In a specific example, target input data input by the target user can be obtained and processed accordingly based on the form of the target input data to generate an image.

[0085] This disclosure does not limit the use of emotion data. It is understood that generating images that better suit the user's needs based on their emotions can improve the relevance of the generated images. For example, if the user is currently happy, a corresponding humorous image can be generated. Optionally, emotion data can be in the form of feature vectors or text; this disclosure does not limit the form of emotion data.

[0086] This disclosure does not limit the method for determining emotion data. Optionally, emotion data representing the target user's emotions can be directly obtained, or emotion data can be determined based on relevant information about the target user. Specifically, emotion data can be determined based on target input data; for example, the target user's emotion data can be determined based on target input data in the form of speech. Alternatively, emotion data can be determined based on the target user's interaction data. For example, in a chat scenario, the target user's emotion data can be determined based on the target user's historical chat records for the day.

[0087] This disclosure does not limit the target virtual image. Optionally, the target virtual image can be any virtual image set by the target user. Images can be generated based on the target virtual image, reducing the difficulty of image generation and improving the efficiency of image generation. This disclosure does not limit the specific content and form of the target virtual image. Optionally, the target virtual image can be a virtual character, a virtual animal, a virtual item, a virtual anime character, etc. The target virtual image can be used to represent the target user or other users. The target virtual image can be a three-dimensional character model or a two-dimensional character, etc. This disclosure does not limit the source of the target virtual image. Optionally, the target virtual image can be set by the target user, set by other users, preset or default, or set by the target user's interaction object, etc.; the target virtual image can be provided by the target user or other users, or generated or constructed by a machine or device, etc. This disclosure does not limit the number of target virtual images. Optionally, one or more target virtual images can be determined. Specifically, the target user can set one or more target virtual images, or multiple target virtual images set by multiple users can be included, facilitating the generation of images of interaction between different virtual images.

[0088] This disclosure does not limit the first image data. Optionally, the first image data can be any image data corresponding to the target virtual image, or any image data containing the target virtual image. For example, virtual images are typically displayed using image data. This disclosure does not limit the form of the first image data; it can be a two-dimensional image or a three-dimensional image. In one specific example, the target virtual image can be a virtual character, and the corresponding first image data can be a three-view drawing of the virtual character, a complete facial image of the virtual character, or an image containing the complete virtual character, etc. In another specific example, the target virtual image can be the image in a user's avatar, thus the image used as the avatar can be determined as the first image data.

[0089] Regarding the three operations of determining target input data, determining emotion data, and determining first image data, the embodiments of this disclosure do not limit the specific execution order. These three operations can be executed in parallel or sequentially. Optionally, the target input data, emotion data, and first image data can be determined in parallel; alternatively, the emotion data can be determined while the target input data is being determined, and the first image data corresponding to the target virtual avatar can be determined while the emotion data is being determined.

[0090] This disclosure does not limit the target image. Optionally, the generated target image may include a target virtual avatar, and the target virtual avatar included in the target image can be used to represent target input data and / or emotional data. Specifically, the target input data and / or emotional data can be represented by the state of the target virtual avatar in the target image. Optionally, the image semantics corresponding to the target virtual avatar in the target image can be used to represent target input data and / or emotional data.

[0091] For ease of understanding, in a specific example, given the target input data "shake hands," the generated target image may contain a virtual target avatar with the "shake hands" gesture, thus representing the target input data. Similarly, given the emotion data "happy," the generated target image may contain a virtual target avatar with a "happy" expression, thus representing the emotion data. Finally, given the target input data "Nice to meet you, shake hands," the generated target image may contain a virtual target avatar with both the "shake hands" gesture and a "happy" expression, thus representing both the target input data and the emotion data.

[0092] This disclosure does not limit the content of the target image. Optionally, the target image may include a target virtual avatar, and may also include other content. Specifically, the target image may include text from the target input data, and may also include image content generated based on at least one of the target input data, emotion data, and first image data. Of course, the target image may also include other image content or image content generated based on other information. Specifically, the corresponding image content in the generated target image can be controlled by prompt words or other methods. In a specific example, for the target input data "Congratulations," the text "Congratulations" can be included in the generated target image (emoticon). Specifically, the image generation can be controlled by the prompt word "Present the target input data in text form in the image."

[0093] This disclosure does not limit the specific form of the target image. Optionally, the target image can be a two-dimensional image, a three-dimensional image, an animated image, etc. In a specific example, for the target input data "handshake", an animated target image can be generated, which includes a dynamic handshake action; for the target input data "congratulations", an animated target image can be generated, which includes a dynamic cupped-hand gesture.

[0094] This disclosure does not limit the specific method of generating the target image. Optionally, the target image can be generated based on a generative large model, an image generation model, or an intelligent agent. Accordingly, the data used to generate the target image may include target input data, emotion data, and first image data. This disclosure does not limit the data used to generate the target image; in addition to the above three types of data, other data can also be combined to generate the target image. For example, the target user's preset image style (e.g., the target user's favorite or commonly used emoji style), the target user's preset expression style (e.g., the target user's preferred humorous, calm, or concise expression style), the target user's historical interaction records, information about the target user's current interaction object, the target user's interaction style with the current interaction object, and other data can be used to generate the target image.

[0095] This disclosure does not limit the specific entity executing the target image generation method. Optionally, the entity executing the above method flow may execute the target image generation method itself to generate the target image locally; alternatively, it may use an external device to execute the target image generation method, sending the data used to generate the target image to the external device and obtaining the target image generated by the external device. The external device may be a server or cloud device, and high-computing-power devices can be used to improve image generation efficiency.

[0096] This disclosure does not limit the display method of the target image. Optionally, the target image can be displayed to the target user, sent to and displayed to the target user, or sent to the client of the logged-in target user for display. Optionally, the target image can be displayed on the target user's interactive interface to facilitate the target user to select the image for interaction. Specifically, the interactive interface can be an interface for the target user to input target input data.

[0097] In a specific example, target input data can be obtained from the target user's chat interface. Specifically, this could be the target input data entered by the target user in a dialog box within the chat interface. The aforementioned method flow is then executed to generate a target image, which is then displayed in the current chat interface. The chat interface can be a chat with a single other user, a group chat, or a chatbot (intelligent customer service, intelligent chatbot, or chat agent, etc.). The target user can further select an image from the target image for interaction. The generated target image could be, for example, an emoticon. The target user can select the generated emoticon for interaction, improving the convenience of image acquisition and interaction, and enhancing the user experience.

[0098] This disclosure does not limit the number of target images generated. Optionally, one or more target images may be generated. When multiple target images are generated, one or more of them may be displayed, or some or all of them may be displayed, to facilitate selection by the target user. Multiple target images also provide the target user with a greater number of choices, improving the ease with which the user can acquire images and interact with them, thus enhancing the user experience.

[0099] It is understandable that one or more target images can be generated based on the above method and process. Specifically, multiple different target images can be generated based on the same input data but different image generation methods; multiple different target images can be generated based on different input data but the same image generation method; or multiple different target images can be generated based on different input data and different image generation methods.

[0100] For example, different generative large models or different agents can be used to generate different target images, or different parameters can be used to generate different target images for the same model, or different target images can be generated based on different input data. Specifically, a target user might have multiple favorite emoji styles, thus allowing different target images to be generated based on different emoji styles.

[0101] This disclosure does not limit the specific application scenario. Specifically, it can generate images for users in interactive scenarios, image generation scenarios, or other scenarios. For example, it can generate images for users in a chat scenario, or it can generate images for users based on image generation software.

[0102] In an optional embodiment, the target input data may include data entered by the target user into the chat interface dialog box. Accordingly, the above method flow may further include: displaying interactive controls; the interactive controls can be used to send the target image in the chat interface. This embodiment can generate an image based on the input data in the chat interface and provide controls to send the generated image, which can improve the efficiency and accuracy of image interaction in chat scenarios, increase the convenience for users to obtain and interact with images in chat scenarios, and improve the user experience of chatting.

[0103] This disclosure does not limit the specific form of the interactive control. Optionally, the interactive control can be a control that displays the generated target image, including a selection component and a send button. Users can select from the displayed target images and send the selected target image, thereby improving the convenience of user operation. For example, a window can pop up in a chat interface, displaying multiple generated target images. Each target image can correspond to a selection component, allowing users to easily select from the selection components. Furthermore, the window can also include a send button, allowing users to click the send button to send the selected target image to the chat interface. Accordingly, the interactive control can be used to select from the generated target images and send the selected target image to the chat interface.

[0104] This disclosure does not limit the content and number of target virtual avatars. Optionally, the target virtual avatar may include the target user's own virtual avatar, such as a preset virtual avatar of the target user, specifically a virtual avatar set by the target user for themselves, or a virtual avatar set by the target user for generating emoticons, or a preset emoticon virtual avatar of the target user. Furthermore, the target virtual avatar may also include the virtual avatars of other users. Optionally, the chat partner in the target user's current chat interface may also have a virtual avatar, thus the virtual avatar of the chat partner can also be used as a target virtual avatar to generate the target image.

[0105] This disclosure does not limit the target user's chat partners; chat partners can include other users, group chats, other users in a group chat, chatbots, intelligent agents, etc. It is understood that chatbots can have virtual avatars, and other users can also have virtual avatars, which can then be used as target virtual avatars to generate target images.

[0106] In one specific example, a virtual avatar can be determined based on the profile picture in the chat interface. Therefore, a target virtual avatar can be determined to generate a target image based on the profile picture of the target user and / or the profile picture of the chat partner.

[0107] Therefore, optionally, the target virtual avatar may include at least one of the following: (1) a preset virtual avatar of the target user; (2) a preset virtual avatar of the target chat object corresponding to the chat interface; (3) a preset emoji virtual avatar of the target user. This embodiment can generate a target image containing the target virtual avatar by setting a target virtual avatar related to the target user, which facilitates improving the degree of fit of the generated image to the user's needs, making it easier to meet the user's needs and reducing the difficulty of image generation. It is understood that the target virtual avatar may also include other virtual avatars, such as the virtual avatar in the target user's avatar, the virtual avatar of other users, the virtual avatar of other users specified by the target user, etc.

[0108] The embodiments of this disclosure do not limit the target chat object. Optionally, the chat interface may correspond to one or more chat objects. For example, for a group chat, the chat interface may correspond to multiple other users in the group chat as chat objects. Optionally, determining the target chat object may specifically involve determining the single chat object as the target chat object when the chat interface corresponds to a single chat object; or determining the target chat object from the multiple chat objects based on a preset determination method when the chat interface corresponds to multiple chat objects. The embodiments of this disclosure do not limit the method of determining the target chat object, nor do they limit the preset determination method. Optionally, the target chat object may be selected by the target user; or the target chat object may be determined based on target input data. For example, the target input data may contain an identifier of the target chat object, thereby determining the corresponding target chat object based on the target input data. In a specific example, the target chat object may be determined based on the target user's selection, or it may be determined based on the chat object targeted in the target input data. The embodiments of this disclosure do not limit the number of target chat objects; one or more target chat objects may be determined.

[0109] The embodiments disclosed herein are not limited to preset virtual emoticons. Specifically, they can be virtual emoticons commonly used by the target user or virtual emoticons set by the target user.

[0110] This disclosure does not limit the method for determining emotion data. Optionally, the current emotion of the target user can be determined based on the target input data, thereby determining emotion data to characterize the target user's current emotion. Furthermore, emotion data can be determined by combining other information, such as the target user's historical interaction data, or the target user's interaction data with other users, etc.

[0111] Therefore, optionally, determining the emotional data used to characterize the target user's emotions can specifically involve: determining the emotional data used to characterize the target user's emotions based on target input data and / or target interaction data. The target interaction data can include interaction data between the target user and other users. The chat interface can be used for interaction between user members within a target group, which can include the target user and other users. The target group can specifically be a set of users, such as a user group chat, or a set of users containing only two interacting users. The chat interface can be used for interaction among user members within the target group, enabling interaction between user members within the target group. This embodiment can combine target input data and target interaction data to determine user emotions, improving the accuracy and comprehensiveness of user emotions and facilitating subsequent improvements in the relevance of the generated image to user needs.

[0112] This disclosure does not limit the target interaction data. Optionally, the target interaction data may include interaction data related to the target user, historical interaction data of the target user (e.g., historical chat logs), interaction data between the target user and other users, or historical interaction data between the target user and other users, etc.

[0113] Understandably, to determine a target user's emotional data, one can combine the target user's input data with their interaction data. Specifically, this could involve combining the target user's recent historical interaction data, such as their interactions within a preset timeframe prior to the current moment, to determine their emotional data and improve the accuracy of the emotional assessment. In a concrete example, the target user's current emotion can be determined based on their current input data and recent chat history.

[0114] This disclosure does not limit the scope to other users. Optionally, other users may specifically include the chat objects corresponding to the target user's chat interface, other users selected by the target user, or the target chat object, etc. Optionally, other users may be users other than the target user in the target group corresponding to the chat interface.

[0115] This disclosure does not limit the method for determining the target interaction data. Specifically, it can be determined from historical interaction data or from interaction data related to the target user.

[0116] Optionally, the method for determining the target interaction data may include: determining historical interaction data between the target user and the interaction object; and determining target interaction data that meets preset interaction conditions from the historical interaction data. The target group may include the interaction object. This embodiment can determine the target interaction data from historical interaction data, which can improve the accuracy and comprehensiveness of user emotions and facilitate subsequent improvements in the degree to which the generated image matches user needs.

[0117] This disclosure does not limit the interaction object; it can specifically be an object interacted with by the target user. Optionally, the interaction object may include an object selected by the target user, a chat object corresponding to the target user's chat interface, or the target chat object, etc.

[0118] This disclosure does not limit the preset interaction conditions. Optionally, the preset interaction conditions may include at least one of the following: (1) the interaction data targeted by the target input data; (2) the interaction data within the target time period; (3) the interaction data between the target user and a specified object in the interaction objects. The target group may include the specified object. This embodiment can improve the efficiency of determining the target interaction data based on the preset interaction conditions.

[0119] This disclosure does not limit the interactive data to which the target input data is targeted. Optionally, the target input data may include response data that responds to a portion of the interactive data, data that references a portion of the interactive data, data that summarizes a portion of the interactive data, and so on. It is understood that the portion of the interactive data to which the target input data is targeted can be defined as the interactive data to which the target input data is targeted.

[0120] This disclosure does not limit the target time period. Optionally, the target time period can be a time period within a preset duration prior to the current time, thereby improving the real-time performance of the target interactive data. Of course, the target time period can also be other time periods.

[0121] This disclosure does not limit the specified object. Optionally, the specified object may be the target chat object, or an object specified by the target user in the interaction object, etc.

[0122] In one optional embodiment, the target user's current mood can be determined based on historical chat logs between the target user and the target chat partner, as well as the target input data. Furthermore, the target image can be generated based on a preset virtual avatar of the target chat partner.

[0123] In an alternative embodiment, the target user's emotional data can also be determined by combining their expression style. For example, if the target user's expression style tends to be "concise and calm," this can help in determining the target user's emotional data.

[0124] Optionally, the emotional data used to characterize the target user's emotions can be determined based on at least one of the following: (1) target input data; (2) target interaction data; and (3) the target user's expression style. The target interaction data may include interaction data between the target user and other users. This embodiment can improve the accuracy and comprehensiveness of user emotions. The interpretation of the target interaction data can be found in the explanations of other embodiments.

[0125] This disclosure does not limit the method for determining the expression style. Optionally, it can be an expression style selected by the target user, or it can be an expression style determined based on the target user's relevant interaction data or historical interaction data. Furthermore, considering that the target user may adopt different expression styles for different interaction objects—for example, a rigorous expression style for work scenarios and an optimistic expression style for friendship scenarios—different expression styles can be determined by combining interaction data with different interaction objects. It is understood that multiple determined expression styles can be used to determine the user's emotion, or different user emotions can be determined separately based on different expression styles to generate different target images. Alternatively, the corresponding expression style can be determined based on the chat object.

[0126] Optionally, the method for determining the target user's expression style includes: determining the target user's expression style based on the target user's historical interaction data. This embodiment can improve the accuracy of user expression style by relying on historical interaction data.

[0127] This disclosure does not limit the use of historical interaction data. Optionally, it could be historical interaction data between the target user and the interaction object, historical interaction data between the target user and a specific object within the interaction object, historical interaction data between the target user and the target chat object, and so on.

[0128] This disclosure does not limit the specific process of generating the target image. Optionally, the information used to generate the target image can be directly input into a generative large model or agent to generate the target image, or the information used to generate the target image can be processed before being input into a generative large model or agent to generate the target image.

[0129] Therefore, optionally, a target image is generated based on the target input data, emotion data, and the first image data. Specifically, this can be achieved by: determining image performance description data based on the target input data and emotion data; and generating a target image to characterize the image performance description data using the first image data as a reference element. This embodiment improves the efficiency and accuracy of target image generation by using the first image data as a reference element to determine the image performance description data based on the target input data and emotion data.

[0130] This disclosure does not limit the image representation description data. Optionally, the image representation description data can specifically be prompts for image generation, which can be used to characterize the target input data and emotion data. For example, based on the emotion data "happy," the image representation description data can be determined as "generate an image that reflects a happy emotion" or "the generated image needs to reflect a happy emotion," etc. Based on the target input data, the text content or image content in the image representation description data can be further determined. For example, the target input data can be used as the text content in the target image, or data that can be used as the image content of the target image can be extracted based on the target input data, specifically items, people, events, or actions in the target input data, etc. For the target input data "reading a book," the item "book" can be extracted. The image representation description data used to determine the image representation description data can be "the generated image needs to contain the text content 'reading a book,'" or "the generated image needs to contain a book," etc. It is understood that the generated target image can conform to the image representation description data.

[0131] This disclosure does not limit the specific manner in which the first image data is used as a reference element. Optionally, for the image generation method, the image can be generated based on the first image data, or the image content or image semantics in the first image data can be extracted and combined with other information to generate the target image, etc.

[0132] This disclosure does not limit the specific timing or triggering conditions for executing the above method flow, nor does it limit the specific timing or triggering conditions for generating the target image.

[0133] Optionally, the above method flow can be executed or a target image can be generated when the target user performs a preset trigger operation or inputs preset characters.

[0134] Optionally, a target image is generated based on the target input data, emotion data, and the first image data. Specifically, this can be achieved by generating the target image based on the target input data, emotion data, and the first image data when a preset trigger condition is detected. This disclosure does not limit the preset trigger condition. Optionally, the preset trigger condition may include at least one of the following: (1) the target input data contains a preset character; (2) a preset trigger operation is performed. This disclosure does not limit the preset character or the preset trigger operation. Optionally, the preset character may be a hash symbol or other special symbols, and the preset trigger operation may be clicking a shortcut key, clicking a virtual button, or performing voice input, etc.

[0135] For ease of understanding, this disclosure also provides an optional embodiment in which the target user's expression style can be determined in advance based on the target user's historical interaction data. Specifically, different expression styles can be determined based on different interaction objects, or the overall expression style of the target user can be determined by combining historical interaction data with different interaction objects.

[0136] Accordingly, the target user's current emotional data can be determined based on the target user's target input data (e.g., the data entered by the target user in the dialog box of the current chat interface), the pre-determined expression style, and at least one of the historical interaction data.

[0137] Historical interaction data may include chat history corresponding to the target user's current chat interface, or historical interaction data between the target user and the chat partner corresponding to the current chat interface, etc.

[0138] Afterwards, a target image can be generated by combining multiple data, such as an emoji. Among them, at least one of the following can be combined to generate the target image: (1) the current emotional data of the determined target user; (2) the target input data; (3) the expression style of the target user; (4) the interactive image style of the target user, such as the emoji style of the target user, etc.; (5) the target virtual image, such as the target user's preset image, the target user's avatar, etc.; (6) the virtual image of the interactive object or chat object; (7) the virtual image in the target user's commonly used interactive images (emojis), etc.

[0139] For ease of understanding, this disclosure also provides an application embodiment.

[0140] In social media chat applications, users typically need to select from existing emojis to send. This embodiment proposes a method to automatically generate emotion-related emojis based on the user's current chat content and automatically add them to the chat window.

[0141] (1) Get the chat history of the user in the current chat interface.

[0142] (2) Combine chat logs with the user's personal knowledge database to extract the user's speaking style, and combine this with prompt words to send the data to a large language model to infer the user's current mood. Specifically, this may include: when selecting chat logs, categorizing them by speaker if it is a group chat; determining which speaker the user will reply to based on the content the user "quotes" or the other users the input content is directed to. Otherwise, it is considered a reply to the last speaker who is not the user; combining the chat logs of the relevant speakers with the user's historical records (the most recent 10 records) obtained from the personal knowledge database with prompt words, allowing the large language model to analyze the user's speaking style towards this speaker.

[0143] For example, inputting "What is the user's historical speaking style (expression style) in the current chat window?" into the large model allows it to combine its personal knowledge database to determine the context of the current chat window and analyze the user's historical speaking style. Furthermore, inputting "The user's input in the current chat window is: 'That's right, hahaha.' Please infer the user's emotion" into the large model allows it to combine the analyzed historical speaking style to determine the user's current emotion.

[0144] (3) Combine the user's personal knowledge database to extract the user's favorite emoji styles, and based on the user's input content, the user's favorite emoji styles and current mood, use text-to-image model or other generative models to generate emoji images (e.g., emoji packs).

[0145] For example, you can input "what kind of emoji style the user prefers" into a large model. The model can then combine this information with emoji information from the user's personal knowledge database to determine the user's preferred emoji style. Furthermore, based on the user's input ("just kidding, hahaha"), their preferred emoji style, and their current mood, it can use a text-to-image model or other generative models to generate emoji images.

[0146] (4) Automatically add emoticons to the chat window for users to use.

[0147] Specifically, this could involve adding generated emoji images to the user's chat input box, allowing the user to select and decide whether to send the emoji image.

[0148] The generation of emoticons in this embodiment is achieved by combining the context and semantic understanding of the text content input by the user.

[0149] For example, when a user types "hahaha," it sometimes indicates genuine happiness, and sometimes it might be a wry smile. This implementation can combine the user's personal speaking style with recent chat history, and then use a large language model to analyze which type of emoji should be used. When the large language model makes its judgment, it can also be asked to provide a confidence value for the judgment. If the confidence value is less than a certain threshold, the default expression is used; for example, "hahaha" generally represents a happy expression.

[0150] The emojis are generated using a generative large-scale model generation method. If the emoji is a person or animal, user input can be incorporated into the emoji to increase its fun factor.

[0151] Figure 3 The diagram illustrates a chat interface according to an embodiment of the present disclosure. The chat interface can display a target image generated according to an embodiment of the present disclosure. Specifically, the target image may be an emoticon.

[0152] A chat interface can include dialog boxes (or input boxes) and an interactive area, which can contain user interactions or interaction logs. Figure 3 The diagram shows the interaction between the target user and user A.

[0153] According to the embodiments of this disclosure, the corresponding method embodiments can be executed to determine the information used to generate the target image (e.g., input data in the dialog box, interaction records in the interaction area, the user's target virtual image, the user's current emotional data, etc.), generate the target image and display it in the chat interface to facilitate user image interaction. Specifically, users can easily send target images.

[0154] Specifically, this can be achieved by displaying multiple generated target images (e.g., in the "Recommended Emoticons" window that pops up in the chat interface) Figure 3 (The four target images shown in the image). The generated target images may contain the target virtual avatar, the target virtual avatar's expression or actions, and text content, such as the user's input data "Nice to meet you".

[0155] Users can interact with target images by clicking or other actions. Specifically, clicking on an emoticon will display the selected emoticon in a dialog box, and users can then delete or keep the selected emoticon within the dialog box as needed. Users can also select multiple target images (emoticons) to be displayed in the dialog box.

[0156] Corresponding to the above method embodiments, this disclosure also provides apparatus embodiments.

[0157] Figure 4 A schematic diagram illustrating the structure of an image display device according to an embodiment of the present disclosure is shown. Figure 4 As shown, the image display device may include: a first determining unit 301, a second determining unit 302, a generating unit 303, and an image display unit 304.

[0158] The first determining unit 301 is used to determine the target input data input by the target user and to determine the emotional data used to characterize the target user's emotions. In one embodiment, the first determining unit 301 can be used to perform the operation S201 described above, which will not be repeated here.

[0159] The second determining unit 302 is used to determine the first image data corresponding to the target virtual image. In one embodiment, the second determining unit 302 can be used to perform the operation S202 described above, which will not be repeated here.

[0160] The generation unit 303 is used to generate a target image based on the target input data, emotion data, and first image data; wherein the target image contains a target virtual image; the image semantics corresponding to the target virtual image in the target image are used to characterize the target input data and / or emotion data. In one embodiment, the generation unit 303 can be used to perform the operation S203 described above, which will not be repeated here.

[0161] The image display unit 304 is used to display the target image. In one embodiment, the image display unit 304 can be used to perform the operation S204 described above, which will not be repeated here.

[0162] Optionally, the target input data includes data entered by the target user into the chat interface dialog box; the image display unit 304 is also used to: display interactive controls; the interactive controls are used to send the target image in the chat interface.

[0163] Optionally, the target virtual avatar includes at least one of the following: a preset virtual avatar of the target user, a preset virtual avatar of the target chat object corresponding to the chat interface, and a preset emoji virtual avatar of the target user.

[0164] Optionally, the first determining unit 301 is used to: determine emotional data to characterize the emotions of the target user based on the target input data and / or target interaction data; wherein the target interaction data includes interaction data between the target user and other users; the chat interface is used for interaction between user members included in the target group, which includes the target user and other users.

[0165] Optionally, the method for determining the target interaction data includes: determining the historical interaction data between the target user and the interaction object; determining the target interaction data that meets the preset interaction conditions from the historical interaction data; and the target group containing the interaction object.

[0166] Optionally, the preset interaction conditions include at least one of the following: the interaction data targeted by the target input data; the interaction data within the target time period; the interaction data between the target user and a specified object in the interaction objects, wherein the target group contains the specified object.

[0167] Optionally, the generation unit 303 is used to: determine image performance description data based on target input data and emotion data; and generate a target image to characterize the image performance description data by using the first image data as a reference element.

[0168] Optionally, the first determining unit 301 is configured to: determine emotional data for characterizing the target user's emotions based on at least one of the following: target input data, target interaction data, and the target user's expression style; wherein the target interaction data includes interaction data between the target user and other users.

[0169] Optionally, the method for determining the target user's expression style includes: determining the target user's expression style based on the target user's historical interaction data.

[0170] For an explanation of this device embodiment, please refer to other embodiments.

[0171] According to embodiments of this disclosure, any plurality of units among the first determining unit 301, the second determining unit 302, the generating unit 303, and the image display unit 304 can be combined into one unit, or any one of these units can be split into multiple units. Alternatively, at least part of the functionality of one or more of these units can be combined with at least part of the functionality of other units and implemented in one unit. According to embodiments of this disclosure, at least one of the first determining unit 301, the second determining unit 302, the generating unit 303, and the image display unit 304 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in a suitable combination of any of these. Alternatively, at least one of the first determining unit 301, the second determining unit 302, the generating unit 303, and the image display unit 304 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0172] Figure 5 A block diagram schematically illustrates an electronic device suitable for implementing an image display method according to an embodiment of the present disclosure.

[0173] like Figure 5 As shown, an electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage portion 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0174] RAM 903 stores various programs and data required for the operation of electronic device 900. Processor 901, ROM 902, and RAM 903 are interconnected via bus 904. Processor 901 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 902 and / or RAM 903. It should be noted that the programs may also be stored in one or more memories other than ROM 902 and RAM 903. Processor 901 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.

[0175] According to embodiments of this disclosure, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to a bus 904. The electronic device 900 may also include one or more of the following components connected to the I / O interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 910 as needed so that computer programs read from it can be installed into the storage section 908 as needed.

[0176] This disclosure also provides an electronic device, including: a first application running on the electronic device, and a display unit; the first application is configured to: parse at least one task based on input, call a target model to execute at least one task, and at least perform: determining target input data input by a target user, and determining emotional data used to characterize the target user's emotions; determining first image data corresponding to a target virtual avatar; generating a target image based on the target input data, emotional data, and first image data, specifically, calling a target model, inputting the target input data, emotional data, and first image data into the target model, and determining the target image generated by the target model; wherein, the target image contains a target virtual avatar; the image semantics corresponding to the target virtual avatar in the target image are used to characterize the target input data and / or emotional data; and calling the display unit to display the target image.

[0177] Optionally, the first application can be used to perform any of the method embodiments in this disclosure.

[0178] Optionally, the first application can be an agent deployed in an electronic device, an application of artificial intelligence technology. It can be based on a target model (such as a large language model, text-based graph model, or generative large model). The agent's behavior can be determined by the target model invoked by the agent based on the current state and external input, and can be used to specifically solve a certain type of problem. The target model provides the decision-making basis through learning and training on a large amount of data; the agent can also invoke tools, plugins, and knowledge bases to provide reasoning, decision-making, and execution capabilities. In short, the target model can provide decision support, and the agent can continuously optimize the performance of its internal target model through data fed back from actual applications or information input by the user, such as prompts. (The prompt can be the text corresponding to the user input information and / or a pre-set second text used to guide the target model in reasoning).

[0179] In a specific example, the agent can receive user input and data from other applications (such as a smart assistant application on the device). The agent can then analyze the received input, further call a target model for processing, and obtain the processing result from the target model. The agent can then display or perform operations based on the processing result of the target model. Specifically, the agent can execute the above method embodiment to determine the data used to generate the target image and call the target model to generate the target image. The target model could be, for example, a text-based image model or a generative large model, etc.

[0180] This disclosure does not limit the first application, the target model, or the display unit. Optionally, the display unit may be a display component in an electronic device, such as the screen of the electronic device.

[0181] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0182] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the methods of the embodiments of this disclosure.

[0183] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure can be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0184] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.

Claims

1. An image display method, comprising: Determine the target input data input by the target user, and determine the emotional data used to characterize the target user's emotions; Determine the first image data corresponding to the target virtual image; A target image is generated based on the target input data, the emotion data, and the first image data; wherein, the target image contains the target virtual image; the image semantics corresponding to the target virtual image in the target image are used to characterize the target input data and / or the emotion data; Display the target image.

2. The method according to claim 1, wherein the target input data includes data input by the target user into the chat interface dialog box; the method further includes: Display interactive controls; the interactive controls are used to send the target image in the chat interface.

3. The method according to claim 2, wherein the target virtual avatar comprises at least one of the following: The preset virtual avatar of the target user, the preset virtual avatar of the target chat object corresponding to the chat interface, and the preset virtual emoticon avatar of the target user.

4. The method according to claim 2, wherein determining the emotional data used to characterize the target user's emotion includes: Based on the target input data and / or target interaction data, determine the emotional data used to characterize the target user's emotions; The target interaction data includes interaction data between the target user and other users; the chat interface is used for interaction between user members in the target group, which includes the target user and the other users.

5. The method according to claim 4, wherein the method for determining the target interactive data includes: Determine the historical interaction data between the target user and the interaction object; From the historical interaction data, target interaction data that meets the preset interaction conditions is identified; The target group contains the interactive object.

6. The method according to claim 5, wherein the preset interaction conditions include at least one of the following: The target input data refers to the interactive data; Interaction data within the target time period; The interaction data between the target user and a specified object in the interaction object, wherein the specified object is included in the target group.

7. The method according to claim 1, wherein generating the target image based on the target input data, the emotion data, and the first image data comprises: Based on the target input data and the emotion data, determine the image expression description data; Using the first image data as a reference element, a target image is generated to characterize the image representation description data.

8. The method according to claim 1, wherein determining the emotional data used to characterize the target user's emotion comprises: Emotional data for characterizing the target user's emotions is determined based on at least one of the following: the target input data, the target interaction data, and the target user's expression style; The target interaction data includes the interaction data between the target user and other users.

9. The method according to claim 8, wherein the method for determining the expression style of the target user includes: Based on the target user's historical interaction data, the target user's expression style is determined.

10. An electronic device, comprising: A first application running on the electronic device, and a display unit; The first application is configured to: parse at least one task based on the input, and invoke the target model to execute at least one task, for at least the following purposes: Determine the target input data input by the target user, and determine the emotional data used to characterize the target user's emotions; Determine the first image data corresponding to the target virtual image; The target model is invoked, and the target input data, the emotion data, and the first image data are input into the target model to determine the target image generated by the target model; wherein, the target image contains the target virtual image; the image semantics corresponding to the target virtual image in the target image are used to characterize the target input data and / or the emotion data; The display unit is invoked to display the target image.