Image generation method, apparatus and display device

By employing multimodal information processing methods and utilizing encoding, fusion, and multi-stage networks in the image generation model, the problem that single-modal information-generated images cannot meet users' personalized needs is solved, thus achieving higher-quality image generation.

CN119520934BActive Publication Date: 2026-05-12HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HISENSE VISUAL TECH CO LTD
Filing Date
2024-11-05
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing single-modal information generation image technology cannot meet users' personalized needs and cannot effectively integrate multiple information sources to generate images that meet users' expectations.

Method used

A multimodal information processing method is adopted, which uses the encoding network, fusion network and multi-stage network in the image generation model to process audio, text, image and video information respectively, and generate images that meet the user's needs.

Benefits of technology

It achieves effective fusion of multimodal information, generating more realistic images that meet user expectations, thus improving user experience and the quality of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119520934B_ABST
    Figure CN119520934B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an image generation method, device and display equipment, the method comprising: acquiring multi-modal information input by a user; processing each modality information in the multi-modal information based on an encoding network in an image generation model to obtain a feature vector corresponding to each modality information; performing fusion processing on the feature vector corresponding to each modality information based on a fusion network in the image generation model to obtain a fusion vector; and processing the fusion vector based on the multi-modal information through a multi-stage network in the image generation model to obtain a target image corresponding to the multi-modal information. The image generation method implemented by the present application can generate a corresponding image based on multi-modal information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of display device technology, and in particular to an image generation method, apparatus and display device. Background Technology

[0002] With the rapid development of image generation technology, artificial intelligence and machine learning technologies can achieve automatic image generation. These technologies can generate images based on single-modal information, such as generating images based on text information or corresponding images based on image information. However, with the increasing demand for personalized images from users, this single-modal information generation method can no longer meet their needs. Summary of the Invention

[0003] This application provides an image generation method, apparatus, and display device that can generate corresponding images based on multimodal information to meet users' personalized needs.

[0004] A first aspect of this application provides an image generation method, comprising: first, acquiring multimodal information input by a user; wherein the multimodal information includes at least two of audio information, text information, image information, and video information. Then, processing each modal information based on an encoding network in an image generation model to obtain a feature vector corresponding to each modal information. Second, fusing the feature vectors corresponding to each modal information based on a fusion network in the image generation model to obtain a fused vector. Finally, based on the multimodal information, processing the fused vector through a multi-stage network in the image generation model to obtain a target image corresponding to the multimodal information.

[0005] The image generation method provided in this application, after acquiring multimodal information input by the user, inputs the multimodal information into an image generation model. The image is then processed sequentially through an encoding network, a fusion network, and a multi-stage network within the image generation model to obtain a target image that conforms to the multimodal information. The image generation method provided in this application can process multimodal information to generate images that meet requirements, satisfying users' personalized needs and improving the user experience.

[0006] In some embodiments, based on multimodal information, the fusion vector is processed by a multi-stage network in the image generation model to obtain a target image corresponding to the multimodal information, including: determining at least one auxiliary feature information corresponding to the multimodal information; processing the fusion vector and at least one auxiliary feature information based on at least one feature sub-network in the multi-stage network to obtain a target vector; and decoding the target vector based on the image decoding sub-network in the multi-stage network to obtain a target image.

[0007] Based on the above scheme, when the multi-stage network processes the fused vector, auxiliary feature information is introduced. The auxiliary feature information enables the multi-stage network to extract more realistic and vivid features, thereby improving the quality of the subsequently generated target image and enabling the image generation model to generate images that better meet the user's needs.

[0008] In some embodiments, processing a fusion vector and at least one auxiliary feature information based on at least one feature subnetwork in a multi-stage network to obtain a target vector includes: processing the fusion vector and first auxiliary feature information based on a first feature subnetwork in at least one feature subnetwork to determine a first output vector; wherein the at least one auxiliary feature information includes first auxiliary feature information, second auxiliary feature information, and third auxiliary feature information; processing the first output vector and second auxiliary feature information based on an intermediate feature subnetwork in at least one feature subnetwork to determine a second output vector; and processing the second output vector and third auxiliary feature information based on a second feature subnetwork in at least one feature subnetwork to determine the target vector.

[0009] Based on the above scheme, in a multi-stage network, the vector output by the previous feature sub-network and the auxiliary feature information corresponding to the current feature sub-network are input into the current feature sub-network. After processing by the current feature sub-network, its output feature vector is then input into the next feature sub-network. By dividing the multi-stage network into hierarchical relationships, each stage receives the output information and auxiliary information from the previous stage, achieving fine-grained control and enabling controllable generation across multiple stages. This improves the output quality of the image and enhances the execution efficiency and performance of the image generation model.

[0010] In some embodiments, the feature vectors corresponding to each modality information are fused based on the fusion network in the image generation model to obtain a fusion vector, including: determining a preset length; and based on the preset length, the feature vectors corresponding to each modality information are fused through the fusion network to obtain a fusion vector of the preset length.

[0011] Based on the above scheme, the fusion network generates a fusion vector of a preset fixed length by fusing the feature vectors corresponding to the multimodal information. Generating a fusion vector of a preset length can improve the operating efficiency while ensuring the processing performance of the encoding network, making it more suitable for real-time processing and large-scale data analysis scenarios.

[0012] In some embodiments, based on a preset length, a fusion network is used to fuse the feature vectors corresponding to each modality information to obtain a fusion vector of the preset length. This includes: performing feature selection processing on the feature vectors corresponding to each modality information through a fusion network to obtain selected feature vectors; mapping the selected feature vectors to a preset feature space to obtain mapped feature vectors; and performing fusion processing on the mapped feature vectors based on the preset length to obtain the fusion vector of the preset length. The fusion processing includes at least one of weighted fusion, feature-level fusion, and decision-level fusion.

[0013] Based on the above scheme, during the fusion process, the most useful features can be selected through feature selection, thereby reducing the dimensionality of the feature space and improving the generalization ability and computational efficiency of the image generation model. Simultaneously, through feature mapping and fusion processing, multiple feature vectors are fused into a new feature vector, realizing the fusion processing of multimodal information and providing a foundation for subsequent image generation based on multimodal information.

[0014] In some embodiments, the modal information is processed based on the encoding network in the image generation model to obtain the feature vector corresponding to each modal information, including at least one of the following: inputting audio information into the audio encoding subnetwork of the encoding network to generate the audio feature vector corresponding to the audio information; inputting text information into the text encoding subnetwork of the encoding network to generate the text feature vector corresponding to the text information; inputting image information into the image encoding subnetwork of the encoding network to generate the image feature vector corresponding to the image information; and inputting video information into the video encoding subnetwork of the encoding network to generate the video feature vector corresponding to the video information.

[0015] Based on the above scheme, the encoding network includes multiple encoding sub-networks for processing different modal information, thereby processing various modal information of user entry and exit, and providing a foundation for the subsequent fusion network to fuse the feature vectors corresponding to multimodal information.

[0016] In some embodiments, the audio information is input into the audio coding subnetwork in the coding network to generate an audio feature vector corresponding to the audio information, including: preprocessing the audio information to obtain preprocessed audio information; performing feature extraction on the preprocessed modal information to obtain an initial feature vector; and encoding the initial feature vector to obtain an audio feature vector.

[0017] Based on the above scheme, the audio coding sub-network in the coding network can perform audio information encoding processing, including preprocessing, feature extraction and encoding processing, to extract meaningful feature representations from the user-input audio information and improve the accuracy of subsequent image generation.

[0018] In some embodiments, the method further includes: acquiring multimodal training information; processing the multimodal training information based on the image generation model to be trained to generate a predicted image; wherein the image generation model to be trained includes a coding network to be trained, a fusion network to be trained, and a multi-stage network to be trained; acquiring sample images; using the predicted image as the initial training output information of the image generation model to be trained and the sample image as the supervision information, iterating the image generation model to be trained to obtain an image generation model.

[0019] Based on the above scheme, this application embodiment also proposes a training process for the image generation model. The image generation model to be trained is trained by acquiring multimodal training information, and the image generation model to be trained is iteratively trained using sample images as supervision information to obtain the trained image generation model. The trained image generation model can be used to convert multimodal information into corresponding target images, realize the processing of multimodal information to generate images that meet the requirements, meet the personalized needs of users, and improve the user experience.

[0020] A second aspect of this application provides an image generation apparatus, including an acquisition module and a determination module. The acquisition module is configured to acquire multimodal information input by a user; wherein the multimodal information includes at least two of audio information, text information, image information, and video information. The determination module is configured to process each modal information based on an encoding network in an image generation model to obtain a feature vector corresponding to each modal information; to fuse the feature vectors corresponding to each modal information based on a fusion network in the image generation model to obtain a fused vector; and to process the fused vector based on the multimodal information through a multi-stage network in the image generation model to obtain a target image corresponding to the multimodal information.

[0021] The image generation apparatus provided in this application, after acquiring multimodal information input by the user, inputs the multimodal information into an image generation model. The image is then processed sequentially through an encoding network, a fusion network, and a multi-stage network within the image generation model to obtain a target image that conforms to the multimodal information. The image generation method provided in this application can process multimodal information to generate images that meet requirements, satisfying users' personalized needs and improving the user experience.

[0022] A third aspect of this application provides a display device, including a display and a controller coupled to the display. The display is configured to display an application interface for an image generation application; the application interface includes a first input area, a second input area, and an image display area. The controller is configured to: respond to receiving image information input by a user in the first input area and text information input in the second input area; process each modal information in the multimodal information based on an encoding network in an image generation model to obtain feature vectors corresponding to each modal information; wherein the multimodal information includes image information and text information; fuse the feature vectors corresponding to each modal information based on a fusion network in the image generation model to obtain a fusion vector; and process the fusion vector based on the multimodal information through a multi-stage network in the image generation model to obtain a target image corresponding to the multimodal information.

[0023] The display device provided in this application embodiment can receive image information input by a user in a first input area and text information input in a second input area, and generate a target image based on the image information, text information, and a trained image generation model. The target image can satisfy the requirements of both image and text information. The display device provided in this application embodiment can process multimodal information to generate images that meet requirements, satisfy personalized user needs, and improve user experience. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in an embodiment of this application.

[0026] Figure 2 This is a schematic diagram of the hardware configuration of the display device provided in the embodiments of this application;

[0027] Figure 3 This is a schematic diagram of the hardware configuration of the control device provided in the embodiments of this application;

[0028] Figure 4 A schematic diagram of an application interface provided in an embodiment of this application;

[0029] Figure 5 A schematic diagram of another application interface provided in an embodiment of this application;

[0030] Figure 6 A schematic diagram illustrating an image display method provided in an embodiment of this application;

[0031] Figure 7 A schematic diagram of an application interface provided in an embodiment of this application;

[0032] Figure 8 A schematic diagram illustrating another image display method provided in an embodiment of this application;

[0033] Figure 9 A schematic diagram of another application interface provided in an embodiment of this application;

[0034] Figure 10 A schematic diagram illustrating yet another image display method provided in an embodiment of this application;

[0035] Figure 11 A schematic diagram illustrating yet another image display method provided in an embodiment of this application;

[0036] Figure 12 A schematic diagram illustrating another application interface provided in an embodiment of this application;

[0037] Figure 13 A schematic diagram illustrating yet another image display method provided in an embodiment of this application;

[0038] Figure 14 A schematic diagram illustrating an image generation method provided in an embodiment of this application;

[0039] Figure 15 A schematic diagram of an image generation model provided in an embodiment of this application;

[0040] Figure 16 A schematic diagram illustrating yet another image generation method provided in an embodiment of this application;

[0041] Figure 17 A schematic diagram illustrating yet another image generation method provided in an embodiment of this application;

[0042] Figure 18 A schematic diagram of a multi-stage network structure provided in an embodiment of this application;

[0043] Figure 19 A flowchart illustrating an image display method provided in an embodiment of this application;

[0044] Figure 20 This is a flowchart illustrating another image display method provided in an embodiment of this application. Detailed Implementation

[0045] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0046] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0047] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0048] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0049] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0050] In this embodiment, the display device 200 generally refers to a device with screen display and data processing capabilities. For example, the display device 200 includes, but is not limited to, smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.

[0051] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application. For example... Figure 1 As shown, a user can operate the display device 200 via touch operation, a mobile terminal 300, and a control device 100. The control device 100 receives user input commands and converts them into control commands that the display device 200 can recognize and respond to. For example, the control device 100 can be a remote control, a stylus, a gamepad, etc.

[0052] In some embodiments, the control device 100 may be a remote control, and the communication between the remote control and the display device includes at least one of infrared protocol communication, Bluetooth protocol communication, and other short-range communication methods, controlling the display device 200 wirelessly or via a wired connection. Users can control the display device 200 by inputting user commands through buttons on the remote control, voice input, control panel input, etc.

[0053] The mobile terminal 300 can function as a control device for human-computer interaction between the user and the display device 200. It can also function as a communication device for establishing a communication connection with the display device 200 and exchanging data. In some embodiments, the mobile terminal 300 can have software applications installed on it and communicate with the display device 200 via network communication protocols to achieve one-to-one control and data communication. Furthermore, it can transmit audio and video content displayed on the mobile terminal 300 to the display device 200 for synchronized display.

[0054] In some embodiments, the mobile terminal 300 or other electronic devices may also simulate the functions of the control device 100 by running an application that controls the display device 200.

[0055] like Figure 1 The diagram also shows that the display device 200 communicates with the server 400 via various communication methods. This allows the display device 200 to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0056] Display device 200 can provide broadcast television reception function, and can also be equipped with intelligent network television function that provides computer support, including but not limited to network television, smart television, Internet Protocol television (IPTV), etc.

[0057] Figure 2 This is a hardware configuration block diagram of a display device 200 provided in some embodiments of this application.

[0058] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0059] In some embodiments, detector 230 is used to acquire signals from the external environment or to interact with the outside world. For example, detector 230 includes a light receiver, a sensor for acquiring ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0060] In some embodiments, device interface 240 may include, but is not limited to, a high-definition multimedia interface (HDMI), an analog or data high-definition component input interface (component), a composite video input interface (CVBS), a USB input interface (USB), an RGB port, etc. Device interface 240 may also be a composite input / output interface formed by multiple of the above interfaces.

[0061] Display device 200 can connect to external device 500 via device interface 240. External device 500 can be a set-top box, game console, PC, or other similar device. When display device 200 connects to external device 500 via device interface 240, a corresponding signal source channel can be established, and input signals can be received through this signal source channel. For example, display device 200 can connect to a game device via an HDMI interface to establish an HDMI channel, and receive audio and video signals sent by the game device based on the HDMI channel.

[0062] In some embodiments, the display 260 includes display function components for presenting images and driving components for driving image display. The display 260 is used to receive and display image signals output from the controller 250. For example, the display 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces, etc.

[0063] In some embodiments, the communication device 220 is a component used to communicate with external devices or the server 400 according to various communication protocol types. The display device 200 may have multiple communication devices 220 depending on the supported communication methods. For example, when the display device 200 supports wireless network communication, it may have a communication device 220 with WiFi functionality. When the display device 200 supports Bluetooth connectivity, it needs to have a communication device 220 with Bluetooth functionality.

[0064] The communication device 220 enables the display device 200 to communicate with external devices or the server 400 via wireless or wired connections. Wired connections utilize data cables, interfaces, or other components to connect the display device 200 to external devices. Wireless connections utilize wireless signals or wireless networks. The display device 200 can directly establish a connection with external devices or indirectly through gateways, routers, or other connection devices.

[0065] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the display device 200.

[0066] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0067] In some embodiments, a user can input user commands through a graphical user interface (GUI) displayed on a display 260, and the user input interface receives user input commands through the graphical user interface (GUI).

[0068] In some embodiments, the audio output device 270 can be a built-in speaker of the display device 200 or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, through which the audio output device can be connected to the display device 200 to output sound from the display device 200.

[0069] In some embodiments, the user input interface 280 can be used to receive instructions from user input.

[0070] To enable user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program used to manage and control the hardware and software resources of the display device 200. The operating system can control the display device to provide a user interface; for example, the operating system can directly control the display device to provide a user interface, or it can provide a user interface by running an application. The operating system also allows users to interact with the display device 200.

[0071] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system that is deeply customized based on a specific operating platform, or an independent operating system specifically developed for display devices.

[0072] An operating system can be divided into different modules or levels based on the functions it implements. For example, in some embodiments, such as Figure 3 As shown, taking the Android system as an example, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the System Library layer, and the Kernel layer. In some embodiments, the operating system of the display device can also be other operating systems.

[0073] In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with the user based on the applications. The application layer may contain at least one application, which may be a built-in Windows program, system settings program, or clock program of the operating system; or it may be an application developed by a third-party developer. In specific implementations, the application packages in the application layer are not limited to the examples above.

[0074] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.

[0075] like Figure 3As shown, the application framework layer in this embodiment includes a view system, managers, and content providers. The view system designs and implements the application's interface and interactions, and includes lists, grids, textboxes, and buttons. The managers include at least one of the following modules: an activity manager for interacting with all running activities in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0076] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling changes to the display window, such as shrinking the display window, shaking the display, or distorting the display.

[0077] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime library layer, such as the C / C++ instruction library, to implement the functions to be performed by the framework layer.

[0078] In some embodiments, the kernel layer is a functional layer situated between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, ... Figure 3 As shown, hardware drivers can be configured in the kernel layer. The drivers included in the kernel layer can be at least one of the following: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0079] It should be noted that the above examples are merely a simple division of operating system functions and do not limit the specific form of the operating system of the display device 200 in this application embodiment. Depending on the function of the display device, the type of operating system, and other factors, the number of levels and the specific level type of the operating system may be expressed in other forms.

[0080] In some embodiments, the display device 200 may run an image generation application to generate the target image. This image generation application may be a system application built into the operating system of the display device 200. Alternatively, the image generation application may be a standalone application or a third-party application developed based on the operating system of the display device 200.

[0081] During operation, the display device 200 can receive control commands input by the user. Some of these control commands are used to control the display device 200 to generate a target image; these commands can be referred to as image generation commands. Depending on the interaction methods supported by the display device 200, the user can input image generation commands in different ways.

[0082] In some embodiments, the display device 200 can obtain image generation instructions by detecting user input via key input on the accompanying control device 100. Specifically, while the display device 200 is displaying the application list interface, the user can use the arrow keys on the control device 100 to move the focus cursor in the application list interface. After the focus cursor moves to the image generation application icon, the user can press the confirmation key on the control device 100 to control the display device 200 to run the image generation application, i.e., input the image generation instruction. The display device 200 can then run the image generation application according to the image generation instruction.

[0083] In some embodiments, when the display device 200 supports interaction methods such as touch, voice, motion sensing, and gestures, the display device 200 can detect that the user can input interaction signals through the corresponding interaction method and obtain image generation instructions based on the interaction signals. For example, when the voice assistant application built into the display device 200 detects that the user inputs a voice command such as "Please generate an image containing a bird, with blue feathers on its back and white feathers on its belly," it determines that the user's intent is an image generation intent based on the keywords in the voice command. That is, it determines that the current voice command is an image generation instruction, and at this time, the display device 200 can run the image generation application in response to the image generation instruction.

[0084] In some embodiments, the image generation instruction can also be generated by the display device 200 by monitoring various runtime events during operation and generating the instruction when a specific runtime event is detected. For example, the display device 200 can run a search engine application such as "×× Search," and the search interface can include a text input box and an image input box. When the user enters search text in the text input box and / or enters a search image in the image input box, the display device 200 can automatically generate an image generation instruction and run the image generation application in response to the instruction.

[0085] In response to an image generation command, the display device 200 can run an image generation application and display the application interface of the image generation application.

[0086] In some embodiments, in response to a received launch operation for an image generation application, the display device 200 launches the image generation application, and the display shows the application interface corresponding to the image generation application. The application interface of the image generation application may include two input areas: a first input area and a second input area. The first input area is used for inputting and displaying image content from the input data; this first input area may also be referred to as a drawing area or image drawing area. The second input area is used for inputting and displaying text content from the input data; this second input area may also be referred to as a text input area.

[0087] For example, the first input area can be used to display an image, which can be an image drawn by the user, an image taken by the user through an image acquisition device on the display device, or an image uploaded through a terminal device.

[0088] For example, when the display device supports touch functionality, users can directly draw images in the first input area, or capture images using the image sensor on the display device, or upload images via a terminal device. When the display device does not support touch functionality, users can capture images using the image sensor on the display device, or upload images via a terminal device.

[0089] For example, the second input area is used to display descriptive text input by the user. The descriptive text displayed in the second input area can be descriptive text input by the user, or it can be descriptive text generated based on the user's input. For example, the descriptive text can be a description of the user's needs.

[0090] In some embodiments, the application interface of the image generation application may further include an image display area for displaying the generated target image. The image display area may be displayed side by side with the first input area to facilitate user comparison of the target image and the image content in the input data.

[0091] For example, the image display area is used to display the image generated after processing by the image generation model. The image display area can also be called the generation area or image generation area. For instance, the image generation model can generate a target image based on the image displayed in the first input area and / or the descriptive text displayed in the second input area, and then display the target image in the image display area. The target image conforms to the requirements of the descriptive image and the descriptive text.

[0092] For example, the image display area can be used to display the image displayed in the first input area after processing by the image generation model, or it can be used to display the image of the descriptive text displayed in the second input area after processing by the image generation model, or it can be used to display the image of the descriptive image displayed in the first input area and the image of the descriptive text displayed in the second input area after processing by the image generation model.

[0093] The first and second input areas can be represented by different shapes and distributed in different locations on the application interface, according to the UI design specifications of the image generation application. In some embodiments, the first input area, the second input area, and the image display area can be three independent areas on the application interface that do not overlap. For example, Figure 4 As shown, the first input area can be located in the upper left corner of the application interface. To facilitate inputting and displaying image content, the first input area can be a rectangular area. The second input area can be located at the bottom of the application interface. Similarly, to facilitate inputting and displaying text content, the second input area can be a long, narrow area. The image display area is located in the upper right corner of the application interface, and to facilitate displaying the target image, the image display area can also be a rectangular area.

[0094] In some embodiments, there may be overlapping areas among the first input area, the second input area, and the image display area. The display device 200 can determine the area to which the content belongs based on the location of the focus marker. For example, when the display device 200 receives a user's request to move the focus cursor to the first input control corresponding to the first input area, the entire screen area of ​​the current application interface can be used to receive the user's input image. That is, the first input area at this time is the entire application interface with the focus marker in the first input control state. Similarly, when the display device 200 receives a user's request to move the focus cursor to the second input control corresponding to the first input area, the first input area is the entire application interface with the focus marker in the second input control state.

[0095] To facilitate user input and image display, in some embodiments, the display device 200 can set a display hierarchy among the first input area, the second input area, and the image display area. For example, when the display device 200 detects that the user is performing image-related operations such as touch drawing, image capture, or image uploading, the display hierarchy of the first input area can be set higher than that of the second input area and the image display area. Similarly, when the display device 200 detects that the user is performing text-related operations such as text input or voice interaction, the display hierarchy of the second input area can be set higher than that of the first input area and the image display area. Likewise, after the display device 200 detects that an image generation model has generated a target image, the display hierarchy of the image display area can be set higher than that of the first and second input areas.

[0096] The display device 200 can dynamically adjust the display content in the first input area, the second input area, and the image display area at different operating stages.

[0097] The first input area can display different content at different interaction stages depending on the interaction method supported by the display device 200. In some embodiments, if the display device 200 supports touch interaction, for example, the display 260 of the display device 200 is connected to a touch component. The touch component can form a touch screen with the display 260 to detect touch signals input by the user based on the touch screen.

[0098] The first input area can be associated with a touch component, which detects user input touch points through changes in capacitance or resistance. When the user moves their finger or stylus on the touchscreen, the display device 200 continuously records the coordinates of the touch points. Touchscreen interaction programs, such as the touchscreen driver, record these coordinate points, forming a series of point data, i.e., a set of trajectory detection points.

[0099] After obtaining the set of detection points for the drawing trajectory, the image drawing application draws the trajectory based on this point data. That is, the image drawing application can draw lines on the first input area based on the line style corresponding to the user-selected drawing control. For example, if the user selects a pen control, the image drawing application will display lines in the first input area with the line style (red, thick line, solid line) that match the shape of the touch trajectory detection points, based on the pen control's current line style: red, thick line, solid line. The image drawing application can be a standalone application or a drawing function unit within an image generation application. Line drawing can be achieved using algorithms such as Bresenham.

[0100] After detecting and generating a set of drawing trajectory detection points, the display device 200 can save the drawn trajectory as an image file. That is, the display device 200 converts the drawing trajectory detection points into a set of trajectory pixels to generate a user-drawn image (also known as a first image). Similarly, the display device 200 can generate a user-drawn image through a standalone image drawing application or through a drawing function unit in an image generation application.

[0101] During the generation of the drawn image (such as the first image), the display device 200 can monitor the touch signals detected by the touch component in real time, determine the touch node based on the touch signal, and generate the drawn image based on the touch trajectory at the appropriate touch node. Therefore, the first input area in the application interface is used to obtain the drawn image input by the user. That is, the display device 200 can respond to receiving a drawing instruction input in the first input area, extract the touch trajectory corresponding to the drawing instruction, generate the drawn image based on the touch trajectory, and control the display to display the drawn image in the first input area.

[0102] In some embodiments, the display device 200 can generate a drawn image according to a preset detection cycle. That is, the display device 200 can obtain a preset detection cycle and extract the touch trajectory of the user input based on the first input area according to the preset detection cycle. The preset detection cycle is set according to the denoising rounds of the image generation model. For example, if the denoising rounds of the image generation model are 30 rounds and the total time is 3 seconds, then the preset detection cycle is 1 second, meaning that detection is performed once every 10 rounds of denoising. The display device 200 can use the touch position when the input time reaches a preset touch response cycle as the touch node, thereby obtaining the touch trajectory of the first input area every 1 second and generating a drawn image based on the touch trajectory.

[0103] In some embodiments, the display device 200 can listen to touch events in the touch trajectory and generate a drawn image based on the listened touch events. The touch events may include touch down events (ACTION_DOWN), touch move events (ACTION_MOVE), and touch up events (ACTION_UP). After the display device 200 listens for a touch down event (ACTION_DOWN) through the touch component, it continuously listens for touch up events (ACTION_UP) and uses the touch up event as an image generation node. That is, after listening for a touch up event, it generates a drawn image in response to the touch up event. The application interface is then refreshed according to a preset detection cycle to display the drawn image in the first input area.

[0104] In some embodiments, if the display device 200 does not support touch interaction, i.e., the display device 200 does not include a touch component, the display device 200 can acquire a descriptive image, wherein the descriptive image is a file in an image format input by the user. The descriptive image can be acquired in real time by an image acquisition device such as a camera, or it can be uploaded by the user via a mobile terminal 300.

[0105] To acquire descriptive images in real time, in some embodiments, the display device 200 may have a built-in or external image acquisition device. Specifically, the display device 200 may include a device interface 240 configured to connect to an image acquisition device. The image acquisition device incorporates an image sensor such as a charge-coupled device (CCD) or complementary metal-oxide-semiconductor (CMOS), which can perform photographing or video recording of the detection area to acquire descriptive images.

[0106] For descriptive images acquired in real time, the display device 200 can display a shooting interaction control for real-time image acquisition in the application interface. When the user inputs an image acquisition command based on the shooting interaction control, the display device 200 can run an image acquisition device through the image acquisition application and acquire a descriptive image through the image acquisition device.

[0107] After the image acquisition device starts running, the display device 200 can display the environmental image captured by the image acquisition device through the first input area. During the display of the environmental image, the display device 200 can also receive a confirmation command input by the user. In response to the confirmation command, it acquires the environmental image corresponding to the moment the confirmation command is input to obtain a descriptive image.

[0108] For example, when the display device 200 does not support touch interaction, but has a built-in or external image acquisition device such as a camera, the user can draw a descriptive image on a drawing medium such as paper, and then take a picture of the drawing medium through the image acquisition device to obtain the descriptive image.

[0109] In order to acquire the descriptive image uploaded by the user, in some embodiments, the display device 200 further includes a communication device 220 configured to establish a communication connection with the mobile terminal 300. The mobile terminal 300 has a camera device capable of scanning barcodes.

[0110] In some embodiments, the application interface of the image generation application may include a QR code scanning control. When the display device 200 receives an image acquisition instruction input by the user based on the QR code scanning control, it can display an upload identification code for uploading the image in the first input area, such as... Figure 5As shown, when a user activates the QR code scanning function on mobile terminal 300 and scans the QR code using the camera device, an image selection interface for uploading image files will be displayed on mobile terminal 300. After the user selects the image to be uploaded and enters a confirmation command, mobile terminal 300 can send the selected image to display device 200 in response to the confirmation command. Upon receiving the image, display device 200 will display the image in the first input area of ​​the application interface.

[0111] In some embodiments, the display device 200 can also obtain the user-uploaded descriptive image through the server 400. That is, the communication device 220 of the display device 200 is configured to establish a communication connection with the server 400. The server 400 also establishes a communication connection with the mobile terminal 300 to receive the descriptive image uploaded by the user through the mobile terminal 300 and store the uploaded descriptive image in a specific Uniform Resource Locator (URL) address. In the step of receiving the descriptive image uploaded by the mobile terminal 300 by scanning the upload identification code, the display device 200 can first obtain the URL address issued by the server. That is, obtain the storage address of the descriptive image sent by the mobile terminal 300 to the server 400 after scanning the upload identification code. An image receiving request is generated based on the URL address, and the image receiving request is sent to the server, i.e., accessing the URL address issued by the server 400, causing the server 400 to respond to the image receiving request and issue the descriptive image. The display device 200 then receives the descriptive image issued by the server 400.

[0112] For example, in response to a user pressing a directional key on the control device 100, the display device 200 controls the focus cursor to move within the image generation application interface. When the focus cursor moves to the QR code scanning control, the display device 200 can receive a confirmation command from the user by pressing the confirmation key on the control device 100. At this time, in response to the confirmation command, the display device 200 controls itself to display the upload identification code, that is, to display a QR code for uploading a descriptive image in the first input area. After the user initiates scanning the QR code using the camera device of the mobile terminal 300, the selected descriptive image (such as the seventh image) can be uploaded to the server 400. After receiving the descriptive image uploaded by the user, the server 400 can send a URL address to the display device 200. After receiving the URL address, the display device 200 cancels the display of the QR code in the first input area and accesses the URL address to download the descriptive image. Furthermore, after downloading the descriptive image, the display device 200 then displays the descriptive image again through the first input area.

[0113] It should be noted that the method of acquiring images through the first input area provided in the above embodiments can be used alone or in combination. That is, for the display device 200 that supports touch interaction, the descriptive image can also be acquired by acquiring images in real time through an image acquisition device, or by uploading images through a mobile terminal 300.

[0114] In some embodiments, the display device 200 may be configured to acquire a drawing image and / or a description image through a first input area. Since both the drawing image and the description image can be displayed in the first input area, the display device 200 can acquire the drawing image or the description image by detecting whether an image is displayed in the first input area.

[0115] In some embodiments, after generating the target image, the display device 200 can also perform image file operations such as saving and sharing the target image in response to user interaction. For this purpose, the application interface of the image generation application can also include an image function area, which includes at least one interactive control. After displaying the target image, the display device 200 can receive image interaction commands input by the user based on the interactive controls. In response to the image interaction commands, it obtains the control identification information of the target interactive control, which is the interactive control selected by the image interaction commands. Then, it queries the function process corresponding to the control identification information and runs the function process based on the function mapping table.

[0116] like Figure 4 As shown, the image function area may include a save control. After displaying the target image, the display device 200 can obtain the image interaction command input by the user for the save control. The display device 200 can then respond to the image interaction command, obtain the recognition information "Save button" of the save control, and invoke the image file saving process, i.e., "ImageFile saving process," based on the recognition information to store the target image in a specific storage space.

[0117] In some embodiments, such as Figure 4 As shown, the image function area may also include a sharing control. The display device 200, in response to a user's image interaction command input based on the sharing control, can invoke the "Image File sharing process" according to the sharing control's identification information "sharebutton". The target image is then sent to the target client according to the file sharing process.

[0118] Display device 200 can also send the generated target image to server 400 via a file sharing process for storage. The target client can download the target image from server 400 by sending a request to server 400.

[0119] In some embodiments, the display device 200 may also support interactive operations on multiple target images. That is, the image function area may further include interactive controls for picture books (such as...). Figure 4 (The illustrated picture book generation control). After receiving the image interaction command input by the user based on the picture book interaction control, the display device 200 can obtain the target images saved within a preset generation period to generate a target image set. Then, it can obtain the descriptive text information input in the second input area during the target image generation process to generate a picture book file based on the target image set and the descriptive text information.

[0120] In some embodiments, after generating the target image, the display device 200 can receive image processing instructions input by the user. These instructions can be input based on image processing controls within the image function area of ​​the application interface. For example, the image function area may include image processing controls such as background enhancement and image quality enhancement. The display device 200 can obtain the image processing instructions by detecting any image processing control selected by the user through touch interaction.

[0121] After receiving an image processing instruction, the display device 200 can respond by invoking an image processing algorithm model. Then, based on the image processing algorithm model, it performs corresponding image processing on the target image. These image processing algorithms include, but are not limited to, color and brightness adjustment, background blurring, edge detection and matting, background replacement, depth estimation, image segmentation, and style transfer. For example, if a user inputs an image processing instruction by clicking the background enhancement control, the display device 200 can determine the foreground and background regions in the image based on the descriptive text entered in the second input area. It then invokes a color and brightness adjustment algorithm to adjust parameters such as color, contrast, and brightness of pixels within the background region of the target image, making the background more harmonious or highlighting the foreground subject.

[0122] Figure 6 This is a schematic diagram illustrating an image display method provided in an embodiment of this application. The following is in conjunction with... Figure 6 When the display device supports touch functionality, the implementation process of the image display method provided in this application embodiment will be described. For example... Figure 6 As shown, the method includes steps 610 to 640 as shown below.

[0123] Step 610: In response to receiving a touch command input by the user in the first input area, control the display to display the first image drawn by the user in the first input area.

[0124] In some examples, the display device can listen for touch commands input by the user in the first input area in real time. When a touch command is detected, the device can display an image drawn by the user (i.e., a first image) in the first input area according to the touch command. The touch command can be input by the user moving their limbs (such as fingers) or a touch device (such as a stylus) on the touchscreen of the first input area. It should be noted that the following embodiments illustrate the example of a user inputting touch commands through a touch device (such as a stylus).

[0125] For example, when a user moves through the first input area using a touch device, the coordinates of the touch point will change. The display device can detect the change in the touch point coordinates, obtain the movement trajectory of the touch point, generate a map image based on the movement trajectory of the touch point, and display the first image drawn by the user in the first input area.

[0126] Figure 7 This is a schematic diagram of an application interface provided in an embodiment of this application. For example... Figure 7 As shown, the application interface 700 includes a first input area 710, a second input area 720, and an image display area 730. The first input area 710 on the left side of the application interface 700 displays a descriptive image drawn by the user (such as the first image).

[0127] Step 620: Obtain the second image generated based on the first image, and control the display to show the second image in the image display area.

[0128] For example, the second image displayed in the first input area can be generated by processing the first image using an image generation model. In practical applications, considering factors such as device performance and product requirements, the image generation model can be deployed either on the display device or on a server.

[0129] In some examples, the image generation model can be a pre-trained model that can convert input multimodal information such as images, text, audio, and video into a corresponding reference image that meets the requirements of multimodal information. It should be noted that the image generation model in the following embodiments ( Figure 14 The embodiments shown will be described in detail.

[0130] For example, after acquiring the second image, the display device can control the monitor to display the second image in the image display area. Figure 7 As shown, the image display area 730 on the right side of the application interface 700 displays a reference image (such as a second image) generated based on the descriptive image (such as the first image) displayed in the first input area 710.

[0131] In other words, the image generation model can realize the function of image-to-image, transforming user-drawn images into more realistic images that meet the user's needs.

[0132] Figure 8 This is a schematic diagram of another image display method provided in the embodiments of this application, which is described below in conjunction with... Figure 8 The implementation process of step 620 is explained when the image generation model is deployed on both the display device and the server. For example... Figure 8 As shown, step 620 above includes steps 621 and 622 as shown below.

[0133] Step 621: Process the first image based on the image generation model to obtain the second image, and control the display to show the second image in the image display area.

[0134] For example, when an image generation model is deployed on a display device, after the display device acquires a first image, it can process the first image based on the image generation model deployed on it to obtain a second image, and then display the second image in the image display area.

[0135] In some embodiments, step 621 may include: processing the first image using an image generation model based on a preset detection period to obtain a second image, and controlling the display to show the second image in the image display area.

[0136] For example, the preset detection period can be set according to user needs; for instance, the preset detection period can be a preset duration. The display device can input the image drawn by the user within the current detection period into the image generation model for processing every preset detection period to obtain the second image corresponding to the current detection period, and display the second image corresponding to the current detection period in the image display area.

[0137] In some examples, the first image displayed in the first input area may or may not change between two adjacent detection periods. That is, the first images corresponding to two adjacent detection periods may be the same or different.

[0138] For example, when the first input area displays the first image corresponding to the first detection cycle, the display device saves the first image corresponding to the first detection cycle. After entering the second detection cycle following the first detection cycle, if the user continues to draw an image in the first input area, the display device acquires the first image displayed in the first input area corresponding to the second detection cycle and compares it with the first image corresponding to the first detection cycle to determine whether the first image displayed in the first input area has changed.

[0139] For example, the display device can compare the trajectory of the first image corresponding to the second detection cycle with the trajectory of the first image corresponding to the first detection cycle. When the trajectory of the first image corresponding to the second detection cycle completely overlaps with the trajectory of the first image corresponding to the first detection cycle, it indicates that the first image corresponding to the second detection cycle is the same as the first image corresponding to the first detection cycle, that is, the first image displayed in the first input area has not changed in the two detection cycles. When the trajectory of the first image corresponding to the second detection cycle does not completely overlap with the trajectory of the first image corresponding to the first detection cycle, it indicates that the first image corresponding to the second detection cycle is different from the first image corresponding to the first detection cycle, that is, the first image displayed in the first input area has not changed in the two detection cycles.

[0140] Based on the above scheme, the display method provided by this application can process the first image to obtain the second image through the image generation model on the display device based on a preset detection cycle, thereby realizing the periodic synchronous display of the first image drawn by the user in the first input area to the first input area, ensuring that the user can check and confirm the drawn image, improving the portability of user interaction and enhancing the user experience.

[0141] In some embodiments, if it is determined, based on a preset detection period, that the first image corresponding to the current detection period is different from the first image corresponding to the previous detection period, then an image generation model is used to process the first image corresponding to the current detection period to obtain the second image corresponding to the current detection period; and the display is controlled to show the second image corresponding to the current detection period in the image display area.

[0142] In some examples, if the display device determines that the first image corresponding to the current detection period is different from the first image corresponding to the previous detection period, that is, the first image displayed in the first input area has changed, the image generation model can be called to process the first image corresponding to the current detection period to obtain the second image corresponding to the current detection period, and the second image can be displayed in the image display area.

[0143] For example, since the first image corresponding to two adjacent detection cycles changes, the corresponding second image will also change. In order to ensure the consistency between the second image displayed in the image display area and the first image displayed in the first input area, when the first image displayed in the first input area changes, the second image displayed in the second input area needs to be updated accordingly.

[0144] Based on the above scheme, the display device compares the first images of two adjacent detection cycles, and when the first image of the current detection cycle changes from the first image of the previous detection cycle, it processes the first image of the current detection cycle through the image generation model, thereby avoiding repeated processing by the image generation model and improving image generation efficiency.

[0145] In some embodiments, if it is determined, based on a preset detection period, that the first image corresponding to the current detection period is the same as the first image corresponding to the previous detection period, then the second image corresponding to the previous detection period continues to be displayed in the image display area.

[0146] In some examples, if the display device determines that the first image corresponding to the current detection cycle is the same as the first image corresponding to the previous detection cycle, that is, the first image displayed in the first input area has not changed, then the second image corresponding to the previous detection cycle will continue to be displayed in the image display area.

[0147] For example, since the first image corresponding to two adjacent detection cycles does not change, the corresponding second image also does not change. Therefore, there is no need to update the second image displayed in the second input area, nor is there a need for the image generation model to repeatedly process the first image corresponding to the current detection cycle, thereby improving the processing efficiency of the display device.

[0148] Based on the above scheme, the display device compares the first images of two adjacent detection cycles, and maintains the display of the first image of the previous detection cycle when the first images of the two adjacent detection cycles have not changed, thereby avoiding repeated processing of the image generation model and improving image generation efficiency.

[0149] In some embodiments, step 621 may further include: listening to the touch event corresponding to the touch command, processing the first image based on the image generation model to obtain the second image according to the touch event, and controlling the display to display the second image in the image display area.

[0150] In some examples, in addition to displaying the second image in the image display area based on a preset detection period, the display device can also determine the second image to be displayed in the image display area by listening to the touch events corresponding to the touch commands entered by the user in the first input area.

[0151] For example, the touch events corresponding to touch commands include touch press events, touch move events, and touch release events. For instance, when a user triggers a touch event by inputting a touch command in the first input area of ​​a touch device, the type of touch event can be determined first.

[0152] For example, when the touch device moves from never touching the touchscreen corresponding to the first input area to touching the touchscreen corresponding to the first input area, the touch event corresponding to the input touch command is a touch press event. When the touch device continuously touches the touchscreen corresponding to the first input area for a preset time, the touch event corresponding to the input touch command is a touch move event. When the touch device moves from touching the touchscreen corresponding to the first input area to leaving the touchscreen corresponding to the first input area, the touch event corresponding to the input touch command is a touch release event.

[0153] In some examples, when a user is drawing an image in the first input area, touch commands input within a drawing cycle can sequentially trigger touch press, touch move, and touch release events. For instance, when a touch command triggers a touch press event, it indicates that the user has begun drawing an image in the first input area, i.e., the current drawing cycle has begun; when a touch command triggers a touch move event, it indicates that the user is drawing an image in the first input area; and when a touch command triggers a touch release event, it indicates that the user has completed drawing an image in the first input area, i.e., the current drawing cycle has ended.

[0154] It should be noted that touch commands input by the user within a drawing cycle can also trigger touch press and touch release events sequentially. In this case, the user may not have adjusted or updated the first image of the first input area within the current drawing cycle.

[0155] Based on the above solution, the display method provided in this application can obtain a second image by processing the first image through the image generation model on the display device based on the touch event corresponding to the touch command. This enables the image to be displayed synchronously in the first input area when the user completes a stage of drawing in the first input, ensuring that the user can check and confirm the drawn image, improving the convenience of user interaction and enhancing the user experience.

[0156] In some embodiments, while listening for touch press events, touch release events are continuously listened for; in response to listening for touch release events, the first image is processed based on an image generation model to obtain a second image, and the display is controlled to display the second image in the image display area.

[0157] In some examples, when the display device detects a touch press event, it indicates that the current drawing cycle has started. The display device can continue to listen for touch events. When it detects a touch release event, it indicates that the current drawing cycle has ended. In response to the touch release event, the display device can process the first image corresponding to the current drawing cycle displayed in the first input area to obtain a second image, and control the display to display the second image in the image display area.

[0158] In other words, at the end of a drawing cycle, the display device can process the first image corresponding to that drawing cycle through an image generation model, thereby updating the second image in the image display area.

[0159] Based on the above scheme, when the touch press event is triggered, the touch release event is triggered, indicating that the user has completed the drawing of one drawing cycle and stopped drawing. Therefore, the currently drawn image can be processed in a timely manner through the image generation model to ensure the timeliness of image display.

[0160] In some embodiments, step 621 may further include: in response to receiving a confirmation instruction input by the user, processing the first image based on an image generation model to obtain a second image, and controlling the display to show the second image in the image display area.

[0161] In some examples, in addition to displaying the second image in the image display area based on a preset detection period, or displaying the second image in the image display area by listening to the touch event corresponding to the touch command entered by the user in the first input area, the display device can also determine the second image displayed in the image display area by the confirmation command entered by the user.

[0162] For example, when the first image is displayed in the first input area, the user can input a command through the control device, or they can input a confirmation command via voice. The confirmation command is used to indicate that the first image currently displayed in the first input area can be used to generate the corresponding second image. Therefore, after receiving the confirmation command, the display device, in response to the confirmation command, processes the first image currently displayed in the first input area through the image generation model to obtain the second image, and displays the second image in the image display area.

[0163] In some examples, the application interface may include a confirmation control, which the user can use to input confirmation commands via a control device (such as a remote control).

[0164] In some examples, users can enter confirmation commands at any point during the first image drawing process, as needed. (See reference...) Figure 7 When a user draws a bird image in the first input area 710, the user can enter a confirmation command after completing the drawing of the entire bird image, or when completing a part of the bird image. This application embodiment does not limit this.

[0165] For example, when a user completes drawing part of a bird image, they can input a confirmation command. The image generation model then generates a corresponding reference image based on that part of the bird image and displays it in the image display area 730. The user continues drawing the bird image in the first input area. After completing the remaining part of the image, the user inputs a confirmation command again. The image generation model then generates a corresponding reference image based on the complete bird image currently displayed in the first input area 710 and displays it in the image display area 730.

[0166] Based on the above solution, the present application provides a display method that can process the first image to obtain the second image based on the user's confirmation command through the image generation model on the display device, thereby realizing the synchronous display of the drawn image to the first input area according to the user's needs and improving the user experience.

[0167] Step 622: Send the first image to the server to receive the second image generated by the server based on the first image, and control the display to show the second image in the image display area.

[0168] For example, when an image generation model is deployed on a server, after determining a first image, the display device can send the first image to the server for processing to obtain a second image. For instance, after receiving the first image from the display device, the server can process the first image using the image generation model deployed on it to obtain the second image. The server then sends the second image to the display device, which controls the monitor to display the second image in the image display area.

[0169] In some embodiments, sending the first image to the server in step 622 above includes: sending the first image to the server based on a preset detection period; or, listening to touch events corresponding to touch commands and sending the first image to the server based on the touch events; or, in response to receiving a confirmation command input by the user, sending the first image to the server.

[0170] In some examples, the process of the display device sending the first image to the server is similar to the process in step 621 above where the display device generates the second image based on the first image. For example, the display device may send the first image to the server based on a preset detection period, or based on a touch instruction corresponding to a detected processing instruction, or based on a confirmation instruction input by the user.

[0171] In some embodiments, when sending a first image to the server based on a preset detection period, if the display device determines that the first image corresponding to the current detection period is different from the first image corresponding to the previous detection period, it indicates that the first image displayed in the first input area has been updated and the second image corresponding to the current detection period needs to be regenerated. In this case, the display device can send the first image corresponding to the current detection period to the server, and the server processes the first image corresponding to the current detection period based on the image generation model to obtain the second image corresponding to the current detection period.

[0172] For example, if the display device determines that the first image corresponding to the current detection cycle is the same as the first image corresponding to the previous detection cycle, it indicates that the first image displayed in the first input area has not been updated from the previous detection cycle to the current detection cycle. Therefore, the second image displayed in the image display area also does not need to be updated, meaning the server does not need to regenerate the second image. In this case, the display device does not need to send the first image corresponding to the current detection cycle to the server, but instead controls the display to continue displaying the second image determined in the previous detection cycle in the image display area.

[0173] In some embodiments, the display device may also continuously listen for touch release events when a touch press event is detected; and in response to a touch release event, send a first image to the server.

[0174] In some examples, when the display device detects a touch lift event, indicating the end of the current drawing cycle, it can send the first image corresponding to the current drawing cycle to the server, and process the first image using an image generation model deployed on the server to generate the corresponding second image.

[0175] For example, when the display device receives a confirmation command from the user, it can also send the first image currently displayed in the first input area to the server in response to the confirmation command, and process the first image through the image generation model deployed on the server to generate a corresponding second image.

[0176] Step 630: In response to receiving the first descriptive text input by the user, control the display to display the first descriptive text in the second input area.

[0177] For example, the first descriptive text is text input by the user or text generated based on the user's input description. The first descriptive text is the user's textual description of the desired image. After determining the first descriptive text, the display device controls the display to show the first descriptive text in the second input area.

[0178] In some examples, the application interface may include text input controls and voice input controls. Users can select the text input control via the control device to input descriptive text; or, users can select the voice input control via the control device to input descriptive speech.

[0179] For example, a user can select the text input control via remote control to bring up the virtual input keyboard, and then input the corresponding text in the second input area 720 to obtain the first descriptive text. Alternatively, the user can select the voice input control to bring up the voice input mode of the display device and input descriptive voice.

[0180] Figure 9 This is a schematic diagram illustrating another application interface provided in an embodiment of this application. For example... Figure 9 As shown in (a), the application interface 700 includes a text input control 910 and a voice input control 920. The user can input descriptive text through the text input control 910, and the descriptive text displayed in the second input area 720 is "Bird, with blue feathers on its back and white feathers on its belly".

[0181] In some embodiments, step 630 includes: in response to receiving a descriptive speech input by a user, processing the descriptive speech to obtain an audio spectrum; acquiring a first descriptive text obtained after processing the audio spectrum based on a speech recognition model; and controlling a display to display the first descriptive text in a second input area.

[0182] For example, when the first descriptive text input by the user is text generated based on the user-input descriptive speech, the display device, after acquiring the user-input descriptive speech, first processes the descriptive speech to obtain an audio spectrum. After acquiring the audio spectrum, a speech recognition model can be used to process the audio spectrum to obtain the first descriptive text. The speech recognition model can be a pre-trained model, and it can be deployed on the display device or on a server.

[0183] In some examples, when the speech recognition model is deployed on a display device, the display device can process the audio spectrum corresponding to the acquired descriptive speech using the speech recognition model deployed on it to obtain the first descriptive text. When the speech recognition model is deployed on a server, after obtaining the audio spectrum corresponding to the descriptive speech, the display device first sends the audio spectrum to the server. The server then processes the audio spectrum using its deployed speech recognition model to obtain the first descriptive text and sends the first descriptive text to the display device. After obtaining the first descriptive text, the display device controls the monitor to display the first descriptive text in the second input area.

[0184] Based on the above solution, the display method provided in this application can convert the user-inputted descriptive speech into descriptive text using a pre-trained speech recognition model, and then display the descriptive text in the second input area, thereby facilitating the subsequent generation of the third image. By converting the descriptive speech into descriptive text, the user can avoid needing to input text in the second input area, improving input efficiency and enhancing the user experience.

[0185] In some embodiments, step 630 above, controlling the display to display the first descriptive text in the second input area, includes: obtaining the number of characters of the third descriptive text currently displayed in the second input area; if the number of characters of the third descriptive text is zero, adding the first descriptive text to the second input area and controlling the display to display the first descriptive text in the second input area; if the number of characters of the third descriptive text is not zero, adding the first descriptive text after the third descriptive text to update the descriptive text in the second input area, and controlling the display to display the updated descriptive text in the second input area.

[0186] In some examples, after receiving the user's input of the first descriptive text, the display device can first determine whether the second input area currently displays descriptive text. For example, the display device can obtain the number of characters in the third descriptive text displayed in the second input area; the number of characters in the third descriptive text may be zero or non-zero.

[0187] For example, when the third descriptive text has zero characters, it indicates that the second input area is currently empty. In this case, the display device can display the first descriptive text entered by the user in the second input area. When the third descriptive text has a non-zero character count, it indicates that the second input area is not currently empty. In this case, the user can add the first descriptive text after the third descriptive text to update the descriptive text displayed in the second input area, and control the display to show the updated descriptive text in the second input area. That is, when the third descriptive text has a non-zero character count, the second input area can display a new descriptive text composed of the first and third descriptive texts.

[0188] Continue to refer to Figure 9 ,like Figure 9 As shown in (b), the second input area 720 currently displays the third descriptive text "bird, blue feathers on the back". In this case, if the user inputs the first descriptive text "white feathers on the belly", the display device obtains the number of characters in the third descriptive text. Since the number of characters in the third descriptive text is not zero, the display device can add the user input first descriptive text after the third descriptive text and display the updated descriptive text "bird, blue feathers on the back; white feathers on the belly" in the second input area 720.

[0189] Based on the above scheme, the display method provided in this application, when receiving the first descriptive text input by the user, obtains the number of characters of the third descriptive text displayed in the second input area. If it is determined that the number of characters of the third descriptive text is not zero, that is, if there is descriptive text in the second input area, the newly input first descriptive text is added after the original third descriptive text, thereby supplementing the descriptive text in the second input area and further improving the convenience of user interaction.

[0190] Step 640: Process the first descriptive text and the first image to obtain a third image that matches the first descriptive text, and control the display to show the third image in the image display area.

[0191] In some examples, after determining the first descriptive text, the display device can process the first descriptive text and the first image using an image generation model to obtain a third image that conforms to the first descriptive text and the first image.

[0192] Figure 10This is a schematic diagram of another image display method provided in the embodiments of this application, which is described below in conjunction with... Figure 10 The implementation process of step 640 is explained when the image generation model is deployed on both the display device and the server. For example... Figure 10 As shown, step 640 includes steps 641 and 642 as described below. Step 641 describes the implementation process of step 640 when the image generation model is deployed on a display device, and step 642 describes the implementation process of step 640 when the image generation model is deployed on a server.

[0193] Step 641: Process the first descriptive text and the first image based on the image generation model to obtain a third image that matches the first descriptive text, and control the display to show the third image in the image display area.

[0194] For example, when displaying a second image, the image display area, in response to user input of first descriptive text, processes the first descriptive text and the first image based on an image generation model deployed in the display device to obtain a third image. The third image meets the requirements of both the first image displayed in the first input area and the first descriptive text displayed in the second input area.

[0195] In some examples, the third image may be different from or the same as the second image. For example, when the first descriptive text has zero characters, the third image is the same as the second image; or, when the content described by the first descriptive text has already been described in the first image, the third image may also be the same as the second image. This application does not limit this aspect.

[0196] In some examples, after a third image is generated based on an image generation model, the display device can control the display to show the third image in the image display area. For example, when the third image is the same as the second image, the display device can control the display to maintain the display of the second image in the image display area; when the third image is different from the second image, the display device can control the display to exit the display of the second image in the image display area and display the third image.

[0197] In some embodiments, after step 641 above, the method further includes: in response to receiving a modification instruction for the first image input by a user in the first input area, determining a fourth image; if the fourth image is different from the first image, processing the first descriptive text and the fourth image based on an image generation model to obtain a fifth image that conforms to the first descriptive text, and controlling the display to display the fifth image in the image display area.

[0198] In some examples, the user can modify the first image displayed in the first input area. After the first image is modified, the first input area can display the modified image (such as the fourth image).

[0199] For example, when a first image is displayed in the first input area, the user can move their limbs (such as fingers) or a touch device (such as a stylus) on the touchscreen corresponding to the first input area to input modification commands. When the display device detects the modification command, it can generate a modified fourth image based on the first image and the modification command. For instance, the display device can generate the fourth image based on the movement trajectory corresponding to the first image and the modification command.

[0200] In some examples, after determining the fourth image, the display device can compare the fourth image with the first image previously displayed in the first input area to determine whether the fourth image is the same as the first image.

[0201] For example, if the first image is the same as the fourth image, it indicates that the user has not modified the first image, or the user's modification operation on the first image was unsuccessful. In this case, the display device controls the display to maintain the display of the currently displayed second image in the image display area.

[0202] For example, if the first image and the fourth image are different, it indicates that the user has modified the first image. In this case, the display device needs to reprocess the fourth image and the first descriptive text using the image generation model to generate a fifth image that matches the fourth image and the first descriptive text. After generating the fifth image, the display device controls the monitor to remove the display of the second image from the image display area and display the fifth image.

[0203] Based on the above scheme, this application can support users to modify the first image drawn in the first input area, and after modification, reprocess the modified fourth image and the first descriptive text displayed in the second input area through an image generation model to ensure that a fifth image that conforms to the modified fourth image and the first descriptive text is generated.

[0204] In some embodiments, after step 641 above, the method further includes: in response to receiving second descriptive text input by a user, controlling the display to display the second descriptive text in a second input area; processing the second descriptive text and the first image based on an image generation model to obtain a sixth image that conforms to the second descriptive text, and controlling the display to display the sixth image in an image display area.

[0205] In some examples, in addition to modifying the first image displayed in the first input area, users can also modify the first descriptive text displayed in the second input area as needed. For example, a user can re-enter second descriptive text in the second input area. This second descriptive text is either the text entered by the user or text generated based on the user's descriptive speech, and it differs from the first descriptive text.

[0206] For example, a user can first delete the first description text displayed in the second input area, and after deleting the first description text, enter the second description text and control the second input area to display the second description text.

[0207] After obtaining the second descriptive text, the display device can process the second descriptive text and the first image through an image generation model to obtain a sixth image that matches the first image and the second descriptive text, and control the display to display the sixth image in the image display area.

[0208] For example, a user can modify both the first image displayed in the first input area and the first descriptive text displayed in the second input area.

[0209] In some examples, the user modifies the first image displayed in the first input area to obtain a fourth image. A timer starts when the fourth image is displayed in the first input area. Before the timer reaches a preset time threshold, if the user re-enters second descriptive text in the second input area, the display device can process the fourth image and the second descriptive text using an image generation model to obtain an image that matches both the fourth image and the second descriptive text, and then display that image in the image display area. The fourth image differs from the first image, and the second descriptive text differs from the first descriptive text. The preset time threshold can be adjusted according to user needs, and this application embodiment does not limit this.

[0210] In other examples, after receiving a user's instruction to modify the first image in the first input area, the display device can display a modified fourth image in the first input area. It then processes the fourth image and the first descriptive text using an image generation model to obtain a modified image, which is then displayed in the image display area. During this process, if the user re-enters the second descriptive text in the second input area, the display device processes the fourth image and the second descriptive text using the image generation model to obtain an image that matches both the fourth image and the second descriptive text, and displays this image in the image display area.

[0211] It should be noted that the display device can process the modifications made by the user to the first image in the first input area and the first description text in the second input area in sequence, or it can process the modifications made to the first image and the first description text within the preset time threshold simultaneously, based on the preset time threshold. This application embodiment does not limit this.

[0212] Based on the above scheme, this application can support users to modify the first descriptive text displayed in the second input area, and after modification, reprocess the first image displayed in the first input area and the modified second descriptive text displayed in the second input area through an image generation model to ensure that a sixth image that conforms to the first image and the modified second descriptive text is generated.

[0213] The display method provided in this application embodiment utilizes an image generation model deployed on the display device. This model processes a first image input by the user in the first input area and / or a first descriptive text input in the second input area to generate a third image that meets the user's needs, which is then displayed in the image display area. Therefore, this application embodiment provides a function capable of generating corresponding images based on multimodal information, improving the accuracy of image generation by the display device and enhancing the user experience. Furthermore, deploying the image generation model on the display device further improves the security and efficiency of the third image generation process. Additionally, since the first input area, second input area, and image display area on the application interface can display different content at different stages, the convenience of user interaction is further enhanced.

[0214] Step 642: Receive a third image generated by the server based on the first description text and the first image, and control the display to show the third image in the image display area.

[0215] In some embodiments, when the image generation model is deployed on a server, after step 630 above, the method further includes sending a first image and a first descriptive text to the server.

[0216] For example, after receiving the first descriptive text, the display device can send the first image and the first descriptive text to the server. After receiving the first image and the first descriptive text, the server processes the first image and the first descriptive text using an image generation model deployed on it to generate a third image, and displays the third image in the image display area.

[0217] In some examples, when the first descriptive text is text entered by the user, the display device can directly send the first descriptive text to the server. When the first descriptive text is text generated from user-inputted speech, if the speech recognition model is deployed on the server, the display device can send the first image and the speech description to the server, where the speech recognition model converts the speech description into the first descriptive text. If the speech recognition model is not on the display device, the display device converts the speech description into the corresponding first descriptive text using the speech recognition model, and then sends the first image and the first descriptive text to the server.

[0218] For example, after receiving a first image and a first descriptive text, the server processes the first image and the first descriptive text using an image generation model deployed on it to obtain a third image that matches the first image and the first descriptive text, and then sends the third image to the display device. After receiving the second image, the display device can display the third image in the image display area.

[0219] For example, if the image generation model is deployed on a server, when a user inputs a modification instruction for the first image in the first input area, the display device determines the modified fourth image. If the fourth image is different from the first image currently displayed in the first input area, the display device sends the first descriptive text and the fourth image to the server. The server processes the first descriptive text and the fourth image using the image generation model deployed on it, generates a fifth image that conforms to the first descriptive text and the fourth image, and sends the fifth image to the display device. The display device receives the fifth image and controls the display to show the fifth image in the image display area.

[0220] Accordingly, if the image generation model is deployed on the server, when the display device receives the second descriptive text input by the user, it can send the second descriptive text and the first image to the server. The server processes the second descriptive text and the first image through the image generation model deployed on it, generates a sixth image that matches the second descriptive text and the first image, and sends the sixth image to the display device. The display device receives the sixth image and controls the display to show the sixth image in the image display area.

[0221] It should be noted that when the image generation model is deployed on the server, the images displayed in the image display area are generated by the server and sent to the display device. The image display process in the image display area is similar to that in the above embodiment where the image generation model is deployed on the display device. To avoid repetition, it will not be described again here.

[0222] The display method provided in this application, after obtaining a first image drawn by the user in the first input area and / or a first descriptive text entered in the second input area, can send the first image and / or the first descriptive text to a server. The server's image generation model then processes the first image and / or the first descriptive text to generate a third image that meets the user's needs. The generated third image is then sent to a display device so that the display device can display the third image in the image display area. Therefore, this application provides a function that can generate corresponding images based on multimodal information, improves the accuracy of image generation by the display device, and enhances the user experience. Furthermore, deploying the image generation model on a server provides higher processing power and more powerful computing resources, making it easier to manage and expand. Additionally, since the first input area, second input area, and image display area on the application interface can display different content at different stages, the convenience of user interaction is further improved.

[0223] Figure 11 A schematic diagram illustrating another image display method provided in an embodiment of this application, as shown below. Figure 11 As shown, the image display method includes steps 1110 to 1140 as shown below.

[0224] Step 1110: Obtain the seventh image, extract the key feature points of the seventh image to obtain the eighth image, and control the display to show the eighth image in the first input area.

[0225] For example, the seventh image is an image captured by an image acquisition device and / or an image uploaded by a mobile terminal. When the display device does not support touch operation or the user cannot draw the first image in the first input area, the display device may also use the image captured by the image acquisition device, or the image uploaded by the user via the mobile terminal, as the seventh image and display the seventh image in the first input area. When the display device supports touch functionality, if the user does not want to draw the first image in the first input area, they may also capture the seventh image by the image acquisition device or upload the seventh image via the mobile terminal.

[0226] In some examples, the display's application interface includes image acquisition controls, which may include camera capture controls and barcode scanning controls. Users can provide a seventh image to the display device through either the camera capture control or the barcode scanning control. For example, continue referring to... Figure 7 The interface on the right side of the second input area 720 includes a camera interaction control 740 and a QR code scanning interaction control 750.

[0227] In some embodiments, obtaining the seventh image in step 1110 includes: in response to receiving an image acquisition instruction input by a shooting interaction control on the application interface, receiving an image acquired by an image acquisition device and displaying the image acquired by the image acquisition device in a first input area; and in response to a shooting instruction input by the user, determining the captured image as the seventh image.

[0228] In some examples, the display's application interface includes a camera control that allows the user to input image input commands. For instance, the user can input image input commands via a control device (such as a remote control), or the user can input image input commands via voice. This application embodiment does not limit this approach.

[0229] For example, an image acquisition command is used to invoke the image acquisition device in the display device. Upon receiving the command, the controller activates the image acquisition device to capture images of the current environment. During image acquisition, the first input area displays the captured image of the current environment, allowing the user to view and confirm the capture. When the user inputs a capture command (confirmation of capture), the image acquisition device acquires the image and sends it to the controller. The controller designates this image as the seventh image and displays it on the first input area.

[0230] For example, such as Figure 7 As shown, when the user moves the focus to the shooting interaction control 740 using the remote control and presses the remote control to select the shooting interaction control, they can input an image acquisition command. The image acquisition command activates the image acquisition device on the display device, enters the shooting mode, and displays the captured image in the first input area.

[0231] In some embodiments, obtaining the seventh image in step 1110 above further includes: in response to receiving an image acquisition instruction input to the QR code scanning interaction control of the application interface, controlling the display to display an identification code in the first input area; and receiving the seventh image uploaded by the mobile terminal through scanning the identification code.

[0232] In some examples, the application interface of the display device may also include a QR code scanning control. The user can input an image acquisition command into the QR code scanning control to retrieve an identification code, which is used to upload the image after scanning by the terminal device. In response to the image acquisition command, the display device controls the screen to display the identification code in the first input area. The user scans the identification code using a mobile terminal (such as a mobile phone), thereby sending the desired image (such as a seventh image) to the display device. The display device receives the seventh image and displays it in the first input area.

[0233] In some examples, after acquiring the seventh image, the first input area can exit the display of the identification code and display the seventh image. The identification code can be a QR code, barcode, RFID tag, or other form; this application embodiment does not limit this. Alternatively, the display device may also exit the display of the identification code in the first input area after detecting that the user has completed scanning the identification code; this application embodiment does not limit this as well.

[0234] For example, such as Figure 7 As shown, when a user moves the focus to the QR code scanning control 750 using the remote control and presses the remote to select the control, they can input an image acquisition command. The image acquisition command can retrieve an identification code, which is then displayed in the first input area. The user scans this identification code using their mobile terminal to upload the seventh image.

[0235] In some embodiments, receiving the seventh image uploaded by the mobile terminal via scanning the identification code includes: obtaining a URL address sent by the server; generating an image receiving request based on the URL address; sending the image receiving request to the server so that the server responds to the image receiving request by sending the seventh image; and receiving the seventh image sent by the server.

[0236] In some examples, the URL address is the storage address of the seventh image that the terminal device sends to the server after scanning the identification code.

[0237] For example, after scanning an identification code, a mobile terminal can send the image to be uploaded (such as the seventh image) to the server. The server stores the seventh image, generates a URL address based on the storage location of the seventh image, and sends the URL address to the display device. In response to the image acquisition command, the display device generates an image receiving request based on the URL address and sends the image receiving request to the server. The server sends the seventh image to the display device based on the URL address in the image receiving request. The display device receives the seventh image and displays it in the first input area.

[0238] In some embodiments, step 1110 includes: acquiring a seventh image and controlling the display to show the seventh image in the first input area; extracting key feature points from the seventh image to obtain an eighth image and controlling the display to show the eighth image in the first input area.

[0239] In some examples, after acquiring the seventh image, the display device controls the display to show the seventh image in the first input area. The display device extracts key feature points from the seventh image displayed in the first input area, uses the extracted feature point map as the eighth image, and controls the display to show the eighth image in the first input area.

[0240] In some embodiments, step 1110 above, which involves extracting key feature points from the seventh image to obtain the eighth image, includes: determining the image type of the seventh image; if the image type of the seventh image is a first type, then extracting key feature points from the seventh image using a skeleton feature extraction algorithm to obtain the eighth image; if the image type of the seventh image is a second type, then extracting key feature points from the seventh image using an edge detection algorithm and a skeleton feature extraction algorithm to obtain the eighth image.

[0241] In some examples, the image type includes a first type and a second type, where the first type image is a line drawing and the second type image is a non-line drawing. Since the source of the seventh image can be a photograph or an upload by a user via a mobile terminal, the seventh image may not contain obvious line features. In order for the image generation model to generate a more accurate image based on the seventh image, the display device can perform image processing and feature extraction operations on the acquired seventh image.

[0242] For example, when the seventh image is a line drawing, the display device can extract key feature points from the seventh image using a skeleton feature extraction algorithm to transform it into the eighth image. When the seventh image is not a line drawing, the display device first transforms the seventh image into a boundary image using an edge detection algorithm, and then transforms the boundary image into the eighth image using a skeleton feature extraction algorithm. This eighth image can be a skeleton structure diagram.

[0243] In some examples, the display device performs grayscale conversion and binarization on the seventh image to generate a binary image, performs edge detection on the binary image using an edge detection algorithm to generate a boundary image, extracts morphological features from the boundary image to generate a skeleton structure map, and displays the generated skeleton structure map as the eighth image in the first input area.

[0244] In some examples, after the eighth image is determined, the first input area can exit the display of the seventh image and display the eighth image; alternatively, the first input area can display both the seventh and eighth images simultaneously; or, the first input area can display either the seventh or the eighth image based on the user's selection.

[0245] In some embodiments, in response to selecting the first option, the display is controlled to display a seventh image in the first input area; in response to switching the focus to the second option, the display is controlled to display an eighth image in the first input area.

[0246] In some examples, the application interface displayed on the monitor may include at least one option control.

[0247] For example, when the application interface includes an option control, such as a first option control, the first option control can simultaneously correspond to the first option and the second option. When the first input area displays the seventh image, it indicates that the first option control currently corresponds to the first option. In this case, when the user moves the focus back to the first option control using the remote control and selects the first option control using the remote control, the first option control corresponds to the second option. In this case, the display can exit the display of the seventh image in the first input area and display the eighth image.

[0248] For example, when the application interface includes two option controls, such as a first option control and a second option control, the first option control can correspond to the first option, and the second option control can correspond to the second option. When the user moves the focus to the first option control using the remote control, in response to the selection of the first option control corresponding to the first option, the seventh image is displayed in the second input area; when the user moves the focus to the second option control using the remote control, in response to the selection of the second option control corresponding to the second option, the eighth image is displayed in the second input area.

[0249] Figure 12 This is a schematic diagram illustrating another application interface provided in an embodiment of this application. For example... Figure 12 As shown in (a), the first input area 710 includes two option controls, a first option control 711 and a second option control 712. When the user selects the first option control 711 via the remote control, the display area of ​​the first input area 710 displays the seventh image. When the user selects the second option control 712 via the remote control, the display area of ​​the first input area 710 displays the eighth image.

[0250] For example, the first input area can also display both the seventh and eighth images simultaneously. For instance, the first input area can be divided into two display areas, where one display area (referred to as the first display area) can display the seventh image, and the other display area (referred to as the second display area) can display the eighth image. Users can view both the original image (such as the seventh image) that they have captured or uploaded, and the skeleton image generated from the original image (such as the eighth image) through the first input area.

[0251] In some examples, the first display area and the second display area may be the same size or different. For example, when the first display area and the second display area are different sizes, the second display area may be larger than the first display area, such as the first display area being a portion of the second display area. This application does not limit the size or distribution of the first and second display areas.

[0252] For example, such as Figure 12 As shown in (b) above, the display area of ​​the first input area 710 includes a display area for displaying the seventh image and a display area for displaying the eighth image. The display area for displaying the seventh image is a portion of the display area for displaying the eighth image.

[0253] Step 1120: Process the eighth image based on the image generation model to obtain the ninth image, and control the display to show the ninth image in the image display area.

[0254] In some examples, after determining the eighth image, the display device can process the eighth image through image generation model processing to obtain the ninth image, and display the ninth image in the image display area.

[0255] It should be noted that step 1120 is similar to step 621 in the above embodiment. In step 1120, the eighth image is equivalent to the first image in step 621, and the ninth image in step 1120 is equivalent to the second image in step 621. The image generation model's processing of the image and the process of displaying the image in the image display area have been described in detail in step 621 above. To avoid repetition, they will not be repeated here.

[0256] Step 1130: In response to receiving the first descriptive text input by the user, control the display to display the first descriptive text in the second input area.

[0257] It should be noted that step 1130 is similar to step 630 in the above embodiment, and will not be repeated here to avoid repetition.

[0258] Step 1140: Process the first descriptive text and the eighth image based on the image generation model to obtain the tenth image that matches the first descriptive text, and control the display to show the tenth image in the image display area.

[0259] In some examples, after obtaining the first descriptive text, the display device can process the first descriptive text and the eighth image based on an image generation model to obtain a tenth image that matches the first descriptive text and the eighth image, and then display the tenth image in the image display area.

[0260] It should be noted that step 1140 is similar to step 640 in the above embodiment, and will not be repeated here to avoid repetition.

[0261] In some embodiments, after step 1140, the method further includes: in response to receiving a modification instruction for the eighth image input by a user in the first input area, determining an eleventh image; if the eleventh image is different from the eighth image, processing the first descriptive text and the eleventh image based on an image generation model to obtain a twelfth image that conforms to the first descriptive text, and controlling the display to display the twelfth image in the image display area.

[0262] In some examples, when the display device supports touch functionality and the first input area displays the eighth image, the user can also modify the eighth image displayed in the first input area. For example, the user can use their limbs (such as fingers) or a touch device (such as a stylus) to move in the first input area to input a modification command. When the display device detects the modification command, it can generate a modified eleventh image based on the eighth image and the modification command. The display device determines whether the eleventh image is the same as the eighth image. If it determines that the eleventh image is different from the eighth image, it processes the eleventh image and the first descriptive text displayed in the second input area using an image generation model to obtain the twelfth image. The twelfth image conforms to the requirements of the eleventh image and the first descriptive text.

[0263] It should be noted that the user's modification of the eighth image is similar to the modification process of the first image in the above embodiments, and will not be repeated here to avoid repetition.

[0264] In some embodiments, based on a preset detection period, an image generation model is used to process the first descriptive text and the eleventh image to obtain the twelfth image, and the display is controlled to display the twelfth image in the image display area; or, a touch event corresponding to a touch command is monitored, and based on the touch event, the first descriptive text and the eleventh image are processed using the image generation model to obtain the twelfth image, and the display is controlled to display the twelfth image in the image display area; or, in response to receiving a confirmation command input by the user, the first descriptive text and the eleventh image are processed using the image generation model to obtain the twelfth image, and the display is controlled to display the twelfth image in the image display area.

[0265] It should be noted that the display device can also generate the twelfth image based on a preset detection cycle, the touch event corresponding to the touch command, or the confirmation command input by the user. In this embodiment, the process of generating the twelfth image based on the first description text and the eleventh image is similar to the process of generating the second image based on the first image in the above embodiment. To avoid repetition, it will not be described again here.

[0266] In some embodiments, after step 1140, the method further includes: in response to receiving second descriptive text input by a user, controlling the display to display the second descriptive text in a second input area; processing the second descriptive text and the eighth image based on an image generation model to obtain a thirteenth image that conforms to the second descriptive text, and controlling the display to display the thirteenth image in an image display area.

[0267] In some examples, the display device may also update the first descriptive text displayed in the second input area to obtain a second descriptive text, and then process the second descriptive text and the eighth image displayed in the first input area using an image generation model to obtain a thirteenth image, which is then displayed in the image display area. The thirteenth image conforms to the requirements of both the second descriptive text and the eighth image.

[0268] The display method provided in this application embodiment deploys an image generation model. This model can process a seventh image input by the user in the first input area via shooting or uploading, and / or a first descriptive text input in the second input area, thereby generating an image that meets the user's needs and displaying it in the image display area. Therefore, this application embodiment provides a function that can generate corresponding images based on multimodal information, and improves the accuracy of image generation by the display device, thus enhancing the user experience. Furthermore, deploying the image generation model on the display device can improve the security and efficiency of the third image generation process. Additionally, since the first input area on the application interface can display images shot and uploaded by the user, the convenience of user interaction is further improved.

[0269] Figure 13 This is a schematic diagram illustrating another image display method provided in an embodiment of this application. It should be noted that... Figure 13 The image display method shown is the same as Figure 11 The difference in the image display methods shown is that, Figure 12 In the image display method shown, the image generation model is deployed on the display device. Figure 13 In the image display method shown, the image generation model is deployed on a server. For example... Figure 13 As shown, the method includes steps 1310 to 1340 as shown below.

[0270] Step 1310: Obtain the seventh image, extract the key feature points of the seventh image to obtain the eighth image, and control the display to show the eighth image in the first input area.

[0271] The seventh image is an image captured by an image acquisition device and / or an image uploaded by a mobile terminal.

[0272] It should be noted that step 1310 is similar to step 1110 in the above embodiments, and will not be repeated here to avoid repetition.

[0273] Step 1320: Send the eighth image to the server to receive the ninth image generated by the server based on the eighth image, and control the display to show the ninth image in the image display area.

[0274] In some examples, step 1320 differs from step 1120 in the above embodiments in that, after determining the eighth image, the display device needs to send the eighth image to the server, and process the eighth image using the image generation model on the server to obtain the ninth image. The server then sends the generated ninth image to the display device, which displays the ninth image in the image display area.

[0275] Step 1330: In response to receiving the first descriptive text input by the user, control the display to show the first descriptive text in the second input area, and send the eighth image and the first descriptive text to the server.

[0276] The first descriptive text is either text entered by the user or text generated based on the user's input description.

[0277] In some examples, step 1330 differs from step 1130 in the above embodiments in that, after determining the first descriptive text, the display device sends the first descriptive text and the eighth image to the server for processing, rather than having the display device process them.

[0278] Step 1340: Receive the tenth image generated by the server based on the first description text and the eighth image, and control the display to show the tenth image in the image display area.

[0279] In some examples, step 1340 differs from step 1140 in the above embodiments in that, after receiving the first descriptive text, the server processes the first descriptive text and the eighth image using an image generation model deployed on the server to obtain the tenth image. The server then sends the tenth image to the display device, which controls the display to show the tenth image in the image display area.

[0280] The display method provided in this application, after acquiring a seventh image entered by the user in the first input area through shooting or uploading, and / or a first descriptive text entered in the second input area, can send the first image and / or the first descriptive text to a server. The image generation model on the server then processes the seventh image and / or the first descriptive text to generate an image that meets the user's requirements. The generated image is then sent to a display device so that the display device can display the image in the image display area. Therefore, this application provides a function that can generate corresponding images based on multimodal information, improves the accuracy of image generation by the display device, and enhances the user experience. Furthermore, deploying the image generation model on a server provides higher processing power and more powerful computing resources, making it easier to manage and expand. Additionally, since the first input area on the application interface can display images captured and uploaded by the user, the convenience of user interaction is further improved.

[0281] The above embodiments provide a detailed description of the display device and the image display method. The image displayed in the image display area of ​​the display device is generated by an image generation model based on the image information displayed in the first input area and the text information displayed in the second input area. The process of the image generation model generating an image based on multimodal information input by the user in the above embodiments is described below.

[0282] In some examples, with the rapid development of image generation technology, artificial intelligence and machine learning techniques can achieve automatic image generation. However, related technologies still focus on image generation based on single-modal information. For example, generating images based on text information or generating corresponding images based on image information. As users' personalized needs increase, this single-modal information-based image generation can no longer meet their requirements.

[0283] Based on this, this application provides an image generation method that processes multimodal information input by the user through a trained image generation model to generate a target image that meets the requirements (i.e., the image displayed in the image display area in the above embodiment).

[0284] The following is combined Figure 14 The process of generating images using an image generation model is described below. It should be noted that since the image generation model can be deployed on a server or a display device, the image generation method provided in the following embodiments can be implemented by either a display device or a server; this application does not limit this approach.

[0285] Figure 14 This is a schematic diagram illustrating an image generation method provided in an embodiment of this application. Figure 14 As shown, the method includes steps 1410 to 1440 as shown below.

[0286] Step 1410: Obtain multimodal information input by the user.

[0287] For example, modal information can be information of different forms (or types) input by the user, and modal information can be audio information, text information, image information, or video information. Specifically, the multimodal information obtained in this embodiment may include at least two modal information types selected from audio information, text information, image information, and video information.

[0288] In some examples, multimodal information may include any two modalities of audio, text, image, and video information; or, multimodal information may include any three modalities of audio, text, image, and video information; or, multimodal information may include audio, text, image, and video information. This application does not limit this. For example, in the above embodiments, the multimodal information received by the display device includes image information input from the first input area and text information input from the second input area.

[0289] It should be noted that modal information may include not only audio information, text information, image information, or video information, but may also include more modal information such as historical style status information and historical viewed content information. This application embodiment does not limit this.

[0290] In some examples, users can input multimodal information through a display device. The display device's screen shows an application interface generated from the image, which may include multiple input areas where the user can input different modal information.

[0291] For example, refer to the above Figure 7 Users can input first modal information in the first input area 710 of the application interface 700 and input second modal information in the second input area 720. The first modal information can be image information (such as the first image in the above embodiment) and the second modal information can be text information (such as the first descriptive text in the above embodiment).

[0292] In other examples, the application interface displayed on the monitor may include a third input area in addition to the first and second input areas. The user can input third modal information through this third input area, which can be audio or video information. It should be noted that the application interface of the monitor may also include a greater number of input areas (such as a fourth input area), but this embodiment does not limit this. That is, the user can input corresponding modal information in multiple input areas separately, or at least in any two of the multiple input areas.

[0293] In some examples, users can input audio information to the display device via a microphone on a remote control, input text information via the remote control, or input image or video information via a camera on the display device. Alternatively, users can input audio information via a sound event sensor, or modal information via an Internet of Things (IoT) device. This application does not specifically limit the input method of modal information.

[0294] For example, when the image generation model is deployed on a display device, the display device obtains multimodal information based on the modal information input by the user in each input area; when the image generation model is deployed on a server, the display device sends the modal information input by the user in each input area to the server so that the server can obtain the multimodal information. This application embodiment uses the deployment of the image generation model on a display device as an example for illustrative purposes.

[0295] Figure 15 This is a schematic diagram of an image generation model provided in an embodiment of this application. Figure 15 As shown, the image generation model 1500 includes an encoding network 1510, a fusion network 1520, and a multi-stage network 1530. After obtaining the multimodal information input by the user, the multimodal information can be input into the encoding network 1510, and after being processed by the encoding network 1510, the fusion network 1520, and the multi-stage network 1530 in sequence, a target image that meets the requirements is generated.

[0296] The processing procedures of the coding network 1510, fusion network 1520, and multi-stage network 1530 in the image generation model 1500 are explained below.

[0297] Step 1420: Process the modal information based on the encoding network in the image generation model to obtain the feature vector corresponding to each modal information.

[0298] For example, the coding network may include multiple coding sub-networks, each of which can process its corresponding modal information separately.

[0299] For example, such as Figure 15 As shown, when the modal information is audio information, text information, image information, or video information, the encoding network 1510 may include an audio encoding sub-network 1511, a text encoding sub-network 1512, an image encoding sub-network 1513, and a video encoding sub-network 1514.

[0300] It should be noted that the encoding network may include a greater number of encoding sub-networks. This application embodiment does not limit the number or type of encoding sub-networks in the encoding network. This application embodiment uses examples where the encoding network may include audio encoding sub-networks, text encoding sub-networks, image encoding sub-networks, and video encoding sub-networks for illustrative purposes.

[0301] In some embodiments, step 1420 includes: inputting audio information into an audio coding subnetwork in an encoding network to generate an audio feature vector corresponding to the audio information; inputting text information into a text coding subnetwork in an encoding network to generate a text feature vector corresponding to the text information; inputting image information into an image coding subnetwork in an encoding network to generate an image feature vector corresponding to the image information; and inputting video information into a video coding subnetwork in an encoding network to generate a video feature vector corresponding to the video information.

[0302] In some examples, the audio coding subnetwork can encode audio information to obtain the corresponding audio feature vector. The text coding subnetwork can encode text information to obtain the corresponding text feature vector. The image coding subnetwork can encode image information to obtain the corresponding image feature vector. The video coding subnetwork can encode video information to obtain the corresponding video feature vector.

[0303] For example, when the user inputs multimodal information including image and text information, the text information can be input into the text coding sub-network for processing, and the image information into the image coding sub-network for processing. In other words, multiple coding sub-networks in the coding network can determine whether to perform coding processing based on the input multimodal information; not all of the multiple coding sub-networks necessarily need to perform coding processing.

[0304] In some examples, the structures of the coding subnetworks can be the same or different. These coding subnetworks can be pre-trained using training data.

[0305] For example, the text encoding subnetwork can convert text information into fixed-length text feature vectors based on a transformer model.

[0306] In some examples, the text encoding subnetwork can transform text information into a fixed-length text feature vector through steps such as word segmentation, vocabulary mapping, word embedding, and encoding. For instance, the text encoding subnetwork can transform input text information into tokens of length 77.

[0307] For example, when text information is input into the text encoding sub-network, it first segments the text information into a series of tokens, which can be words, phrases, or other linguistic units. Then, based on a predefined vocabulary, each token is mapped to a unique ID; the vocabulary is constructed during training based on a large amount of text data, and each word or token can correspond to an index ID. Next, each token is converted into a corresponding word embedding vector, which is obtained by looking up the corresponding row in the embedding matrix. Since the Transformer does not have built-in order information, a positional encoding needs to be added to each token to represent its position in the sentence. The positional encoding can be a predefined fixed vector or trainable parameters, and these positional encoding vectors are added to the word embedding vector of each token. Finally, the text encoding sub-network uses a model based on Transformer architecture to process these embedding vectors, capturing the contextual relationships in the text information through mechanisms such as multi-head self-attention, and generating a series of text feature vectors.

[0308] For example, the image coding subnetwork uses a Convolutional Neural Network (CNN) or a Vision Transformer to transform image information into a fixed-length image feature vector. For instance, the image coding subnetwork can convert input image information into tokens (i.e., image feature vectors) of length 196. Specifically, the CNN extracts image features through convolutional operations and flattens the feature maps into one-dimensional vectors as tokens; the Vision Transformer divides the image into small blocks and transforms each block into tokens containing features and contextual information through linear embedding and the Transformer coding subnetwork.

[0309] It should be noted that the video coding subnetwork and the image coding subnetwork have similar structures and processing procedures. The difference lies in the fact that the video information processing process needs to consider the spatiotemporal correlation between frames. To avoid repetition, this application uses the image coding subnetwork as an example for illustration.

[0310] In some examples, the image coding subnetwork can use a CNN to extract image features from image information and transform them into a fixed-length image feature vector through processes such as convolution, feature extraction, and token transformation.

[0311] For example, when image information is input into the image encoding subnetwork, the CNN performs convolution operations on the input image using convolutional layers based on convolutional kernels to generate feature maps. Different convolutional kernels can extract different features, such as edges, textures, and shapes. As the convolutional layers deepen, the CNN can capture higher-level abstract features and construct complex feature representations by stacking convolutional layers, pooling layers (such as max pooling), and fully connected layers. Finally, the output of the CNN (such as the feature map of the last convolutional layer) is usually flattened into a one-dimensional vector. This one-dimensional vector can be regarded as image tokens, where each token contains feature information of a local region in the image.

[0312] In other examples, the image coding subnetwork can also use the Vision Transformer to extract image features from image information and transform them into a fixed-length image feature vector through processes such as image segmentation, linear embedding, positional encoding, and Vision Transformer encoding.

[0313] For example, the Vision Transformer first divides the input image into a series of patches. Each patch is treated as an independent token, containing pixel information of a local region in the image. Each patch is transformed into a fixed-length vector (i.e., the embedding representation of the token) through a linear embedding layer. This process is typically implemented using a convolutional layer where the kernel size is the same as the patch size, and the number of output channels equals the length of the embedding vector. Next, a positional encoding is added to each token. This positional encoding can be fixed or trainable parameters. After linear embedding and positional encoding, all tokens are fed into the Transformer encoding sub-network for processing. The Transformer encoding sub-network captures the contextual relationships between tokens through mechanisms such as multi-head self-attention. Finally, the output representation of each token contains the features and contextual information of the corresponding region in the image. The output of the Vision Transformer is a series of encoded tokens, each containing a feature representation of a local region in the image.

[0314] In some embodiments, the above-mentioned input of audio information into the audio coding sub-network in the coding network to generate an audio feature vector corresponding to the audio information includes: preprocessing the audio information to obtain preprocessed audio information; performing feature extraction on the preprocessed modal information to obtain an initial feature vector; and encoding the initial feature vector to obtain an audio feature vector.

[0315] For example, the audio coding subnetwork can use an autoregressive model (such as WaveNet or Transformer architecture) to process sound waveforms. The audio coding subnetwork learns the intrinsic features of the audio signal through the model and encodes them into low-dimensional vectors. For instance, the audio coding subnetwork can transform input audio information into tokens of length 42 (i.e., audio feature vectors).

[0316] In some examples, the preprocessing of the audio coding subnetwork may include processes such as sampling, quantization, noise reduction, and normalization. For example, the audio information (such as an audio signal) may first be sampled, that is, converted into a discrete time series. Then, noise and interference in the audio signal may be removed by filters, and the audio signal may be normalized to give it a uniform amplitude range, thus obtaining the preprocessed audio information.

[0317] In some examples, the feature extraction process of the audio coding subnetwork extracts meaningful features from the preprocessed audio information. These features can be used in the subsequent tokenization process. For example, Fourier transform, Mel Frequency Cepstral Coefficients (MFCC), and Linear Predictive Coding Coefficients (LPC) can be used to extract useful features from the preprocessed audio information. These extracted features are then combined into a feature vector, known as the initial feature vector. Processing the feature vector can reflect the temporal structure and frequency characteristics of the audio information.

[0318] In some examples, after determining the initial feature vector, the initial feature vector can be transformed into fixed-length tokens. This process can also be called the token generation process, which can further transform the extracted features into a series of discrete and finite tokens.

[0319] For example, through step 1420 above, the multimodal information can be encoded by a network to output fixed-length feature vectors corresponding to each modality, such as... Figure 15 As shown, the fixed-length feature vectors output by the encoding network 1510 can be input into the fusion network for further processing.

[0320] Step 1430: Based on the fusion network in the image generation model, the feature vectors corresponding to each modality information are fused to obtain the fused vector.

[0321] In some examples, fusion networks can also be called unified tokens encoding networks. Fusion networks can provide a common encoding space for different modalities (such as text, images, audio, and video), enabling multimodal information to be processed and analyzed within a unified framework. For example, fusion networks can handle the different characteristics of various modalities, such as the sequential nature of text, the two-dimensional nature of images, and the temporal sequence nature of audio.

[0322] For example, the fusion network provided in this application embodiment possesses multimodal compatibility, feature extraction capabilities, and scalability. First, the fusion network can process data from different modalities and map them to a common space. Second, the fusion network's powerful feature extraction capability can extract meaningful feature representations from the raw data. Simultaneously, the fusion network is easily extensible to adapt to new data types and features. The fusion network provided in this application embodiment can prioritize operational efficiency while ensuring performance, making it suitable for real-time processing and large-scale data analysis scenarios.

[0323] In some embodiments, step 1430 includes: determining a preset length; and based on the preset length, performing fusion processing on the feature vectors corresponding to each modality information through a fusion network to obtain a fusion vector of the preset length.

[0324] In some examples, the preset length can be set according to requirements. This application embodiment does not limit the length of the fusion vector (i.e., the preset length). For example, the preset length can be 256, or it can be 1024. That is to say, after the feature vectors corresponding to each modality information are input into the fusion network, a fusion vector with a length of 256 can be generated after processing by the fusion network.

[0325] For example, when multimodal information includes image information, text information, and audio information, after processing by the encoding network, the encoding network can input text feature vectors of length 77, image feature vectors of length 196, and audio feature information of length 42 into the fusion network. After the three feature vectors of length are fused by the fusion network, a new vector of length 256 (i.e., the fused vector) is generated.

[0326] In some embodiments, a fusion network is used to perform feature selection processing on the feature vectors corresponding to each modality information to obtain selected feature vectors; the selected feature vectors are mapped to a preset feature space to obtain mapped feature vectors; and the mapped feature vectors are fused based on a preset length to obtain a fused vector of a preset length.

[0327] For example, the fusion process includes at least one of the following fusion methods: weighted fusion, feature-level fusion, and decision-level fusion.

[0328] In some examples, firstly, after obtaining the feature vectors (hereinafter also referred to as modal feature vectors) corresponding to each modality, feature selection processing can be performed on each modal feature vector to select the most useful features from each modality feature vector, resulting in selected feature vectors for each modality. Feature selection processing can reduce the dimensionality of the feature space, improving the model's generalization ability and computational efficiency. Secondly, after obtaining the selected feature vectors for each modality, feature mapping is performed on the selected feature vectors for each modality, mapping selected feature vectors of different lengths to a common feature space, resulting in mapped feature vectors. Finally, based on one or more combinations of feature fusion methods such as superposition, weighted averaging, feature-level fusion, or decision-level fusion, the mapped feature vectors can be fused into a single fused vector.

[0329] In some examples, the feature vector fusion process is illustrated using a feature vector in the form of [B, N, C]. Here, B represents the batch size, which is the number of batches during training or inference, N is the feature dimension, and C is the number of channel channels.

[0330] For example, multiple feature vectors can be fused into a single fusion vector using a weighted fusion method. This involves calculating the vector weights of text, audio, and image feature vectors using the correlation coefficient method. Then, a weighted operation is performed on these vector weights to obtain weighted text, audio, and image vectors. Finally, these weighted text, audio, and image vectors are summed to obtain the fusion vector.

[0331] For example, when there are text feature vectors [B1, N1, C1], speech feature vectors [B2, N2, C2], and image feature vectors [B3, N3, C3], multiple feature vectors can be concatenated or weighted concatenation to fuse them into a fused vector. That is:

[0332] [B0,N0,C0]=w1[B1,N1,C1]+w2[B2,N2,C2]+w3[B3,N3,C3];

[0333] Where [B0, N0, C0] are the fusion vectors; w1, w2, w3 are the learnable coefficients.

[0334] For example, multiple feature vectors can be fused based on a cross-attention mechanism to generate a fused vector. This involves defining a query sequence, a key sequence, and a value sequence using text feature vectors, audio feature vectors, and image feature vectors respectively. Then, the similarity between the query sequence and the key sequence is calculated, and the similarity weights are normalized using the Softmax function. Finally, a weighted sum is performed based on the normalized weights to obtain a weighted value sequence. The weighted value sequence is then fused with the query sequence and / or key sequence to obtain the fused vector.

[0335] For example, when there are text feature vectors [B1, N1, C1], speech feature vectors [B2, N2, C2], and image feature vectors [B3, N3, C3], feature vector fusion can be performed based on a cross-attention mechanism, i.e.:

[0336] [B0,N0,C0]=Trans_Attention([B1,N1,C1],[B2,N2,C2],[B3,N3,C3]);

[0337] Where [B0, N0, C0] is the fusion vector.

[0338] Step 1440: Based on multimodal information, the fusion vector is processed through a multi-stage network in the image generation model to obtain the target image corresponding to the multimodal information.

[0339] For example, the target image obtained by processing the fused vector through a multi-stage network meets the requirements of multimodal information. A multi-stage network may include at least one feature sub-network. The number of at least one feature sub-network can be one or more; when there are multiple at least one feature sub-networks, these multiple feature sub-networks are sequentially connected end-to-end.

[0340] For example, such as Figure 15 As shown, the multi-stage network 1530 may include N feature subnetworks, such as the first feature subnetwork, the second feature subnetwork, ..., and the Nth feature subnetwork (i.e. the last feature subnetwork), where N is any integer greater than or equal to 3, and the output of the previous feature subnetwork serves as the input of the next feature subnetwork connected to it.

[0341] It should be noted that the embodiments of this application do not limit the number of feature subnetworks in a multi-stage process. The embodiments of this application use an example of three or more feature subnetworks for illustrative purposes.

[0342] Figure 16 A schematic diagram of another image generation method provided in the embodiments of this application, as shown below. Figure 16 As shown, step 1440 above includes steps 1441 to 1443 as shown below.

[0343] Step 1441: Determine at least one auxiliary feature information corresponding to the multimodal information.

[0344] For example, auxiliary feature information assists the multi-stage network in extracting more accurate features from the fused vector. The multi-stage network can extract the target vector from the fused vector based on at least one auxiliary feature information, facilitating subsequent generation of the target image. This auxiliary feature information can also be called conditional guidance information. Auxiliary feature information is related to multimodal information; when the multimodal information input to the image generation model differs, the corresponding auxiliary feature information may also differ. For example, auxiliary feature information includes feature information corresponding to the guiding auxiliary image, which can include images of different styles, content, structures, etc. For instance, auxiliary feature information can be obtained by extracting image features from the guiding auxiliary image through image encoding.

[0345] In some examples, guiding maps may include human skeleton maps, depth maps, and edge mapping maps. The human skeleton map provides the basic skeletal structure of a person, guiding the poses and movements of the person in the generated image, ensuring that the generated human posture is natural and reasonable. The depth map represents the depth information of the image and can be used to create realistic 3D effects, making the image appear more layered. Especially when generating complex scenes such as landscapes and cityscapes, the depth map helps the model better understand the depth relationships within the scene, thus generating more realistic images. The edge mapping emphasizes the contours and boundaries in the image, helping to maintain the sharpness and detail of the generated image. When generating intricate patterns, text, or objects that need to maintain a specific shape, the edge map ensures that the generated image has clear outlines and accurate boundaries.

[0346] In other examples, the guiding auxiliary diagram may also include color maps, texture maps, normal maps, height maps, reflection maps, lighting maps, and skeleton maps of other animals (such as birds), etc. The embodiments of this application do not limit the guiding auxiliary diagram.

[0347] For example, auxiliary feature information can be determined based on the guidance auxiliary map. For instance, feature extraction can be performed on the guidance auxiliary map to obtain the auxiliary feature information corresponding to each guidance auxiliary map. After determining the auxiliary feature information, multiple auxiliary feature information entries are pre-stored in a guidance feature table. The guidance feature table can be pre-stored on a display device or server; this embodiment does not limit this. Alternatively, the image generation model may also include a generation sub-network to generate auxiliary feature information corresponding to each feature sub-network and input it to the corresponding feature sub-network. This embodiment does not limit the generation and determination of auxiliary feature information.

[0348] In some examples, at least one auxiliary feature can be determined based on the multimodal information. For instance, when the multimodal information involves people, the corresponding auxiliary feature includes the feature information of the human skeleton map; when the multimodal information involves scenery, the corresponding auxiliary feature includes the feature information corresponding to the edge map and the depth map.

[0349] For example, such as Figure 9 As shown in (a), when the multimodal information includes the image information displayed in the first input area (i.e., the image of the bird drawn by the user) and the text information displayed in the second input area (such as the first descriptive text: bird, blue feathers on the back, white feathers on the belly), the corresponding auxiliary feature information can be determined based on the image information and the text information, including at least the feature information corresponding to the bird skeleton structure map, the feature information corresponding to the edge map, the feature information corresponding to the depth map, and the feature information corresponding to the color map.

[0350] In some examples, the number of at least one auxiliary feature in the multimodal information can be one or more. For example, the number of auxiliary features corresponding to the multimodal information may be M, where M is an integer greater than or equal to 1. The number M of auxiliary features corresponding to the multimodal information can be less than the number N of feature subnetworks in the multi-stage network (i.e., M > N), or it can be greater than or equal to the number N of feature subnetworks.

[0351] For example, in multiple feature subnetworks, each feature subnetwork may correspond to auxiliary feature information, or some feature subnetworks may correspond to auxiliary feature information, while others may not. The auxiliary feature information corresponding to each feature subnetwork is different.

[0352] For example, when a feature subnetwork corresponds to multiple auxiliary feature information, the multiple auxiliary feature information can be weighted and then input into the feature subnetwork for processing.

[0353] Step 1442: Process the fusion vector and at least one auxiliary feature information based on at least one feature subnetwork in the multi-stage network to obtain the target vector.

[0354] For example, after determining at least one auxiliary feature information, the at least one auxiliary feature information and the fusion vector determined in step 1443 can be input into at least one feature subnetwork in the multi-stage network to generate a target vector. The target vector can reflect the features of the multimodal information and the features of the guiding auxiliary graph.

[0355] In some examples, after determining the M auxiliary feature information corresponding to the multimodal information, the corresponding auxiliary feature information can be assigned to the N feature network sub-networks of the multi-stage network.

[0356] In some examples, each feature subnetwork includes multiple processing steps, forming a chain structure to execute processing tasks of varying complexity corresponding to the feature subnetwork. The execution steps for each feature subnetwork can be the same or different. For example, when the complexity of the processing task corresponding to the feature subnetwork is high, it will have more steps.

[0357] like Figure 17 As shown, step 1442 above includes steps 14421 to 14422 as shown below. It should be noted that... Figure 17 An example is given using at least one auxiliary feature information, including a first auxiliary feature information, a second auxiliary feature information, and a third auxiliary feature information.

[0358] Step 14421: Process the fusion vector and the first auxiliary feature information based on the first feature subnetwork in at least one feature subnetwork to determine the first output vector.

[0359] In some examples, the first feature subnetwork is the first feature subnetwork in a multi-stage network, i.e. Figure 15 The first feature subnetwork in the multi-stage network 1530. After determining the auxiliary feature information corresponding to the first feature subnetwork as the first auxiliary feature information, the fusion vector output by the fusion network and the first auxiliary feature information can be input into the first feature subnetwork.

[0360] For example, the fusion vector (i.e., the token vector) can be aligned and random noise can be added to obtain the input vector. The input vector and the first auxiliary feature information can then be input into the first feature sub-network to obtain the first input vector. The random noise added to the aligned token vector can be Gaussian noise or other types of noise.

[0361] Step 14422: Process the first output vector and the second auxiliary feature information based on the intermediate feature subnetwork in at least one feature subnetwork to determine the second output vector.

[0362] In some examples, the number of intermediate feature subnetworks can be one or more. For example, an intermediate subnetwork may include the second feature subnetwork among multiple feature subnetworks, or it may include the second and third feature subnetworks among multiple feature subnetworks, or it may include more feature subnetworks. That is, an intermediate feature subnetwork includes at least one feature subnetwork between the first and last feature subnetworks in a multi-stage network. The embodiments of this application do not limit this.

[0363] For example, after obtaining the first output vector from the first feature sub-network, the first output vector and the auxiliary feature information corresponding to the intermediate feature sub-network can be input into the intermediate feature sub-network. After processing by the intermediate sub-network, a second output vector is output. The second output vector is the feature vector output by the last intermediate feature sub-network.

[0364] For example, when the intermediate subnetwork can include the second feature subnetwork up to the (N-1)th feature subnetwork, the second output vector is the vector output by the (N-1)th feature subnetwork.

[0365] In some examples, when the intermediate feature subnetwork includes the second to the (N-1)th feature subnetwork among multiple feature subnetworks, each intermediate feature subnetwork may correspond to an auxiliary feature information (i.e., the second auxiliary feature information); or, some intermediate feature subnetworks may correspond to auxiliary feature information, while other intermediate feature vectors may not have corresponding auxiliary feature information.

[0366] For example, when N is 5, the intermediate subnetwork includes a second feature subnetwork, a third feature subnetwork, and a fourth feature subnetwork. The second feature subnetwork corresponds to one auxiliary feature (such as a second auxiliary feature), while the third and fourth feature subnetworks do not have corresponding auxiliary feature information. Therefore, the second feature subnetwork can process the first output vector and the second auxiliary feature information to output a first intermediate vector; the first intermediate vector is input to the third feature subnetwork, and after processing by the third feature subnetwork, a second intermediate vector is generated; the second intermediate vector is input to the fourth feature subnetwork, and after processing by the fourth feature subnetwork, a second output vector is generated.

[0367] Step 14423: Process the second output vector and the third auxiliary feature information based on the second feature subnetwork in at least one feature subnetwork to determine the target vector.

[0368] In some examples, the second feature subnetwork is the last feature subnetwork in the multi-stage network, i.e., the Nth feature subnetwork. After determining the second output vector of the intermediate feature subnetwork (i.e., the (N-1)th feature subnetwork) and the third auxiliary information corresponding to the Nth feature subnetwork, the second output vector and the third auxiliary feature information are input into the Nth feature subnetwork to obtain the target vector.

[0369] Step 1443: Decode the target vector based on the image decoding subnetwork in the multi-stage network to obtain the target image.

[0370] In some examples, the multi-stage network also includes an image decoding subnetwork (also called an image decoder), which is used to convert the target vector into a target image that conforms to the multimodal information of the user input. For example, image decoding processing may include feature vector decoding, inverse quantization, inverse transform, and pixel reconstruction. The image decoding subnetwork can recover pixel data close to the original image from the target vector.

[0371] Figure 18 This is a schematic diagram of a multi-stage network structure provided in an embodiment of this application.

[0372] like Figure 18 As shown, the multi-stage network 1530 can obtain the input vector Z0 based on the fusion vector and random noise, and then input the input vector Z0 and the first auxiliary feature information into the first feature sub-network. After processing by the first feature sub-network, the first output vector Z1 is obtained; then the first output vector Z1 and the second auxiliary feature information are input into the second feature sub-network to obtain the second output vector Z2.

[0373] Similarly, processing can continue through other feature subnetworks until the Nth feature subnetwork. The output vector from the previous feature subnetwork and the third auxiliary feature information can be input into the Nth feature subnetwork. After processing by the Nth feature subnetwork, the target vector Zn is output. Finally, the target vector Zn is input into the image decoding subnetwork for decoding to obtain the target image.

[0374] It's important to note that each feature subnetwork in the multi-stage network performs different tasks, such as style, content, and structure. The multi-stage network adopts an adapter-based model, adding elements like edge contours and color (i.e., auxiliary feature information) to enable individual and combined applications of multiple tasks. Furthermore, the multi-stage network defines a clear hierarchical relationship, with each layer (e.g., each feature subnetwork) receiving guidance from the upper layer, enabling fine-grained control. Within each feature subnetwork, a chain-like structure allows for generation at different stages, catering to tasks of varying complexity. This chain-like structure is primarily designed for interactive conditional input and supports controllable multi-stage generation. The intermediate control condition output is the implicit feature representation of the previous stage, and the dimensionality of the feature output remains consistent across all stages.

[0375] In some embodiments, the image generation method further includes: acquiring multimodal training information; processing the multimodal training information based on the image generation model to be trained to generate a predicted image; acquiring sample images; using the predicted image as the initial training output information of the image generation model to be trained and the sample image as the supervision information, iterating the image generation model to be trained to obtain an image generation model.

[0376] For example, the image generation model to be trained includes a coding network to be trained, a fusion network to be trained, and a multi-stage network to be trained. The image generation model obtained after iterative training includes the coding network, the fusion network, and the multi-stage network. That is to say, during the training process of the image generation model, it can be trained as a whole, or each network in the image generation model to be trained can be trained separately. This application embodiment does not limit this.

[0377] In some examples, when training the image generation model to be trained using multimodal training information, the loss value of the image generation model to be trained can be determined based on the sample image and the predicted image, and the model parameters of the image generation model to be trained can be adjusted according to the loss value until a preset number of iterations is reached or the image generation model converges, thereby obtaining the trained image generation model (i.e., the image generation model in the above embodiment). The trained image generation model can be used to convert multimodal information into the corresponding target image.

[0378] The image generation method provided in this application, after acquiring multimodal information input by the user, inputs the multimodal information into an image generation model. The image is then processed sequentially through an encoding network, a fusion network, and a multi-stage network within the image generation model to obtain a target image that conforms to the multimodal information. The image generation method provided in this application can process multimodal information to generate images that meet requirements, satisfying users' personalized needs and improving the user experience.

[0379] Figure 19This is a schematic flowchart illustrating an image display method provided in an embodiment of this application. The following is in conjunction with... Figure 19 The image display method provided in the embodiments of this application will be described.

[0380] like Figure 19 As shown, the image display method provided in this application embodiment can be applied to a display device, which may include an application control module, a signal receiving module, and an algorithm module. It should be noted that... Figure 19 This example illustrates the deployment of an image generation model on a display device.

[0381] In some examples, the signal receiving module includes a voice receiving module, a text input module, and an image receiving module. When the display device supports touch functionality, the signal receiving module may also include a touch module. Figure 19 The following is an example illustrating the signal receiving module, which includes a touch module.

[0382] In some examples, the algorithm modules include a speech recognition model, an image preprocessing model, and an image generation model (such as the image generation model 1500 in the above embodiments); wherein, the image generation model includes a text feature extractor (such as the text encoding sub-network 1512 in the encoding network 1510 in the above embodiments), an image feature extractor (such as the image encoding sub-network 1513 in the encoding network 1510 in the above embodiments), a feature fusion layer (such as the fusion network 1520 in the above embodiments), and a generation layer (such as the multi-stage network 1530 in the above embodiments).

[0383] like Figure 19 As shown, when a user opens an application (such as an image generation application) through a display device, the application control module can call various modules in the signal receiving module to perform monitoring. For example, the application control module can call the voice receiving module for voice monitoring, the touch module for touch monitoring, and the text input module for typing monitoring.

[0384] For example, such as Figure 7 As shown, when a user triggers the shooting interaction control 740 or the scanning interaction control 750 in the image generation application, an image can be captured by an image acquisition device or uploaded via a mobile terminal. The image receiving module can detect the image and send it to the algorithm module. The image preprocessing model in the algorithm module performs edge or skeleton extraction processing on the image (such as key point feature extraction processing) and sends the processed image to the application control module. The application control module draws the image in the application interface drawing area (such as...). Figure 7 The first input area 710 in the drawing area displays the image, or the application control module updates the image currently displayed in the drawing area based on the processed image.

[0385] For example, such as Figure 9 As shown in (a), taking the drawing area as the first input area 710 as an example, when the user draws an image in the first input area 710 of the application interface, the touch receiving module can listen to the user's drawing operation and send the drawing information to the application control module. The application control module displays the drawn image in the first input area 710, or updates the image currently displayed in the first input area 710 based on the drawing information. The application control module sends the image signal corresponding to the image displayed in the first input area 710 to the algorithm module. The image feature extractor in the algorithm module extracts features from the image signal to obtain image features, and inputs the image features to the feature fusion layer.

[0386] like Figure 19 As shown, when a user inputs voice through far-field or near-field means, the voice receiving module can listen to the voice signal and transmit it to the algorithm module. The audio feature extractor in the algorithm module extracts features from the voice signal to obtain voice features, and then inputs the voice features to the feature fusion layer.

[0387] For example, the voice receiving module can directly send the monitored voice signal to the speech recognition model in the algorithm module. The speech recognition model converts the voice signal into a text signal and sends the text signal to the application control module. Alternatively, when a user types text on the display device, the text input module can monitor the typed text information and send it to the application control module. After receiving the text information, the application control module displays the text information in the text input area of ​​the application interface (such as the second input area in the above embodiment), or updates the text currently displayed in the text input area based on the text information. The application control module sends the text signal corresponding to the text displayed in the text input area to the algorithm module. The text feature extractor in the algorithm module extracts features from the text signal to obtain text features. These text features are then input to the feature fusion layer.

[0388] like Figure 19 As shown, after receiving image features, speech features, and text features, the feature fusion layer in the algorithm module performs feature fusion processing on the image features, speech features, and text features to obtain fused features. These fused features are then input into the generation layer in the algorithm module. After processing by the generation layer, a generated image (such as the third image in the above embodiment) conforming to each modal information (speech signal, image, and text signal) is obtained. The generation layer feeds back the generated image to the application control module, which controls the generation area of ​​the application interface (such as the image display area in the above embodiment) to display the generated image.

[0389] like Figure 19As shown, when the user closes the application, the application control module can disconnect the monitoring actions of each module in the signal receiving module. For example, the application control module can disconnect the voice monitoring of the voice receiving module, disconnect the touch monitoring of the touch module, and disconnect the typing monitoring of the text input module.

[0390] Figure 20 This is a schematic flowchart illustrating another image display method provided in an embodiment of this application. Figure 20 As shown, the display device also includes an application module, which includes a saving module, a sharing module, a background beautification module, a picture book generation module, and an image quality enhancement module.

[0391] like Figure 20 As shown, after the generated image is displayed in the generation area of ​​the application interface, it can be processed accordingly through the save module, share module, background beautification module, picture book generation module, and image quality enhancement module.

[0392] For example, combining Figure 4 As shown, in response to the user triggering the image quality enhancement control in the image generation application, the application control module calls the image quality enhancement interface to perform image quality enhancement processing on the generated image through the image quality enhancement module in the application module; the image quality enhancement module can then feed back the processed high-definition image to the application control module, which then controls the generation area to display the high-definition image. Combined with... Figure 4 As shown, in response to the user triggering the picture book generation control in the image generation application, the application control module calls the picture book generation interface to process the generated image through the picture book generation module in the application module, generating the corresponding picture book information. Combined with... Figure 4 As shown, in response to a user triggering the background enhancement control in the image generation application, the application control module calls the background enhancement interface to enhance the generated image through the background enhancement module within the application module. The background enhancement module can then feed back the processed enhanced image to the application control module, which in turn controls the generation area to display the enhanced image. Combined with... Figure 4 As shown, in response to the user triggering the sharing control in the image generation application, the application control module calls the sharing interface to share the generated image through the sharing module in the application module. Combined with... Figure 4 As shown, in response to the user triggering the save control in the image generation application, the application control module calls the save interface to save the generated image through the save module in the application module.

[0393] like Figure 20As shown, when the user closes the application, the sharing module in the application module sends a window switching command to the application control module. Based on the window switching command, the application control module switches the current sharing application window to the background. The picture book generation module sends a close application command to the application control module, and the application control module closes the current application.

[0394] The same or similar parts among the various embodiments in this specification can be referred to mutually, and will not be repeated here.

[0395] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or certain parts of the embodiments of the present invention.

[0396] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0397] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.

Claims

1. An image generation method, characterized in that, The method includes: Obtain multimodal information from user input; The encoding network in the image generation model is used to process each modal information in the multimodal information to obtain the feature vector corresponding to each modal information; Based on the fusion network in the image generation model, the feature vectors corresponding to each modality information are fused to obtain a fusion vector; Determine at least one auxiliary feature information corresponding to the multimodal information, wherein the auxiliary feature information is the feature information extracted from the fusion vector, and the auxiliary feature information includes at least skeleton features, depth features, and edge structure features; Based on the first feature subnetwork in the multi-stage network of the image generation model, the fusion vector and the first auxiliary feature information are processed to determine the first output vector; wherein, the multi-stage network includes the first feature subnetwork, the intermediate feature subnetwork and the second feature subnetwork connected in sequence, and the at least one auxiliary feature information includes the first auxiliary feature information, the second auxiliary feature information and the third auxiliary feature information; The first output vector and the second auxiliary feature information are processed based on the intermediate feature subnetwork to determine the second output vector; The second output vector and the third auxiliary feature information are processed based on the second feature sub-network to obtain the target vector; The target vector is decoded based on the image decoding subnetwork in the multi-stage network to obtain the target image corresponding to the multimodal information.

2. The method according to claim 1, characterized in that, The fusion network in the image generation model fuses the feature vectors corresponding to each modality information to obtain a fused vector, including: Determine the preset length; Based on the preset length, the feature vectors corresponding to each modality information are fused through the fusion network to obtain the fusion vector of the preset length.

3. The method according to claim 2, characterized in that, The step of fusing the feature vectors corresponding to each modality information through the fusion network based on the preset length to obtain the fusion vector of the preset length includes: The fusion network performs feature selection processing on the feature vectors corresponding to each modality information to obtain selected feature vectors. The selected feature vector is mapped to a preset feature space to obtain the mapped feature vector; Based on the preset length, the mapped feature vector is fused to obtain the fused vector of the preset length; wherein, the fusion process includes at least one of weighted fusion, feature-level fusion and decision-level fusion.

4. The method according to claim 1, characterized in that, The multimodal information includes at least two of the following: audio information, text information, image information, and video information; the encoding network in the image generation model processes each modal information to obtain the feature vector corresponding to each modal information, including at least two of the following: The audio information is input into the audio coding sub-network of the coding network to generate an audio feature vector corresponding to the audio information; The text information is input into the text encoding subnetwork of the encoding network to generate a text feature vector corresponding to the text information; The image information is input into the image coding subnetwork of the coding network to generate an image feature vector corresponding to the image information; The video information is input into the video coding subnetwork in the coding network to generate a video feature vector corresponding to the video information.

5. The method according to claim 4, characterized in that, The step of inputting the audio information into the audio coding sub-network of the coding network to generate the audio feature vector corresponding to the audio information includes: The audio information is preprocessed to obtain preprocessed audio information; The preprocessed modal information is subjected to feature extraction to obtain an initial feature vector; The initial feature vector is encoded to obtain the audio feature vector.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: Obtain multimodal training information; The multimodal training information is processed based on the image generation model to be trained to generate a predicted image; wherein the image generation model to be trained includes an encoding network to be trained, a fusion network to be trained, and a multi-stage network to be trained. Acquire sample images; Using the predicted image as the initial training output information of the image generation model to be trained, and the sample image as the supervision information, the image generation model is iterated to obtain the image generation model.

7. An image generation device based on multimodal information, characterized in that, include: The acquisition module is configured to acquire multimodal information input by the user. The module is defined as follows: The image generation model processes the modal information based on the encoding network to obtain the feature vector corresponding to each modal information. Based on the fusion network in the image generation model, the feature vectors corresponding to each modality information are fused to obtain a fusion vector; Determine at least one auxiliary feature information corresponding to the multimodal information, wherein the auxiliary feature information is the feature information extracted from the fusion vector, and the auxiliary feature information includes at least skeleton features, depth features, and edge structure features; Based on the first feature subnetwork in the multi-stage network of the image generation model, the fusion vector and the first auxiliary feature information are processed to determine the first output vector; wherein, the multi-stage network includes the first feature subnetwork, the intermediate feature subnetwork and the second feature subnetwork connected in sequence, and the at least one auxiliary feature information includes the first auxiliary feature information, the second auxiliary feature information and the third auxiliary feature information; The first output vector and the second auxiliary feature information are processed based on the intermediate feature subnetwork to determine the second output vector; The second output vector and the third auxiliary feature information are processed based on the second feature sub-network to obtain the target vector; The target vector is decoded based on the image decoding subnetwork in the multi-stage network to obtain the target image corresponding to the multimodal information.

8. A display device, characterized in that, include: The display is configured to show the application interface of an image generation application; wherein the application interface includes a first input area, a second input area, and an image display area; The controller, coupled to the display, is configured to: In response to receiving image information input by the user in the first input area and text information input in the second input area; The encoding network in the image generation model processes each modal information in the multimodal information to obtain the feature vector corresponding to each modal information; wherein, the multimodal information includes the image information and the text information; Based on the fusion network in the image generation model, the feature vectors corresponding to each modality information are fused to obtain a fusion vector; Determine at least one auxiliary feature information corresponding to the multimodal information, wherein the auxiliary feature information is the feature information extracted from the fusion vector, and the auxiliary feature information includes at least skeleton features, depth features, and edge structure features; Based on the first feature subnetwork in the multi-stage network of the image generation model, the fusion vector and the first auxiliary feature information are processed to determine the first output vector; wherein, the multi-stage network includes the first feature subnetwork, the intermediate feature subnetwork and the second feature subnetwork connected in sequence, and the at least one auxiliary feature information includes the first auxiliary feature information, the second auxiliary feature information and the third auxiliary feature information; The first output vector and the second auxiliary feature information are processed based on the intermediate feature subnetwork to determine the second output vector; The second output vector and the third auxiliary feature information are processed based on the second feature sub-network to obtain the target vector; The target vector is decoded based on the image decoding subnetwork in the multi-stage network to obtain the target image corresponding to the multimodal information.