Display device and image display method

By setting up an image generation model with a multimodal training sample set on a display device and combining audio, text, image, and video information, the problem of discrepancy between the image generated by the display device and the user's expectations is solved, thereby improving the accuracy of image generation and the user experience.

WO2026098174A1PCT designated stage Publication Date: 2026-05-15HISENSE VISUAL TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HISENSE VISUAL TECH CO LTD
Filing Date
2025-10-15
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

When existing display devices generate images, the text-based image model has limitations in understanding user semantics, resulting in a deviation between the output image and the user's expectations, which affects the user experience.

Method used

A display device and image display method are provided. By setting a first input area, a second input area and an image display area on the display, and combining an image generation model with a multimodal training sample set, a target image is generated using audio, text, image and video information. This supports users to draw images and input descriptive text, thereby improving the accuracy of image generation.

Benefits of technology

By training the image generation model with a multimodal training sample set, the accuracy of images generated by display devices and the user experience are enhanced, ensuring that the generated images meet the user's text descriptions and image input requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025127927_15052026_PF_FP_ABST
    Figure CN2025127927_15052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present disclosure are a display device and an image display method. The display device comprises a display and a controller. The display displays an application interface, wherein the application interface comprises a first input area, a second input area and an image display area. The controller is configured to: in response to a touch instruction in the first input area, control the display to display in the first input area a first image drawn by a user; input the first image into a trained image generation model to obtain a second image, and control the display to display the second image in the image display area; in response to an acquisition instruction for first description text, control the display to display the first description text in the second input area; and input the first description text and the first image into the image generation model to obtain a third image matching the first description text, and control the display to display the third image in the image display area. The present disclosure can solve the problem of an image generated by a display device having an excessively large deviation from the actual expectation of a user.
Need to check novelty before this filing date? Find Prior Art

Description

Display devices and image display methods

[0001] Cross-reference to related applications

[0002] This disclosure claims priority to Chinese patent applications filed on November 5, 2024, No. 202411570616.2 and November 5, 2024, No. 202411570694.2, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to the field of display device technology, and more particularly to a display device and an image display method. Background Technology

[0004] A display device is a device capable of displaying image information. Through semantic recognition and image generation technologies, it can generate target images based on user-input text. For example, if a user inputs the descriptive text "Please generate an image of a girl with braids" into the display device, the device can generate an image of "a girl with braids" corresponding to that text. However, since the display device generates different images based on the text input into a text-based image model, and this model has limitations in understanding user semantics, the output image may deviate from the user's actual expectations, impacting the user experience. Summary of the Invention

[0005] In a first aspect, a display device is provided, comprising: a display configured to display an application interface of an image generation application; the application interface including a first input area, a second input area, and an image display area; and a controller coupled to the display and configured to: in response to a touch command in the first input area, control the display to display a first image drawn by a user in the first input area; input the first image into a trained image generation model to obtain a second image, and control the display to display the second image in the image display area; in response to a command to obtain first descriptive text, control the display to display the first descriptive text in the second input area; the first descriptive text characterizes a textual description of an image to be generated by a user; input the first descriptive text and the first image into the image generation model to obtain a third image conforming to the first descriptive text, and control the display to display the third image in the image display area; wherein the image generation model is trained based on a multimodal training sample set; each multimodal training sample corresponds to a sample label, the sample label being a target image to be generated by the desired model; each multimodal training sample includes at least one modal information among audio, text, image, and video, the audio and the text characterizing the generation requirements of the target image corresponding to the multimodal training sample.

[0006] In a second aspect, a display device is provided, comprising: a display configured to display an application interface of an image generation application; the application interface including a first input area, a second input area, and an image display area; and a controller coupled to the display and configured to: acquire a seventh image, extract key feature points from the seventh image to obtain an eighth image, and control the display to display the eighth image in the first input area; wherein the seventh image is an image captured by an image acquisition device and / or an image uploaded by a mobile terminal; input the eighth image into a trained image generation model to obtain a ninth image, and control the display to display the ninth image in the image display area; and, in response to an instruction to acquire first descriptive text, control the display to... The display shows the first descriptive text in the second input area; the first descriptive text represents the user's textual description of the image to be generated; the first descriptive text and the eighth image are input into the image generation model to obtain a tenth image that conforms to the first descriptive text, and the display is controlled to display the tenth image in the image display area; wherein, the image generation model is trained based on a multimodal training sample set; each multimodal training sample corresponds to a sample label, the sample label being a target image to be generated by the desired model; each multimodal training sample includes at least one modal information from audio, text, image, and video, the audio and text representing the generation requirements of the target image corresponding to the multimodal training sample.

[0007] Thirdly, an image display method is provided, comprising: responding to a touch command in a first input area, displaying a first image drawn by a user in the first input area; inputting the first image into a trained image generation model to obtain a second image, and displaying the second image in an image display area; responding to a command to obtain first descriptive text, displaying the first descriptive text in a second input area; the first descriptive text representing a textual description of an image to be generated by a user; inputting the first descriptive text and the first image into the image generation model to obtain a third image conforming to the first descriptive text, and displaying the third image in the image display area; wherein the image generation model is trained based on a multimodal training sample set; each multimodal training sample corresponds to a sample label, the sample label representing the target image to be generated by the desired model; each multimodal training sample includes at least one modal information among audio, text, image, and video, the audio and the text representing the generation requirements of the target image corresponding to the multimodal training sample.

[0008] Fourthly, an image display method is provided, comprising: acquiring a seventh image; extracting key feature points from the seventh image to obtain an eighth image; and displaying the eighth image in a first input area; wherein the seventh image is an image captured by an image acquisition device and / or an image uploaded by a mobile terminal; inputting the eighth image into a trained image generation model to obtain a ninth image; and displaying the ninth image in an image display area; in response to an instruction to acquire first descriptive text, displaying the first descriptive text in a second input area; wherein the first descriptive text represents a textual description of an image to be generated by a user; inputting the first descriptive text and the eighth image into the image generation model to obtain a tenth image conforming to the first descriptive text; and controlling the display to display the tenth image in the image display area; wherein the image generation model is trained based on a multimodal training sample set; each multimodal training sample corresponds to a sample label, the sample label representing the target image to be generated by the desired model; each multimodal training sample includes at least one modal information among audio, text, image, and video, the audio and the text representing the generation requirements of the target image corresponding to the multimodal training sample. Attached Figure Description

[0009] Figure 1 is a schematic diagram of an application interface according to some embodiments;

[0010] Figure 2 is a schematic diagram of another application interface according to some embodiments;

[0011] Figure 3 is a schematic diagram of an image display method according to some embodiments;

[0012] Figure 4 is a schematic diagram of an application interface according to some embodiments;

[0013] Figure 5 is a schematic diagram of another application interface according to some embodiments;

[0014] Figure 6 is a schematic diagram of another image display method according to some embodiments;

[0015] Figure 7 is a schematic diagram of another application interface according to some embodiments;

[0016] Figure 8 is a schematic diagram of an image generation method according to some embodiments;

[0017] Figure 9 is a schematic diagram of an image generation model according to some embodiments;

[0018] Figure 10 is a schematic diagram of a multi-stage network structure according to some embodiments. Detailed Implementation

[0019] In this embodiment of the disclosure, display device 200 generally refers to a device with screen display and data processing capabilities. For example, display device 200 includes, but is not limited to, smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc. Users can operate display device 200 through touch operation, mobile terminal 300, and control device 100. Control device 100 can be a remote control, stylus, gamepad, etc.

[0020] To enable user interaction, the display device 200 may run an operating system. The operating system is a computer program used to manage and control the hardware and software resources of the display device 200. The operating system can control the display device to provide a user interface; for example, the operating system can directly control the display device to provide a user interface, or it can provide a user interface by running applications. The operating system also allows users to interact with the display device 200. It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system deeply customized for a specific operating platform, or an independent operating system specifically developed for the display device.

[0021] In some embodiments, the display device 200 may run an image generation application to generate the target image. This image generation application may be a system application built into the operating system of the display device 200. Alternatively, the image generation application may be a standalone application or a third-party application developed based on the operating system of the display device 200.

[0022] During operation, the display device 200 can receive control commands input by the user. Some of these control commands are used to control the display device 200 to generate a target image; these commands can be referred to as image generation commands. Depending on the interaction methods supported by the display device 200, the user can input image generation commands in different ways.

[0023] In some embodiments, the display device 200 can obtain image generation instructions by detecting user input via key input on the accompanying control device 100. Specifically, while the display device 200 is displaying the application list interface, the user can use the arrow keys on the control device 100 to move the focus cursor in the application list interface. After the focus cursor moves to the image generation application icon, the user can press the confirmation key on the control device 100 to control the display device 200 to run the image generation application, i.e., input the image generation instruction. The display device 200 can then run the image generation application according to the image generation instruction.

[0024] In some embodiments, when the display device 200 supports interaction methods such as touch, voice, motion sensing, and gestures, the display device 200 can detect that the user can input interaction signals through the corresponding interaction method and obtain image generation instructions based on the interaction signals. For example, when the voice assistant application built into the display device 200 detects that the user inputs a voice command such as "Please generate an image containing a bird, with blue feathers on its back and white feathers on its belly," it determines that the user's intent is an image generation intent based on the keywords in the voice command. That is, it determines that the current voice command is an image generation instruction, and at this time, the display device 200 can run the image generation application in response to the image generation instruction.

[0025] In some embodiments, the image generation instruction can also be generated by the display device 200 by monitoring various runtime events during operation and generating the instruction when a specific runtime event is detected. For example, the display device 200 can run a search engine application such as "×× Search," and the search interface can include a text input box and an image input box. When the user enters search text in the text input box and / or enters a search image in the image input box, the display device 200 can automatically generate an image generation instruction and run the image generation application in response to the instruction.

[0026] In response to an image generation command, the display device 200 can run an image generation application and display the application interface of the image generation application.

[0027] In some embodiments, in response to a launch operation of an image generation application, the display device 200 launches the image generation application, and the display shows the application interface corresponding to the image generation application. The application interface of the image generation application may include two input areas, namely a first input area and a second input area. The first input area is used for inputting and displaying image content from the input data; the first input area may also be called a drawing area or image drawing area. The second input area is used for inputting and displaying text content from the input data; the second input area may also be called a text input area.

[0028] In some embodiments, the first input area can be used to display an image, which may be an image drawn by the user, an image captured by the user through an image acquisition device on the display device, or an image uploaded through a terminal device.

[0029] For example, when the display device supports touch functionality, users can directly draw images in the first input area, or capture images using the image sensor on the display device, or upload images via a terminal device. When the display device does not support touch functionality, users can capture images using the image sensor on the display device, or upload images via a terminal device.

[0030] In some embodiments, the second input area is used to display descriptive text input by the user. The descriptive text displayed in the second input area can be descriptive text input by the user, or it can be descriptive text generated based on the user's input. For example, the descriptive text can be a description of the user's needs.

[0031] In some embodiments, the application interface of the image generation application may further include an image display area for displaying the generated target image. The image display area may be displayed side by side with the first input area to facilitate user comparison of the target image and the image content in the input data.

[0032] In some embodiments, the image display area displays the image generated after processing by the image generation model. This image display area may also be referred to as the generation area or image generation area. For example, the image generation model can generate a target image based on the image displayed in the first input area and / or the descriptive text displayed in the second input area, and display the target image in the image display area. The target image conforms to the requirements of the descriptive image and the descriptive text.

[0033] For example, the image display area can be used to display the image displayed in the first input area after processing by the image generation model, or it can be used to display the image of the descriptive text displayed in the second input area after processing by the image generation model, or it can be used to display the image of the descriptive image displayed in the first input area and the image of the descriptive text displayed in the second input area after processing by the image generation model.

[0034] The first and second input areas can be designed according to the UI design specifications of the image generation application, exhibiting different shapes and distributed in different locations on the application interface. In some embodiments, the first input area, the second input area, and the image display area can be three independent, non-overlapping areas on the application interface. For example, as shown in Figure 1, the first input area can be located in the upper left corner of the application interface. To facilitate the input and display of image content, the first input area can be a rectangular area. The second input area can be located at the bottom of the application interface. Similarly, to facilitate the input and display of text content, the second input area can be a long, narrow area. The image display area is located in the upper right corner of the application interface, and to facilitate the display of the target image, the image display area can also be a rectangular area.

[0035] In some embodiments, there may be overlapping areas among the first input area, the second input area, and the image display area. The display device 200 can determine the area to which the content belongs based on the location of the focus marker. For example, when the display device 200 receives a user's request to move the focus cursor to the first input control corresponding to the first input area, the entire screen area of ​​the current application interface can be used to receive the user's input image. That is, the first input area at this time is the entire application interface with the focus marker in the first input control state. Similarly, when the display device 200 receives a user's request to move the focus cursor to the second input control corresponding to the first input area, the first input area is the entire application interface with the focus marker in the second input control state.

[0036] To facilitate user input and image display, the display device 200 can configure the display hierarchy among the first input area, the second input area, and the image display area. For example, when the display device 200 detects that the user is performing image-related operations such as touch drawing, image capture, or image uploading, the display hierarchy of the first input area can be set higher than that of the second input area and the image display area. Conversely, when the display device 200 detects that the user is performing text-related operations such as text input or voice interaction, the display hierarchy of the second input area can be set higher than that of the first input area and the image display area. Similarly, after the display device 200 detects that an image generation model has generated a target image, the display hierarchy of the image display area can be set higher than that of the first and second input areas.

[0037] The display device 200 can dynamically adjust the displayed content in the first input area, the second input area, and the image display area at different operating stages. Specifically, the first input area can display different content at different interaction stages depending on the interaction method supported by the display device 200. In some embodiments, if the display device 200 supports touch interaction, for example, the display 260 of the display device 200 is connected to a touch component. The touch component can form a touchscreen with the display 260 for detecting touch signals input by the user based on the touchscreen.

[0038] The first input area can be associated with a touch component, which detects user input touch points through changes in capacitance or resistance. When the user moves their finger or stylus on the touchscreen, the display device 200 continuously records the coordinates of the touch points. Touchscreen interaction programs, such as the touchscreen driver, record these coordinate points, forming a series of point data, i.e., a set of trajectory detection points.

[0039] After obtaining the set of detection points for the drawing trajectory, the image drawing application draws the trajectory based on this point data. That is, the image drawing application can draw lines on the first input area based on the line style corresponding to the user-selected drawing control. For example, if the user selects a pen control, the image drawing application will display lines in the first input area with the line style (red, thick line, solid line) that match the shape of the touch trajectory detection points, based on the pen control's current line style: red, thick line, solid line. The image drawing application can be a standalone application or a drawing function unit within an image generation application. Line drawing can be achieved using algorithms such as Bresenham.

[0040] After detecting and generating a set of drawing trajectory detection points, the display device 200 can save the drawn trajectory as an image file. That is, the display device 200 converts the drawing trajectory detection points into a set of trajectory pixels to generate a user-drawn image (also known as a first image). Similarly, the display device 200 can generate a user-drawn image through a standalone image drawing application or through a drawing function unit in an image generation application.

[0041] During the generation of the drawn image (such as the first image), the display device 200 can monitor the touch signals detected by the touch component in real time, determine the touch node based on the touch signal, and generate the drawn image based on the touch trajectory at the appropriate touch node. Therefore, the first input area in the application interface is used to obtain the drawn image input by the user. That is, the display device 200 can respond to the drawing command input in the first input area, extract the touch trajectory corresponding to the drawing command, generate the drawn image based on the touch trajectory, and control the display to display the drawn image in the first input area.

[0042] In some embodiments, the display device 200 can generate a drawn image according to a preset detection cycle. That is, the display device 200 can obtain a preset detection cycle and extract the touch trajectory of the user input based on the first input area according to the preset detection cycle. The preset detection cycle is set according to the denoising rounds of the image generation model. For example, if the denoising rounds of the image generation model are 30 rounds and the total time is 3 seconds, then the preset detection cycle is 1 second, meaning that detection is performed once every 10 rounds of denoising. The display device 200 can use the touch position when the input time reaches a preset touch response cycle as the touch node, thereby obtaining the touch trajectory of the first input area every 1 second and generating a drawn image based on the touch trajectory.

[0043] In some embodiments, the display device 200 can listen to touch events in the touch trajectory and generate a drawn image based on the listened touch events. The touch events may include touch down events (ACTION_DOWN), touch move events (ACTION_MOVE), and touch up events (ACTION_UP). After the display device 200 listens for a touch down event (ACTION_DOWN) through the touch component, it continuously listens for touch up events (ACTION_UP) and uses the touch up event as an image generation node. That is, after listening for a touch up event, it generates a drawn image in response to the touch up event. The application interface is then refreshed according to a preset detection cycle to display the drawn image in the first input area.

[0044] In some embodiments, if the display device 200 does not support touch interaction, i.e., the display device 200 does not include a touch component, the display device 200 can acquire a descriptive image, wherein the descriptive image is a file in an image format input by the user. The descriptive image can be acquired in real time by an image acquisition device such as a camera, or it can be uploaded by the user via a mobile terminal 300.

[0045] To acquire descriptive images in real time, in some embodiments, the display device 200 may have a built-in or external image acquisition device. Specifically, the display device 200 may include a device interface 240 configured to connect to an image acquisition device. The image acquisition device incorporates an image sensor such as a charge-coupled device (CCD) or complementary metal-oxide-semiconductor (CMOS), which can perform photographing or video recording of the detection area to acquire descriptive images.

[0046] For descriptive images acquired in real time, the display device 200 can display a shooting interaction control for real-time image acquisition in the application interface. When the user inputs an image acquisition command based on the shooting interaction control, the display device 200 can run an image acquisition device through the image acquisition application and acquire a descriptive image through the image acquisition device.

[0047] After the image acquisition device starts running, the display device 200 can display the environmental image captured by the image acquisition device through the first input area. During the display of the environmental image, the display device 200 can also receive a confirmation command input by the user. In response to the confirmation command, it acquires the environmental image corresponding to the moment the confirmation command is input to obtain a descriptive image.

[0048] For example, when the display device 200 does not support touch interaction, but has a built-in or external image acquisition device such as a camera, the user can draw a descriptive image on a drawing medium such as paper, and then take a picture of the drawing medium through the image acquisition device to obtain the descriptive image.

[0049] In order to acquire the descriptive image uploaded by the user, in some embodiments, the display device 200 further includes a communication device 220 configured to establish a communication connection with the mobile terminal 300. The mobile terminal 300 has a camera device capable of scanning barcodes.

[0050] In some embodiments, the application interface of the image generation application may include a QR code scanning control. When the display device 200 receives an image acquisition instruction input by the user based on the QR code scanning control, it can display an upload identification code for uploading the image in the first input area, as shown in Figure 2. The user operates the mobile terminal 300 to activate the QR code scanning function. After scanning the QR code with the camera device, an image selection interface for uploading image files can be displayed on the mobile terminal 300. After the user selects the description image to be uploaded and inputs a confirmation instruction, the mobile terminal 300 can send the selected description image to the display device 200 in response to the confirmation instruction. Upon receiving the description image, the display device 200 displays the description image in the first input area of ​​the application interface.

[0051] In some embodiments, the display device 200 can also obtain the user-uploaded descriptive image through the server 400. That is, the communication device 220 of the display device 200 is configured to establish a communication connection with the server 400. The server 400 also establishes a communication connection with the mobile terminal 300 to receive the descriptive image uploaded by the user through the mobile terminal 300 and store the uploaded descriptive image in a specific Uniform Resource Locator (URL) address. In the step of receiving the descriptive image uploaded by the mobile terminal 300 by scanning the upload identification code, the display device 200 can first obtain the URL address issued by the server. That is, obtain the storage address of the descriptive image sent by the mobile terminal 300 to the server 400 after scanning the upload identification code. An image receiving request is generated based on the URL address, and the image receiving request is sent to the server, i.e., accessing the URL address issued by the server 400, causing the server 400 to respond to the image receiving request and issue the descriptive image. The display device 200 then receives the descriptive image issued by the server 400.

[0052] For example, in response to a user pressing a directional key on the control device 100, the display device 200 controls the focus cursor to move within the image generation application interface. When the focus cursor moves to the QR code scanning control, the display device 200 can receive a confirmation command from the user by pressing the confirmation key on the control device 100. At this time, in response to the confirmation command, the display device 200 controls itself to display the upload identification code, that is, to display a QR code for uploading a descriptive image in the first input area. After the user initiates scanning the QR code using the camera device of the mobile terminal 300, the selected descriptive image (such as the seventh image) can be uploaded to the server 400. After receiving the descriptive image uploaded by the user, the server 400 can send a URL address to the display device 200. After receiving the URL address, the display device 200 cancels the display of the QR code in the first input area and accesses the URL address to download the descriptive image. Furthermore, after downloading the descriptive image, the display device 200 then displays the descriptive image again through the first input area.

[0053] It should be noted that the method of acquiring images through the first input area provided in the above embodiments can be used alone or in combination. That is, for the display device 200 that supports touch interaction, the descriptive image can also be acquired by acquiring images in real time through an image acquisition device, or by uploading images through a mobile terminal 300.

[0054] In some embodiments, the display device 200 may be configured to acquire a drawing image and / or a description image through a first input area. Since both the drawing image and the description image can be displayed in the first input area, the display device 200 can acquire the drawing image or the description image by detecting whether an image is displayed in the first input area.

[0055] In some embodiments, after generating the target image, the display device 200 can also perform image file operations such as saving and sharing the target image in response to user interaction. For this purpose, the application interface of the image generation application can also include an image function area, which includes at least one interactive control. After displaying the target image, the display device 200 can receive image interaction commands input by the user based on the interactive controls. In response to the image interaction commands, it obtains the control identification information of the target interactive control, which is the interactive control selected by the image interaction commands. Then, it queries the function process corresponding to the control identification information and runs the function process based on the function mapping table.

[0056] As shown in Figure 1, the image function area may include a save control. After displaying the target image, the display device 200 can obtain the image interaction command input by the user for the save control. The display device 200 can then respond to the image interaction command, obtain the recognition information "Save button" of the save control, and invoke the image file saving process based on the recognition information to store the target image in a specific storage space.

[0057] In some embodiments, as shown in FIG1, the image function area may further include a sharing control. Upon receiving an image interaction command input by the user based on the sharing control, the display device 200 may invoke an "Image File sharing process" according to the "share button" identification information of the sharing control. The target image is then sent to the target client according to the file sharing process. The display device 200 may also send the generated target image to the server 400 via the file sharing process for storage. The target client can download the target image from the server 400 by sending a request to the server 400 to obtain the target image.

[0058] In some embodiments, the display device 200 can also support interactive operations on multiple target images. That is, the image function area can also include a picture book interaction control (as shown in Figure 1, a picture book generation control). After receiving an image interaction command input by the user based on the picture book interaction control, the display device 200 can acquire the target images saved within a preset generation period to generate a target image set. Then, it acquires the descriptive text information input in the second input area during the target image generation process to generate a picture book file based on the target image set and the descriptive text information.

[0059] In some embodiments, after generating the target image, the display device 200 can receive image processing instructions input by the user. These instructions can be input based on image processing controls within the image function area of ​​the application interface. For example, the image function area may include image processing controls such as background enhancement and image quality enhancement. The display device 200 can obtain the image processing instructions by detecting any image processing control selected by the user through touch interaction.

[0060] After receiving an image processing instruction, the display device 200 can respond by invoking an image processing algorithm model. Then, based on the image processing algorithm model, it performs corresponding image processing on the target image. These image processing algorithms include, but are not limited to, color and brightness adjustment, background blurring, edge detection and matting, background replacement, depth estimation, image segmentation, and style transfer. For example, if a user inputs an image processing instruction by clicking the background enhancement control, the display device 200 can determine the foreground and background regions in the image based on the descriptive text entered in the second input area. It then invokes a color and brightness adjustment algorithm to adjust parameters such as color, contrast, and brightness of pixels within the background region of the target image, making the background more harmonious or highlighting the foreground subject.

[0061] Figure 3 is a schematic diagram of an image display method according to some embodiments. The implementation process of the image display method according to some embodiments when the display device supports touch functionality will be described below with reference to Figure 3. As shown in Figure 3, the method includes steps 610 to 640 as follows.

[0062] Step 610: In response to the touch command of the first input area, control the display to display the first image drawn by the user in the first input area.

[0063] The display device can monitor touch commands input by the user in the first input area in real time. When a touch command is detected, the device can display an image drawn by the user (i.e., a first image) in the first input area according to the touch command. The touch command can be input by the user moving their limbs (such as fingers) or a touch device (such as a stylus) on the touchscreen of the first input area. It should be noted that the following embodiment illustrates the user inputting touch commands through a touch device (such as a stylus).

[0064] For example, when a user moves the touch point in the first input area, the coordinates of the touch point will change. The display device can detect the change in the coordinates of the touch point, obtain the movement trajectory of the touch point, generate a first image based on the movement trajectory of the touch point, and display the first image drawn by the user in the first input area.

[0065] Figure 4 is a schematic diagram of an application interface according to some embodiments. As shown in Figure 4, the application interface 700 includes a first input area 710, a second input area 720, and an image display area 730. The first input area 710 on the left side of the application interface 700 displays a descriptive image (such as a first image) drawn by the user.

[0066] Step 620: Obtain the second image generated based on the first image, and control the display to show the second image in the image display area.

[0067] In some embodiments, the second image displayed in the first input area can be generated by processing the first image using an image generation model. Considering factors such as device performance and product requirements, in practical applications, the image generation model can be deployed on the display device or on a server.

[0068] The image generation model of this disclosure is trained based on a multimodal training sample set. Each multimodal training sample corresponds to a sample label, which represents the target image to be generated by the desired model. Each multimodal training sample includes at least one modal information among audio, text, image, and video, with audio and text representing the generation requirements of the target image corresponding to the multimodal training sample. The image generation model can convert the input multimodal information such as image, text, audio, and video into a corresponding reference image, which can meet the requirements of the multimodal information. It should be noted that the image generation model is described in detail in the following embodiments (the embodiment shown in Figure 8).

[0069] The implementation process of step 620 is described below. Step 620 includes steps 621 and 622 as shown below. These correspond to the cases where the image generation model is deployed on the display device and the server, respectively.

[0070] Step 621: Input the first image into the trained image generation model to obtain the second image, and control the display to show the second image in the image display area.

[0071] In some embodiments, when an image generation model is deployed on a display device, after acquiring a first image, the display device can process the first image based on the image generation model deployed thereon to obtain a second image, and then display the second image in the image display area.

[0072] In some embodiments, step 621 may include: processing the first image using an image generation model based on a preset detection period to obtain a second image, and controlling the display to show the second image in the image display area.

[0073] In some embodiments, the preset detection period can be set according to user needs; for example, the preset detection period can be a preset duration. The display device can input the image drawn by the user within the current detection period into the image generation model for processing every preset detection period to obtain a second image corresponding to the current detection period, and then display the second image corresponding to the current detection period in the image display area.

[0074] In some embodiments, the first image displayed in the first input area may change or may not change between two adjacent detection cycles. That is, the first images corresponding to two adjacent detection cycles may be the same or different.

[0075] Based on the above scheme, the display method provided by this disclosure can process the first image to obtain the second image through the image generation model on the display device based on a preset detection cycle, thereby realizing the periodic synchronous display of the first image drawn by the user in the first input area to the first input area, ensuring that the user can check and confirm the drawn image, improving the portability of user interaction and enhancing the user experience.

[0076] In some embodiments, if it is determined, based on a preset detection period, that the first image corresponding to the current detection period is different from the first image corresponding to the previous detection period, then an image generation model is used to process the first image corresponding to the current detection period to obtain the second image corresponding to the current detection period; and the display is controlled to show the second image corresponding to the current detection period in the image display area.

[0077] In some embodiments, if the display device determines that the first image corresponding to the current detection cycle is different from the first image corresponding to the previous detection cycle, that is, the first image displayed in the first input area has changed, the image generation model can be invoked to process the first image corresponding to the current detection cycle to obtain the second image corresponding to the current detection cycle, and the second image can be displayed in the image display area.

[0078] For example, since the first image corresponding to two adjacent detection cycles changes, the corresponding second image will also change. In order to ensure the consistency between the second image displayed in the image display area and the first image displayed in the first input area, when the first image displayed in the first input area changes, the second image displayed in the second input area needs to be updated accordingly.

[0079] Based on the above scheme, the display device compares the first images of two adjacent detection cycles, and when the first image of the current detection cycle changes from the first image of the previous detection cycle, it processes the first image of the current detection cycle through the image generation model, thereby avoiding repeated processing by the image generation model and improving image generation efficiency.

[0080] In some embodiments, if, based on a preset detection period, it is determined that the first image corresponding to the current detection period is the same as the first image corresponding to the previous detection period, then the second image corresponding to the previous detection period continues to be displayed in the image display area. If the display device determines that the first image corresponding to the current detection period is the same as the first image corresponding to the previous detection period, that is, the first image displayed in the first input area has not changed, then the second image corresponding to the previous detection period continues to be displayed in the image display area.

[0081] For example, since the first image corresponding to two adjacent detection cycles does not change, the corresponding second image also does not change. Therefore, there is no need to update the second image displayed in the second input area, nor is there a need for the image generation model to repeatedly process the first image corresponding to the current detection cycle, thereby improving the processing efficiency of the display device.

[0082] Based on the above scheme, the display device compares the first images of two adjacent detection cycles, and maintains the display of the first image of the previous detection cycle when the first images of the two adjacent detection cycles have not changed, thereby avoiding repeated processing of the image generation model and improving image generation efficiency.

[0083] In some embodiments, step 621 may further include: listening to the touch event corresponding to the touch command, processing the first image based on the image generation model to obtain the second image according to the touch event, and controlling the display to display the second image in the image display area.

[0084] In some embodiments, in addition to displaying the second image in the image display area based on a preset detection period as described above, the display device can also determine the second image displayed in the image display area by listening to the touch events corresponding to the touch commands entered by the user in the first input area.

[0085] In some embodiments, the touch events corresponding to touch commands include touch press events, touch move events, and touch release events. For example, after a user triggers a touch event by inputting a touch command in the first input area of ​​a touch device, the type of touch event can be determined first.

[0086] For example, when the touch device moves from never touching the touchscreen corresponding to the first input area to touching the touchscreen corresponding to the first input area, the touch event corresponding to the input touch command is a touch press event. When the touch device continuously touches the touchscreen corresponding to the first input area for a preset time, the touch event corresponding to the input touch command is a touch move event. When the touch device moves from touching the touchscreen corresponding to the first input area to leaving the touchscreen corresponding to the first input area, the touch event corresponding to the input touch command is a touch release event.

[0087] In some embodiments, when a user draws an image in the first input area, touch commands input within a drawing cycle can sequentially trigger touch press events, touch move events, and touch release events. For example, when a touch command triggers a touch press event, it indicates that the user has begun drawing an image in the first input area, i.e., the current drawing cycle has begun; when a touch command triggers a touch move event, it indicates that the user is drawing an image in the first input area; when a touch command triggers a touch release event, it indicates that the user has completed drawing the image in the first input area, i.e., the current drawing cycle has been completed.

[0088] It should be noted that touch commands input by the user within a drawing cycle can also trigger touch press and touch release events sequentially. In this case, the user may not have adjusted or updated the first image of the first input area within the current drawing cycle.

[0089] Based on the above scheme, the display method provided in this disclosure can obtain a second image by processing the first image through the image generation model on the display device based on the touch event corresponding to the touch command. This enables the image to be displayed synchronously in the first input area when the user completes a stage of drawing in the first input, ensuring that the user can check and confirm the drawn image, improving the convenience of user interaction and enhancing the user experience.

[0090] In some embodiments, while listening for touch press events, touch release events are continuously listened for; in response to listening for touch release events, the first image is processed based on an image generation model to obtain a second image, and the display is controlled to display the second image in the image display area.

[0091] In some embodiments, when the display device detects a touch press event, it indicates the start of the current rendering cycle. The display device can continuously monitor touch events, and when it detects a touch release event, it indicates the end of the current rendering cycle. In response to the touch release event, the display device can process the first image corresponding to the current rendering cycle displayed in the first input area to obtain a second image, and control the display to display the second image in the image display area. That is, the display device can process the first image corresponding to the rendering cycle through an image generation model at the end of a rendering cycle, thereby updating the second image in the image display area.

[0092] Based on the above scheme, when the touch press event is triggered, the touch release event is triggered, indicating that the user has completed the drawing of one drawing cycle and stopped drawing. Therefore, the currently drawn image can be processed in a timely manner through the image generation model to ensure the timeliness of image display.

[0093] In some embodiments, step 621 may further include: in response to a confirmation command input by the user, processing the first image based on an image generation model to obtain a second image, and controlling the display to show the second image in the image display area.

[0094] In some embodiments, in addition to displaying the second image in the image display area based on a preset detection period, or displaying the second image in the image display area by listening to the touch event corresponding to the touch command input by the user in the first input area, the display device may also determine the second image displayed in the image display area by the confirmation command input by the user.

[0095] In some embodiments, when a user displays a first image in the first input area, the user can input a command through a control device, or a confirmation command via voice. The confirmation command indicates that the first image currently displayed in the first input area can be used to generate a corresponding second image. Therefore, upon receiving the confirmation command, the display device, in response to the confirmation command, processes the first image currently displayed in the first input area using an image generation model to obtain a second image, and displays the second image in the image display area.

[0096] In some embodiments, the application interface may include a confirmation control, which allows the user to input confirmation commands via a control device (such as a remote control).

[0097] In some embodiments, the user can input a confirmation command at any time during the first image drawing process as needed. Referring to FIG4, when the user draws an image of a bird in the first input area 710, the user can input a confirmation command after completing the drawing of the entire bird image, or when completing a portion of the bird image. This disclosure does not limit this.

[0098] For example, when a user completes drawing part of a bird image, they can input a confirmation command. The image generation model then generates a corresponding reference image based on that part of the bird image and displays it in the image display area 730. The user continues drawing the bird image in the first input area. After completing the remaining part of the image, the user inputs a confirmation command again. The image generation model then generates a corresponding reference image based on the complete bird image currently displayed in the first input area 710 and displays it in the image display area 730.

[0099] Based on the above scheme, the display method provided in this disclosure can process the first image to obtain the second image based on the confirmation command input by the user through the image generation model on the display device, thereby realizing the synchronous display of the drawn image to the first input area according to the user's needs, and improving the user experience.

[0100] Step 622: Send the first image to the server to receive the second image generated by the server based on the first image, and control the display to show the second image in the image display area.

[0101] Step 630: In response to the instruction to acquire the first descriptive text, control the display to display the first descriptive text in the second input area.

[0102] In some embodiments, the first descriptive text is text entered by the user or text generated based on the user's input description. Specifically, the first descriptive text is a textual description of the desired image by the user. After determining the first descriptive text, the display device controls the display to show the first descriptive text in the second input area.

[0103] In some embodiments, the application interface may include a text input control and a voice input control. Users can select the text input control through the control device to input descriptive text; or, users can select the voice input control through the control device to input descriptive voice.

[0104] For example, a user can select the text input control via remote control to bring up the virtual input keyboard, and then input the corresponding text in the second input area 720 to obtain the first descriptive text. Alternatively, the user can select the voice input control to bring up the voice input mode of the display device and input descriptive voice.

[0105] Figure 5 is a schematic diagram of another application interface according to some embodiments. As shown in Figure 5(a), the application interface 700 includes a text input control 910 and a voice input control 920. The user can input descriptive text through the text input control 910, and the descriptive text displayed in the second input area 720 is "bird, with blue feathers on its back and white feathers on its belly".

[0106] In some embodiments, step 630 includes: processing the descriptive speech in response to user input to obtain an audio spectrum; acquiring a first descriptive text obtained after processing the audio spectrum based on a speech recognition model; and controlling a display to display the first descriptive text in a second input area.

[0107] In some embodiments, when the first descriptive text input by the user is text generated based on the user-input descriptive speech, the display device, after acquiring the user-input descriptive speech, first processes the descriptive speech to obtain an audio spectrum. After acquiring the audio spectrum, a speech recognition model can be used to process the audio spectrum to obtain the first descriptive text. The speech recognition model can be a pre-trained model, and it can be deployed on the display device or a server.

[0108] In some embodiments, when the speech recognition model is deployed on a display device, the display device can process the audio spectrum corresponding to the acquired descriptive speech using the speech recognition model deployed on it to obtain a first descriptive text. When the speech recognition model is deployed on a server, after obtaining the audio spectrum corresponding to the descriptive speech, the display device first sends the audio spectrum to the server. The server then processes the audio spectrum using the speech recognition model deployed on it to obtain the first descriptive text and sends the first descriptive text to the display device. After obtaining the first descriptive text, the display device controls the display to show the first descriptive text in the second input area.

[0109] Based on the above scheme, the display method provided in this disclosure can convert the user-inputted descriptive speech into descriptive text using a pre-trained speech recognition model, and then display the descriptive text in the second input area, thereby facilitating the subsequent generation of the third image. By converting the descriptive speech into descriptive text, the user can avoid needing to input text in the second input area, improving input efficiency and enhancing the user experience.

[0110] In some embodiments, step 630 above, controlling the display to display the first descriptive text in the second input area, includes: obtaining the number of characters of the third descriptive text currently displayed in the second input area; if the number of characters of the third descriptive text is zero, then adding the first descriptive text to the second input area and controlling the display to display the first descriptive text in the second input area; if the number of characters of the third descriptive text is not zero, then after adding the first descriptive text to the third descriptive text, updating the descriptive text in the second input area and controlling the display to display the updated descriptive text in the second input area.

[0111] In some embodiments, after receiving input from the user of the first descriptive text, the display device may first determine whether the second input area currently displays descriptive text. For example, the display device may obtain the number of characters in the third descriptive text displayed in the second input area; the number of characters in the third descriptive text may be zero or non-zero.

[0112] For example, when the third descriptive text has zero characters, it indicates that the second input area is currently empty. In this case, the display device can display the first descriptive text entered by the user in the second input area. When the third descriptive text has a non-zero character count, it indicates that the second input area is not currently empty. In this case, the user can add the first descriptive text after the third descriptive text to update the descriptive text displayed in the second input area, and control the display to show the updated descriptive text in the second input area. That is, when the third descriptive text has a non-zero character count, the second input area can display a new descriptive text composed of the first and third descriptive texts.

[0113] Referring to Figure 5, as shown in Figure 5(b), the second input area 720 currently displays the third descriptive text "bird, blue feathers on the back". In this case, if the user inputs the first descriptive text "white feathers on the belly", the display device obtains the number of characters in the third descriptive text. Since the number of characters in the third descriptive text is not zero, the display device can add the user input first descriptive text after the third descriptive text and display the updated descriptive text "bird, blue feathers on the back; white feathers on the belly" in the second input area 720.

[0114] Based on the above scheme, the display method provided in this disclosure, when receiving the first descriptive text input by the user, obtains the number of characters of the third descriptive text displayed in the second input area. If it is determined that the number of characters of the third descriptive text is not zero, that is, if there is descriptive text in the second input area, the newly input first descriptive text is added after the original third descriptive text, thereby supplementing the descriptive text in the second input area and further improving the convenience of user interaction.

[0115] Step 640: Process the first descriptive text and the first image to obtain a third image that matches the first descriptive text, and control the display to show the third image in the image display area.

[0116] In some embodiments, after determining the first descriptive text, the display device may process the first descriptive text and the first image through an image generation model to obtain a third image that conforms to the first descriptive text and the first image.

[0117] In some embodiments, step 640 includes steps 641 and 642 as described below. Step 641 is the implementation process of step 640 when the image generation model is deployed on a display device, and step 642 is the implementation process of step 640 when the model is deployed on a server.

[0118] Step 641: Input the first descriptive text and the first image into the image generation model to obtain a third image that matches the first descriptive text, and control the display to show the third image in the image display area.

[0119] In some embodiments, when displaying a second image, the image display area, in response to user input of first descriptive text, processes the first descriptive text and the first image based on an image generation model deployed in the display device to obtain a third image. The third image conforms to the requirements of both the first image displayed in the first input area and the first descriptive text displayed in the second input area.

[0120] In some embodiments, the third image may be different from or the same as the second image. For example, when the number of characters in the first descriptive text is zero, the third image is the same as the second image; or, when the content described by the first descriptive text has already been described in the first image, the third image may also be the same as the second image. This application does not limit this aspect.

[0121] In some embodiments, after generating a third image based on an image generation model, the display device can control the display to show the third image in the image display area. For example, when the third image is the same as the second image, the display device can control the display to maintain the display of the second image in the image display area; when the third image is different from the second image, the display device can control the display to exit the display of the second image in the image display area and display the third image.

[0122] In some embodiments, after step 641 above, the method further includes: in response to receiving a modification instruction for the first image input by a user in the first input area, determining a fourth image; if the fourth image is different from the first image, processing the first descriptive text and the fourth image based on an image generation model to obtain a fifth image that conforms to the first descriptive text, and controlling the display to display the fifth image in the image display area.

[0123] In some embodiments, a user can modify a first image displayed in the first input area. After modification, the first input area can display a modified image (such as a fourth image). For example, when the first image is displayed in the first input area, the user can move their limbs (such as fingers) or a touch device (such as a stylus) on the touchscreen corresponding to the first input area to input a modification command. When the display device detects the modification command, it can generate a modified fourth image based on the first image and the modification command. For example, the display device can generate the fourth image based on the movement trajectory corresponding to the first image and the modification command.

[0124] In some embodiments, after determining the fourth image, the display device can compare the fourth image with the first image previously displayed in the first input area to determine whether the fourth image and the first image are the same. For example, if the first image and the fourth image are the same, it indicates that the user has not modified the first image, or the user's modification operation on the first image was unsuccessful. In this case, the display device controls the display to maintain the currently displayed second image in the image display area. If the first image and the fourth image are different, it indicates that the user has modified the first image. In this case, the display device needs to reprocess the fourth image and the first descriptive text using an image generation model to generate a fifth image that conforms to the fourth image and the first descriptive text. After generating the fifth image, the display device controls the display to exit the display of the second image in the image display area and display the fifth image. Based on the above scheme, this application can support the user's modification of the first image drawn in the first input area, and after modification, reprocess the modified fourth image and the first descriptive text displayed in the second input area using an image generation model to ensure that a fifth image conforming to the modified fourth image and the first descriptive text is generated.

[0125] In some embodiments, after step 641 above, the method further includes: in response to receiving second descriptive text input by a user, controlling the display to display the second descriptive text in a second input area; processing the second descriptive text and the first image based on an image generation model to obtain a sixth image that conforms to the second descriptive text, and controlling the display to display the sixth image in an image display area.

[0126] In some embodiments, in addition to modifying the first image displayed in the first input area, the user can also modify the first descriptive text displayed in the second input area as needed. For example, the user can re-enter second descriptive text in the second input area. The second descriptive text is either text entered by the user or text generated based on the user's input description, and it differs from the first descriptive text.

[0127] For example, a user can first delete the first description text displayed in the second input area, and after deleting the first description text, enter the second description text and control the second input area to display the second description text.

[0128] After obtaining the second descriptive text, the display device can process the second descriptive text and the first image through an image generation model to obtain a sixth image that matches the first image and the second descriptive text, and control the display to display the sixth image in the image display area.

[0129] In some embodiments, a user may modify both the first image displayed in the first input area and the first descriptive text displayed in the second input area.

[0130] In some embodiments, the user modifies the first image displayed in the first input area to obtain a fourth image. A timer starts when the fourth image is displayed in the first input area. Before the timer reaches a preset time threshold, if the user re-enters second descriptive text in the second input area, the display device can process the fourth image and the second descriptive text using an image generation model to obtain an image that matches the fourth image and the second descriptive text, and then display this image in the image display area. The fourth image is different from the first image, and the second descriptive text is different from the first descriptive text. The preset time threshold can be adjusted according to user needs, and this embodiment does not limit this.

[0131] In some embodiments, after receiving a modification instruction for the first image input by a user in the first input area, the display device can display a modified fourth image in the first input area, process the fourth image and the first descriptive text using an image generation model to obtain a modified image, and then display the modified image in the display area. During this process, if the user re-enters the second descriptive text in the second input area, the display device processes the fourth image and the second descriptive text using the image generation model to obtain an image that conforms to both the fourth image and the second descriptive text, and then displays this image in the image display area.

[0132] It should be noted that the display device can process the modifications made by the user to the first image in the first input area and the first descriptive text in the second input area sequentially, or it can process the modifications to the first image and the first descriptive text within a preset time threshold simultaneously, based on a preset time threshold. This application does not limit this approach. Based on the above scheme, it can support users in modifying the first descriptive text displayed in the second input area, and after modification, the first image displayed in the first input area and the modified second descriptive text displayed in the second input area can be reprocessed by an image generation model to ensure that a sixth image that conforms to the first image and the modified second descriptive text is generated.

[0133] The display method provided in this application embodiment utilizes an image generation model deployed on the display device. This model processes a first image input by the user in the first input area and / or a first descriptive text input in the second input area to generate a third image that meets the user's needs, which is then displayed in the image display area. Therefore, this application embodiment provides a function capable of generating corresponding images based on multimodal information, improving the accuracy of image generation by the display device and enhancing the user experience. Furthermore, deploying the image generation model on the display device further improves the security and efficiency of the third image generation process. Additionally, since the first input area, second input area, and image display area on the application interface can display different content at different stages, the convenience of user interaction is further enhanced.

[0134] Step 642: Receive a third image generated by the server based on the first description text and the first image, and control the display to show the third image in the image display area.

[0135] It should be noted that when the image generation model is deployed on a server, the images displayed in the image display area are generated by the server and sent to the display device. The image display process in the image display area is similar to that of the image generation model deployed on the display device, and will not be described again here.

[0136] The above process enables the generation of corresponding images based on multimodal information, improving the accuracy of images generated by display devices and enhancing the user experience. Furthermore, deploying the image generation model on a server provides higher processing power and more robust computing resources, making it easier to manage and expand. Additionally, the ability to display different content at different stages in the application interface's first input area, second input area, and image display area further enhances the convenience of user interaction.

[0137] Figure 6 is a schematic diagram of another image display method according to some embodiments. As shown in Figure 6, the image display method includes:

[0138] Step 910: Obtain the seventh image, extract the key feature points of the seventh image to obtain the eighth image, and control the display to show the eighth image in the first input area.

[0139] In some embodiments, the seventh image is an image captured by an image acquisition device and / or an image uploaded by a mobile terminal. When the display device does not support touch operation or the user cannot draw the first image in the first input area, the display device may also use the image captured by the image acquisition device, or the image uploaded by the user via a mobile terminal, as the seventh image and display the seventh image in the first input area. When the display device supports touch functionality, if the user does not want to draw the first image in the first input area, they may also capture the seventh image by an image acquisition device or upload the seventh image via a mobile terminal.

[0140] In some embodiments, the application interface of the display includes an image acquisition control, which may include a camera interaction control and a barcode scanning interaction control. The user can provide a seventh image to the display device through the camera interaction control or the barcode scanning interaction control. For example, continuing to refer to FIG4, the interface on the right side of the second input area 720 includes a camera interaction control 740 and a barcode scanning interaction control 750.

[0141] In some embodiments, obtaining the seventh image in step 910 includes: receiving an image captured by an image capture device in response to an image acquisition instruction input by a shooting interaction control on the application interface, and displaying the image captured by the image capture device in a first input area; and determining the captured image as the seventh image in response to a shooting instruction input by the user.

[0142] In some embodiments, the display's application interface includes a camera interaction control, through which the user can input image acquisition commands. For example, the user can input image acquisition commands through a control device (such as a remote control), or the user can input image acquisition commands through voice. This disclosure does not limit the scope of the embodiments.

[0143] For example, as shown in Figure 4, when the user moves the focus to the shooting interaction control 740 using the remote control and presses the remote control to select the shooting interaction control, they can input an image acquisition command. The image acquisition command activates the image acquisition device on the display device, enters the shooting mode, and displays the captured image in the first input area.

[0144] In some embodiments, obtaining the seventh image in step 910 above further includes: responding to an image acquisition instruction input by the scanning interaction control of the application interface, controlling the display to display an identification code in the first input area; and receiving the seventh image uploaded by the mobile terminal through scanning the identification code.

[0145] In some embodiments, the application interface of the display device may further include a QR code scanning control. The user can input an image acquisition command into the QR code scanning control to retrieve an identification code, which is used to upload the image after scanning by a terminal device. In response to the image acquisition command, the display device controls the display to show the identification code in the first input area. The user scans the identification code using a mobile terminal (such as a mobile phone) to send the desired image (such as a seventh image) to the display device. The display device receives the seventh image and displays it in the first input area.

[0146] In some embodiments, after acquiring the seventh image, the first input area may exit the display of the identification code and display the seventh image. The identification code may be a QR code, barcode, RFID tag, or other form, and this disclosure is not limited thereto. Alternatively, the display device may also exit the display of the identification code in the first input area after detecting that the user has completed scanning the identification code; this disclosure is not limited in its embodiments.

[0147] For example, as shown in Figure 4, when a user moves the focus to the QR code scanning control 750 using the remote control and presses the remote to select the control, they can input an image acquisition command. This command retrieves an identification code, which is then displayed in the first input area. The user scans this code using their mobile terminal to upload the seventh image.

[0148] In some embodiments, receiving the seventh image uploaded by the mobile terminal via scanning the identification code includes: obtaining a URL address sent by the server; generating an image receiving request based on the URL address; sending the image receiving request to the server so that the server responds to the image receiving request by sending the seventh image; and receiving the seventh image sent by the server.

[0149] In some embodiments, the URL address is the storage address of the seventh image sent by the terminal device to the server after scanning the identification code.

[0150] In some embodiments, step 910 includes: acquiring a seventh image and controlling the display to display the seventh image in the first input area; extracting key feature points from the seventh image to obtain an eighth image and controlling the display to display the eighth image in the first input area.

[0151] In some embodiments, after acquiring the seventh image, the display device controls the display to display the seventh image in the first input area. The display device extracts key feature points from the seventh image displayed in the first input area, uses the extracted feature point map as the eighth image, and controls the display to display the eighth image in the first input area.

[0152] In some embodiments, step 910 above, which involves extracting key feature points from the seventh image to obtain the eighth image, includes: determining the image type of the seventh image; if the image type of the seventh image is a first type, then extracting key feature points from the seventh image using a skeleton feature extraction algorithm to obtain the eighth image; if the image type of the seventh image is a second type, then extracting key feature points from the seventh image using an edge detection algorithm and a skeleton feature extraction algorithm to obtain the eighth image.

[0153] In some embodiments, the image type includes a first type and a second type, wherein the first type of image is a line drawing, and the second type of image is a non-line drawing. Since the seventh image can be captured by a camera or uploaded by a user via a mobile terminal, it may not contain obvious line features. To enable the image generation model to generate a more accurate image based on the seventh image, the display device can perform image processing and feature extraction operations on the acquired seventh image. For example, when the seventh image is a line drawing, the display device can extract key feature points from the seventh image using a skeleton feature extraction algorithm to transform the seventh image into an eighth image. When the seventh image is a non-line drawing, the display device first transforms the seventh image into a boundary image using an edge detection algorithm, and then transforms the boundary image into an eighth image using a skeleton feature extraction algorithm. This eighth image can be a skeleton structure diagram.

[0154] In some embodiments, the display device performs grayscale conversion and binarization on the seventh image to generate a binary image, performs edge detection on the binary image using an edge detection algorithm to generate a boundary image, extracts morphological features from the boundary image to generate a skeleton structure map, and displays the generated skeleton structure map as the eighth image in the first input area.

[0155] In some embodiments, after the eighth image is determined, the first input area may exit the display of the seventh image and display the eighth image; alternatively, the first input area may display both the seventh and eighth images simultaneously; or, the first input area may display either the seventh or the eighth image according to the user's selection.

[0156] In some embodiments, in response to selecting the first option, the display is controlled to display a seventh image in the first input area; in response to switching the focus to the second option, the display is controlled to display an eighth image in the first input area.

[0157] In some embodiments, the application interface displayed on the monitor may include at least one option control.

[0158] For example, when the application interface includes two option controls, such as a first option control and a second option control, the first option control can correspond to the first option, and the second option control can correspond to the second option. When the user moves the focus to the first option control using the remote control, in response to the selection of the first option control corresponding to the first option, the seventh image is displayed in the second input area; when the user moves the focus to the second option control using the remote control, in response to the selection of the second option control corresponding to the second option, the eighth image is displayed in the second input area.

[0159] Figure 7 is a schematic diagram of another application interface according to some embodiments. As shown in Figure 7(a), the first input area 710 includes two option controls, a first option control 711 and a second option control 712. When the user selects the first option control 711 via a remote control, the display area of ​​the first input area 710 displays a seventh image. When the user selects the second option control 712 via a remote control, the display area of ​​the first input area 710 displays an eighth image.

[0160] In some embodiments, the first input area may simultaneously display both the seventh and eighth images. For example, the first input area may be divided into two display areas, one of which (as referred to as the first display area) can display the seventh image, and the other (as referred to as the second display area) can display the eighth image. Users can view both the original image (such as the seventh image) captured or uploaded via the first input area and the skeleton diagram generated from the original image (such as the eighth image).

[0161] In some embodiments, the first display area and the second display area may be the same size or different. For example, when the first display area and the second display area are different sizes, the second display area may be larger than the first display area, such as the first display area being a portion of the second display area. This disclosure does not limit the size or distribution of the first and second display areas.

[0162] For example, as shown in Figure 7(b), the display area of ​​the first input area 710 includes a display area for displaying the seventh image and a display area for displaying the eighth image. The display area for displaying the seventh image is a portion of the display area for displaying the eighth image.

[0163] Step 920: Input the eighth image into the trained image generation model to obtain the ninth image, and control the display to show the ninth image in the image display area.

[0164] In some embodiments, after determining the eighth image, the display device can process the eighth image using an image generation model to obtain a ninth image, and then display the ninth image in the image display area. It should be noted that step 920 is similar to step 621 in the above embodiments, wherein the eighth image in step 920 is equivalent to the first image in step 621, and the ninth image in step 920 is equivalent to the second image in step 621. The image processing procedure of the image generation model and the process of displaying the image in the image display area have been described in detail in step 621 above, and will not be repeated here.

[0165] Step 930: In response to the instruction to acquire the first descriptive text, control the display to display the first descriptive text in the second input area.

[0166] It should be noted that step 930 is similar to step 630 in the above embodiment, and will not be described again here.

[0167] Step 940: Input the first descriptive text and the eighth image into the image generation model to obtain the tenth image that matches the first descriptive text, and control the display to show the tenth image in the image display area.

[0168] In some embodiments, after obtaining the first descriptive text, the display device can process the first descriptive text and the eighth image based on an image generation model to obtain a tenth image that conforms to the first descriptive text and the eighth image, and display the tenth image in the image display area.

[0169] It should be noted that step 940 is similar to step 640 in the above embodiment, and will not be described again here.

[0170] In some embodiments, after step 940, the method further includes: in response to a user's modification instruction for the eighth image input in the first input area, obtaining an eleventh image; if the eleventh image is different from the eighth image, processing the first descriptive text and the eleventh image based on an image generation model to obtain a twelfth image that conforms to the first descriptive text, and controlling the display to display the twelfth image in the image display area.

[0171] In some embodiments, when the display device supports touch functionality and the first input area displays an eighth image, the user can also modify the eighth image displayed in the first input area. For example, the user can move their limbs (such as fingers) or a touch device (such as a stylus) in the first input area to input a modification command. When the display device detects a modification command, it can generate a modified eleventh image based on the eighth image and the modification command. The display device determines whether the eleventh image is the same as the eighth image. If the eleventh image is different from the eighth image, it processes the eleventh image and the first descriptive text displayed in the second input area using an image generation model to obtain a twelfth image. The twelfth image conforms to the requirements of the eleventh image and the first descriptive text.

[0172] It should be noted that the user's modification of the eighth image is similar to the modification process of the first image in the above embodiments, and will not be repeated here.

[0173] In some embodiments, the eleventh image in the first input area can be periodically detected. If the eleventh image in the current detection period is different from the eleventh image in the previous detection period, the eleventh image in the current detection period is input into the image generation model to obtain the twelfth image in the current detection period. The display is then controlled to display the twelfth image corresponding to the current detection period in the image display area. Alternatively, after detecting that the touch event corresponding to the touch instruction is a touch press event, if an adjacent touch event is detected as a touch release event, the eleventh image is input into the image generation model to obtain the twelfth image, and the display is controlled to display the twelfth image in the image display area. Alternatively, in response to a confirmation instruction input by the user, the eleventh image is input into the image generation model to obtain the twelfth image, and the display is controlled to display the twelfth image in the image display area.

[0174] It should be noted that the display device can also generate the twelfth image based on a preset detection period, the touch event corresponding to the touch command, or the confirmation command input by the user. In this embodiment, the process of generating the twelfth image based on the first descriptive text and the eleventh image is similar to the process of generating the second image based on the first image in the above embodiment, and will not be described again here.

[0175] In some embodiments, after step 940, the method further includes: in response to an instruction to acquire the second descriptive text, controlling the display to display the second descriptive text in the second input area; processing the second descriptive text and the eighth image based on an image generation model to obtain a thirteenth image that conforms to the second descriptive text, and controlling the display to display the thirteenth image in the image display area.

[0176] In some embodiments, the display device may also update the first descriptive text displayed in the second input area to obtain a second descriptive text, and process the second descriptive text and the eighth image displayed in the first input area using an image generation model to obtain a thirteenth image, which is then displayed in the image display area. The thirteenth image conforms to the requirements of both the second descriptive text and the eighth image.

[0177] The above process enables the generation of corresponding images based on multimodal information, improves the accuracy of images generated by the display device, and enhances the user experience. Furthermore, deploying the image generation model on the display device improves the security and efficiency of the third image generation process. Since the first input area on the application interface can display images captured and uploaded by the user, the convenience of user interaction is further enhanced.

[0178] The process of image generation by the image generation model is described below with reference to Figure 8. It should be noted that since the image generation model can be deployed on a server or a display device, the image generation method provided in the following embodiments can be implemented by a display device or a server; this disclosure does not limit this. As shown in Figure 8, the method includes:

[0179] Step 1110: Obtain multimodal information input by the user.

[0180] In some embodiments, modal information can be information of different forms (or types) input by the user, and the modal information can be audio information, text information, image information, or video information. Specifically, the multimodal information obtained in this embodiment may include at least two modal information selected from audio information, text information, image information, and video information.

[0181] In some embodiments, multimodal information may include any two modalities selected from audio, text, image, and video information; or, multimodal information may include any three modalities selected from audio, text, image, and video information; or, multimodal information may include audio, text, image, and video information. This disclosure does not limit this. For example, in the above embodiments, the multimodal information received by the display device includes image information input from the first input area and text information input from the second input area. It should be noted that, in addition to audio, text, image, or video information, modal information may also include more modal information such as historical style status information and historically viewed content information; this disclosure does not limit this.

[0182] In some embodiments, a user can input multimodal information through a display device. The display device's screen shows an application interface generated from the image, which may include multiple input areas, allowing the user to input different modal information in each area.

[0183] For example, referring to Figure 4 above, a user can input first modal information in the first input area 710 of the application interface 700 and input second modal information in the second input area 720. The first modal information can be image information (such as the first image in the above embodiment), and the second modal information can be text information (such as the first descriptive text in the above embodiment).

[0184] In some embodiments, the application interface displayed on the monitor may include a third input area in addition to the first and second input areas. The user can input third modal information through the third input area, which may be audio or video information. It should be noted that the application interface of the monitor may also include a greater number of input areas (such as a fourth input area), and this embodiment does not limit this. That is, the user can input corresponding modal information in multiple input areas respectively, or can input corresponding modal information in any at least two of the multiple input areas.

[0185] In some embodiments, when the image generation model is deployed on a display device, the display device obtains multimodal information based on the modal information input by the user in each input area; when the image generation model is deployed on a server, the display device sends the modal information input by the user in each input area to the server, enabling the server to obtain the multimodal information. This disclosure uses the deployment of the image generation model on a display device as an example for illustrative purposes.

[0186] Figure 9 is a schematic diagram of an image generation model according to some embodiments. As shown in Figure 9, the image generation model 1500 includes an encoding network 1510, a fusion network 1520, and a multi-stage network 1530. After obtaining the multimodal information input by the user, the multimodal information can be input into the encoding network 1510, and after being processed by the encoding network 1510, the fusion network 1520, and the multi-stage network 1530 in sequence, a target image that meets the requirements is generated.

[0187] The processing procedures of coding network 1510, fusion network 1520 and multi-stage network 1530 are explained below.

[0188] Step 1120: Process the modal information based on the encoding network in the image generation model to obtain the feature vector corresponding to each modal information.

[0189] In some embodiments, the encoding network may include multiple encoding sub-networks, each of which can process its corresponding modal information separately. For example, as shown in FIG9, when the modal information is audio information, text information, image information, or video information, the encoding network 1510 may include an audio encoding sub-network 1511, a text encoding sub-network 1512, an image encoding sub-network 1513, and a video encoding sub-network 1514.

[0190] It should be noted that the encoding network may include a greater number of encoding sub-networks, and this disclosure does not limit the number or type of encoding sub-networks in the encoding network. This disclosure provides an illustrative example using an encoding network that may include audio encoding sub-networks, text encoding sub-networks, image encoding sub-networks, and video encoding sub-networks.

[0191] In some embodiments, step 1120 includes: inputting audio information into an audio encoding subnetwork in an encoding network to generate an audio feature vector corresponding to the audio information; inputting text information into a text encoding subnetwork in an encoding network to generate a text feature vector corresponding to the text information; inputting image information into an image encoding subnetwork in an encoding network to generate an image feature vector corresponding to the image information; and inputting video information into a video encoding subnetwork in an encoding network to generate a video feature vector corresponding to the video information.

[0192] In some embodiments, the audio coding subnetwork can encode audio information to obtain an audio feature vector. The text coding subnetwork can encode text information to obtain a text feature vector. The image coding subnetwork can encode image information to obtain an image feature vector. The video coding subnetwork can encode video information to obtain a video feature vector. For example, when the multimodal information input by the user includes both image and text information, the text information can be input into the text coding subnetwork for processing, and the image information can be input into the image coding subnetwork for processing. That is, multiple coding subnetworks in the coding network can determine whether to perform coding processing based on the input multimodal information; not all multiple coding subnetworks necessarily need to perform coding processing.

[0193] In some embodiments, the structures of the various encoding sub-networks can be the same or different. Each encoding sub-network can be pre-trained based on training data. The text encoding sub-network can convert text information into fixed-length text feature vectors based on a transformer model. The text encoding sub-network can convert text information into fixed-length text feature vectors through steps such as word segmentation, vocabulary mapping, word embedding, and encoding processing. For example, the text encoding sub-network can convert input text information into tokens of length 77.

[0194] When text information is input into the text encoding sub-network, it first segments the text information into a series of tokens, which can be words, phrases, or other linguistic units. Then, based on a predefined vocabulary, each token is mapped to a unique ID; this vocabulary is constructed during training using a large amount of text data, and each word or token can correspond to an index ID. Next, each token is converted into a corresponding word embedding vector, obtained by looking up the corresponding row in the embedding matrix. Since the Transformer does not have built-in order information, a positional encoding is added to each token to represent its position in the sentence. The positional encoding can be a predefined fixed vector or trainable parameters, and these positional encoding vectors are added to the word embedding vector of each token. Finally, the text encoding sub-network uses a model based on Transformer architecture to process these embedding vectors, capturing the contextual relationships in the text information through mechanisms such as multi-head self-attention and generating a series of text feature vectors.

[0195] In some embodiments, the image coding subnetwork uses a Convolutional Neural Network (CNN) or a Vision Transformer to convert image information into a fixed-length image feature vector. For example, the image coding subnetwork can convert the input image information into tokens (i.e., image feature vectors) of length 196. Specifically, the CNN extracts image features through convolutional operations and flattens the feature maps into one-dimensional vectors as tokens; the Vision Transformer divides the image into small blocks and converts each block into tokens containing features and contextual information through linear embedding and a Transformer coding subnetwork. It should be noted that the structure and processing of the video coding subnetwork are similar to those of the image coding subnetwork. The difference lies in the fact that the video information processing process needs to consider the spatiotemporal correlation between frames, etc. This disclosure uses the image coding subnetwork as an example for illustration.

[0196] In some embodiments, the image coding subnetwork can use a CNN to extract image features from image information through processes such as convolution, feature extraction, and token transformation, and transform them into a fixed-length image feature vector. For example, when image information is input into the image coding subnetwork, the CNN performs convolution operations on the input image based on convolutional kernels through convolutional layers to generate a feature map; different convolutional kernels can extract different features, such as edges, textures, and shapes. As the convolutional layers deepen, the CNN can capture higher-level abstract features and construct complex feature representations by stacking convolutional layers, pooling layers (such as max pooling), and fully connected layers. Finally, the output of the CNN (such as the feature map of the last convolutional layer) is usually flattened into a one-dimensional vector, which can be regarded as image tokens, each token containing feature information of a local region in the image.

[0197] In some embodiments, the image coding subnetwork may also use Vision Transformer to extract image features from image information and transform them into fixed-length image feature vectors through processes such as image segmentation, linear embedding, positional encoding, and Vision Transformer encoding.

[0198] In some embodiments, the above-mentioned input of audio information into the audio coding sub-network in the coding network to generate an audio feature vector corresponding to the audio information includes: preprocessing the audio information to obtain preprocessed audio information; performing feature extraction on the preprocessed modal information to obtain an initial feature vector; and encoding the initial feature vector to obtain an audio feature vector.

[0199] In some embodiments, the audio coding subnetwork can employ an autoregressive model (such as WaveNet or Transformer architecture) to process sound waveforms. The audio coding subnetwork learns the intrinsic features of the audio signal through the model and encodes them into low-dimensional vectors. For example, the audio coding subnetwork can transform input audio information into tokens of length 42 (i.e., audio feature vectors).

[0200] In some embodiments, the preprocessing of the audio coding subnetwork may include processes such as sampling, quantization, noise reduction, and normalization. For example, the audio information (such as an audio signal) may first be sampled, that is, converted into a discrete time series. Then, noise and interference in the audio signal may be removed by filters, and the audio signal may be normalized to give it a uniform amplitude range, thus obtaining the preprocessed audio information.

[0201] In some embodiments, the feature extraction process of the audio coding subnetwork extracts meaningful features from the preprocessed audio information. These features can be used in the subsequent tokenization process. For example, Fourier transform, Mel Frequency Cepstral Coefficients (MFCC), and Linear Predictive Coding Coefficients (LPC) can be used to extract useful features from the preprocessed audio information. The extracted features are then combined into a feature vector, known as the initial feature vector. Processing the feature vector can reflect the temporal structure and frequency characteristics of the audio information.

[0202] In some embodiments, after determining the initial feature vector, the initial feature vector can be converted into fixed-length tokens. This process can also be called a token generation process, which converts the extracted features into discrete and finite tokens. For example, through step 1120 above, the multimodal information can be processed by the encoding network to output fixed-length feature vectors corresponding to each modality, as shown in Figure 9. The fixed-length feature vectors output by the encoding network 1510 can be input into the fusion network for further processing.

[0203] Step 1130: Based on the fusion network in the image generation model, the feature vectors corresponding to each modality information are fused to obtain the fused vector.

[0204] In some embodiments, a fusion network can also be referred to as a unified tokens encoding network. A fusion network can provide a common encoding space for different modalities of information (such as text, image, audio, and video), enabling multimodal information to be processed and analyzed within a unified framework. For example, a fusion network can handle the different characteristics of various modalities, such as the sequential nature of text, the two-dimensional nature of images, and the temporal sequence nature of audio.

[0205] For example, fusion networks according to some embodiments possess multimodal compatibility, feature extraction capabilities, and scalability. First, fusion networks can process data from different modalities and map them to a common space. Second, the powerful feature extraction capabilities of fusion networks can extract meaningful feature representations from raw data. Simultaneously, fusion networks are easily extensible to adapt to new data types and features. Fusion networks according to some embodiments can prioritize operational efficiency while ensuring performance, making them suitable for real-time processing and large-scale data analysis scenarios.

[0206] In some embodiments, step 1130 includes: determining a preset length; and based on the preset length, performing fusion processing on the feature vectors corresponding to each modality information through a fusion network to obtain a fusion vector of the preset length.

[0207] In some embodiments, the preset length can be set according to requirements, and the present disclosure does not limit the length of the fusion vector (i.e., the preset length). For example, the preset length can be 256, or it can be 1024. That is, after the feature vectors corresponding to each modality information are input into the fusion network, a fusion vector with a length of 256 can be generated after processing by the fusion network.

[0208] In some embodiments, a fusion network is used to perform feature selection processing on the feature vectors corresponding to each modality information to obtain selected feature vectors; the selected feature vectors are mapped to a preset feature space to obtain mapped feature vectors; and the mapped feature vectors are fused based on a preset length to obtain a fused vector of a preset length.

[0209] In some embodiments, the fusion process includes at least one of weighted fusion, feature-level fusion, and decision-level fusion.

[0210] In some embodiments, firstly, after obtaining the feature vectors (hereinafter also referred to as modal feature vectors) corresponding to each modality, feature selection processing can be performed on each modal feature vector to select the most useful features from each modality feature vector, resulting in selected feature vectors for each modality. Feature selection processing can reduce the dimensionality of the feature space, improving the model's generalization ability and computational efficiency. Secondly, after obtaining the selected feature vectors corresponding to each modality, feature mapping is performed on the selected feature vectors corresponding to each modality, mapping selected feature vectors of different lengths to a common feature space, resulting in mapped feature vectors. Finally, based on one or more combinations of feature fusion methods such as superposition, weighted averaging, feature-level fusion, or decision-level fusion, the mapped feature vectors can be fused into a single fused vector.

[0211] In some embodiments, the feature vector fusion process is illustrated using a feature vector in the form of [B, N, C]. Here, B represents the batch size, which is the number of batches during training or inference; N is the feature dimension; and C is the number of channel channels.

[0212] For example, when there are text feature vectors [B1, N1, C1], speech feature vectors [B2, N2, C2], and image feature vectors [B3, N3, C3], multiple feature vectors can be concatenated or weighted concatenation to merge them into a fused vector. That is: [B0, N0, C0] = w1[B1, N1, C1] + w2[B2, N2, C2] + w3[B3, N3, C3];

[0213] Where [B0, N0, C0] are the fusion vectors; w1, w2, w3 are the learnable coefficients.

[0214] For example, multiple feature vectors can be fused based on a cross-attention mechanism to generate a fused vector. This involves defining a query sequence, a key sequence, and a value sequence using text feature vectors, audio feature vectors, and image feature vectors respectively. Then, the similarity between the query sequence and the key sequence is calculated, and the similarity weights are normalized using the Softmax function. Finally, a weighted sum is performed based on the normalized weights to obtain a weighted value sequence. The weighted value sequence is then fused with the query sequence and / or key sequence to obtain the fused vector.

[0215] When the text feature vector is [B1, N1, C1], the speech feature vector is [B2, N2, C2], and the image feature vector is [B3, N3, C3], feature vector fusion can be performed based on the cross-attention mechanism, that is: [B0, N0, C0] = Trans_Attention([B1, N1, C1], [B2, N2, C2], [B3, N3, C3]);

[0216] Where [B0, N0, C0] is the fusion vector.

[0217] Step 1140: Based on multimodal information, the fusion vector is processed through a multi-stage network in the image generation model to obtain the target image corresponding to the multimodal information.

[0218] In some embodiments, the target image obtained by processing the fusion vector by the multi-stage network meets the requirements of multimodal information. The multi-stage network may include at least one feature subnetwork. The number of at least one feature subnetwork can be one or more, and when the number of at least one feature subnetwork is more than one, these feature subnetworks are connected sequentially.

[0219] For example, as shown in Figure 9, the multi-stage network 1530 may include N feature subnetworks, such as the first feature subnetwork, the second feature subnetwork, ..., and the Nth feature subnetwork (i.e. the last feature subnetwork), where N is any integer greater than or equal to 3, and the output of the previous feature subnetwork is used as the input of the next feature subnetwork connected to it.

[0220] In some embodiments, step 1140 may include steps 1141 to 1143 as shown below.

[0221] Step 1141: Determine at least one auxiliary feature information corresponding to the multimodal information.

[0222] In some embodiments, auxiliary feature information is information that assists the multi-stage network in extracting more accurate features from the fusion vector. The multi-stage network can extract the target vector from the fusion vector based on at least one auxiliary feature information, facilitating subsequent generation of the target image. This auxiliary feature information can also be called conditional guidance information. Auxiliary feature information is related to multimodal information; when the multimodal information input to the image generation model is different, the corresponding auxiliary feature information may also be different. For example, auxiliary feature information includes feature information corresponding to the guidance map, which can include images of different styles, content, structures, etc. For example, auxiliary feature information can be obtained by extracting image features from the guidance map through image encoding.

[0223] In some embodiments, the guiding auxiliary maps may include a human skeleton map, a depth map, and an edge structure map. The human skeleton map provides the basic skeletal structure of a person, guiding the posture and movements of the person in the generated image, ensuring that the generated human posture is natural and reasonable. The depth map represents the depth information of the image and can be used to create realistic 3D effects, making the image appear more layered. Especially when generating complex scenes such as landscapes and cityscapes, the depth map helps the model better understand the depth relationships of the scene, thereby generating more realistic images. The edge map emphasizes the contours and boundaries in the image, helping to maintain the sharpness and detail of the generated image. When generating intricate patterns, text, or objects that need to maintain a specific shape, the edge map ensures that the generated image has clear contours and accurate boundaries.

[0224] In some embodiments, auxiliary feature information can be determined based on the guidance auxiliary map. For example, feature extraction can be performed on the guidance auxiliary map to obtain auxiliary feature information corresponding to each guidance auxiliary map. After determining the auxiliary feature information, multiple auxiliary feature information are pre-stored in a guidance feature table. The guidance feature table can be pre-stored on a display device or server; this embodiment of the present disclosure does not limit this. Alternatively, the image generation model may also include a generation subnetwork to generate auxiliary feature information corresponding to each feature subnetwork and input it to the corresponding feature subnetwork. This embodiment of the present disclosure does not limit the generation and determination of auxiliary feature information.

[0225] In some embodiments, at least one auxiliary feature information corresponding to the multimodal information can be determined based on the multimodal information. For example, when the multimodal information involves a person, its corresponding auxiliary feature information includes the feature information of the human skeleton map; when the multimodal information involves a landscape, its corresponding auxiliary feature information includes the feature information corresponding to the edge map and the depth map. For example, as shown in Figure 5(a), when the multimodal information includes the image information displayed in the first input area (i.e., the image of a bird drawn by the user) and the text information displayed in the second input area (such as the first descriptive text: bird, blue feathers on the back, white feathers on the belly), the corresponding auxiliary feature information can be determined based on the image information and the text information to include at least the feature information corresponding to the bird skeleton structure map, the feature information corresponding to the edge map, the feature information corresponding to the depth map, and the feature information corresponding to the color map.

[0226] In some embodiments, the number of at least one auxiliary feature information in the multimodal information can be one or more. For example, the number of auxiliary feature information corresponding to the multimodal information is M, where M is an integer greater than or equal to 1. The number M of auxiliary feature information corresponding to the multimodal information can be less than the number N of feature subnetworks in the multi-stage network (i.e., M > N), or it can be greater than or equal to the number N of feature subnetworks.

[0227] In some embodiments, when a feature subnetwork corresponds to multiple auxiliary feature information, the multiple auxiliary feature information can be weighted and then input into the feature subnetwork for processing.

[0228] Step 1142: Process the fusion vector and at least one auxiliary feature information based on at least one feature subnetwork in the multi-stage network to obtain the target vector.

[0229] In some embodiments, after determining at least one auxiliary feature information, the at least one auxiliary feature information and the fusion vector determined in step 1143 can be input into at least one feature subnetwork in the multi-stage network to generate a target vector. The target vector can reflect the features of the multimodal information and the features of the guiding auxiliary graph.

[0230] In some embodiments, after determining the M auxiliary feature information corresponding to the multimodal information, the corresponding auxiliary feature information can be assigned to the N feature network sub-networks of the multi-stage network.

[0231] In some embodiments, each feature subnetwork includes a processing procedure with multiple steps, forming a chain structure to execute processing tasks of varying complexity corresponding to the feature subnetwork. The execution steps for each feature subnetwork can be the same or different. For example, when the complexity of the processing task corresponding to the feature subnetwork is high, it will have more steps.

[0232] Step 1142 above includes steps 11421 to 11422 as shown below. For ease of understanding, steps 11421 to 11422 are described using at least one auxiliary feature information including first auxiliary feature information, second auxiliary feature information, and third auxiliary feature information as an example.

[0233] Step 11421: Process the fusion vector and the first auxiliary feature information based on the first feature subnetwork in at least one feature subnetwork to determine the first output vector.

[0234] In some embodiments, the first feature subnetwork is the first feature subnetwork in a multi-stage network, that is, the first feature subnetwork in the multi-stage network 1530 in Figure 9. After determining that the auxiliary feature information corresponding to the first feature subnetwork is the first auxiliary feature information, the fusion vector output by the fusion network and the first auxiliary feature information can be input into the first feature subnetwork.

[0235] For example, the fusion vector (i.e., the token vector) can be aligned and random noise can be added to obtain the input vector. The input vector and the first auxiliary feature information can then be input into the first feature sub-network to obtain the first input vector. The random noise added to the aligned token vector can be Gaussian noise or other types of noise.

[0236] Step 11422: Process the first output vector and the second auxiliary feature information based on the intermediate feature subnetwork in at least one feature subnetwork to determine the second output vector.

[0237] In some embodiments, the number of intermediate feature subnetworks can be one or more. For example, an intermediate subnetwork may include a second feature subnetwork among multiple feature subnetworks, or it may include a second and a third feature subnetwork among multiple feature subnetworks, or it may include more feature subnetworks. That is, an intermediate feature subnetwork includes at least one feature subnetwork between the first and last feature subnetworks in a multi-stage network. This disclosure does not limit this aspect.

[0238] In some embodiments, after obtaining the first output vector from the first feature subnetwork, the first output vector and the auxiliary feature information corresponding to the intermediate feature subnetwork can be input into the intermediate feature subnetwork. After processing by the intermediate subnetwork, a second output vector is output. The second output vector is the feature vector output by the last intermediate feature subnetwork.

[0239] For example, when the intermediate subnetwork can include the second feature subnetwork up to the (N-1)th feature subnetwork, the second output vector is the vector output by the (N-1)th feature subnetwork.

[0240] In some embodiments, when the intermediate feature subnetwork includes the second to the (N-1)th feature subnetworks, each intermediate feature subnetwork may correspond to an auxiliary feature information (i.e., the second auxiliary feature information); or, some intermediate feature subnetworks may correspond to auxiliary feature information, while others may not have corresponding auxiliary feature information. For example, when N is 5, the intermediate subnetwork includes a second, a third, and a fourth feature subnetwork. The second feature subnetwork corresponds to one auxiliary feature information (such as the second auxiliary feature information), while the third and fourth feature subnetworks do not have corresponding auxiliary feature information. Therefore, the second feature subnetwork can process the first output vector and the second auxiliary feature information to output a first intermediate vector; the first intermediate vector is input to the third feature subnetwork, and after processing by the third feature subnetwork, a second intermediate vector is generated; the second intermediate vector is input to the fourth feature subnetwork, and after processing by the fourth feature subnetwork, a second output vector is generated.

[0241] Step 11423: Process the second output vector and the third auxiliary feature information based on the second feature subnetwork in at least one feature subnetwork to determine the target vector.

[0242] In some embodiments, the second feature subnetwork is the last feature subnetwork in the multi-stage network, i.e., the Nth feature subnetwork. After determining the second output vector output by the intermediate feature subnetwork (i.e., the (N-1)th feature subnetwork) and the third auxiliary information corresponding to the Nth feature subnetwork, the second output vector and the third auxiliary feature information are input into the Nth feature subnetwork to obtain the target vector.

[0243] Step 1143: Decode the target vector based on the image decoding subnetwork in the multi-stage network to obtain the target image.

[0244] In some embodiments, the multi-stage network further includes an image decoding sub-network (also known as an image decoder), which is used to convert the target vector into a target image that conforms to the multimodal information input by the user. For example, image decoding processing may include feature vector decoding, inverse quantization, inverse transform, and pixel reconstruction. The image decoding sub-network can recover pixel data close to the original image from the target vector.

[0245] Figure 10 is a schematic diagram of a multi-stage network structure according to some embodiments.

[0246] As shown in Figure 10, the multi-stage network 1530 can obtain the input vector Z0 based on the fusion vector and random noise, and then input the input vector Z0 and the first auxiliary feature information into the first feature sub-network. After processing by the first feature sub-network, the first output vector Z1 is obtained; then the first output vector Z1 and the second auxiliary feature information are input into the second feature sub-network to obtain the second output vector Z2.

[0247] Similarly, processing can continue through other feature subnetworks until the Nth feature subnetwork. The output vector from the previous feature subnetwork and the third auxiliary feature information can be input into the Nth feature subnetwork. After processing by the Nth feature subnetwork, the target vector Zn is output. Finally, the target vector Zn is input into the image decoding subnetwork for decoding to obtain the target image.

[0248] It's important to note that each feature subnetwork in the multi-stage network performs different tasks, such as style, content, and structure. The multi-stage network adopts an adapter-based model, adding elements like edge contours and color (i.e., auxiliary feature information) to enable individual and combined applications of multiple tasks. Furthermore, the multi-stage network defines a clear hierarchical relationship, with each layer (e.g., each feature subnetwork) receiving guidance from the upper layer, enabling fine-grained control. Within each feature subnetwork, a chain-like structure allows for generation at different stages, catering to tasks of varying complexity. This chain-like structure is primarily designed for interactive conditional input and supports controllable multi-stage generation. The intermediate control condition output is the implicit feature representation of the previous stage, and the dimensionality of the feature output remains consistent across all stages.

[0249] In some embodiments, the image generation method further includes: acquiring multimodal training information; processing the multimodal training information based on the image generation model to be trained to generate a predicted image; acquiring sample images; using the predicted image as the initial training output information of the image generation model to be trained and the sample image as the supervision information, iterating the image generation model to be trained to obtain an image generation model.

[0250] In some embodiments, the image generation model to be trained includes a coding network to be trained, a fusion network to be trained, and a multi-stage network to be trained. The image generation model obtained after iterative training includes the coding network, the fusion network, and the multi-stage network. That is, during the training process of the image generation model, it can be trained as a whole, or each network in the image generation model to be trained can be trained separately. This disclosure does not limit this approach.

[0251] In some embodiments, when training the image generation model to be trained using multimodal training information, the loss value of the image generation model to be trained can be determined based on the sample image and the predicted image, and the model parameters of the image generation model to be trained can be adjusted according to the loss value until a preset number of iterations is reached or the image generation model converges, thereby obtaining the trained image generation model (i.e., the image generation model in the above embodiments). The trained image generation model can be used to convert multimodal information into the corresponding target image.

[0252] Through the above process, after obtaining the multimodal information input by the user, the multimodal information is fed into the image generation model. It then passes through the encoding network, fusion network, and multi-stage network within the image generation model to obtain the target image that conforms to the multimodal information. This method enables the processing of multimodal information to generate images that meet the requirements, satisfying personalized user needs and improving the user experience.

Claims

1. A display device, comprising: The display is configured to show the application interface of an image generation application; wherein the application interface includes a first input area, a second input area, and an image display area; The controller, coupled to the display, is configured to: In response to a touch command in the first input area, the display is controlled to show a first image drawn by the user in the first input area; The first image is input into a trained image generation model to obtain a second image, and the display is controlled to show the second image in the image display area; In response to a command to acquire the first descriptive text, the display is controlled to show the first descriptive text in the second input area; the first descriptive text represents a textual description of the image to be generated by the user. The first descriptive text and the first image are input into the image generation model to obtain a third image that conforms to the first descriptive text, and the display is controlled to display the third image in the image display area; wherein, the image generation model is trained based on a multimodal training sample set; each multimodal training sample corresponds to a sample label, and the sample label is a representation of the target image to be generated by the expected model; each multimodal training sample includes at least one modal information among audio, text, image and video, and the audio and the text represent the generation requirements of the target image corresponding to the multimodal training sample.

2. The display device according to claim 1, wherein the controller executes the step of inputting the first image into a trained image generation model to obtain a second image, and controls the display to display the second image in the image display area, specifically configured as follows: The first image within the first input area is periodically detected. If the first image in the current detection period is different from the first image in the previous detection period, the first image in the current detection period is input into the image generation model to obtain the second image for the current detection period. The display is then controlled to show the second image corresponding to the current detection period in the image display area; or... After detecting that the touch event corresponding to the touch command is a touch press event, if an adjacent touch event is detected as a touch release event, then the first image is input into the image generation model to obtain the second image, and the display is controlled to display the second image in the image display area; or, In response to a user's confirmation command, the first image is input into the image generation model to obtain the second image, and the display is controlled to show the second image in the image display area.

3. The display device according to claim 1 or 2, wherein after the controller executes the command to control the display to show the third image in the image display area, it is further configured to: In response to a user's instruction to modify the first image input in the first input area, a fourth image is obtained; If the fourth image is different from the first image, the first descriptive text and the fourth image are input into the image generation model to obtain a fifth image that conforms to the first descriptive text, and the display is controlled to display the fifth image in the image display area.

4. The display device according to claim 1 or 2, wherein after the controller executes the command to control the display to show the third image in the image display area, it is further configured to: In response to a command to acquire the second descriptive text, the display is controlled to show the second descriptive text in the second input area; The second descriptive text and the first image are input into the image generation model to obtain a sixth image that conforms to the second descriptive text, and the display is controlled to display the sixth image in the image display area.

5. The display device according to claim 1 or 2, wherein the controller executes a retrieving instruction in response to the first descriptive text, controlling the display to display the first descriptive text in the second input area, specifically configured as follows: In response to the user's input of a descriptive speech, the descriptive speech is processed to obtain an audio spectrum; The system obtains the first descriptive text obtained by processing the audio spectrum based on a speech recognition model, and controls the display to show the first descriptive text in the second input area.

6. The display device according to claim 5, wherein the controller executes the command to control the display to show the first descriptive text in the second input area, specifically configured as follows: Get the number of characters in the third description text currently displayed in the second input area; If the third description text has zero characters, then the first description text is added to the second input area, and the display is controlled to show the first description text in the second input area; If the number of characters in the third description text is not zero, the first description text is added after the third description text to update the description text in the second input area, and the display is controlled to display the updated description text in the second input area.

7. A display device, comprising: The display is configured to show the application interface of an image generation application; wherein the application interface includes a first input area, a second input area, and an image display area; The controller, coupled to the display, is configured to: A seventh image is acquired, key feature points of the seventh image are extracted to obtain an eighth image, and the display is controlled to display the eighth image in the first input area; wherein, the seventh image is an image captured by an image acquisition device and / or an image uploaded by a mobile terminal; The eighth image is input into the trained image generation model to obtain the ninth image, and the display is controlled to display the ninth image in the image display area; In response to a command to acquire the first descriptive text, the display is controlled to show the first descriptive text in the second input area; the first descriptive text represents a textual description of the image to be generated by the user. The first descriptive text and the eighth image are input into the image generation model to obtain a tenth image that conforms to the first descriptive text, and the display is controlled to display the tenth image in the image display area; wherein, the image generation model is trained based on a multimodal training sample set; each multimodal training sample corresponds to a sample label, and the sample label is a representation of the target image to be generated by the expected model; each multimodal training sample includes at least one modal information among audio, text, image and video, and the audio and text represent the generation requirements of the target image corresponding to the multimodal training sample.

8. The display device according to claim 7, wherein the controller performing the acquisition of the seventh image is specifically configured as follows: In response to an image acquisition command triggered by a shooting interaction control in the application interface, the system receives an image captured by an image acquisition device and displays the image captured by the image acquisition device in the first input area. In response to a shooting command, the image captured by the user is acquired and used as the seventh image.

9. The display device according to claim 7, wherein the controller performing the acquisition of the seventh image is specifically configured as follows: In response to an image acquisition command input via the QR code scanning interaction control on the application interface, the display is controlled to show the identification code in the first input area; wherein, The identification code is used to upload images after scanning the code on the terminal device; The seventh image is received by the mobile terminal by scanning the identification code.

10. The display device according to claim 9, wherein the controller is specifically configured to receive the seventh image uploaded by the terminal device by scanning the identification code, wherein the controller performs the action of receiving the seventh image uploaded by the terminal device by scanning the identification code. Obtain the URL address sent by the server; where, The URL address is the storage address of the seventh image sent by the terminal device to the server after scanning the identification code; A request to receive an image is generated based on the URL address; The image receiving request is sent to the server, so that the server responds to the image receiving request by sending the seventh image; Receive the seventh image sent by the server.

11. The display device according to claim 7, wherein the controller performs the acquisition of the seventh image, extracts key feature points from the seventh image to obtain an eighth image, and controls the display to display the eighth image in the first input area, specifically configured as follows: Acquire the seventh image and control the display to show the seventh image in the first input area; The key feature points of the seventh image are extracted to obtain the eighth image, and the display is controlled to display the eighth image in the first input area.

12. The display device of claim 11, wherein the first input area includes a first option for displaying the seventh image and a second option for the user to display the eighth image; the controller is further configured to: In response to the focus selection of the first option, the display is controlled to show the seventh image in the first input area; In response to the focus switching to the second option, the display is controlled to show the eighth image in the first input area.

13. The display device according to claim 7, wherein the controller performs the extraction of key feature points from the seventh image to obtain an eighth image, specifically configured as follows: Determine the image type of the seventh image; If the image type of the seventh image is the first type, then the key feature points of the seventh image are extracted according to the skeleton feature extraction algorithm to obtain the eighth image; If the image type of the seventh image is the second type, then the key feature points of the seventh image are extracted according to the edge detection algorithm and the skeleton feature extraction algorithm to obtain the eighth image.

14. The display device according to any one of claims 7-13, wherein after the controller executes the command to control the display to show the tenth image in the image display area, it is further configured to: In response to the user's instruction to modify the eighth image input in the first input area, the eleventh image is obtained; If the eleventh image is different from the eighth image, the first descriptive text and the eleventh image are input into the image generation model to obtain a twelfth image that conforms to the first descriptive text, and the display is controlled to display the twelfth image in the image display area.

15. The display device according to any one of claims 7-13, wherein after the controller executes the command to control the display to show the tenth image in the image display area, it is further configured to: In response to a command to acquire the second descriptive text, the display is controlled to show the second descriptive text in the second input area; The second descriptive text and the eighth image are input into the image generation model to obtain a thirteenth image that matches the second descriptive text, and the display is controlled to display the thirteenth image in the image display area.

16. The display device according to any one of claims 7-13, wherein the controller executes the acquisition instruction in response to the first descriptive text and controls the display to display the first descriptive text in the second input area, specifically configured as follows: In response to the user's input of a descriptive speech, the descriptive speech is processed to obtain an audio spectrum; The audio spectrum is processed based on a speech recognition model to obtain the first descriptive text, and the display is controlled to show the first descriptive text in the second input area.

17. The display device according to claim 16, wherein the controller executes the command to control the display to display the first descriptive text in the second input area, specifically configured as follows: Get the number of characters in the third description text currently displayed in the second input area; If the third description text has zero characters, then the first description text is added to the second input area, and the display is controlled to show the first description text in the second input area; If the number of characters in the third description text is not zero, the first description text is added after the third description text to update the description text in the second input area, and the display is controlled to display the updated description text in the second input area.

18. An image display method, applied to the display device according to any one of claims 1-6, comprising: In response to a touch command in the first input area, a first image drawn by the user is displayed in the first input area; The first image is input into the trained image generation model to obtain the second image, which is then displayed in the image display area. In response to the instruction to acquire the first descriptive text, the first descriptive text is displayed in the second input area; the first descriptive text represents the user's textual description of the image to be generated. The first descriptive text and the first image are input into the image generation model to obtain a third image that conforms to the first descriptive text, and the third image is displayed in the image display area; wherein, the image generation model is trained based on a multimodal training sample set; each multimodal training sample corresponds to a sample label, and the sample label represents the target image to be generated by the expected model; each multimodal training sample includes at least one modal information among audio, text, image and video, and the audio and the text represent the generation requirements of the target image corresponding to the multimodal training sample.

19. An image display method, applied to the display device according to any one of claims 7-17, comprising: A seventh image is acquired, and key feature points of the seventh image are extracted to obtain an eighth image, which is then displayed in the first input area; wherein the seventh image is an image captured by an image acquisition device and / or an image uploaded by a mobile terminal; The eighth image is input into the trained image generation model to obtain the ninth image, which is then displayed in the image display area. In response to the instruction to acquire the first descriptive text, the first descriptive text is displayed in the second input area; the first descriptive text represents the user's textual description of the image to be generated. The first descriptive text and the eighth image are input into the image generation model to obtain a tenth image that conforms to the first descriptive text, and the display is controlled to display the tenth image in the image display area; wherein, the image generation model is trained based on a multimodal training sample set; each multimodal training sample corresponds to a sample label, and the sample label is a representation of the target image to be generated by the expected model; each multimodal training sample includes at least one modal information among audio, text, image and video, and the audio and text represent the generation requirements of the target image corresponding to the multimodal training sample.