Electronic device and method for controlling image content displayed thereon
The electronic device uses a remote coordinating procedure with AI models to control ChLC displays via speech commands, addressing slow response and inefficient interfaces, offering hands-free operation and reduced power consumption.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2026-04-02
AI Technical Summary
Existing cholesteric liquid crystal (ChLC) displays face challenges in displaying arbitrary content instantaneously and have slow response speeds, with inefficient user interfaces like pen input or software keyboards complicating control mechanisms, particularly in devices like e-paper and digital photo frames.
An electronic device equipped with a microphone, processor, and display controller performs a remote coordinating procedure using a large language model, speech recognition model, and image generation model to generate control signals for displaying images based on speech commands, with the display turning on during updates and off after completion, allowing power-saving operation.
Enables efficient, hands-free control of image content on ChLC displays with reduced power consumption by turning the display on only during updates and maintaining the displayed image even when powered off, enhancing user experience and reducing energy usage.
Smart Images

Figure CN2024120954_02042026_PF_FP_ABST
Abstract
Description
ELECTRONIC DEVICE AND METHOD FOR CONTROLLING IMAGE CONTENT DISPLAYED THEREONTECHNICAL FIELD
[0001] The present disclosure relates to display devices, and, in particular, to an electronic device and a method for controlling image content displayed thereon.DESCRIPTION OF THE RELATED ART
[0002] A cholesteric liquid crystal (ChLC) display exhibits bi-stable characteristics, allowing it to conserve power by maintaining the display or information without the need for a continuous electric field. ChLC technology can be utilized in various applications, including temperature display boards, digital picture books, e-books, e-paper, and electronic whiteboards.
[0003] Digital picture books and digital photo frames are anticipated to be prominent applications for e-paper technology, primarily due to its ultra-low power consumption. However, there has been no straightforward method to display arbitrary content instantaneously. Furthermore, the slow response speed of e-paper presents challenges. For example, pen input or software keyboard input is not only inefficient as a user interface for e-paper, which is predominantly used for displaying static images, but also complicates the control mechanisms required to drive the e-paper.SUMMARY
[0004] An aspect of the present disclosure provides an electronic device, which includes a microphone, a processor, a display controller, and a display device. The microphone is configured to receive a speech signal from a user of the electronic device. The processor is configured to perform a remote coordinating procedure based on the speech signal between a large language model, a speech recognition model, and an image generation model. The display controller is configured to generate control signals based on an image signal obtained from the remote coordinating procedure. The display device is configured to render a display image based on the control signals and the image signal. The display device is turned on during updating the display image and turned off upon completion of updating the display image.
[0005] Another aspect of the present disclosure provides a method for controlling image content displayed on an electronic device, which includes a microphone, a processor, a display controller, and a display device. The method includes the following steps: utilizing the microphone to receive a speech signal from a user of the electronic device; utilizing the processor to perform a remote coordinating procedure, based on the speech signal, between a large language model (LLM) , a speech recognition model, and an image generation model; utilizing the display controller to generate control signals based on an image signal obtained from the remote coordinating procedure; and utilizing the display device to render a display image based on the control signals and the image signal. The display device is turned on during updating the display image, and turned off upon completion of updating the display image.BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Aspects of the present disclosure are best understood from the following detailed description when read with the accompanying figures. It is noted that in accordance with the standard practice in the industry, various features are not drawn to scale. In fact, the dimensions of the various features may be arbitrarily increased or reduced for clarity of discussion.
[0007] FIG. 1 is a block diagram of an electronic device in accordance with some embodiments of the present disclosure.
[0008] FIG. 2 is a flowchart illustrating the remote coordinating procedure between different AI models in accordance with some embodiments of the present disclosure.
[0009] FIG. 3 is a diagram illustrating a candidate image in accordance with some embodiments of the present disclosure.
[0010] FIG. 4 is a diagram illustrating the procedure for displaying a display image on the ChLC display panel in a particular display order in accordance with some embodiments of the present disclosure.
[0011] FIG. 5 is a flowchart of a method for controlling image content displayed on an electronic device in accordance with some embodiments of the present disclosure.DETAILED DESCRIPTION
[0012] The following disclosure provides many different embodiments, or examples, for implementing different features of the provided subject matter. Specific examples of operations, components, and arrangements are described below to simplify the present disclosure. These are, of course, merely examples and are not intended to be limiting. For example, a first operation performed before or after a second operation in the description may include embodiments in which the first and second operations are performed together, and may also include embodiments in which additional operations may be performed between the first and second operations. For example, the formation of a first feature over, on or in a second feature in the description that follows may include embodiments in which the first and second features are formed in direct contact, and may also include embodiments in which additional features may be formed between the first and second features, such that the first and second features may not be in direct contact. In addition, the present disclosure may repeat reference numerals and / or letters in the various examples. This repetition is for the purpose of simplicity and clarity and does not in itself dictate a relationship between the various embodiments and / or configurations discussed.
[0013] Time relative terms, such as "prior to, " "before, " "posterior to, " "after" and the like, may be used herein for ease of description to describe the relationship of one operation or feature to another operation (s) or feature (s) as illustrated in the figures. Such time relative terms are intended to encompass different sequences of the operations depicted in the figures. Further, spatially relative terms, such as "beneath, " "below, " "lower, " "above, " "upper" and the like, may be used herein for ease of description to describe the relationship of one element or feature to another element (s) or feature (s) as illustrated in the figures. Such spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. The apparatus may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein may likewise be interpreted accordingly. Relative terms for connections, such as "connect, " "connected, " "connection, " "couple, " "coupled, " "in communication, " and the like, may be used herein for ease of description to describe an operational connection, coupling, or linking one between two elements or features. The relative terms for connections are intended to encompass different connections, couplings, or links of the devices or components. The devices or components may be directly or indirectly connected, coupled, or linked to one another through, for example, another set of components. The devices or components may be connected, coupled, or linked with each other by wire and / or wirelessly.
[0014] As used herein, the singular terms "a, " "an, " and "the" may include plural referents unless the context clearly indicates otherwise. For example, reference to a device may include multiple devices unless the context clearly indicates otherwise. The terms "comprising" and "including" may indicate the existences of the described features, integers, steps, operations, elements, and / or components, but may not exclude the existence of combinations of one or more of the features, integers, steps, operations, elements, and / or components. The term "and / or" may include any or all combinations of one or more listed items.
[0015] Additionally, amounts, ratios, and other numerical values are sometimes presented herein in a range format. It is to be understood that such range format is used for convenience and brevity and should be understood flexibly to include numerical values explicitly specified as limits of a range, but also to include all individual numerical values or sub-ranges encompassed within that range as if each numerical value and sub-range is explicitly specified.
[0016] FIG. 1 is a block diagram of an electronic device in accordance with some embodiments of the present disclosure.
[0017] In some embodiments, the electronic device 1 may be a portable device, such as an e-paper, e-book, digital picture book, etc., equipped with a cholesteric liquid crystal (ChLC) display device, and capable of generating an image signal (e.g., a color image signal or a grayscale image signal) based on a text command or a speech command by performing a remote coordinating procedure between a plurality of remote artificial intelligence (AI) models of different types. The electronic device 1 includes a processor 102, a display controller 104, a memory 106, a microphone 108, a network interface 110, one or more peripheral device 112, a display device 120, and a storage device 130 connected to each other through bus 111, as depicted in FIG. 1.
[0018] In some embodiments, the processor 102 is configured to control the overall operation of the electronic device 1. The processor 102 may be implemented using various types of processors, such as a central processing unit (CPU) , a microprocessor, a multimedia processor, or a graphics processor, but the present disclosure is not limited thereto. In an embodiment, the processor 102 may be implemented using as an integrated circuit, a mobile application processor (AP) , or a system on chip (SoC) . In some embodiments, the processor 102 may be regarded as an artificial intelligence (AI) coordinator that is configured to execute the coordinating program 132 to perform a remote coordinating procedure, based on an input command, between a large language model (LLM) 141, a speech recognition model 142, and an image generation model 143.
[0019] In some embodiments, the display controller 104 may control the overall operation of the display device 120. For example, the display controller 104 may be a general-purpose processor, a microcontroller, or a micro control unit (MCU) that is configured to generate control signals based on an image signal from the processor 102 obtained from the remote coordinating procedure between the AI models. In some embodiments, the display controller 104 may be a standalone controller disposed in the electronic device 1. Alternatively, the display controller 104 may be integrated into the display device 120.
[0020] In some embodiments, the display controller 104 may include a timing controller (TCON) which may receive the image signal from the processor 102. For example, the image signal may be in the form of a stream and may include timing information. The timing controller may generate, based on the timing information, control signals for controlling the driving circuit 121.
[0021] In some embodiments, memory 106 may include a non-volatile memory (NVM) and a volatile memory (both not shown) that are configured to store intermediate data generated by the processor 102 during execution of the application program 131 and the coordinating program 132.
[0022] In some embodiments, the microphone 108 may be configured to detect an acoustic signal around the electronic device 1, where the acoustic signal may include a speech signal indicating an input command to the electronic device. The microphone 108 may transmit the received acoustic signal and / or speech signal to the processor 102 for subsequent processes.
[0023] In some embodiments, the electronic device 1 may communicate with network 140 through the network interface 110. The network interface 110 may employ a plurality of wired and / or wireless communication protocols and / or technologies. Examples of various generations (e.g., third (3G) , fourth (4G) , or fifth (5G) ) of communication protocols and / or technologies that may be employed by the network may include, but are not limited to, Global System for Mobile communication (GSM) , General Packet Radio Services (GPRS) , Enhanced Data GSM Environment (EDGE) , Code Division Multiple Access (CDMA) , Wideband Code Division Multiple Access (W-CDMA) , Code Division Multiple Access 2000 (CDMA2000) , High Speed Downlink Packet Access (HSDPA) , Long Term Evolution (LTE) , Universal Mobile Telecommunications System (UMTS) , Evolution-Data Optimized (Ev-DO) , Worldwide Interoperability for Microwave Access (WIMAX) , time division multiple access (TDMA) , Orthogonal frequency-division multiplexing (OFDM) , ultra-wide band (UWB) , Wireless Application Protocol (WAP) , user datagram protocol (UDP) , transmission control protocol / Internet protocol (TCP / IP) , various portions of the Open Systems Interconnection (OSI) model protocols, session initiated protocol / real-time transport protocol (SIP / RTP) , short message service (SMS) , multimedia messaging service (MMS) , or various ones of a variety of other communication protocols and / or technologies.
[0024] In some embodiments, one or more peripheral devices 112 may be electrically coupled to the processor 102 through bus 111. For example, the one or more peripheral devices 112 may include a stylus, a keyboard, a touch-sensitive device, etc., but the present disclosure is not limited thereto. For example, a user can operate the electronic device 1 using one or more gestures on the user interface rendered by the ChLC display panel 122 through the one or more peripheral devices 112.
[0025] In some embodiments, the display device 120 may be a ChLC display device, which includes a driving circuit 121, a ChLC display panel 122, and a touch module 123. The ChLC display panel 122 may be a ChLC color display panel which includes a blue ChLC display panel, a green ChLC display panel, and a red ChLC display panel (not shown in FIG. 1) in a stacked structure.
[0026] In some embodiments, the driving circuit 121 may receive control signals and the image signal from the display controller 104 to control the ChLC display panel to render a display image based on the control signals and the image signal. For example, the driving circuit 121 may provide voltages to gate lines and data lines of the ChLC display panel 122 in response to the control signals. The touch module 123 may be a touch-sensitive surface disposed on or integrated into the ChLC display panel 122, and it is configured to detect the location and movement of a user’s finger or stylus on the touch-sensitive surface. The touch module 123 may support various touch gestures, such as tapping, swiping, and pinching, to perform different functions. In some embodiments, the technology used by the touch module 123 can vary, including resistive, capacitive, infrared, surface acoustic wave methods, each offering different levels of sensitivity, accuracy, and durability. In some embodiments, the ChLC display panel 122 and the touch module 123 can be integrated into a touch display panel 124.
[0027] In some embodiments, the storage device 130 may be a non-volatile memory, such as a hard disk drive, a solid-state disk, a flash memory, a read-only memory, etc., but the present disclosure is not limited thereto. The storage device 130 may store an application program 131, a coordinating program 132, and an image database 133. When the processor 102 executes the application program 131, the electronic device 1 may serve as a digital picture book or a digital photo frame. When the processor 102 executes the coordinating program 132, the processor 102 may act as the AI controller that coordinates between the LLM 141, speech recognition model 142, and the image generation model 143. Additionally, the image database 133 may be configured to store one or more candidate images that are generated by processor 102 and / or the display controller 104 during the remote coordinating procedure, the details of which will be described later.
[0028] In some embodiments, the speech recognition model 142 is an artificial intelligence model designed to convert spoken language into text. The speech recognition model 142 employs advanced algorithms and machine learning techniques to analyze and interpret audio signals. For example, the speech recognition model 142 may receive the speech command from the processor 102 via network 140, and recognize the textual information within the speech command. In some embodiments, the LLM 141 is an artificial intelligence model designed to understand and generate human language. The LLM 141 is trained on vast amounts of text data, enabling it to predict and product coherent and contextually relevant text based on the text input it receives. For example, the LLM 141 may identify the title string, the display-image instruction, and display-position instruction from the textual information associated with the speech command.
[0029] In some embodiments, the image generation model 143 is an artificial intelligence model designed to create visual content from textual description or other input data. Utilizing advanced machine learning techniques, particularly deep learning, the image generation model 143 is trained on vast datasets comprising various images and corresponding annotations. Through this training, the image generation model 143 learns to understand and replicate intricate patterns, textures, and structures found in the visual data. It employs neural networks, often Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs) , to generate high-quality, realistic images. In some embodiments, the image generation model 143 could be any existing image generation model, such as Stable Diffusion, Dall-E, etc., but the present disclosure is not limited thereto.
[0030] FIG. 2 is a flowchart illustrating the remote coordinating procedure between different AI models in accordance with some embodiments of the present disclosure. Please refer to both FIG. 1 and FIG. 2.
[0031] In some embodiments, a user can instruct the electronic device 1 to generate a display image with a designated display position using a speech command. The speech command may include commands for generating a display image, a display position of the display image, and title indication of the display image. In some embodiments, the speech command may further include image correction options of the display image and / or the display order and display interval of the playlist while displaying the display image on the ChLC display panel 122. The microphone 108 of the electronic device 1 may receive the speech command, and then forward the received speech command to the processor 102 (step 202) . Here, the processor 102 may act as an AI controller that performs the remote coordinating procedure.
[0032] For example, the processor 102 may transmit the speech command to the speech recognition model 142 via network 140 (step 204) . Upon inputting the speech command to the speech recognition model 142, the speech recognition model 142 may recognize textual information within the speech command, and return the recognized textual information back to the processor 102 (step 206) . Afterwards, the processor 102 may forward the recognized textual information to the LLM 141 via network 140 (step 208) to recognize a title string, a display-image instruction, and a display-position instruction from the textual information. The LLM 141 may then return the recognized title string, display-image instruction, and display-position instruction back to the processor 102 (step 210) . Thereafter, the processor 102 may generate display coordinate data based on the display-position instruction. For example, the display coordinate data may include position information regarding the selected candidate image (e.g., the target image) or the display image, such as at the upper middle, upper left, upper right, bottom middle, bottom left, bottom right, centralized, full screen, etc. In some cases, the display coordinate data may also include the orientation of the display image, such as portrait or landscape.
[0033] In some embodiments, the processor 102 may forward the display-image instruction to the image generation model 143 via network 140 (step 212) , allowing the image generation model 143 to generate one or more candidate images (e.g., images 401 to 40N shown in FIG. 4) based on the display-image instruction. For example, the display-image instruction may include the string “Acat sitting in the street” , and the image generation model 143 may generate one or more candidate images associated with the content or scene indicated by the string of the display-image instruction. FIG. 3 illustrates that image 300 represents one of the candidate images. The image generation model 143 may then transmit the one or more candidate images to the processor 102 via network 140 (step 214) . The user may select all or a portion of the one or more candidate images via the peripheral device 112 (or by another speech command) , and the processor 102 may store the selected candidate images (e.g., selected candidate images 1 to M shown in FIG. 4) in a particular album of the image database 133.
[0034] In some embodiments, the image parameters, such as color balance (e.g., hue, saturation, value) , sharpness, color temperature, etc., of the candidate images can be changed via image correction options set by a speech command or by the peripheral device 112 operating on the user interface displayed on the ChLC display panel 122. In the first example, the user may send a speech command “show warm colors” to the electronic device 1, and the processor 102 may perform an image correction process on the target image or the candidate images with a lower color temperature. In the second example, the user may send a speech command “show a clearer image” to the electronic device 1, and the processor 102 may perform an image correction process on the target image or the candidate images with a higher sharpness.
[0035] In some embodiments, when the electronic device 1 acts as a digital photo frame, the processor 102 may select one of the candidate images stored in the particular album of the image database 133 as a target image according to a particular playlist or a default playlist. The processor 102 may send the target image to the display controller 104 (step 216) , allowing the display controller 104 to generate the display image (e.g., image 430 shown in FIG. 4) by synthesizing the target image (e.g., image 300 shown in FIG. 3) with a title image (e.g., image 420 show in FIG. 4) including the title string indicated in the speech command. Specifically, the display controller 104 may generate a title image (e.g., image 420 show in FIG. 4) based on the title string indicated in the speech command, and overlay the title image on the target image (e.g., image 300 shown in FIG. 3) according to the position information of the target image to generate the display image (e.g., image 430 shown in FIG. 4) . In some embodiments, the title string may include a title and its color and font of the one or more candidate images. In case that the color and font of the title string are not available, a default color (e.g., white) and a default font (e.g., Calibri) can be used to generate the title string within the title image, but the present disclosure is not limited thereto.
[0036] Subsequently, the display device 120 may display the display image on the ChLC display panel 122 according to the control signals and the image signal from the display controller 104 (step 218) . For example, the playlist may include a display order and a display interval associated with the display image displayed on the ChLC display panel 122. For example, the processor 102 or the display controller 104 may periodically update the target image using the candidate image with the designated image number, such as image number 08, 10, 06, 07, and 05 shown in the playlist 410 in FIG. 4, indicated in the playlist per display interval (e.g., 30 minutes) . The display order of the candidate images stored in the image database 133 and the display interval in the playlist can be changed via a speech command or by the peripheral device 112 operating on the user interface displayed on the ChLC display panel 122.
[0037] FIG. 5 is a flowchart of a method for controlling image content displayed on an electronic device in accordance with some embodiments of the present disclosure. Please refer to both FIG. 1 and FIG. 5. The method 500 shown in FIG. 5 includes steps 510 to 540.
[0038] Step 510: Utilizing the microphone 108 to receive a speech signal from a user of the electronic device 1. In some embodiments, the microphone 108 may be configured to detect an acoustic signal around the electronic device 1, where the acoustic signal may include a speech signal indicating an input command to the electronic device. The microphone 108 may transmit the received acoustic signal and / or speech signal to the processor 102 for subsequent processes.
[0039] Step 520: Utilizing the processor 102 to perform a remote coordinating procedure, based on the speech signal, between a large language model (LLM) 141, a speech recognition model 142, and an image generation model 143. In some embodiments, the processor 102 may act as an AI controller that performs the remote coordinating procedure between the LLM 141, the speech recognition model 142, and the image generation model 143, the details of which can be referred to as the embodiment of FIG. 2.
[0040] Step 530: Utilizing the display controller 104 to generate control signals based on an image signal obtained from the remote coordinating procedure between the AI models. In some embodiments, the display controller 104 may be configured to receive the target image (e.g., the selected candidate image) from the processor 102, and generate an image signal and its corresponding control signals based on the target image. The display controller 104 may further send the image signal and its corresponding control signals to the driving circuit 121 for rendering the target image on the ChLC display panel 122.
[0041] Step 540: Utilizing the cholesteric liquid crystal display panel 122 to render a display image based on the control signals and the image signal. In some embodiments, the driving circuit 121 may receive control signals and the image signal from the display controller 104 to control the ChLC display panel to render a display image based on the control signals and the image signal. For example, the driving circuit 121 may provide voltages to gate lines and data lines of the ChLC display panel 122 in response to the control signals. It should be noted that the display device 120 is turned on during updating the display image, and turned off upon completion of updating the display image. Additionally, the displaying of the display image on the ChLC display panel 122 of the display device 120 is maintained when the display device 120 is turned off.
[0042] Accordingly, an electronic device and a method for controlling image content displayed thereon are provided. The electronic device, such as an e-paper, is capable of receiving a speech command and perform a remote coordinating procedure between different types of AI models, thereby obtaining candidate images associated with the scenario indicated in the speech command and periodically displaying the candidate images on the ChLC display panel of the electronic device in a particular display order. Additionally, the image parameters of the candidate image displayed on the electronic device can also adjusted via another speech command, thereby achieving improved hand-free user experience. Furthermore, the display device within the electronic device is turned on when receiving a speech command for updating the display screen on the ChLC display panel, and turned on upon completion of the updating of the display screen, and the displaying of the updated display screen is maintained even though the power to the display device is turned off, thereby reducing power consumption of the electronic device.
[0043] The nature and use of the embodiments are discussed in detail as follows. It should be appreciated, however, that the present disclosure provides many applicable inventive concepts that can be embodied in a wide variety of specific contexts. The specific embodiments discussed are merely illustrative of specific ways to embody and use the disclosure, without limiting the scope thereof.
Claims
1.An electronic device, comprising:a microphone, configured to receive a speech signal from a user of the electronic device;a processor, configured to perform a remote coordinating procedure, based on the speech signal, between a large language model (LLM) , a speech recognition model, and an image generation model;a display controller, configured to generate control signals based on an image signal obtained from the remote coordinating procedure; anda display device configured to render a display image based on the control signals and the image signal,wherein the display device is turned on during updating the display image, and turned off upon completion of updating the display image.2.The electronic device of Claim 1, wherein:the display device comprises a cholesteric liquid crystal display panel to render the display image; andduring the remote coordinating procedure, the processor is configured to perform the following operations:inputting the speech signal to the speech recognition model to recognize textual information within the speech signal;forwarding the textual information to the large language model (LLM) to recognize a title string, a display-image instruction, and a display-position instruction from the textual information;generating display coordinate data based on the display-position instruction;forwarding the display-image instruction to the image generation model to generate one or more candidate images based on the display-image instruction;transmitting the title string, the display coordinate data, and the one or more candidate images to the display controller; andgenerating the image signal based on the one or more candidate images.3.The electronic device of Claim 2, wherein the title string comprises a title and its color and font of the one or more candidate images, and the display coordinate data comprises position information about the display image, which is selected from the one or more candidate images, displayed on the display panel.4.The electronic device of Claim 3, wherein the image generation model is a text-to-image artificial-intelligence model.5.The electronic device of Claim 4, wherein the display-image instruction comprises a description of the one or more candidate images.6.The electronic device of Claim 2, wherein the display controller stores all or a portion of the one or more candidate images in an image database of the electronic device.7.The electronic device of Claim 6, wherein the display controller generates a title image based on the title string, and overlays the title image on one of the one or more candidate images selected from the image database to generate the display image.8.The electronic device of Claim 7, wherein the display controller performs an image-correction process on the display image to adjust one or more image parameters of the display image based on an image-correction instruction recognized from the speech signal by the speech recognition model and the LLM.9.The electronic device of Claim 8, wherein the one or more image parameters comprises any one of an HSV (hue-saturation-value) color space, a sharpness, a contrast, and a color temperature of the display image.10.The electronic device of Claim 7, wherein the display controller builds a playlist of the one or more candidate images stored in the image database based on a speech command recognized from the speech signal by the speech recognition model and the LLM to periodically display each candidate image stored in the image database on the display panel.11.A method for controlling image content displayed on an electronic device, wherein the electronic device comprises a microphone, a processor, a display controller, and a display device, the method comprising:utilizing the microphone to receive a speech signal from a user of the electronic device;utilizing the processor to perform a remote coordinating procedure, based on the speech signal, between a large language model (LLM) , a speech recognition model, and an image generation model;utilizing the display controller to generate control signals based on an image signal obtained from the remote coordinating procedure; andutilizing the display device to render a display image based on the control signals and the image signal,wherein the display device is turned on during updating the display image, and turned off upon completion of updating the display image.12.The method of Claim 11, wherein:the display device comprises a cholesteric liquid crystal display panel to render the display image; andduring the remote coordinating procedure, the method further comprises:inputting, by the processor, the speech signal to the speech recognition model to recognize textual information within the speech signal;forwarding, by the processor, the textual information to the large language model (LLM) to recognize a title string, a display-image instruction, and a display-position instruction from the textual information;generating, by the processor, display coordinate data based on the display-position instruction;forwarding, by the processor, the display-image instruction to the image generation model to generate one or more candidate images based on the display-image instruction;transmitting, by the processor, the title string, the display coordinate data, and the one or more candidate images to the display controller; andgenerating, by the processor, the image signal based on the one or more candidate images.13.The method of Claim 12, wherein the title string comprises a title of the one or more candidate images, and the display coordinate data comprises position information about the display image, which is selected from the one or more candidate images, displayed on the display panel.14.The method of Claim 13, wherein the image generation model is a text-to-image artificial-intelligence model.15.The method of Claim 14, wherein the display-image instruction comprises a description of the one or more candidate images.16.The method of Claim 12, further comprising: utilizing the display controller to store all or a portion of the one or more candidate images in an image database of the electronic device.17.The method of Claim 16, further comprising: utilizing the display controller to generate a title image based on the title string, and to overlay the title image on one of the one or more candidate images selected from the image database to generate the display image.18.The method of Claim 17, further comprising: utilizing the display controller to perform an image-correction process on the display image to adjust one or more image parameters of the display image based on an image-correction instruction recognized from the speech signal by the speech recognition model and the LLM.19.The method of Claim 18, wherein the one or more image parameters comprises any one of an HSV (hue-saturation-value) color space, a sharpness, a contrast, and a color temperature of the display image.20.The method of Claim 17, further comprising: utilizing the display controller to build a playlist of the one or more candidate images stored in the image database based on a speech command recognized from the speech signal by the speech recognition model and the LLM to periodically display each candidate image stored in the image database on the display panel.
Citation Information
Patent Citations
Multi-modal image color divider and editor
CN115249273A
Image generation method and device, computer equipment and storage medium
CN116363242A
Artificial intelligence poem artistic conception image generation and construction method based on semantic ontology
CN118537445A
Multi-modal interaction method and device
CN118585637A