Display device, server and picture book generation method

By presetting picture book applications in the display device, the automatic generation and editing of picture books is realized, which solves the problem that traditional display devices cannot meet users' personalized needs and realizes user-defined picture book display.

CN119991868APending Publication Date: 2025-05-13HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411980675.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Traditional display devices cannot automatically generate and modify picture books, and cannot meet users' personalized needs.

Method used

By presetting picture book applications in the display device, the automatic generation and editing of picture books can be realized. Users can generate and modify picture books by entering text and images, display differentiate text and image display areas, edit controls to obtain user input content, generate target images and insert them into picture books.

Benefits of technology

It realizes the automatic generation and editing of picture books based on user needs, meets user personalized needs, and displays picture books content that meets users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991868A_ABST
    Figure CN119991868A_ABST
Patent Text Reader

Abstract

The invention discloses a display device, a server and a picture book generation method, and the method comprises the steps: the display device displays a function page of a picture book application on a display in response to an instruction of opening the picture book application; in response to a picture book generation instruction triggered on the functional page, generating a target picture book, the target picture book comprising at least one picture book page with a page order; displaying a picture book playing page on the display, wherein a picture book display area of the picture book playing page is used for displaying at least one picture book page according to a page sequence; in response to a trigger operation on a picture book editing control in the picture book playing page, acquiring newly-added text content input by a user, and generating a target image corresponding to the newly-added text content based on the newly-added text content and a picture book style corresponding to the target picture book; forming a newly added picture book page by the newly added text content and the target image; and inserting the newly added picture book page into the target picture book to obtain an updated picture book, and displaying the updated picture book in the picture book display area. According to the method, the picture book meeting the personalized requirements of the user can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of display devices, and in particular to a display device, a server, and a picture book generation method. Background Art

[0002] With the rapid development of display devices, the functions that display devices can provide to users are becoming more and more abundant. At present, display devices include smart TVs, smart set-top boxes, smart boxes, and products with smart display screens. Taking smart TVs as an example, smart TVs enable more and more scenarios. They are not only used as devices for watching TV programs at home, but also for playing games, playing electronic photo albums, and displaying information.

[0003] Users can view preset images or videos through the display device. For example, users can view preset picture books through the display device, but cannot view non-preset picture books, let alone modify the picture books. Therefore, traditional display devices have the problem of being unable to display picture books that meet the user's personalized needs. Summary of the invention

[0004] The present application provides a display device, a server and a picture book generation method, which can automatically generate a target picture book according to the user's picture book generation requirements, and rewrite or continue the generated target picture book according to the user's picture book editing requirements, so as to display a picture book that meets the user's personalized needs on the display device.

[0005] In a first aspect, a display device is provided, comprising:

[0006] a display configured to display content from a broadcast system or network and / or a user interface;

[0007] and at least one processor connected to the display and configured to execute instructions to cause the display device to:

[0008] In response to an instruction to open the picture book application, displaying a function page of the picture book application on the display;

[0009] In response to a picture book generation instruction triggered on a function page, a target picture book is generated, the target picture book includes at least one picture book page with a page sequence, each picture book page includes a picture book text and a picture book image; a picture book playback page is displayed on a display, the picture book playback page includes a picture book display area and a picture book editing control; the picture book display area is used to display at least one picture book page according to the page sequence;

[0010] In response to a trigger operation on a picture book editing control, newly added text content input by the user is obtained, and based on the newly added text content and the picture book style corresponding to the target picture book, a target image corresponding to the newly added text content is generated; the newly added text content and the target image are combined into a newly added picture book page; the newly added picture book page is inserted into the target picture book to obtain an updated picture book, and the updated picture book is displayed in a picture book display area.

[0011] In some embodiments, the picture book display area includes a text display area and an image display area, the text display area is used to display the picture book text in the picture book page, and the image display area is used to display the picture book image in the picture book page;

[0012] When at least one processor executes a trigger operation on the picture book editing control to obtain the newly added text content input by the user, the processor is further configured to:

[0013] In response to the triggering operation on the picture book editing control, the target picture book text currently displayed in the text display area is cleared, the target text input by the user in the text display area is received, and the target text is used as the newly added text content.

[0014] In some embodiments, when at least one processor executes a trigger operation on a picture book editing control to obtain the newly added text content input by the user, the processor is further configured to:

[0015] In response to the triggering operation of the picture book editing control, a voice input prompt is displayed on the picture book playback page; the interactive voice input by the user is received, the interactive voice is converted into text, and the text is used as the newly added text content.

[0016] In some embodiments, the picture book playback page further includes a forward expansion control; when at least one processor executes inserting a new picture book page into a target picture book, it is further configured to:

[0017] In response to the triggering operation of the forward expansion control, the page number of the target picture book page currently displayed in the picture book display area is used as the page number of the newly added picture book page, and the page number of the target picture book page and each picture book page after the target picture book page is increased by one.

[0018] In some embodiments, the picture book playback page further includes a backward expansion control; when at least one processor executes inserting a new picture book page into a target picture book, it is further configured to:

[0019] In response to the triggering operation of the backward expansion control, the page number of the next picture book page of the target picture book page currently displayed in the picture book display area is used as the page number of the newly added picture book page, and the page number of the next picture book page and each picture book page after the next picture book page is increased by one.

[0020] In some embodiments, when at least one processor generates a target image corresponding to the newly added text content based on the newly added text content and the picture book style corresponding to the target picture book, the processor is further configured to:

[0021] Based on the newly added text content and the picture book style corresponding to the target picture book, a picture book editing instruction is constructed; the picture book editing instruction carries the newly added text content, the picture book style corresponding to the target picture book, and the identification data of the target picture book;

[0022] Sending the picture book editing instruction to the server to instruct the server to generate a target image corresponding to the newly added text content based on the picture book editing instruction;

[0023] Receive the target image fed back by the server.

[0024] In a second aspect, a picture book generation method is provided, which is applied to the display device described in the first aspect, and the method includes:

[0025] In response to an instruction to open the picture book application, displaying a function page of the picture book application on the display;

[0026] In response to a picture book generation instruction triggered on a function page, a target picture book is generated, the target picture book includes at least one picture book page with a page sequence, each picture book page includes a picture book text and a picture book image; a picture book playback page is displayed on a display, the picture book playback page includes a picture book display area and a picture book editing control; the picture book display area is used to display at least one picture book page according to the page sequence;

[0027] In response to a trigger operation on a picture book editing control, newly added text content input by the user is obtained, and based on the newly added text content and the picture book style corresponding to the target picture book, a target image corresponding to the newly added text content is generated; the newly added text content and the target image are combined into a newly added picture book page; the newly added picture book page is inserted into the target picture book to obtain an updated picture book, and the updated picture book is displayed in a picture book display area.

[0028] In the above embodiment, a display device and a picture book generation method are provided. The display device is pre-installed with a picture book application. In response to an instruction to open the picture book application, the picture book application can be opened and a function page of the picture book application can be displayed. In response to a picture book generation instruction for triggering on the function page, a target picture book can be automatically generated and a picture book playback page can be displayed. The picture book playback page can display at least one picture book page with a page sequence in the target picture book, thereby realizing automatic generation and display of a picture book according to a user's picture book generation requirement. In response to a triggering operation of a picture book editing control in the picture book playback page, newly added text content input by a user is obtained. Based on the newly added text content and a picture book style corresponding to the target picture book, a target image corresponding to the newly added text content is generated. The newly added text content and the target image can form a newly added picture book page. The newly added picture book page is inserted into the target picture book and displayed. Since the newly added picture book page has text content that meets the user's requirements, an image that matches the text content, and has a picture book style consistent with the target picture book, it is realized that the generated target picture book is rewritten or continued according to the user's picture book editing requirements, and a picture book that meets the user's personalized requirements is displayed on the display device.

[0029] In a third aspect, a picture book generation method is provided, which is applied to a server, and the server communicates data with the display device described in the first aspect; the method comprises:

[0030] Receive a picture book editing instruction sent by a display device, where the picture book editing instruction carries newly added text content, a picture book style corresponding to a target picture book, and identification data of the target picture book;

[0031] Generate new picture book texts through the language model based on the newly added text content and identification data; Generate target images through the visual generation model based on the newly added picture book texts and picture book styles;

[0032] Send the target image to the display device.

[0033] In some embodiments, generating a new picture book text through a language model according to the new text content and identification data includes:

[0034] According to the identification data, the picture book text of the target picture book is obtained; the picture book text includes story character information, story outline and story content;

[0035] Assembling a first input text based on the newly added text content, the picture book text and the preset first prompt words; the first input text is used to guide the first language big model to generate a text that meets the requirements of the role and the outline; the first input text is input into the first language big model to generate a story text;

[0036] Based on the story text and the preset second prompt word, a second input text is assembled; the second input text is used to guide the second language large model to generate text that meets the image description requirements; the second input text is input into the second language large model to generate a newly added picture book text.

[0037] In some embodiments, based on the newly added picture book text and picture book style, a large model is generated by visual means to generate a target image, including:

[0038] Encode the newly added picture book text and picture book style respectively to obtain the prompt word vector and style feature vector;

[0039] Integrate the style feature vector and the prompt word vector to obtain the conditional vector;

[0040] The conditional vector is input into the visual generative model, and the visual generative model generates the target image based on the conditional vector.

[0041] In the above embodiment, a picture book generation method is applied to a server, in which the server can receive a picture book editing instruction sent by a display device, the picture book editing instruction carrying the newly added text content, the picture book style corresponding to the target picture book and the identification data of the target picture book, and generate the newly added text content according to the newly added text content and the identification data by using the text-generating-text capability of the language big model, and then generate a target image matching the newly added text content and the picture book style by using the text-generating-image capability of the visual generation big model, and finally send the target image to the display device. This picture book editing method based on the user's picture book editing needs, combined with the text-generating-text and text-generating-image capabilities of the big model, can generate newly added text content that meets the user's picture book editing needs and a target image with a consistent style, which is conducive to realizing the function of rewriting or continuing the target picture book, thereby displaying a picture book that meets the user's personalized needs on the display device. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1 An operation scenario between a display device and a control device according to some embodiments is shown;

[0044] Figure 2 shows a hardware configuration block diagram of a control device 100 according to some embodiments;

[0045] Figure 3shows a hardware configuration block diagram of a display device 200 according to some embodiments;

[0046] Figure 4 shows a software configuration diagram in a display device 200 according to some embodiments;

[0047] Figure 5 A flowchart of a picture book generation method according to some embodiments is exemplarily shown;

[0048] Figure 6 A schematic diagram exemplarily shows a function page according to some embodiments;

[0049] Figure 7 A schematic diagram exemplarily shows a picture book waiting page according to some embodiments;

[0050] Figure 8 A schematic diagram of a picture book playback page according to some embodiments is exemplarily shown;

[0051] Fig. 9 A flowchart of a picture book generation method according to some other embodiments is exemplified;

[0052] Fig.10 A flowchart of a data processing method for visually generating a large model according to some embodiments is exemplified;

[0053] Fig.11 A timing diagram of a picture book generation method according to some embodiments is exemplarily shown;

[0054] Fig.12 A flowchart of a picture book generation method according to some other embodiments is exemplified;

[0055] Fig.13 The figure exemplarily shows an internal structure diagram of a computer device according to some embodiments. DETAILED DESCRIPTION

[0056] In order to make the purpose and implementation method of the present application clearer, the exemplary implementation method of the present application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only part of the embodiments of the present application, rather than all the embodiments.

[0057] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.

[0058] The terms "first", "second", "third", etc. in the specification and claims of this application and the above drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances.

[0059] The terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0060] The display device provided in the embodiments of the present application may have various implementation forms, for example, it may be a television, a smart television, a laser projection device, a monitor, an electronic bulletin board, an electronic table, etc. Figure 1 and Figure 2 This is a specific implementation of the display device of the present application.

[0061] Figure 1 FIG. 1 is a schematic diagram of an operation scenario between a display device and a control device according to an embodiment. Figure 1 As shown, the user can operate the display device 200 through the smart device 300 or the control apparatus 100 .

[0062] In some embodiments, the control device 100 may be a remote controller, and the communication between the remote controller and the display device includes infrared protocol communication or Bluetooth protocol communication, and other short-range communication methods, and the display device 200 is controlled wirelessly or wired. The user may control the display device 200 by inputting user commands through buttons on the remote controller, voice input, control panel input, etc.

[0063] In some embodiments, a smart device 300 (such as a mobile terminal, a tablet computer, a computer, a laptop computer, etc.) may also be used to control the display device 200. For example, the display device 200 is controlled using an application running on the smart device.

[0064] In some embodiments, the display device may not use the above-mentioned smart device or control device to receive instructions, but may receive user control through touch or gestures.

[0065] In some embodiments, the display device 200 can also be controlled in a manner other than the control device 100 and the smart device 300. For example, the user's voice command control can be directly received through a module for obtaining voice commands configured inside the display device 200, or the user's voice command control can be received through a voice control device set outside the display device 200.

[0066] In some embodiments, the display device 200 also communicates data with the server 400. The display device 200 may be allowed to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks. The server 400 may provide various content and interactions to the display device 200. The server 400 may be one cluster or multiple clusters, and may include one or more types of servers.

[0067] Figure 2 Schematically shows a block diagram of a configuration of the control device 100 according to an exemplary embodiment. Figure 2 As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, a memory, and a power supply. The control device 100 can receive input operation instructions from the user, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, playing the role of an interactive intermediary between the user and the display device 200.

[0068] like Figure 3 The display device 200 includes at least one of a tuner and demodulator 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface.

[0069] In some embodiments, the controller includes a processor, a video processor, an audio processor, a graphics processor, a RAM, a ROM, and a first interface to an nth interface for input / output.

[0070] The display 260 includes a display screen component for presenting images, and a driving component for driving image display, which is used to receive image signals output from the controller, and display video content, image content, and menu control interface components and user control UI interface.

[0071] The display 260 may be a liquid crystal display, an OLED display, or a projection display, and may also be a projection device and a projection screen.

[0072] The communicator 220 is a component for communicating with an external device or server according to various communication protocol types. For example, the communicator may include at least one of a Wifi module, a Bluetooth module, a wired Ethernet module, and other network communication protocol chips or near field communication protocol chips, and an infrared receiver. The display device 200 can establish transmission and reception of control signals and data signals with the external control device 100 or the server 400 through the communicator 220.

[0073] The user interface can be used to receive control signals from the control device 100 (such as an infrared remote controller, etc.).

[0074] The detector 230 is used to collect signals from the external environment or the external interaction. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of ambient light; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environment scenes, user attributes or user interaction gestures; or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.

[0075] The external device interface 240 may include, but is not limited to, any one or more of the following: a high-definition multimedia interface (HDMI), an analog or digital high-definition component input interface (component), a composite video input interface (CVBS), a USB input interface (USB), an RGB port, etc. It may also be a composite input / output interface formed by the above multiple interfaces.

[0076] The tuner-demodulator 210 receives broadcast television signals via wired or wireless reception, and demodulates audio and video signals, such as EPG data signals, from a plurality of wireless or wired broadcast television signals.

[0077] In some embodiments, the controller 250 and the tuner-demodulator 210 may be located in different separate devices, that is, the tuner-demodulator 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0078] The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200. For example, in response to receiving a user command for selecting a UI object to be displayed on the display 260, the controller 250 can perform operations related to the object selected by the user command.

[0079] In some embodiments, the controller includes a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM Random Access Memory (RAM), ROM (Read-Only Memory, ROM), a first interface to an nth interface for input / output, a communication bus (Bus), etc.

[0080] The user may input a user command through a graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input command through the graphical user interface (GUI). Alternatively, the user may input a user command through a specific sound or gesture, and the user input interface recognizes the sound or gesture through a sensor to receive the user input command.

[0081] "User interface" is the medium interface for interaction and information exchange between applications or operating systems and users. It realizes the conversion between the internal form of information and the form acceptable to users. The commonly used form of user interface is the Graphical User Interface (GUI), which refers to the user interface related to computer operation displayed in a graphical way. It can be an interface element such as an icon, window, control, etc. displayed on the display screen of an electronic device, where the control can include icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, widgets, etc.

[0082] See also Figure 4 In some embodiments, the system is divided into four layers, from top to bottom, namely, the application layer (Applications) layer (referred to as "application layer"), the application framework layer (Application Framework) layer (referred to as "framework layer"), the Android runtime (Android runtime) and system library layer (referred to as "system runtime library layer"), and the kernel layer.

[0083] In some embodiments, at least one application is running in the application layer, and these applications can be window programs, system settings programs, clock programs, etc. provided by the operating system, or applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the above examples.

[0084] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications. The application framework layer includes some predefined functions. The application framework layer is equivalent to a processing center that determines the actions that applications in the application layer take. Through the API interface, applications can access system resources and obtain system services during execution.

[0085] like Figure 4 As shown, in the embodiment of the present application, the application framework layer includes managers, content providers, etc., wherein the manager includes at least one of the following modules: an activity manager (ActivityManager) is used to interact with all activities running in the system; a location manager (Location Manager) is used to provide system services or applications with access to system location services; a package manager (Package Manager) is used to retrieve various information related to the application package currently installed on the device; a notification manager (NotificationManager) is used to control the display and clearing of notification messages; a window manager (Window Manager) is used to manage icons, windows, toolbars, wallpapers, and desktop components on the user interface.

[0086] In some embodiments, the activity manager is used to manage the life cycle of each application and the common navigation back function, such as controlling the exit, opening, and back of the application. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, capturing the screen, and controlling the display window changes (for example, reducing the display window, shaking the display, distorting the display, etc.).

[0087] In some embodiments, the system runtime layer provides support for the upper layer, namely the framework layer. When the framework layer is used, the Android operating system will run the C / C++ library contained in the system runtime layer to implement the functions to be implemented by the framework layer.

[0088] In some embodiments, the kernel layer is a layer between hardware and software. Figure 4 As shown, the kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0089] Traditional picture books are books that are mainly paintings with a small amount of text. With the rapid development of display devices, the functions that display devices can provide to users are becoming more and more abundant. Users can watch preset images or videos through display devices. For example, they can watch the contents of picture books through display devices. Picture books can be preset in the display device, and each picture book contains multiple images or videos. The display device can display multiple images in the picture book or display the videos in the picture book in sequence, so that users can watch the preset picture books through the display device. However, if the user needs to watch a picture book with a specific theme, but the display device does not have the picture book with the theme preset, the user cannot watch it; at the same time, traditional display devices cannot modify the picture book. Therefore, traditional display devices have the problem of being unable to display picture books that meet the personalized needs of users.

[0090] In order to solve the problem that traditional display devices cannot display picture books that meet the personalized needs of users, an embodiment of the present application provides a picture book generation method. In this method, a picture book application is pre-installed on the display device. In response to an instruction to open the picture book application, the picture book application can be opened and a function page of the picture book application can be displayed. In response to a picture book generation instruction triggered on the function page, a target picture book can be automatically generated and a picture book playback page can be displayed. The picture book playback page can display at least one picture book page with a page sequence in the target picture book, thereby realizing automatic generation and display of picture books according to the user's picture book generation needs; in response to a picture book editing instruction on the picture book playback page, a picture book page can be automatically generated and displayed. The trigger operation of the editing control is used to obtain the newly added text content input by the user, and based on the newly added text content and the picture book style corresponding to the target picture book, a target image corresponding to the newly added text content is generated. The newly added text content and the target image can form a newly added picture book page, which is inserted into the target picture book and displayed. Since the newly added picture book page has text content that meets user needs, an image that matches the text content, and has a picture book style consistent with the target picture book, it is possible to rewrite or continue the generated target picture book according to the user's picture book editing needs, and display the picture book that meets the user's personalized needs on the display device.

[0091] like Figure 5 As shown, Figure 5 A flowchart of a picture book generation method according to some embodiments is exemplarily shown. The method is applied to a display device, and the method includes:

[0092] S100: In response to an instruction to open a picture book application, display a function page of the picture book application on a display.

[0093] The display device includes a display and at least one processor, wherein the display is configured to display content from a broadcast system or network and / or a user interface; and at least one processor is connected to the display and configured to execute the picture book generation method. A plurality of applications are installed in the display device, and the picture book application is an application installed in the display device for generating, editing and displaying picture books. The function page refers to a page for displaying the functions of the picture book application. Exemplarily, the function page displays prompts for generating and editing picture books according to user needs.

[0094] In some embodiments, application icons of multiple applications are displayed on the display. After the user triggers the icon of the picture book application, the processor generates an instruction to open the picture book application in response to the triggering operation of the picture book application, opens the picture book application in response to the instruction to open the picture book application, and displays a function page of the picture book application on the display.

[0095] In some embodiments, the function page further includes a picture book collection control and a fine work collection control. In response to the triggering operation of the picture book collection control in the function page, the picture book collection page is displayed on the display, and the picture book collection page includes at least one picture book; in response to the triggering operation of any picture book in the picture book collection page, the corresponding picture book content is played on the display. Similarly, in response to the triggering operation of the fine work collection control in the function page, the fine work collection page is displayed on the display, and the picture book collection page includes at least one fine picture book, and the fine picture book is a picture book selected by preset rules in the picture book collection; in response to the triggering operation of any picture book in the fine picture book collection page, the corresponding picture book content is played on the display.

[0096] S200. In response to a picture book generation instruction triggered on a function page, a target picture book is generated, wherein the target picture book includes at least one picture book page with a page sequence, and each picture book page includes a picture book text and a picture book image; a picture book playback page is displayed on a display, wherein the picture book playback page includes a picture book display area and a picture book editing control; the picture book display area is used to display at least one picture book page in accordance with the page sequence.

[0097] The function page has a picture book generation control, for example, the picture book generation control can be a button, which generates a picture book generation instruction in response to a trigger operation of the picture book generation control, and generates a target picture book in response to the picture book generation instruction. In some embodiments, the function page can also include at least one of a story character input control, a story theme input control, and a style selection control, such as Figure 6A schematic diagram of a function page according to some embodiments is shown. In response to a trigger operation on a story role input control, the story role input by the user is obtained, and in response to a trigger operation on a story theme input control, the story theme input by the user is obtained. The function page may also include a story role selection control, and in response to a selection operation on the story role selection control, a target role selected by the user from a plurality of preset roles is obtained and used as the story role. In addition, the user may input the story role, story theme, and picture book style in the form of voice or text. The processor may obtain the story role and story theme input by the user in response to a triggered picture book generation instruction, and generate a target picture book using a large model based on the story role and story theme. The processor may also obtain the target style selected by the user from a plurality of preset styles in response to a selection operation on the style selection control, and use it as the picture book style.

[0098] In other embodiments, the processor may respond to a triggered picture book generation instruction, obtain interactive voice input by the user, identify interactive text corresponding to the interactive voice, and generate a target picture book through a large model based on the interactive text.

[0099] The target picture book includes at least one picture book page with a page sequence, each picture book page includes a picture book text and a picture book image, and the picture book texts in multiple picture book pages can be combined into a picture book story with a contextual logical order according to the page sequence. The picture book image can display the story characters, story themes, and story contents indicated by the picture book text.

[0100] The picture book playback page refers to the page for playing the picture book. The picture book playback page can be displayed immediately after the picture book generation instruction is triggered, or after the picture book generation instruction is triggered, the picture book waiting page is displayed first to instruct the user to wait for the picture book generation to be completed; after the picture book generation is completed, the picture book playback page is displayed. Figure 7 A schematic diagram of a picture book waiting page according to some embodiments is shown. A prompt text is displayed on the picture book waiting page, and the prompt text may be: picture book creation in progress. Exemplarily, after the picture book text is generated, the picture book text may be displayed on the picture book waiting page until all picture book images are generated, and then the picture book waiting page is exited.

[0101] like Figure 8 A schematic diagram of a picture book playback page according to some embodiments is shown. The picture book playback page includes a picture book display area and a picture book editing control, wherein the picture book display area is used to display at least one picture book page in page order, and the picture book display area can display one picture book page at a time, or can display multiple picture book pages at a time.

[0102] The picture book editing control is used to respond to the user's trigger operation to instruct the processor to edit the target picture book. The picture book editing control can be single or multiple. For example, the picture book playback page has a picture book editing control, which is used to instruct the processor to edit all the picture book texts of the target picture book. For another example, the picture book playback page has multiple picture book editing controls, and each picture book editing control can be used to instruct the processor to edit the corresponding picture book page.

[0103] S300. In response to a trigger operation on a picture book editing control, newly added text content input by a user is obtained, and based on the newly added text content and the picture book style corresponding to the target picture book, a target image corresponding to the newly added text content is generated; the newly added text content and the target image are combined into a newly added picture book page; the newly added picture book page is inserted into the target picture book to obtain an updated picture book, and the updated picture book is displayed in a picture book display area.

[0104] The newly added text content is the text content input by the user, which is used to rewrite or continue the target picture book. In some embodiments, the processor displays a text input control on the picture book playback page in response to a trigger operation on the picture book editing control, and the user receives the newly added text content input by the user.

[0105] The picture book style refers to the style type corresponding to the picture book image of the target picture book, for example, it can be a cartoon style, a comic style, an illustration style, etc. In some embodiments, reference Figure 6 , the user can actively select the desired picture book style before the target picture book is generated. In other embodiments, the user may not specify the picture book style, and the large model automatically selects an appropriate picture book style according to the picture book text.

[0106] The target image is generated by the large model based on the newly added text content and the picture book style corresponding to the target picture book. The target image can show the story characters, story themes, and story content indicated by the newly added text content, and the style of the target image is consistent with the picture book style corresponding to the target picture book.

[0107] The processor uses the newly added text content as the picture book text and the target image as the picture book image to form a newly added picture book page. The newly added picture book page can be inserted into the target picture book as an independent picture book page. For example, the insertion position can be the position indicated by the user, or the position indicated by the large model according to the contextual semantic logic. After the newly added picture book page is inserted, an updated picture book is obtained, thereby realizing the rewriting of the picture book. For another example, the insertion position can also be the end of the picture book. After the newly added picture book page is inserted, an updated picture book is obtained, thereby realizing the continuation of the picture book.

[0108] The picture book display area displays updated picture books, so that the picture books displayed in the display device meet the personalized needs of users.

[0109] In an embodiment of the present application, a picture book application is pre-installed on the display device. In response to an instruction to open the picture book application, the picture book application can be opened and a function page of the picture book application can be displayed. In response to a picture book generation instruction for triggering on the function page, a target picture book can be automatically generated and a picture book playback page can be displayed. The picture book playback page can display at least one picture book page with a page sequence in the target picture book, thereby realizing automatic generation and display of a picture book according to the user's picture book generation requirements. In response to a triggering operation of a picture book editing control in the picture book playback page, the newly added text content input by the user is obtained, and based on the newly added text content and the picture book style corresponding to the target picture book, a target image corresponding to the newly added text content is generated. The newly added text content and the target image can form a newly added picture book page, which is inserted into the target picture book and displayed. Since the newly added picture book page has text content that meets the user's requirements, an image that matches the text content, and has a picture book style consistent with the target picture book, it is realized that the generated target picture book is rewritten or continued according to the user's picture book editing requirements, and a picture book that meets the user's personalized needs is displayed on the display device.

[0110] In some embodiments, the picture book display area includes a text display area and an image display area, the text display area is used to display the picture book text in the picture book page, and the image display area is used to display the picture book image in the picture book page; in response to the triggering operation of the picture book editing control, the newly added text content input by the user is obtained, including: in response to the triggering operation of the picture book editing control, the target picture book text currently displayed in the text display area is cleared, the target text input by the user in the text display area is received, and the target text is used as the newly added text content.

[0111] Among them, reference Figure 8 , the picture book display area includes a text display area and an image display area. The text display area can be used to display the picture book text in at least one picture book page according to the page sequence, and the image display area is used to display the picture book image in at least one picture book page according to the page sequence. It can be understood that the picture book text and picture book image belonging to the same page sequence are displayed in the same picture book display area.

[0112] In response to the triggering operation on the picture book editing control, the processor can clear the target picture book text currently displayed in the text display area to facilitate receiving the target text input by the user in the text display area.

[0113] Exemplarily, in response to a trigger operation on a picture book editing control, the processor may display a text input control on the picture book playback page and receive a target text entered by the user through the text input control, then clear the target picture book text currently displayed in the text display area and display the target text in the text display area, where the target text is the newly added text content.

[0114] The target text input by the user is received by displaying a virtual keyboard on the display, displaying multiple characters on the virtual keyboard, and determining the multiple characters selected by the user in response to the movement operation of the remote controller, and determining the target text through the multiple characters.

[0115] In the embodiment of the present application, the text display area of ​​the picture book display area displays the picture book text, and the image display area displays the picture book image. The user can input new text content in the form of text input. The new text content can be displayed in the text display area and can be used as the story content for rewriting or continuing the picture book, which is conducive to rewriting or continuing the target picture book, and then generating a picture book that meets the user's personalized needs.

[0116] In some embodiments, in response to a trigger operation on a picture book editing control, newly added text content input by a user is obtained, including: in response to a trigger operation on a picture book editing control, a voice input prompt is displayed on a picture book playback page; interactive voice input by the user is received, the interactive voice is converted into text, and the text is used as newly added text content.

[0117] Among them, the processor can also obtain new text content through voice input. Specifically, the processor responds to the trigger operation of the picture book editing control and displays a voice input prompt on the picture book playback page. The voice input prompt is used to indicate that the user can input the picture book content to be edited in the form of voice.

[0118] In response to the triggering operation of the picture book editing control, the processor can also start the voice module, receive the interactive voice input by the user through the voice module, perform voice recognition on the interactive voice to obtain the text corresponding to the interactive voice, and use the text as the newly added text content. The voice module can be a module configured inside the processor, or it can be an external voice control device.

[0119] The newly added text content is usually a rewrite or continuation of the target picture book. For example, the story theme of the target picture book is: Baibai is a little white cat with blue eyes. It explores the changing seasons and experiences the joy of growing up. The target picture book has a total of 7 pages. The newly added text content can be: This year, Baibai made many friends in each season. It hopes to bring this joy of growth to more friends next year, and everyone can grow together.

[0120] Exemplarily, the processor can display the newly added text content in the currently displayed picture book page, and can also display a picture book page without picture book text and picture book image in the picture book playback page before the currently displayed picture book page, after the currently displayed picture book page, before the last picture book page, or after the last picture book page, and display the newly added text content in the text display area of ​​the picture book page.

[0121] In an embodiment of the present application, the user can input new text content through voice input, and the new text content can be used as the story content of the rewritten or continued picture book, which is conducive to rewriting or continuing the target picture book, and then generating a picture book that meets the user's personalized needs.

[0122] In some embodiments, the picture book playback page also includes a forward expansion control; inserting a new picture book page into the target picture book includes: in response to a trigger operation on the forward expansion control, using the number of pages of the target picture book page currently displayed in the picture book display area as the number of pages of the newly added picture book page, and increasing the number of pages of the target picture book page and each picture book page after the target picture book page by one.

[0123] The forward expansion control refers to a triggerable control in the picture book playback page. For example, the forward expansion control may be a button. When the button is clicked, the processor may respond to the triggering operation of the forward expansion control. Figure 8 The forward expansion control may be a triggerable option in a drop-down list. In response to a trigger operation of the drop-down list, the forward expansion option is displayed. When the forward expansion option is clicked, the processor may respond to the trigger operation of the forward expansion control.

[0124] In response to the triggering operation of the forward expansion control, the processor can use the page number of the currently displayed target picture book page as the page number of the newly added picture book page, and increase the page number of the target picture book page and each picture book page after the target picture book page by one, that is, insert the newly added picture book page before the target picture book page. For example, if the target picture book page is page 7, after the newly added picture book page is constructed, the newly added picture book page can be used as the new page 7, and the page numbers of the original page 7, page 8, page 9, etc. are all increased by one, becoming the new page 8, page 9, page 10, etc.

[0125] In an embodiment of the present application, a triggerable forward expansion control is pre-set in the picture book playback page. According to the user's triggering operation on the forward expansion control, a new picture book page can be inserted before the currently displayed target picture book page, which is conducive to meeting the user's needs for rewriting the picture book.

[0126] In some embodiments, the picture book playback page also includes a backward expansion control; inserting a newly added picture book page into the target picture book includes: in response to a trigger operation on the backward expansion control, using the page number of the next picture book page of the target picture book page currently displayed in the picture book display area as the page number of the newly added picture book page, and increasing the page number of the next picture book page and each picture book page after the next picture book page by one.

[0127] The backward expansion control refers to a triggerable control in the picture book playback page. For example, the backward expansion control can be a button. When the button is clicked, the processor can respond to the triggering operation of the backward expansion control. Figure 8The backward expansion control may be a triggerable option in the drop-down list. In response to the trigger operation of the drop-down list, the backward expansion option is displayed. When the backward expansion option is clicked, the processor may respond to the trigger operation of the backward expansion control.

[0128] In response to the triggering operation of the backward expansion control, the processor can use the page number of the next picture book page of the currently displayed target picture book page as the page number of the newly added picture book page, and increase the page number of the next picture book page and each picture book page after the next picture book page by one, that is, insert the newly added picture book page after the target picture book page. For example, if the target picture book page is page 7, after the newly added picture book page is constructed, the newly added picture book page can be used as the new page 8, and the page numbers of the original page 8, page 9, page 10, etc. are all increased by one, becoming the new page 9, page 10, page 11, etc.

[0129] In an embodiment of the present application, a triggerable backward expansion control is pre-set in the picture book playback page. According to the user's triggering operation on the backward expansion control, a new picture book page can be inserted after the currently displayed target picture book page, which is conducive to meeting the user's demand for rewriting the picture book.

[0130] In some embodiments, a target image corresponding to the newly added text content is generated based on the newly added text content and the picture book style corresponding to the target picture book, including: constructing a picture book editing instruction based on the newly added text content and the picture book style corresponding to the target picture book; the picture book editing instruction carries the newly added text content, the picture book style corresponding to the target picture book, and the identification data of the target picture book; sending the picture book editing instruction to a server to instruct the server to generate a target image corresponding to the newly added text content based on the picture book editing instruction; and receiving the target image feedback from the server.

[0131] Among them, the newly added text content input by the user is used as the picture book text, and the text-generating image capability of the large model needs to be used to generate a target image that matches the newly added text content and the picture book style. In some embodiments, the server has a large model, and the server and the display communicate through a preset protocol. Based on the preset protocol, the processor converts the newly added text content and the picture book style corresponding to the target picture book into a picture book editing instruction that can be recognized by the server and sends it to the server.

[0132] At least one picture book is stored in the processor, and the identification data is an identification for distinguishing each picture book, and each picture book has unique corresponding identification data, for example, the identification data can be a number, etc. In some embodiments, the display device has a device identification, the user has a user identification, the picture book has a story identification, and the picture book page has a page identification. According to actual needs, the device identification, user identification and story identification can be combined to form the identification data of the picture book, and the device identification, user identification, story identification and page identification can be combined to form the identification data of the picture book page in the picture book.

[0133] Since the picture book editing instruction carries the newly added text content, the picture book style corresponding to the target picture book, and the identification data of the target picture book, the picture book editing instruction can instruct the server to generate a target image corresponding to the newly added text based on the picture book editing instruction.

[0134] The processor receives the target image fed back by the server. The target image can display the picture indicated by the newly added text content and is consistent with the picture book style of the target picture book.

[0135] In an embodiment of the present application, the processor constructs a picture book editing instruction based on the newly added text content and the picture book style corresponding to the target picture book, and sends it to the server to instruct the server to generate a target image corresponding to the newly added text content based on the picture book editing instruction, which is conducive to the automatic generation of new picture book pages, and realizes the function of rewriting or continuing the picture book based on the user's picture book editing needs, thereby improving the flexibility of picture book generation and enhancing the user experience.

[0136] like Fig. 9 As shown, Fig. 9 A flowchart of a picture book generation method according to some embodiments is exemplarily shown. The method is applied to a server, and the server communicates data with the aforementioned display device. The method includes:

[0137] S920: Receive a picture book editing instruction sent by the display device, where the picture book editing instruction carries newly added text content, a picture book style corresponding to a target picture book, and identification data of the target picture book.

[0138] S940. Generate new picture book text through the language big model according to the new text content and identification data; generate a target image through the visual big model based on the new picture book text and picture book style.

[0139] S960: Send the target image to the display device.

[0140] Among them, the server has a language big model and a visual generation big model. The language big model has the ability to generate text from text, and the visual generation big model has the ability to generate images from text.

[0141] The server receives the picture book editing instruction sent by the display device, and generates a new picture book text for describing the screen content through the language model based on the new text content and identification data. Since the new text content input by the user often has abstract story content, using only the new text content for text generation may cause the generated image to not meet the user's needs. Therefore, the text generation capability of the language model can be used to convert the new text content into specific story content to ensure that the generated target image meets the user's needs. For example, the new text content is a swallow flying low, and the new picture book text can be a swallow flying low over the lake, gently touching the water surface, stirring up circles of ripples.

[0142] The server can integrate the newly added picture book text and picture book style into a conditional vector, and the visual generation model generates a target image that matches the newly added text content and the picture book style based on the conditional vector. The server sends the generated target image to the display device.

[0143] In the embodiment of the present application, the server can generate new text content and a target image with the same style that meet the user's picture book editing needs based on the user's picture book editing needs and combined with the text-generated text and text-generated image capabilities of the large model, and send them to the display device, which is conducive to realizing the function of rewriting or continuing the target picture book, thereby displaying the picture book that meets the user's personalized needs on the display device.

[0144] In some embodiments, new picture book text is generated through a language macro model based on new text content and identification data, including: obtaining a picture book text of a target picture book based on the identification data; the picture book text includes story character information, story outline and story content; assembling a first input text based on the new text content, the picture book text and a pre-set first prompt word; the first input text is used to guide the first language macro model to generate a text that meets the role and outline requirements; the first input text is input into the first language macro model to generate a story text; assembling a second input text based on the story text and a pre-set second prompt word; the second input text is used to guide the second language macro model to generate a text that meets the image description requirements; the second input text is input into the second language macro model to generate new picture book text.

[0145] The server stores the picture book texts of various picture books, and the picture book text of the target picture book can be found according to the identification data. The picture book text usually includes story character information, story outline and story content.

[0146] The first prompt word is a prompt word template pre-set in the server. The server assembles the newly added text content, the picture book text and the pre-set first prompt word into the first input text. The first input text is used as the prompt word of the first language large model to guide the first language large model to generate text that meets the role and summary requirements. The first input text is input into the first language large model to generate a story text. The story roles and story summary in the story text are consistent with the story role information and story summary in the picture book text. At the same time, the story content in the story text matches the newly added text content. For example, the story content is based on the story role information, story summary, and newly added text content. Expand the character name, skin color, hairstyle, clothing, decoration, etc.

[0147] The second prompt word is a prompt word template pre-set in the server. The server assembles the story text and the pre-set second prompt word into the second input text. The second input text is used as the prompt word of the second language large model to guide the second language large model to generate text that meets the image description requirements. The second input text is input into the second language large model to generate a new picture book text. The new picture book text matches the new text content, and adds description information such as picture book background description, action description, and noun description to enhance the text's ability to describe the picture.

[0148] In the embodiment of the present application, the first language model is guided to generate a story text that meets the role and outline requirements through the newly added text content, the picture book text of the target picture book and the first prompt word, and then the second language model is guided to generate a newly added picture book text that meets the image description requirements through the story text and the second prompt word. After two language models expand the content of the newly added text, the content of the newly added picture book text is continuously enriched, and the picture description ability of the newly added picture book text is improved, which is conducive to generating a picture book text that meets user needs.

[0149] In some embodiments, a target image is generated based on newly added picture book text and picture book style through a visual generation model, including: encoding the newly added picture book text and picture book style respectively to obtain a prompt word vector and a style feature vector; integrating the style feature vector and the prompt word vector to obtain a conditional vector; inputting the conditional vector into the visual generation model, and the visual generation model generates a target image based on the conditional vector.

[0150] Among them, the prompt word vector is a numerical representation obtained by natural language processing encoding of the newly added picture book text, which captures the key description or instructions of the newly added picture book text. The style feature vector is a numerical representation obtained by encoding the picture book style, which comprehensively reflects the style information of the picture book image. The conditional vector is a vector formed by integrating the style feature vector and the prompt word vector, which is used to guide the image generation process of the visual generation model. The conditional vector contains all the key information required to generate the target image, such as visual elements such as character image and style type, as well as semantic information in the text description. The conditional vector enables the generated target image to conform to the changes in the story content while maintaining the original visual consistency. For example, if the original picture book adopts a specific picture book style, this information will be retained in the style feature vector and guide the process of generating the target image by the visual generation model. If a specific scene or character is described in the newly added picture book text, the prompt word vector will guide the visual generation model to generate the corresponding picture content.

[0151] like Fig.10The flowchart of the data processing method of the visual generation big model according to some embodiments is shown as an example. The server may encode the picture book style through a pre-trained encoder and convert it into a style feature vector, which may be a set of numerical feature vectors. The newly added picture book text is encoded according to the pre-trained language model and encoded into a prompt word vector. Subsequently, the style feature vector and the prompt word vector are integrated to obtain a conditional vector. Specifically, the vector integration may be performed by splicing or weighted summation. Next, the conditional vector is formatted into a format that can be recognized and processed by the visual generation big model, and the conditional vector is passed to the visual generation big model through an API or a direct call. The conditional vector is used as a conditional input to guide the visual generation big model to generate a target image. Specifically, the visual generation big model will always be guided by the conditional vector, and in the process of generating the target image, visual elements such as character image characteristics and style types will be considered, while the content of the newly added picture book text will be combined to ensure that the generated target image not only conforms to the newly added text content input by the user, but also maintains the original picture book style.

[0152] Among them, the visual generation model can be a diffusion model based on the cross-attention mechanism. The diffusion model is a generation model, which generates high-quality data samples by gradually adding noise and then learning to denoise. The cross-attention mechanism is an attention mechanism that allows the model to pay attention to information of different modalities during the generation process. In this embodiment, the cross-attention mechanism makes corresponding adjustments, specifically for the query variable (Query, Q variable), the key variable (Key, K variable) and the value variable (Value, V variable). The adjustment is made, wherein the Q variable is the content feature of the current image in the current generation process, and the K variable and the V variable are used to capture the degree of association between the context and different parts of the current image. Specifically, the K variable is a pre-calculated image feature related to the target picture book, which provides a reference for the current generated image. The model evaluates the similarity of the Q variable and the K variable to find which queried image features best match the needs of the current generated image content. The V variable is used to provide detailed information when the match is successful. Once it is determined which image features are most suitable for the needs of the current generated image, the V variable will give specific visual details or style guidance to enrich the quality of the generated image.

[0153] Specifically, unlike the traditional attention mechanism in which the Q variable, K variable and V variable all use the feature vector of the currently generated image, in this embodiment, the image feature data of the currently generated image is used as the query variable (Q variable), while the K variable and V variable use the style feature vector in the conditional vector.

[0154] The server integrates the style feature vector and the prompt word vector to obtain the conditional vector, and inputs the conditional vector into the diffusion model based on the cross-attention mechanism. The diffusion model starts with a random noise image, and gradually iteratively reduces the noise as the time step increases until a clear image is generated. In each round of iterative denoising, the cross-attention mechanism allows the model to adjust the generation strategy according to the information in the conditional vector. Specifically, the image feature data of the current generated image is used as the query variable (Q variable), and the visual feature data in the conditional vector is used as the K variable and the V variable. The model calculates the similarity between the Q variable and the K variable, and weights and aggregates the V variable according to this similarity, thereby guiding the update of the current image feature. The diffusion model learns how to recover meaningful image features from the noise, and gradually removes the noise until the final clear image is generated. Furthermore, the server can also post-process the generated picture book image, such as color correction, resolution adjustment, etc., to improve the quality of the picture book image.

[0155] In other embodiments, if the user wants to adjust the style of the picture book, such as changing the style from "comic" style to "ancient style", the server can encode the adjusted picture book style to obtain the adjusted style feature vector, thereby updating the condition vector and guiding the visual generation model to generate a target image that conforms to the adjusted picture book style.

[0156] In an embodiment of the present application, by integrating the style feature vector and the prompt word vector into a conditional vector and using the conditional vector to guide the visual generation model, the target image generated by the visual generation model can maintain style consistency with the target picture book, and the picture content of the target image is consistent with the newly added text content input by the user, which is conducive to improving the user experience.

[0157] To explain the display device, server, picture book generation method and effects in detail, a most detailed embodiment is described below:

[0158] The display device includes: a display configured to display content from a broadcast system or a network and / or a user interface; and at least one processor connected to the display and configured to execute instructions to perform the picture book generation method. The server includes: a picture book generation service, a first language large model, a second language large model and a visual generation large model.

[0159] like Fig.11The timing diagram of the picture book generation method according to some embodiments is shown as an example. The user triggers the icon of the picture book application, and the processor of the display device generates an instruction to open the picture book application in response to the triggering operation of the picture book application icon. In response to the instruction to open the picture book application, the function page of the picture book application is displayed on the display of the display device; in response to the picture book generation instruction triggered on the function page, a target picture book is generated, and the target picture book includes at least one picture book page with a page sequence, and each picture book page includes picture book text and picture book images; a picture book playback page is displayed on the display, and the picture book playback page includes a picture book display area and a picture book editing control; the picture book display area is used to display at least one picture book page according to the page sequence; in response to the triggering operation of the picture book editing control, the newly added text content input by the user is obtained, and the newly added text content can be obtained by obtaining the interactive voice input by the user and performing voice recognition on the interactive voice. Based on the newly added text content and the picture book style corresponding to the target picture book, a picture book editing instruction is constructed, and the picture book editing instruction carries the newly added text content, the picture book style corresponding to the target picture book, and the identification data of the target picture book; the picture book editing instruction is sent to the picture book generation service in the server to instruct the server to generate the newly added picture book text through the language big model according to the newly added text content and the identification data, and to generate the target image through the visual big model based on the newly added picture book text and the picture book style, and the target image is fed back through the picture book generation service. The display device receives the target image fed back by the server, and combines the newly added text content and the target image into a newly added picture book page; the newly added picture book page is inserted into the target picture book to obtain an updated picture book, and the updated picture book is displayed in the picture book display area.

[0160] Among them, the server generates a new picture book text through the language big model according to the newly added text content and identification data, including: according to the identification data, obtaining the picture book text of the target picture book; the picture book text includes story character information, story outline and story content; based on the newly added text content, the picture book text and a pre-set first prompt word, assembling to obtain a first input text; the first input text is used to guide the first language big model to generate a text that meets the role and outline requirements; the first input text is input into the first language big model to generate a story text; based on the story text and a pre-set second prompt word, assembling to obtain a second input text; the second input text is used to guide the second language big model to generate a text that meets the image description requirements; the second input text is input into the second language big model to generate a new picture book text.

[0161] In some embodiments, a target image is generated based on newly added picture book text and picture book style through a visual generation model, including: encoding the newly added picture book text and picture book style respectively to obtain a prompt word vector and a style feature vector; integrating the style feature vector and the prompt word vector to obtain a conditional vector; inputting the conditional vector into the visual generation model, and the visual generation model generates a target image based on the conditional vector.

[0162] In some embodiments, the picture book display area includes a text display area and an image display area, the text display area is used to display the picture book text in the picture book page, and the image display area is used to display the picture book image in the picture book page; in response to the triggering operation of the picture book editing control, the newly added text content input by the user is obtained, including: in response to the triggering operation of the picture book editing control, the target picture book text currently displayed in the text display area is cleared, the target text input by the user in the text display area is received, and the target text is used as the newly added text content.

[0163] In some embodiments, in response to a trigger operation on a picture book editing control, newly added text content input by a user is obtained, including: in response to a trigger operation on a picture book editing control, a voice input prompt is displayed on a picture book playback page; interactive voice input by the user is received, the interactive voice is converted into text, and the text is used as newly added text content.

[0164] In some embodiments, the picture book playback page also includes a forward expansion control; inserting a new picture book page into the target picture book includes: in response to a trigger operation on the forward expansion control, using the number of pages of the target picture book page currently displayed in the picture book display area as the number of pages of the newly added picture book page, and increasing the number of pages of the target picture book page and each picture book page after the target picture book page by one.

[0165] In some embodiments, the picture book playback page also includes a backward expansion control; inserting a newly added picture book page into the target picture book includes: in response to a trigger operation on the backward expansion control, using the page number of the next picture book page of the target picture book page currently displayed in the picture book display area as the page number of the newly added picture book page, and increasing the page number of the next picture book page and each picture book page after the next picture book page by one.

[0166] In some embodiments, a target image corresponding to the newly added text content is generated based on the newly added text content and the picture book style corresponding to the target picture book, including: constructing a picture book editing instruction based on the newly added text content and the picture book style corresponding to the target picture book; the picture book editing instruction carries the newly added text content, the picture book style corresponding to the target picture book, and the identification data of the target picture book; sending the picture book editing instruction to a server to instruct the server to generate a target image corresponding to the newly added text content based on the picture book editing instruction; and receiving the target image feedback from the server.

[0167] like Fig.12The flowchart of a picture book generation method according to some embodiments is shown as an example. In response to the user's trigger instruction for the picture book editing control, the newly added text content input by the user is obtained, and the identification data is used to obtain the picture book text of the target picture book from the character feature library. The newly added picture book content, the picture book text and the first prompt word are assembled into a first input text, and the first input text is used as a prompt word of the first language large model to guide the first language large model to generate a story text; based on the story text and the second prompt word, the second input text is assembled into a second input text, and the second input text is used as a prompt word of the second language large model to guide the second language large model to generate the newly added picture book text; based on the newly added picture book text and the picture book style, a large model is generated through visual generation to generate a target image; the newly added text content and the target image constitute a newly added picture book page, and are inserted into the target picture book for display.

[0168] In the above embodiment, the display device is pre-installed with a picture book application. In response to an instruction to open the picture book application, the picture book application can be opened and a function page of the picture book application can be displayed. In response to a picture book generation instruction for triggering on the function page, a target picture book can be automatically generated and a picture book playback page can be displayed. The picture book playback page can display at least one picture book page with a page sequence in the target picture book, thereby realizing automatic generation and display of a picture book according to the user's picture book generation requirements. In response to a triggering operation of a picture book editing control in the picture book playback page, the newly added text content input by the user is obtained, and based on the newly added text content and the picture book style corresponding to the target picture book, a target image corresponding to the newly added text content is generated. The newly added text content and the target image can form a newly added picture book page, which is inserted into the target picture book and displayed. Since the newly added picture book page has text content that meets the user's requirements, an image that matches the text content, and has a picture book style consistent with the target picture book, it is realized that the generated target picture book is rewritten or continued according to the user's picture book editing requirements, and a picture book that meets the user's personalized needs is displayed on the display device. The server can receive picture book editing instructions sent by the display device. The picture book editing instructions carry the newly added text content, the picture book style corresponding to the target picture book, and the identification data of the target picture book. According to the newly added text content and the identification data, the language model's text-generating ability is used to generate the newly added text content, and then the visual generation model's text-generating image ability is used to generate a target image that matches the newly added text content and the picture book style. Finally, the target image is sent to the display device. This picture book editing requirement based on the user, combined with the text-generating text and text-generating image capabilities of the large model, can generate newly added text content that meets the user's picture book editing requirements and a target image with a consistent style, which is conducive to realizing the function of rewriting or continuing the target picture book, thereby displaying a picture book that meets the user's personalized needs on the display device.

[0169] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0170] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Fig.13 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (Near Field Communication, NFC) or other technologies. When the computer program is executed by the processor, a picture book generation method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.

[0171] Those skilled in the art will understand that Fig.13 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0172] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0173] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0174] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0175] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0176] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0177] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0178] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A display device, characterized in that: include: a display configured to display content from a broadcast system or network and / or a user interface; and at least one processor connected to the display and configured to execute instructions to cause the display device to: In response to an instruction to open a picture book application, displaying a function page of the picture book application on the display; In response to a picture book generation instruction triggered on the function page, a target picture book is generated, wherein the target picture book includes at least one picture book page with a page sequence, and each picture book page includes a picture book text and a picture book image; a picture book playback page is displayed on the display, wherein the picture book playback page includes a picture book display area and a picture book editing control; the picture book display area is used to display the at least one picture book page according to the page sequence; In response to a triggering operation on the picture book editing control, newly added text content input by the user is obtained, and based on the newly added text content and the picture book style corresponding to the target picture book, a target image corresponding to the newly added text content is generated; the newly added text content and the target image are combined into a newly added picture book page; the newly added picture book page is inserted into the target picture book to obtain an updated picture book, and the updated picture book is displayed in the picture book display area.

2. The display device according to claim 1, characterized in that The picture book display area includes a text display area and an image display area, the text display area is used to display the picture book text in the picture book page, and the image display area is used to display the picture book image in the picture book page; When the at least one processor executes the triggering operation on the picture book editing control to obtain the newly added text content input by the user, the at least one processor is further configured to: In response to a triggering operation on the picture book editing control, the target picture book text currently displayed in the text display area is cleared, the target text input by the user in the text display area is received, and the target text is used as the newly added text content.

3. The display device according to claim 1, characterized in that When the at least one processor executes the triggering operation on the picture book editing control to obtain the newly added text content input by the user, the at least one processor is further configured to: In response to a triggering operation on the picture book editing control, a voice input prompt is displayed on the picture book playback page; an interactive voice input by a user is received, the interactive voice is converted into text, and the text is used as newly added text content.

4. The display device according to claim 1, characterized in that The picture book playback page also includes a forward expansion control; when the at least one processor executes inserting the newly added picture book page into the target picture book, it is further configured to: In response to the triggering operation of the forward expansion control, the page number of the target picture book page currently displayed in the picture book display area is used as the page number of the newly added picture book page, and the page number of the target picture book page and each picture book page after the target picture book page is increased by one.

5. The display device according to claim 1, characterized in that The picture book playback page also includes a backward expansion control; when the at least one processor executes inserting the newly added picture book page into the target picture book, it is further configured to: In response to the triggering operation of the backward expansion control, the page number of the next picture book page of the target picture book page currently displayed in the picture book display area is used as the page number of the newly added picture book page, and the page number of the next picture book page and each picture book page after the next picture book page is increased by one.

6. The display device according to any one of claims 1 to 5, characterized in that: When the at least one processor generates a target image corresponding to the newly added text content based on the newly added text content and the picture book style corresponding to the target picture book, the at least one processor is further configured to: Based on the newly added text content and the picture book style corresponding to the target picture book, construct a picture book editing instruction; the picture book editing instruction carries the newly added text content, the picture book style corresponding to the target picture book, and the identification data of the target picture book; Sending the picture book editing instruction to a server to instruct the server to generate a target image corresponding to the newly added text content based on the picture book editing instruction; The target image fed back by the server is received.

7. A picture book generation method, characterized in that: Applied to the display device according to any one of claims 1 to 6, the method comprises: In response to an instruction to open a picture book application, displaying a function page of the picture book application on the display; In response to a picture book generation instruction triggered on the function page, a target picture book is generated, wherein the target picture book includes at least one picture book page with a page sequence, and each picture book page includes a picture book text and a picture book image; a picture book playback page is displayed on the display, wherein the picture book playback page includes a picture book display area and a picture book editing control; the picture book display area is used to display the at least one picture book page according to the page sequence; In response to a triggering operation on the picture book editing control, newly added text content input by the user is obtained, and based on the newly added text content and the picture book style corresponding to the target picture book, a target image corresponding to the newly added text content is generated; the newly added text content and the target image are combined into a newly added picture book page; the newly added picture book page is inserted into the target picture book to obtain an updated picture book, and the updated picture book is displayed in the picture book display area.

8. A picture book generation method, characterized in that: Applied to a server, the server performs data communication with a display device according to any one of claims 1 to 6; the method comprises: Receiving a picture book editing instruction sent by the display device, the picture book editing instruction carrying newly added text content, a picture book style corresponding to a target picture book, and identification data of the target picture book; Generate a new picture book text through a language macro model according to the new text content and the identification data; generate a target image through a visual macro model based on the new picture book text and the picture book style; The target image is sent to the display device.

9. The method according to claim 8, characterized in that The step of generating a new picture book text by using a language model according to the new text content and the identification data includes: According to the identification data, the picture book text of the target picture book is obtained; the picture book text includes story character information, story summary and story content; Assembling a first input text based on the newly added text content, the picture book text and the preset first prompt word; the first input text is used to guide the first language model to generate a text that meets the requirements of the role and the outline; the first input text is input into the first language model to generate a story text; Based on the story text and the preset second prompt word, a second input text is assembled; the second input text is used to guide the second language model to generate text that meets the image description requirements; the second input text is input into the second language model to generate a new picture book text.

10. The method according to claim 8, characterized in that The step of generating a target image by visually generating a large model based on the newly added picture book text and the picture book style includes: Encoding the newly added picture book text and the picture book style respectively to obtain a prompt word vector and a style feature vector; Integrate the style feature vector and the prompt word vector to obtain a conditional vector; The conditional vector is input into the visual generative model, and the visual generative model generates a target image based on the conditional vector.

Citation Information

Cited By

  • Display device

    WO2026144399A1