Story picture book editing method, display device and server
By adjusting the text data of the story picture book on the server side and calling the visual generation large model to generate a consistent new image, the problem of low quality of the rewritten story picture book is solved, and the user's interactive experience and personalized needs are improved.
Patent Information
- Application Number
- CN202411980710.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
AI Technical Summary
The quality of rewritten story picture books in the prior art is not high, resulting in poor user interaction experience.
The user's picture book editing content is received through the server, the text data of the original story picture book is adjusted, and the prompt words used to guide the visual generation of the big model to generate images are determined based on the adjusted text data. At the same time, the visual feature data of the original story picture book is queried, and the prompt words and visual feature data are used as inputs, and the visual generation model is called to generate a new image to ensure that the newly generated story picture book is visually and text consistent with the original picture book.
It realizes that when users edit story picture books, the newly generated content maintains the consistency with the original picture books visually and text as much as possible, improving the user's interactive experience and personalized needs satisfaction.
Smart Images

Figure CN119991869A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of display devices, and in particular to a story picture book editing method applied to a server, a story picture book editing method applied to a display device, a display device, and a server. Background Art
[0002] With the rapid development of display devices and the increasing diversification of user needs, people's demand for the intelligence of display devices such as smart TVs is getting higher and higher, and the functions of display devices are becoming more and more abundant.
[0003] Currently, users can input their picture book creation requirements through a series of interactions with the display device, and the display device will display the story picture book, and the user can directly read the story picture book on the display device. In addition, in order to meet the user's personalized customization needs, users are also given the right to edit and rewrite the story picture book, that is, when the user thinks that the generated story picture book does not meet their expectations, they can rewrite or continue the story according to their needs.
[0004] However, usually, it is still difficult to maintain the consistency of the content of the story picture books rewritten based on the picture book rewriting requirements input by the user, and the quality is low, which greatly reduces the user's interactive experience. Summary of the invention
[0005] The present application provides a story picture book editing method applied to a server, a story picture book editing method applied to a display device, and a display device and a server, so as to solve the problem that the quality of the rewritten story picture book is not high, resulting in a poor user interaction experience.
[0006] In a first aspect, some embodiments further provide a story picture book editing method, which is applied to a server, wherein the server includes a communication module and a processor. The method includes:
[0007] The solutions of the above embodiments have the following advantages or beneficial effects:
[0008] After receiving the picture book creation service call request carrying the identification data of the story picture book and the editing content of the picture book, the server adjusts the text data of the original story picture book according to the identification data of the story picture book and the editing content provided by the user, and then determines the prompt words used to guide the visual generation model to generate images based on the adjusted text data. These prompt words not only reflect the changes in the text, but also inherit the style and other visual characteristics of the original story picture book, which can guide the visual generation model to maintain consistency when generating images. On the other hand, the visual feature data (such as image, color, element layout) of the original story picture book is queried, and the visual generation model is called with the prompt words and the original visual feature data as input, so that the newly generated image of the visual generation model can be visually consistent with the original story picture book. Finally, the adjusted text data and picture book image are packaged into an edited story picture book and fed back to the display device so that the user can view the edited story picture book in time. Throughout the entire process, users only need to input the corresponding picture book editing instructions on the display device to edit and modify the story picture book. Moreover, no matter how the user edits the story picture book, the newly generated story picture book can be kept as consistent as possible with the original story picture book in terms of vision and text, meeting the personalized needs of users, saving users from repeatedly adjusting the picture book creation needs, and greatly improving the user's interactive experience. Throughout the entire process, users only need to input interactive content on the user interface of the display device to directly read the generated high-quality story picture book on the display device without other tedious operations. Moreover, since the story picture book generated by the server is more in line with the characters and storyline, it can better meet user expectations, saving users from repeatedly adjusting the picture book creation needs, and greatly improving the user's interactive experience.
[0009] In a second aspect, some embodiments further provide a story picture book editing method, which is applied to a display device, wherein the display device comprises: a display and a controller, wherein the method comprises: receiving picture book editing content input by a user for a story picture book, calling a picture book creation service based on the picture book editing content and identification data of the story picture book, so that the picture book creation service adjusts the text data in the story picture book according to the identification data and the picture book editing content, determines new prompt words based on the adjusted text data, wherein the prompt words are used to guide the visual generation large model to generate images, queries the visual feature data of the story picture book according to the identification data, calls the visual generation large model with the visual feature data and the new prompt words as input, generates an adjusted picture book image, packages the adjusted text data and the adjusted picture book image into an edited story picture book for feedback, receives the edited story picture book fed back by the picture book creation service, and displays the edited story picture book.
[0010] The solutions of the above embodiments have the following advantages or beneficial effects:
[0011] After the display device receives the picture book editing content input by the user for the story picture book, the picture book creation service is called based on the picture book editing content and the identification data of the story picture book, so that the picture book creation service adjusts the original text data according to the identification data of the story picture book and the editing content provided by the user, and then determines new prompt words for guiding the visual generation model to generate images based on the adjusted text data. These prompt words not only reflect the changes in the text, but also inherit the style and other visual characteristics of the original story picture book, and can guide the visual generation model to maintain consistency when generating images. On the other hand, the original visual feature data (such as image, color, element layout) will be queried, and the visual generation model will be called with the new prompt words and the original visual feature data as input, so that the newly generated image of the visual generation model can be visually consistent with the original story picture book. Finally, the service call module packages the adjusted text data and the picture book image into an edited story picture book and feeds it back to the controller, and the controller controls the display to display the edited story picture book. During the entire process, users only need to input the corresponding picture book editing instructions on the display device to edit and modify the story picture book. Moreover, no matter how the user edits the story picture book, the newly generated story picture book can maintain consistency with the original story picture book as much as possible in terms of vision and text, meeting the user's personalized needs and saving the user from repeatedly adjusting the picture book creation needs, greatly improving the user's interactive experience.
[0012] In a third aspect, some embodiments provide a display device, including: a display and a controller. The display is configured to display a story picture book; the input interface is configured to receive interactive content input by a user, and send the interactive content to the controller, wherein the interactive content includes picture book editing content; the controller is configured to: construct a picture book editing instruction based on the picture book editing content, wherein the picture book editing instruction carries identification data of the picture book editing content and the story picture book, send the picture book editing instruction to a service call module, receive the edited story picture book fed back by the service call module, and control the display to display the edited story picture book;
[0013] The service calling module is configured to: respond to the picture book editing instruction, call the picture book creation service based on the identification data and the picture book editing content, so that the picture book creation service adjusts the text data in the story picture book according to the identification data and the picture book editing content, determines prompt words based on the adjusted text data, and the prompt words are used to guide the visual generation model to generate images; query the visual feature data of the story picture book according to the identification data, call the visual generation model with the visual feature data and the prompt words as input, generate an adjusted picture book image, and package the adjusted text data and the adjusted picture book image into an edited story picture book and feed it back to the controller.
[0014] The solutions of the above embodiments have the following advantages or beneficial effects:
[0015] The display device displays the story picture book for the user to read. After receiving the picture book editing content input by the user, the display device generates a picture book editing instruction carrying the specific picture book editing content and identification data, and forwards the picture book editing instruction to the service calling module. On the one hand, the service calling module calls the picture book creation service to adjust the text data of the original story picture book according to the identification data of the story picture book and the editing content provided by the user, and then determines the prompt words used to guide the visual generation model to generate images based on the adjusted text data. These prompt words not only reflect the changes in the text, but also inherit the style and other visual characteristics of the original story picture book, which can guide the visual generation model to maintain consistency when generating images; on the other hand, the visual feature data (such as image, color, element layout) of the original story picture book is also queried, and the visual generation model is called with the prompt words and the original visual feature data as input, so that the newly generated image of the visual generation model can be visually consistent with the original story picture book. Finally, the service calling module packages the adjusted text data and the picture book image into an edited story picture book, and feeds it back to the controller, and the controller controls the display to display the edited story picture book. During the entire process, the user only needs to input the corresponding picture book editing instructions on the display device to edit and modify the story picture book. Moreover, no matter how the user edits the story picture book, the newly generated story picture book can maintain consistency with the original story picture book as much as possible in terms of vision and text, meeting the user's personalized needs and saving the user from repeatedly adjusting the picture book creation needs, greatly improving the user's interactive experience.
[0016] In a fourth aspect, some embodiments further provide a server, comprising: a communication module and a processor. The communication module is configured to establish a communication connection with a display device; the processor is configured to: receive a picture book creation service call request sent by the display device, the picture book creation service call request carries the identification data of the story picture book and the picture book editing content, adjust the text data in the story picture book according to the identification data and the picture book editing content, determine a prompt word based on the adjusted text data, the prompt word is used to guide the visual generation model to generate an image, query the visual feature data of the story picture book according to the identification data, and call the visual generation model with the visual feature data and the prompt word as input to generate an adjusted picture book image. The adjusted text data and the adjusted picture book image are packaged as an edited story picture book and fed back to the display device.
[0017] The solutions of the above embodiments have the following advantages or beneficial effects:
[0018] A picture book creation service is deployed on the server. After receiving the picture book creation service call request sent by the display device, on the one hand, the text data of the original story picture book is adjusted according to the identification data of the story picture book and the editing content provided by the user, and then based on the adjusted text data, the prompt words used to guide the visual generation model to generate images are determined. These prompt words not only reflect the changes in the text, but also inherit the style and other visual characteristics of the original story picture book, which can guide the visual generation model to maintain consistency when generating images. On the other hand, the visual feature data (such as image, color, element layout) of the original story picture book is queried, and the visual generation model is called with the prompt words and the original visual feature data as input, so that the newly generated image of the visual generation model can be visually consistent with the original story picture book. Finally, the adjusted text data and picture book image are packaged into an edited story picture book and fed back to the display device so that the user can view the edited story picture book in time. During the entire process, users only need to input the corresponding picture book editing instructions on the display device to edit and modify the story picture book. Moreover, no matter how the user edits the story picture book, the newly generated story picture book can maintain consistency with the original story picture book as much as possible in terms of vision and text, meeting the user's personalized needs and saving the user from repeatedly adjusting the picture book creation needs, greatly improving the user's interactive experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0020] Figure 1 A schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application;
[0021] Figure 2 A schematic diagram of the hardware configuration of a display device provided in some embodiments of the present application;
[0022] Figure 3 A schematic diagram of the hardware configuration of a control device provided in some embodiments of the present application;
[0023] Figure 4 A schematic diagram of software configuration of a display device provided in some embodiments of the present application;
[0024] Figure 5 A schematic diagram of a user interface provided for some embodiments of the present application;
[0025] Figure 6 A schematic diagram of an interface for a picture book editing interface provided in some embodiments of the present application;
[0026] Figure 7 A schematic diagram of an interface for displaying content in a story picture book provided in some embodiments of the present application;
[0027] Figure 8 A schematic diagram of a flow chart of a story picture book editing method executed by a server processor provided in some embodiments of the present application;
[0028] Fig. 9 An interactive sequence diagram of a story picture book editing method provided in some embodiments of the present application;
[0029] Fig.10 A schematic diagram of a flow chart of a story picture book editing method performed by a display device provided in some embodiments of the present application;
[0030] Fig.11 A schematic diagram of a flow chart of a story picture book editing method performed by a display device provided in some other embodiments of the present application;
[0031] Fig.12 A schematic diagram of a process flow of a story picture book editing method executed by a server provided in some embodiments of the present application;
[0032] Fig.13A schematic diagram of a flow chart of a story picture book editing method executed by a server provided in some other embodiments of the present application;
[0033] Fig.14 A detailed flowchart of a story picture book editing method executed by a server provided in some embodiments of the present application;
[0034] Fig.15 A detailed flowchart of a story picture book editing method executed by a server provided in some other embodiments of the present application. DETAILED DESCRIPTION
[0035] The following embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following embodiments do not represent all implementations consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application as detailed in the claims.
[0036] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.
[0037] The terms "first", "second", "third", etc. in the specification and claims of this application and the above drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances.
[0038] The terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.
[0039] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0040] In the embodiment of the present application, the display device 200 generally refers to a device with image display and data processing capabilities. For example, the display device 200 includes but is not limited to a smart TV, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.
[0041] Figure 1This is a schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application. Figure 1 As shown in FIG. 1 , the user can operate the display device 200 through touch operation, the mobile terminal 300 and the control device 100. For example, the control device 100 can be a remote controller, a stylus pen, a handle, etc.
[0042] The mobile terminal 300 can be used as a control device for performing human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device for establishing a communication connection with the display device 200 and performing data interaction. In some embodiments, the mobile terminal 300 can install software applications with the display device 200, and achieve connection and communication through a network communication protocol to achieve the purpose of one-to-one control operation and data communication. The audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 to achieve a synchronous display function.
[0043] like Figure 1 As also shown in FIG. 4 , the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 may be allowed to communicate and connect through a local area network (LAN), a wireless local area network (WLAN) and other networks.
[0044] The display device 200 may provide a broadcast receiving television function, and may also additionally provide an intelligent network television function of a computer support function, including but not limited to network television, smart television, Internet Protocol television (IPTV), etc.
[0045] Figure 2 Some embodiments of the present application provide Figure 1 2 is a block diagram of the hardware configuration of the display device 200.
[0046] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.
[0047] In some embodiments, the detector 230 is used to collect signals of the external environment or external interaction. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of ambient light; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environment scenes, user attributes or user interaction gestures; or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.
[0048] In some embodiments, the display 260 includes a display function component for presenting a screen, a touch component for receiving a user touch operation, and a drive component for driving an image display. The display 260 is used to receive an image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, and components of a menu control interface and a user control UI interface.
[0049] In some embodiments, the communication device 220 is a component for communicating with an external device or server 400 according to various communication protocol types. The display device 200 may be provided with a plurality of communication devices 220 according to different supported communication modes. For example, when the display device 200 supports wireless network communication, the display device 200 may be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.
[0050] The communication device 220 can enable the display device 200 to communicate with the external device or server 400 by wireless or wired connection. Among them, the wired connection can connect the display device 200 with the external device through components such as data cables and interfaces. The wireless connection can connect the display device 200 with the external device through wireless signals or wireless networks. The display device 200 can establish a connection relationship with the external device directly, or indirectly establish a connection relationship through a gateway, a router, a connection device, etc.
[0051] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first interface to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to the user's operation through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.
[0052] In some embodiments, the controller 250 and the tuner-demodulator 210 may be located in different separate devices, that is, the tuner-demodulator 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.
[0053] In some embodiments, the user may input a user command through a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).
[0054] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or may be an external audio output device of the display device 200. In particular, for the external audio output device of the display device 200, the display device 200 may also be provided with an external audio output terminal, and the audio output device may be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.
[0055] In some embodiments, the user input interface 280 may be used to receive instructions from a user.
[0056] Figure 3 Some embodiments of the present application provide Figure 1 The hardware configuration diagram of the control device in the figure. Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.
[0057] The control device 100 is configured to control the display device 200 , and can receive user input operation instructions, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, playing the role of an interactive intermediary between the user and the display device 200 .
[0058] In some embodiments, the control device 100 may be a smart device, for example, the control device 100 may be installed with various applications for controlling the display device 200 according to user needs.
[0059] In some embodiments, Figure 1 As shown, the mobile terminal 300 or other intelligent electronic devices can play a similar function as the control device 100 after installing the application for controlling the display device 200 .
[0060] The controller 110 includes a processor 112, a RAM 113, a ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between the internal components and the external and internal data processing functions.
[0061] The communication interface 130 implements communication of control signals and data signals with the display device 200 under the control of the controller 110. The communication interface 130 may include at least one of other near field communication modules such as a WiFi chip 131, a Bluetooth module 132, and an NFC module 133.
[0062] The user input / output interface 140 , wherein the input interface includes at least one of other input interfaces such as a microphone 141 , a touch panel 142 , a sensor 143 , and a button 144 .
[0063] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with a communication interface 130, such as a WiFi, Bluetooth, NFC or other module, and can encode the user input command through the WiFi protocol, Bluetooth protocol, or NFC protocol and send it to the display device 200.
[0064] The memory 190 is used to store various operating programs, data and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.
[0065] The power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller.
[0066] In order to perform user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system may provide a user interface (control the display device), allow the user to interact with the display device 200, and support the running of various application programs.
[0067] It should be noted that the operating system may be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.
[0068] The operating system can be divided into different modules or layers according to the functions implemented, such as Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom, namely, the application layer (Applications layer) (referred to as "application layer"), the application framework layer (Application Framework layer) (referred to as "framework layer"), the system library layer and the kernel layer.
[0069] In some embodiments, the application layer is used to provide services and interfaces for applications so that the display device 200 can run applications and interact with users based on applications. At least one application can be run in the application layer, and these applications can be window programs, system settings programs, clock programs, etc. that come with the operating system; they can also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the above examples.
[0070] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications. The application framework layer includes some predefined functions. The application framework layer is equivalent to a processing center that determines the actions of applications in the application layer. Applications can access system resources and obtain system services during execution through the API interface.
[0071] like Figure 4 As shown, in the embodiment of the present application, the application framework layer includes a view system, managers, content providers, etc., wherein the view system can design and implement the interface and interaction of the application, and the view system includes lists, grids, text boxes, buttons, etc. The manager includes at least one of the following modules: an activity manager is used to interact with all activities running in the system; a location manager is used to provide system services or applications with access to system location services; a package manager is used to retrieve various information related to the application package currently installed on the device; a notification manager is used to control the display and clearing of notification messages; and a window manager is used to manage icons, windows, toolbars, wallpapers, and desktop components on the user interface.
[0072] In some embodiments, the activity manager is used to manage the life cycle of each application and the usual navigation back function, such as controlling the exit, opening, and back of the application. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, capturing the screen, and controlling the display window changes, for example, reducing the display window, shaking the display, distorting the display, etc.
[0073] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction library contained in the system runtime layer, such as the C / C++ instruction library, to implement the functions to be implemented by the framework layer.
[0074] In some embodiments, the kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. Figure 4As shown, the kernel layer may be configured with hardware drivers, and the drivers included in the kernel layer may be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.
[0075] It should be noted that the above example is only a simple division of the operating system functions and does not constitute a limitation on the specific operating system form of the display device 200 in the embodiment of the present application. Depending on factors such as the function of the display device and the type of operating system, the number of levels and specific level types contained in the operating system may be expressed in other forms.
[0076] With the rapid development of display devices and the increasing diversification of user needs, people's demand for the intelligence of display devices such as smart TVs is also increasing, and the functions of display devices are becoming more and more abundant. At present, users can input picture book creation requirements through a series of interactions with display devices, and the display device will display the story picture book, and users can read the story picture book directly on the display device. In addition, in order to meet the personalized customization needs of users, users are also given the right to edit and rewrite story picture books, that is, when users think that the generated story picture book does not meet their expectations, they can rewrite or continue the story according to their needs.
[0077] However, usually, it is still difficult to maintain the consistency of the content of the story picture books rewritten based on the picture book rewriting requirements input by the user, and the quality is low, which greatly reduces the user's interactive experience.
[0078] In order to solve the above technical problems, the embodiment of the present application provides a story picture book editing method, which is described by taking the application of the method to a server as an example. Figure 5 As shown, the method includes the following steps S802 to S812:
[0079] S802, receiving a picture book creation service call request, where the picture book creation service call request carries identification data of the story picture book and picture book editing content.
[0080] The identification data of a story picture book is information used to uniquely identify the story picture book, including but not limited to the picture book ID (identity), page ID (page ID), title name, and version number, etc. The picture book editing content refers to the specific content that users modify in an existing story picture book, which may involve text, images, typesetting layout, and other aspects.
[0081] Specifically, the user's editing rights include but are not limited to fine-tuning the character name, character image, story summary, and each story content. For example, the user clicks the edit button to enter the story editing page, and can edit the character name, character image, story summary, and each story according to needs. In addition, in addition to editing and rewriting the original picture book data, the user can also expand (i.e. continue) the original picture book data. The specific editing interface can be seen as follows: Figure 6 shown.
[0082] In actual applications, the user opens the picture book creation function by interacting with the display device, and the display of the display device displays the user interface of the picture book creation. The user interface can refer to Figure 7 . Users can input some data for generating story picture books on the user interface according to their own expectations and needs, including inputting or setting data such as character description data, story theme and story summary. Among them, the built-in protagonist name can be selected in the protagonist setting, or the user can define it by himself in the form of text description, or upload a character reference image. After the user enters the corresponding picture book creation requirements, the display device calls the picture book creation service to generate a story picture book based on the picture book creation requirements entered by the user. Specifically, the controller of the display device can call the picture book creation service of the server based on the picture book creation requirements, generate corresponding prompt words for guiding the generation of picture books, and then input the prompt words into the trained large language model to generate a story picture book, and give the story picture book a unique identifier. Finally, the display is controlled to display the generated story picture book for the user to read. It can be understood that the story picture book includes multiple pages, each of which contains corresponding images and text.
[0083] If the user needs to rewrite or continue an existing story picture book, he can select a page of the story picture book he wants to edit and enter the specific editing content in the picture book editing interface. During this process, the user can select the page, paragraph or element to be edited and provide detailed modification instructions. The controller of the display device monitors the actions on the user interface and sends a picture book creation service call request carrying the picture book editing content and the identification data of the story picture book to the server.
[0084] The server receives the picture book creation service call request from the display device through the API gateway, parses the parameters in the request header and request body, and checks the legitimacy of the request, including but not limited to API key verification, user permission verification, etc., to ensure that only authorized devices and users can access the picture book creation service. Subsequently, the identification data of the story picture book (such as the picture book ID) and the picture book editing content (such as text modification, image replacement, etc.) are extracted from the request body.
[0085] S804, adjusting the text data in the story picture book according to the identification data and the picture book editing content.
[0086] The story picture book includes two contents: text data and picture book image. Usually, there may be corresponding text data as a caption next to the picture book image, or the corresponding text data may be included on the picture book image. In this embodiment, the identification data is taken as an example of picture book ID. After extracting the picture book ID, the server can query the original text data of the story picture book from the database or cache according to the picture book ID. Then, in combination with the picture book editing content, the original text is parsed to identify other places such as paragraphs or sentences that need to be modified, and mark them out. Then, the picture book editing content provided by the user is applied to the marked text part to ensure that the modification is in place. Furthermore, the adjusted text data can be checked for consistency and logic to ensure that the modified text maintains the consistency and logic of the story.
[0087] For example, if the user changes the "blue feathers" of the story character bird in the original story picture book to "rainbow feathers", the story character is changed from ['[bird] A bird with bright feathers, with unique blue feathers on its wings'] to ['[bird] A bird with bright feathers, with unique rainbow feathers on its wings'].
[0088] If a story picture book has 10 pages, and the user selects the content of page 6 to be rewritten, when submitting a request to regenerate the picture book, the user can choose to replace only the content of the current page, or to make adaptive adjustments to all pages. If the user chooses to replace only "page 6", the server will query the text data of the page based on the page identifier of page 6, such as pag ID, and adjust the text data of page 6 based on the edited content and the contextual information of page 6. If the user chooses to make adaptive adjustments to all pages, based on the edited content and the contextual information of page 6, the text data of page 6 will be adjusted with emphasis, and the contents of the remaining other pages will be adaptively adjusted.
[0089] S806, determining prompt words based on the adjusted text data, where the prompt words are used to guide the visual generation model to generate images.
[0090] Prompts are keywords or phrases used to guide the work of the visual generation model, which describe the content, style or other characteristics of the desired generated image, and help the model understand the needs of the user. The visual generation model is an algorithmic model based on artificial intelligence that can receive descriptive text input and other parameters (such as visual features) and generate high-quality images accordingly. Specifically, the visual generation model can also be understood as an image generation model, which may include but is not limited to a diffusion model. Prompts are keywords or phrases used to guide the work of the visual generation model, which describe the content, style or other characteristics of the desired generated image, and help the model understand the needs of the user. Exemplarily, the prompts can be "Snow White", "Hug", "Seven Dwarfs". Based on the above prompts, the visual generation model can generate a corresponding image in which Snow White and the seven dwarfs are hugged.
[0091] After the server adjusts the text data of the story picture book according to the editing content of the picture book, it can extract keywords from the adjusted text through natural language processing. These keywords will be used to generate prompt words. At the same time, it can also analyze the overall semantics of the text to infer the emotional color or style elements that may affect the visual expression. Subsequently, the prompt words are generated by combining the keyword and semantic analysis results. These prompt words can accurately guide the visual generation model to generate images that match the text description.
[0092] S808, querying visual feature data of the story picture book according to the identification data.
[0093] Visual feature data is information used to describe the overall visual style of the original story picture book, including but not limited to the color scheme, scene features, typesetting style, and image features of each character. The image features of the character include image features and text description features. For example, the text description feature may be: "Little Dwarf Lele: Male Dwarf, Yellow Curly Hair, Blue Coat, Brown Trousers, Green Hat". The above visual feature data can be used to guide the visual generation model to generate images that are consistent with the original story picture book.
[0094] In actual applications, a visual feature library is maintained in the database of the server, which stores relevant data of each story picture book, including visual feature data such as character image features, color scheme, layout style, theme style, etc. After the server generates a story picture book through the picture book creation service, it will save the visual feature data of the story picture book generated this time to the above visual feature library, that is, the visual feature library stores the visual feature data of the story picture book generated last time.
[0095] After extracting the identification data of the picture book, the server can query the latest visual feature data of the story picture book in the visual feature library based on the identification data.
[0096] S810, taking the visual feature data and the prompt word as input, calling the visual generation model to generate an adjusted picture book image.
[0097] Vision generative models are models for generating images that can receive descriptive text input and other parameters (such as visual features) and generate high-quality images based on them.
[0098] After querying the visual feature data, the server formats the visual feature data and the newly generated prompt words into input data that can be recognized and processed by the visual generation model. Subsequently, the input data is passed to the deployed visual generation model through an API or a direct call. The visual generation model regenerates the adjusted picture book image based on the input visual feature data and prompt words. Specifically, in the process of generating images by the visual generation model, the input visual feature data and prompt words are considered at the same time. The visual feature data and prompt words are input as conditions. In the process of generating each image, information can be dynamically selected and combined to ensure that the generated new image is consistent with the original picture book image in image and style.
[0099] Similarly, if the user chooses to replace only "Page 6", the server will adjust the text data of Page 6 based on the edited content and the contextual information of Page 6, and then extract the prompt words of the adjusted text data of Page 6, and call the visual generation model with the prompt words and visual feature data of Page 6 as input to guide the visual generation model to regenerate the picture book image of Page 6. If the user chooses to make adaptive adjustments to all pages, the server will adjust the text data of Page 6 based on the edited content and the contextual information of Page 6, and make adaptive adjustments to the contents of the remaining other pages, and then extract the prompt words of each page from the adjusted text data of all pages, and call the visual generation model with the prompt words and visual feature data of each page as input to guide the visual generation model to regenerate the picture book images of all pages. It can be understood that if the text data or prompt words of a page have not changed, the picture book image corresponding to the page may not be adjusted.
[0100] Furthermore, after generating a new picture book image, the server may also perform necessary post-processing on the generated image, such as resolution adjustment, color correction, etc., to improve the image quality.
[0101] S812, integrating the adjusted text data and the adjusted picture book image to obtain an edited story picture book.
[0102] After obtaining the adjusted text data and picture book image, the server can integrate the adjusted text data and the newly generated picture book image into a complete story picture book (data file package), the format of which meets the requirements of the real device, and then transmit the story picture book to the display device so that the display device can display the edited story picture book for the user to read. If the user believes that the edited story picture book still needs to be modified, the edited content can be input again, and the display device will send a picture book creation service call request to the server again. The server will feed back the edited story picture book to the display device according to the above picture book editing process, and the display device will re-display the edited story picture book on the user interface until the displayed story picture book meets the user's expectations.
[0103] The solutions of the above embodiments have the following advantages or beneficial effects:
[0104] After receiving the picture book creation service call request, on the one hand, the original text data is adjusted according to the identification data of the story picture book and the editing content provided by the user, and then based on the adjusted text data, the prompt words used to guide the visual generation model to generate images are determined. These prompt words not only reflect the changes in the text, but also inherit the style and other visual characteristics of the original story picture book, which can guide the visual generation model to maintain consistency when generating images. On the other hand, the original visual feature data (such as image, color, element layout) is queried, and the visual generation model is called with the prompt words and the original visual feature data as input, so that the newly generated image of the visual generation model can be visually consistent with the original story picture book. Finally, the adjusted text data and picture book image are packaged into an edited story picture book and fed back to the display device so that the user can view the edited story picture book in time. During the entire process, users only need to input the corresponding picture book editing instructions on the display device to edit and modify the story picture book. Moreover, no matter how the user edits the story picture book, the newly generated story picture book can maintain consistency with the original story picture book as much as possible in terms of vision and text, meeting the user's personalized needs and saving the user from repeatedly adjusting the picture book creation needs, greatly improving the user's interactive experience.
[0105] like Figure 8 As shown, in some embodiments, S804 includes: S824, based on the picture book editing content, updating the text prompt words, taking the updated text prompt words as input, calling the first language model, generating the text data of the adjusted story picture book, and the first language model is used to generate the text data of the story text.
[0106] Text prompt words are key descriptive words used to guide the large language model to generate texts with specific content. Prompt words usually contain key elements such as characters, actions and positions in the story. Large Language Model (LLM) is a large-scale language model specially trained to generate story texts, which can generate coherent and expected texts based on given prompt words. Large language models may include but are not limited to large language models obtained by training and fine-tuning large language models such as GPT, Wenxin XX or Tongyi XX, and the large language model has functions such as intelligent customer service, machine translation, content creation, and image recognition. In this embodiment, in order to distinguish from other large language models, "first" and "second" are used to distinguish. In practical applications, since the first large language model needs to be given corresponding text prompt words to generate story texts, each story text will have its corresponding text prompt words. After generating the story text, the server can associate and store the story text and the text prompt words.
[0107] When implementing it, Fig. 9 As shown, after obtaining the story ID, the server can search for the original text prompt words of the corresponding story picture book in the database or cache based on the story ID, and then parse the picture book editing content entered by the user, determine which text prompt words need to be modified, and apply the picture book editing content to the text prompt words to update the text prompt words. Furthermore, it can also be verified whether the updated text prompt words maintain the consistency and logic of the original story text. Subsequently, the updated text prompt words are converted into a format that can be recognized and processed by the first language model, and then the converted text prompt words are input into the first language model, and the story text is generated again by the first language model.
[0108] If the user continues writing based on the original story text, the text prompt words for the continued content can be determined again based on the continued content input by the user, and the text prompt words for the continued content can be input into the first largest language model. The continued text is generated by the first largest language model, and the original story text and the continued text are integrated to obtain a new story text.
[0109] The solution of the above-mentioned embodiment has the following advantages or beneficial effects: through the large language model, it is possible to efficiently complete the tasks from obtaining prompt words to generating new story texts, and the edited story picture book can not only meet the user's modification intentions, but also maintain the original text quality and style consistency.
[0110] like Figure 8 As shown, in some embodiments, S806 includes: S826, taking the adjusted text data as input, calling the second largest language model, extracting prompt words from the adjusted text data, and obtaining new prompt words.
[0111] In this embodiment, the second largest language model can also be a large language model obtained by training and fine-tuning on a large language model such as GPT, Wenxin XX or Tongyi XX. The trained second largest language model can extract prompt words from text data to guide the visual generation large model to generate images.
[0112] In the traditional solution of generating story picture books based on a large language model, after the large language model generates a new story text, the generated story text is usually directly input into the visual generation large model, so that the visual generation large model generates the corresponding picture book image. However, due to the lack of guidance from the image features of the original story picture book, it is usually difficult to maintain the consistency of the regenerated image with the original story picture book image. In this embodiment, a second large language model is introduced, and the prompt word is extracted from the story text output by the first large language model through the second large language model. The prompt word retains the visual characteristics of the original story text and also reflects the changes in the text, which can guide the visual generation large model to maintain consistency with the original story picture book when generating images.
[0113] When implementing it, Fig. 9 As shown, after obtaining the adjusted text data, the server formats the adjusted text data into an input format that can be recognized and processed by the second largest language model, and then inputs the adjusted text data into the second largest language model through an API or a direct call, and the second largest language model extracts important prompt words from the adjusted text data. Specifically, the second largest language model can extract prompt words from the adjusted text data by natural language processing, and the extracted prompt words are standardized words that meet the input requirements of the visual generation large model.
[0114] For example, if the text data is "Snow White and the dwarfs are running in the forest, the sunlight shines through the treetops onto the ground, forming mottled light and shadows", the extracted prompt words may include "Snow White, dwarfs, forest, running, sunlight, treetops, light and shadows.
[0115] After extracting the prompt word, the server uses the prompt word and visual feature data as input to call the visual generation model, so that the visual generation model generates a picture book image that maintains a high degree of consistency with the original picture book image. The server feeds back the adjusted text data and picture book image to the display device so that the display device can display the edited story picture book.
[0116] The solution of the above embodiment has the following advantages or beneficial effects: the second largest language model can effectively extract high-quality prompt words from the adjusted text data, provide accurate guidance for subsequent image generation tasks, so that the generated image remains consistent with the original picture book image.
[0117] like Figure 8 As shown, in some embodiments, S810 includes: S820, encoding the visual feature data and the prompt word respectively to obtain a visual feature vector and a prompt word vector, integrating the visual feature vector and the prompt word vector to obtain a conditional vector, inputting the conditional vector into the visual generation model, using the conditional vector as a conditional input to guide the visual generation model, and generating an adjusted picture book image.
[0118] The visual feature vector is a numerical representation obtained by encoding the visual feature data, which comprehensively reflects the color, layout, style and other information of the image. The prompt word vector is a numerical representation obtained by encoding the prompt word through natural language processing, which captures the key description or instructions of the text content. The conditional vector is a vector formed by integrating the visual feature vector and the prompt word vector, which is used to guide the image generation process of the visual generation model. The conditional vector contains all the key information required to generate the target image, such as visual elements such as character image, style, layout, color scheme, and semantic information in the text description. The conditional vector can make the generated image conform to the changes in the story content while maintaining the original visual consistency. For example, if the original picture book adopts a specific color scheme or artistic style, this information will be retained in the conditional vector and guide the process of generating images by the visual generation model. If the story describes a specific scene or character, the prompt word will guide the visual generation model to generate the corresponding visual elements.
[0119] In the specific implementation, the visual feature data can be encoded by a pre-trained encoder and converted into a visual feature vector, which can be a set of numerical feature vectors. The prompt word is encoded according to the pre-trained language model and encoded as a prompt word vector. Subsequently, the visual feature vector and the prompt word vector are integrated to obtain a conditional vector. Specifically, the visual feature vector and the prompt word vector can be integrated by splicing or weighted summation to obtain a conditional vector. Then, the conditional vector is formatted into a format that can be recognized and processed by the visual generation model, and the conditional vector is passed to the visual generation model through an API or a direct call. The conditional vector is used as a conditional input to guide the visual generation model to generate an adjusted picture book image. Specifically, the visual generation model will always take the guidance provided by the conditional vector into account the image's character image features, style, layout, color and other visual elements in the image generation process, and combine the content of the text prompt word to ensure that the generated image is consistent with the edited story text and maintains the original visual style. Furthermore, if multiple images need to be generated, the server can also integrate the picture book images generated by the visual generation model in the order of pages to obtain a continuous story picture book image.
[0120] The solution of the above-mentioned embodiment has the following advantages or beneficial effects: by integrating the visual feature data and the prompt words into a conditional vector and using the conditional vector to guide the visual generation model, the picture book image generated by the visual generation model can be made consistent with the original picture book image.
[0121] like Fig.10 As shown, in some embodiments, S810 includes: S830, encoding the visual feature data and the prompt word respectively to obtain a visual feature vector and a prompt word vector, integrating the visual feature vector and the prompt word vector to obtain a conditional vector, inputting the conditional vector into a diffusion model based on a cross-attention mechanism, and using the conditional vector as a condition to guide the diffusion model to generate an adjusted picture book image, wherein the query variable in the cross-attention mechanism is the feature vector of the currently generated image, and the key variable and the value variable are the visual feature vector.
[0122] In this embodiment, the large model of visual generation is taken as a diffusion model based on the cross-attention mechanism as an example. The diffusion model is a generative model that generates high-quality data samples by gradually adding noise and then learning to denoise. The cross-attention mechanism is an attention mechanism that allows the model to pay attention to information of different modalities during the generation process. In this embodiment, the cross-attention mechanism makes corresponding adjustments, specifically for the query variable (Query, Q variable), key variable (Key, K variable) and value variable (Value, V variable), where the Q variable is the content feature of the current image in the current generation process, and the K variable and V variable are used to capture the degree of association between the context and different parts of the current image. Specifically, the K variable is a pre-calculated image feature related to the original picture book, which provides a reference for the current generated image. The model evaluates the similarity of the Q variable and the K variable to find which queried image features best match the needs of the current generated image content. The V variable is used to provide detailed information when the match is successful. Once it is determined which image features are most suitable for the needs of the current generated image, the V variable will give specific visual details or style guidance to enrich the quality of the generated image.
[0123] Specifically, unlike the traditional attention mechanism in which the Q variable, K variable and V variable all use the feature vector of the currently generated image, in this embodiment, the image feature data of the currently generated image is used as the query variable (Q variable), while the K variable and V variable use the visual feature vector in the conditional vector.
[0124] Similar to the previous embodiment, the server integrates the visual feature vector and the prompt word vector to obtain the conditional vector, and inputs the conditional vector into the diffusion model based on the cross-attention mechanism. The diffusion model starts with a random noise image, and gradually iteratively reduces the noise as the time step increases until a clear image is generated. In each round of iterative denoising, the cross-attention mechanism allows the model to adjust the generation strategy according to the information in the conditional vector. Specifically, the image feature data of the current generated image is used as the query variable (Q variable), and the visual feature data in the conditional vector is used as the K variable and the V variable. The model calculates the similarity between the Q variable and the K variable, and weights and aggregates the V variable according to this similarity, thereby guiding the update of the current image feature. The diffusion model learns how to recover meaningful image features from noise, and gradually removes noise until the final clear image is generated. Furthermore, the server can also post-process the generated picture book image, such as color correction, resolution adjustment, etc., to improve the quality of the picture book image.
[0125] In other embodiments, if the user wants to adjust the style of the picture book, such as adjusting the style from "comic" style to "ancient style", the picture book editing content will include style prompts such as "ancient style". The server can then input the style prompts, prompt words and visual features into the visual generation model to guide the visual generation model to generate the image style adjusted to the "ancient style". The specific picture book image generation process can refer to the specific process in the above embodiment, which will not be repeated here.
[0126] The technical solution of the above embodiment has the following advantages or beneficial effects: by generating picture book images through a diffusion model based on a cross-attention mechanism, the model is allowed to generate images by combining visual features and text prompt information, so that the generated picture book images can not only conform to the changes in the story content, but also maintain the original visual consistency.
[0127] like Fig.11 As shown, in some embodiments, S810 also includes: S840, when the picture book editing content includes a reference image, extracting visual features in the reference image to obtain reference image features, updating visual feature data according to the reference image features, and using the updated visual feature data and prompt words as input, calling the visual generation model to generate an adjusted picture book image.
[0128] The reference image is an example image provided by the user to guide the visual generation model to generate images. It contains the visual style, layout or specific elements of the picture book image that the user expects. The reference image feature is a numerical representation obtained by encoding the reference image, reflecting the visual features of the image such as color, texture, shape, etc.
[0129] In actual applications, if the user uploads a corresponding reference image and edits the text in the picture book editing interface, the server receives the picture book editing content input by the user, including text input and reference images, and then preprocesses the reference image such as resizing and normalization to ensure that it meets the input requirements of the encoder. Subsequently, the reference image is encoded through the pre-trained visual feature extraction encoder, and the visual features therein are extracted to obtain the reference image features. Based on the reference image features, the queried visual feature data is then updated, including corresponding replacement of some of the visual features based on the reference image features. For example, if the user only uploads a character reference image of a certain character, the character reference image features are extracted from the character reference image, and then the image features of the character are queried based on the character identifier such as the character name and the character ID. Subsequently, the image features of the queried character are replaced with the reference image features of the character, and the remaining visual feature data that is not involved in the change can be left unprocessed.
[0130] After the visual feature data is updated, the updated visual feature data and the prompt word are used as input to call the visual generation model to generate the adjusted picture book image. The specific image generation process can be found in the specific description of the above embodiment, which will not be repeated here.
[0131] The technical solution of the above embodiment has the following advantages or beneficial effects: by extracting the image features of the reference image and updating the visual feature data based on the reference image features, the generated picture book image can be made consistent with the effect expected by the user, thereby achieving personalized customization.
[0132] In some embodiments, the present application further provides a picture book story generation method, which is applied to a display device, wherein the display device includes a display, a controller, and a service calling module. Fig.12 As shown, the method comprises the following steps:
[0133] S702, receiving picture book editing content input by the user for the story picture book.
[0134] Picture book editing content refers to the specific content of the user's modification of the existing story picture book, which may involve text, images, typesetting layout and other aspects.
[0135] In specific implementation, the user previews the generated initial story text through the user interface on the display device. Then, the user selects a page of the story picture book that he wants to edit and enters the specific editing content. In this process, the user can select the page, paragraph or element to be edited and provide detailed modification instructions. After the user enters the editing content, he can click the "Regenerate" button. At this time, the display device receives and records the editing content entered by the user in real time.
[0136] S704, based on the picture book editing content and the identification data of the story picture book, call the picture book creation service, so that the picture book creation service adjusts the text data in the story picture book according to the identification data and the picture book editing content, determines the prompt words used to guide the visual generation large model to generate images based on the adjusted text data, queries the visual feature data of the story picture book according to the identification data, calls the visual generation large model with the visual feature data and the prompt words as input, generates the adjusted picture book image, and packages the adjusted text data and the adjusted picture book image into an edited story picture book for feedback.
[0137] In specific implementation, the controller of the display device may monitor actions on the user interface, and construct a picture book editing instruction based on the picture book editing content input by the user. The instruction includes the editing content provided by the user (such as new text, image file, etc.) and the identification data of the selected picture book (such as picture book ID, version number, etc.), and pass the picture book editing instruction to the service call module. After receiving the picture book editing instruction, the service call module parses the instruction content and extracts important parameters, such as identification data (picture book ID), editing content (new / modified text), etc., and then executes corresponding operations to generate a new story picture book, and then feeds the new story picture book back to the controller, which controls the display to display the edited story picture book.
[0138] Specifically, the service call module may extract the identification data and picture book editing content after receiving the picture book editing instruction containing identification data and picture book editing content, and construct an API request or RPC (Remote Procedure Call) call to send a picture book creation service request to the picture book creation service. The picture book creation service uses the identification data to locate the story picture book selected by the user, and applies the picture book editing content to update the text data in the story picture book. Subsequently, new prompt words are regenerated based on the adjusted text data. These prompt words will be used to guide the subsequent image generation process. At the same time, the picture book creation service uses the identification data to query the visual feature data related to the picture book data from the database or cache, and uses the queried visual feature data and the newly generated prompt words as input, and calls the visual generation model through the API or direct call to generate the adjusted picture book image. Specifically, in the process of generating images using the visual generation model, the input visual feature data and prompt words are considered simultaneously. The visual feature data and prompt words are input as conditions, and information can be dynamically selected and combined in the generation process of each image, thereby ensuring that the generated new image is consistent with the original picture book image in image and style.
[0139] Next, the picture book creation service integrates the adjusted text data and the newly generated picture book image into a complete edited story picture book, and returns the edited story picture book to the service call module of the display device, which then passes it back to the controller, and the controller display displays the edited story picture book on the user interface. The user can read the edited story picture book, and if the edited story picture book still needs to be modified, the edited content can be input again, and the controller can send a picture book editing request to the service call module again, and the service call module, according to the picture book editing process in the above embodiment, re-feeds back the edited story picture book to the controller.
[0140] It is understandable that if the picture book creation service is a service deployed on a remote server, the generation process of the new story picture book will be executed on the server, and the server will send the packaged edited story picture book to the service call module of the display device, and the service call module will pass it back to the controller, and the controller display will display the edited story picture book. If the picture book creation service is a local service of the display device, the above interaction is the interaction between different modules inside the display device.
[0141] S706, receiving the edited story picture book fed back by the picture book creation service, and displaying the edited story picture book.
[0142] The display device receives the edited story picture book feedback from the picture book creation service, and then controls the display to re-display the edited story picture book on the user interface for the user to read. If the user has modification requirements, the user can continue to submit editing instructions until the displayed story picture book meets the user's expectations.
[0143] like Fig.13 As shown, in some embodiments, the method also includes: S701, obtaining voice data sent by the user, performing intent recognition on the voice data, determining the user's interaction intention, and when the interaction intention represents editing a picture book, displaying a picture book editing interface for the user to input picture book editing content.
[0144] Intent recognition refers to determining the user's specific intention or needs by analyzing the interactive content input by the user. Interaction intent is the result of intention analysis of interactive content, and it reflects the goal that the user hopes to achieve in a specific situation.
[0145] In this embodiment, the display device supports multi-modal interaction, that is, the user can interact with the display device in a variety of interaction modes and input interaction content. Specifically, the interaction modes include touch interaction, voice interaction, gesture interaction, auditory interaction, and visual interaction.
[0146] Taking voice interaction as an example, the display device may have a built-in or external microphone, which captures the user's voice data and sends the voice data to the controller. The controller converts the user's voice into text or commands through voice recognition, analyzes the user's intention and then performs the corresponding operation. The recognition operation of the user's voice data can refer to the relevant technology, and the embodiments of this application will not be described one by one.
[0147] Optionally, the user can control the display device to enter the voice control mode by operating a designated button of the remote controller, or can control the display device to enter the voice control mode by voice.
[0148] Optionally, when the display device is triggered to enter the voice control mode, the user can also send instructions to the display device in text form through a mobile phone, remote control or other device to prevent the display device from being unable to receive the user's voice commands when there is a problem with the microphone.
[0149] In specific implementation, the microphone collects the voice data input by the user and sends the collected voice data to the controller, which converts the user's voice into text or commands through voice recognition. When the user is recognized to say a specific wake-up word such as "XiaoXiaoX", the display device is controlled to turn on the voice interaction mode. The microphone collects the voice data input by the user in real time, and the controller performs intent recognition on the voice data and performs corresponding operations. For example, if the user says: "Edit story picture book" or "Modify story picture book", the microphone collects the voice data input by the user and sends it to the controller. The controller can call a trained intent recognition model (such as a model based on deep learning) to perform intent recognition on the voice data, convert the voice data into text or commands, and recognize that the user's interaction intention is "Edit picture book". Then, the display is controlled to display the picture book editing interface, allowing the user to select the specific picture book they want to edit, and enter the corresponding picture book editing content on the picture book editing interface.
[0150] Optionally, the display device can play an audio prompt or display text instructions on the screen to tell the user how to proceed, for example: "Please select the picture book you want to edit and the specific picture book page you want to edit", or provide more detailed instructions to help the user understand the available editing options. The controller is ready to receive more voice commands, allowing the user to select picture books, specify editing areas, add content, etc. by speaking. For example, the user can say: "Open the first picture book", or "Add a picture to the second page." Once the user selects a specific picture book, the controller will display the content of the picture book and provide intuitive editing tools, such as entering text, dragging and dropping elements, resizing, changing colors, etc. After receiving the picture book editing content entered by the user, the picture book editing instructions are constructed based on the picture book editing content, and the picture book editing instructions are sent to the service call module to generate a new story picture book.
[0151] The above technical solution has the following advantages or beneficial effects: the display device supports voice interaction, can accurately understand the user's intention, and can also guide the user to easily complete the entire editing process. The user can interact with the display device through natural voice input without complicated operation steps, which simplifies the interaction process and greatly improves the interaction experience.
[0152] In some embodiments, the method further includes: in response to a user's editing operation on the story content, setting the editing status of the story content to an editable state, and updating the editing status of the story characters and story outline to a non-editable state, and the picture book editing content includes the edited story content.
[0153] Following the above embodiment, the user's editing rights include story roles (including role names and role images), story summaries, and adjustments to each story content. In actual applications, after the user modifies the story roles and summaries, the story content is regenerated by default, and the story content is corrected according to the user's modifications. After the user modifies the story content, in principle, the user is no longer allowed to modify the story roles and story summaries.
[0154] In specific implementation, if the user selects the story content to be modified, the controller responds to the user's editing operation on the story character, sets the editing status of the story content to an editable state, and controls the display to normally display the area where the story content is located on the picture book editing interface, that is, the user is allowed to edit the story content of the story picture book. At the same time, the editing status of the story character and the story summary is set to a non-editable state, and the control display sets the area where the story character and the story summary are located on the picture book editing interface to gray display (disabled), that is, the user is not allowed to modify the story character and the story summary. The input interface receives the story content rewrite data input by the user in real time and sends it to the controller. If the controller detects that the user attempts to modify a disabled story character or story summary part, the controller may pop up a prompt box or push a prompt message on the picture book editing interface to remind the user that these parts are not editable. In addition, in order to maintain traceability and support the undo function, the controller may record detailed information of this editing operation, including the specific modified content and location.
[0155] Regarding the modification of the character name, the user is only allowed to modify it from the first sentence of the character name, and the modification of character names other than the first sentence of the character name is not allowed. Specifically, if the user wants to modify the character name, the controller responds to the editing operation for the character name, determines whether the area of the character name selected by the user is the first sentence of the character name in the story picture book, and if so, allows the user to modify it from the first sentence of the character name. If it is not the first sentence of the character name, a prompt box will pop up or a prompt message will be pushed in the picture book editing interface to remind the user to edit at the first sentence of the character name. After the user makes the modification, the character names in other locations will be automatically replaced with the modified first sentence of the character name.
[0156] In some embodiments, if the user selects the story content to be modified, after the controller sets the editing status of the story content to an editable state, the display can be controlled to highlight the area where the story content is located on the picture book editing interface, and at the same time, the area where the story characters and story outline are located on the picture book editing interface is set to gray display, so that the user can intuitively understand which areas on the picture book editing interface are editable and which areas are not editable.
[0157] The above technical solution has the following advantages or beneficial effects: when the user modifies the story content, the controller sets the story characters and story outline to an uneditable state, so that the user can only edit the parts that are allowed to be modified without affecting the content of other areas. It also provides a clear user interface and instant feedback, which improves the user experience while maintaining the consistency and integrity of the story.
[0158] In some embodiments, the method also includes: based on the identification data, obtaining the layout data of the story picture book, based on the layout data, layout the adjusted text data and the adjusted picture book image, and controlling the display to respectively display the layout text data and picture book image on the user interface.
[0159] In actual applications, since a story picture book contains multiple pages, each page has text and a picture book image. The controller needs to layout these pages and control the display to display the layout pages on the user interface in sequence to display the entire story picture book.
[0160] In this embodiment, the layout mode includes but is not limited to single-column layout, double-column layout or grid layout. In the single-column layout, each storyboard occupies a whole line, with the text on the top (or bottom) and the picture book image on the bottom (or top). Fig.14 As shown. In other embodiments, the text may be on the left (or on the right) and the picture book image may be on the right (or on the left). It is understandable that the grid layout is more suitable for large-screen displays. The above layout method can automatically adjust the layout according to the screen size, or it can be selected by the user, and is not limited here.
[0161] In specific implementation, since the typesetting data, prompt words, content and other information related to each story picture book are associated and stored in the memory or database. After receiving the edited story picture book fed back by the service call module, the controller can obtain the typesetting data of the story picture book based on the identification data in the story picture book, such as the picture book ID, and then, based on the typesetting data, the adjusted text data and the adjusted picture book image are typeset and laid out. After the controller processes the text and picture book image of each page according to the set typesetting layout, it can control the display to directly display the text and picture book image of each page according to the established typesetting layout. Users can view and manage pages (such as turning pages, scrolling, zooming in or out, etc.) through touch operations or voice operations to enhance the interactive experience.
[0162] The above technical solution has the following advantages or beneficial effects: the controller processes the layout of the storyboard according to the preset layout template and style specifications, ensuring the consistency of the visual effect of each storyboard, improving the overall look and feel and the user's interactive experience.
[0163] In some embodiments, the method also includes: receiving an edited story picture book fed back by a picture book creation service, and if the resolution of the adjusted picture book image does not meet the preset resolution requirement, converting the picture book image into a picture book image that meets the preset resolution requirement, displaying the converted picture book image and adjusting it.
[0164] Specifically, since different models of display devices may have different resolution requirements, if the resolution of a picture book image does not meet the preset resolution requirement (such as the preset resolution is 1080p), the controller will automatically convert the picture book image into an image that meets the preset resolution requirement, and then control the display to display the converted picture book image.
[0165] For example, if the image resolution of one of the storyboards is low, i.e., 720p, and the preset resolution requirement is 1080p, the controller may convert the picture book image to a resolution of 1080p before performing typesetting, layout, and display.
[0166] Optionally, the conversion of low-resolution images to high-resolution images can be achieved through bilinear interpolation, bicubic interpolation, super-resolution reconstruction, wavelet transform, etc. The specific method can be determined according to the actual situation, which will not be discussed here. Super-resolution reconstruction uses deep learning technology to generate higher-resolution images by training models. Wavelet transform can decompose images into sub-bands of different frequencies, then amplify the high-frequency sub-bands, and finally synthesize high-resolution images.
[0167] The above technical solution has the following advantages or beneficial effects: by converting all picture book images into a preset resolution, the display quality of all images on the user interface is ensured to be consistent, display problems caused by inconsistent resolution are avoided, and when users browse different pages, they will not feel abrupt due to changes in image resolution, and the viewing process of the entire story is smoother and more natural.
[0168] In order to solve the above technical problems, an embodiment of the present application provides a display device, which includes a display, a controller and a service calling module.
[0169] The display is configured to display the story picture book on the user interface.
[0170] During specific implementation, the user previews the generated initial story text through a user interface on a display device.
[0171] The controller is configured to: identify the picture book editing content input by the user for the story picture book, construct a picture book editing instruction, and send the picture book editing instruction to the service calling module, the service calling module calls the picture book creation service based on the identification data and the picture book editing content, so that the picture book creation service adjusts the text data in the story picture book according to the identification data and the picture book editing content, determines the prompt words used to guide the visual generation large model to generate images based on the adjusted text data, queries the visual feature data of the story picture book according to the identification data, calls the visual generation large model with the visual feature data and the prompt words as input, generates an adjusted picture book image, packages the adjusted text data and the adjusted picture book image into an edited story picture book and feeds it back to the controller, receives the edited story picture book fed back by the service calling module, and controls the display to display the edited story picture book.
[0172] The service call module acts as a bridge for communication between different parts of the application. Specifically, it is responsible for processing requests from the controller, which may involve calling an API (Application Programming Interface) on a remote server or a local service to complete a specific task. The picture book creation service is the core service for creating and editing digital picture books. It can be a service local to the display device or a service deployed on a remote server.
[0173] In actual applications, if the user has editing needs, he can select a page of the story picture book he wants to edit and enter the specific editing content. After the user enters the editing content, he can click the "Regenerate" button. At this time, the controller recognizes the editing content entered by the user in real time. The controller can monitor the action on the user interface and construct a picture book editing instruction based on the picture book editing content entered by the user. The instruction contains the editing content provided by the user (such as new text, image file, etc.) and the identification data of the selected picture book (such as picture book ID, version number, etc.), and passes the picture book editing instruction to the service call module. After receiving the picture book editing instruction, the service call module parses the instruction content and extracts important parameters, such as identification data (picture book ID), editing content (new / modified text), etc. Then, it performs the corresponding operation to generate a new story picture book, and then feeds the new story picture book back to the controller, and the controller controls the display to display the edited story picture book.
[0174] Specifically, the service call module may extract the identification data and picture book editing content after receiving the picture book editing instruction containing identification data and picture book editing content, and construct an API request or RPC (Remote Procedure Call) call to send a picture book creation service request to the picture book creation service. The picture book creation service uses the identification data to locate the story picture book selected by the user, and applies the picture book editing content to update the text data in the story picture book. Subsequently, the prompt words are regenerated based on the adjusted text data. These prompt words will be used to guide the subsequent image generation process. At the same time, the picture book creation service uses the identification data to query the visual feature data related to the picture book data from the database or cache, and uses the queried visual feature data and the newly generated prompt words as input, and calls the visual generation model through the API or direct call to generate the adjusted picture book image. Specifically, in the process of generating images using the visual generation model, the input visual feature data and prompt words are considered simultaneously. The visual feature data and prompt words are input as conditions, and information can be dynamically selected and combined in the generation process of each image, thereby ensuring that the generated new image is consistent with the original picture book image in image and style.
[0175] Next, the picture book creation service integrates the adjusted text data and the newly generated picture book image into a complete edited story picture book, and returns the edited story picture book to the service call module of the display device, which then passes it back to the controller, and the controller display displays the edited story picture book on the user interface. The user can read the edited story picture book, and if the edited story picture book still needs to be modified, the edited content can be input again, and the controller can send a picture book editing request to the service call module again. The service call module, according to the picture book editing process in the above embodiment, re-feeds back the edited story picture book to the controller, and the controller controls the display to re-display the edited story picture book on the user interface until the displayed story picture book meets the user's expectations.
[0176] It is understandable that if the picture book creation service is a service deployed on a remote server, the generation process of the new story picture book will be executed on the server, and the server will send the packaged edited story picture book to the service call module of the display device, and the service call module will pass it back to the controller, and the controller display will display the edited story picture book. If the picture book creation service is a local service of the display device, the above interaction is the interaction between different modules inside the display device.
[0177] In some embodiments, the display device further includes: a sound collector configured to: obtain voice data sent by a user, and send the voice data to the controller.
[0178] The controller is further configured to: perform intent recognition on voice data, determine the user's interaction intention, and when the interaction intention represents editing a picture book, control the display to display a picture book editing interface for the user to input picture book editing content, construct picture book editing instructions based on the picture book editing content, and send the picture book editing instructions to the service call module.
[0179] In this embodiment, the sound collector can be a microphone or other device capable of collecting sound signals. The display device supports multimodal interaction, that is, the user can interact with the display device in a variety of interactive ways and input interactive content. Specifically, the interactive methods include touch interaction, voice interaction, gesture interaction, auditory interaction, and visual interaction. The specific intention recognition and voice interaction process refer to the above embodiment, which will not be repeated here.
[0180] The above technical solution has the following advantages or beneficial effects: the display device supports voice interaction, can accurately understand the user's intention, and can also guide the user to easily complete the entire editing process. The user can interact with the display device through natural voice input without complicated operation steps, which simplifies the interaction process and greatly improves the interaction experience.
[0181] In some embodiments, the controller is further configured to: in response to the user's editing operation on the story content, control the display to display the story content on the user interface, set the editing status of the story content to editable, and update the editing status of the story characters and story outline to a non-editable state, and the picture book editing content includes the edited story content.
[0182] In some embodiments, the controller is further configured to: obtain layout data of the story picture book based on the identification data, layout the adjusted text data and the adjusted picture book image based on the layout data, and control the display to respectively display the layout text data and picture book image on the user interface.
[0183] In some embodiments, the controller is further configured to: when the resolution of the adjusted picture book image does not meet the preset resolution requirement, convert the picture book image into a picture book image that meets the preset resolution requirement, and control the display to display the converted picture book image.
[0184] Since the above-mentioned display device embodiment is the display device of the above-mentioned story picture book editing method embodiment that executes the display device, the specific limitations in one or more of the display device embodiments provided above can refer to the above limitations on the display device that executes the above-mentioned story picture book editing method embodiment, and will not be repeated here.
[0185] In some embodiments, the embodiments of the present application further provide a server, the server comprising: a communication module and a processor, the communication module being configured to establish a communication connection with a display device.
[0186] like Fig.15 As shown, the processor is configured to perform the following steps S602 to S612, wherein:
[0187] S602, receiving a picture book creation service call request sent by a display device, where the picture book creation service call request carries identification data of a story picture book and picture book editing content.
[0188] S604, adjusting the text data in the story picture book according to the identification data and the picture book editing content.
[0189] S606, determining prompt words based on the adjusted text data, where the prompt words are used to guide the visual generation model to generate images.
[0190] S608, querying visual feature data of the story picture book according to the identification data.
[0191] S610, using the visual feature data and the new prompt word as input, calling the visual generation model to generate an adjusted picture book image.
[0192] S612, packaging the adjusted text data and the adjusted picture book image into an edited story picture book, and feeding back the edited story picture book to a display device.
[0193] In some embodiments, when the processor executes the step of adjusting the text data in the story picture book according to the identification data and the picture book editing content, the processor is further configured to: obtain the text prompt words of the story picture book based on the story identification in the identification data, update the text prompt words based on the picture book editing content, and use the updated text prompt words as input to call the first language model to generate adjusted text data of the story picture book, and the first language model is used to generate text data of the story text.
[0194] In some embodiments, when the processor executes the step of determining new prompt words based on the adjusted text data, the processor is further configured to: take the adjusted text data as input, call the second largest language model, extract prompt words from the adjusted text data, and obtain new prompt words. The second largest language model is used to generate prompt words that guide the visual generation model to generate images.
[0195] In some embodiments, when the processor executes the steps of calling the visual generation model with visual feature data and new prompt words as input to generate an adjusted picture book image, the processor is further configured to: encode the visual feature data and the new prompt words respectively to obtain a visual feature vector and a prompt word vector, integrate the visual feature vector and the prompt word vector to obtain a conditional vector, input the conditional vector into the visual generation model, use the conditional vector as a conditional input to guide the visual generation model, and generate an adjusted picture book image.
[0196] In some embodiments, when the processor executes the steps of inputting the conditional vector into the visual generation large model, guiding the visual generation large model with the conditional vector as the conditional input, and generating an adjusted picture book image, the processor is further configured to: respectively encode the visual feature data and the new prompt word to obtain the visual feature vector and the prompt word vector, integrate the visual feature vector and the prompt word vector to obtain the conditional vector, input the conditional vector into the diffusion model based on the cross-attention mechanism, and guide the diffusion model to generate the adjusted picture book image with the conditional vector as the condition; wherein, the query variable in the cross-attention mechanism is the feature vector of the currently generated image, and the key variable and the value variable are the visual feature vector.
[0197] In some embodiments, the processor is further configured to: when the picture book editing content includes a reference image, extract visual features from the reference image to obtain reference image features, update visual feature data based on the reference image features, use the updated visual feature data and new prompt words as input, call the visual generation model, and generate an adjusted picture book image.
[0198] Since the above-mentioned server embodiment is a server that executes any one of the above-mentioned story picture book editing method embodiments, the specific limitations in the one or more server embodiments provided above can be referred to the above limitations on the server that executes any one of the above-mentioned story picture book editing method embodiments, and will not be repeated here.
[0199] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the method of the above embodiment is implemented when the processor executes the computer program.
[0200] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method of the above embodiment is implemented.
[0201] In one embodiment, a computer program product is provided, including a computer program, which implements the method of the above embodiment when executed by a processor.
[0202] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis such as picture book editing data, stored data, displayed data such as story picture books, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0203] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0204] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0205] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A story picture book editing method, characterized in that: The method comprises: Receive a picture book creation service call request, wherein the picture book creation service call request carries identification data of a story picture book and picture book editing content; Adjusting text data in the story picture book according to the identification data and the picture book editing content; Determine a prompt word based on the adjusted text data, wherein the prompt word is used to guide the visual generation model to generate an image; Querying visual feature data of the story picture book according to the identification data; Taking the visual feature data and the prompt word as input, calling the visual generation model to generate an adjusted picture book image; The adjusted text data and the adjusted picture book image are integrated to obtain an edited story picture book.
2. The method according to claim 1, characterized in that The step of adjusting text data in the story picture book according to the identification data and the picture book editing content includes: Based on the story identifier in the identification data, obtaining text prompt words of the story picture book; Based on the edited content of the picture book, updating the text prompt word; The updated text prompt words are used as input to call the first language model to generate text data of the adjusted story picture book, where the first language model is used to generate text data of the story text.
3. The method according to claim 1, characterized in that The step of determining the prompt word based on the adjusted text data includes: The adjusted text data is used as input, and the second largest language model is called to extract prompt words from the adjusted text data to obtain prompt words. The second largest language model is used to generate prompt words to guide the visual generation model to generate images.
4. The method according to claim 1, characterized in that The method uses the visual feature data and the prompt word as input, calls the visual generation model, and generates an adjusted picture book image, including: Encoding the visual feature data and the prompt word respectively to obtain a visual feature vector and a prompt word vector; Integrate the visual feature vector and the prompt word vector to obtain a conditional vector; The conditional vector is input into the visual generation model, and the conditional vector is used as the conditional input to guide the visual generation model to generate an adjusted picture book image.
5. The method according to claim 4, characterized in that The step of inputting the conditional vector into the visual generation model, using the conditional vector as a conditional input to guide the visual generation model, and generating an adjusted picture book image includes: Encoding the visual feature data and the prompt word respectively to obtain a visual feature vector and a prompt word vector; Integrate the visual feature vector and the prompt word vector to obtain a conditional vector; Inputting the conditional vector into a diffusion model based on a cross-attention mechanism, and using the conditional vector as a condition to guide the diffusion model to generate an adjusted picture book image; Among them, the query variable in the cross-attention mechanism is the feature vector of the current generated image, and the key variable and the value variable are the visual feature vector.
6. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: When the picture book editing content includes a reference image, extracting visual features in the reference image to obtain reference image features; updating the visual feature data according to the reference image feature; Taking the updated visual feature data and prompt words as input, the visual generation model is called to generate the adjusted picture book image.
7. A story picture book editing method, characterized in that: Applied to a display device, the method comprises: Receiving picture book editing content input by users for story picture books; Based on the picture book editing content and the identification data of the story picture book, the picture book creation service is called, so that the picture book creation service adjusts the text data in the story picture book according to the identification data and the picture book editing content, determines the prompt words for guiding the visual generation big model to generate images based on the adjusted text data, queries the visual feature data of the story picture book according to the identification data, calls the visual generation big model with the visual feature data and the prompt words as input, generates an adjusted picture book image, and packages the adjusted text data and the adjusted picture book image into an edited story picture book for feedback; Receive the edited story picture book fed back by the picture book creation service, and display the edited story picture book.
8. A display device, characterized in that: include: A display configured to display the story picture book on a user interface; The controller is configured as: Identify the picture book editing content input by the user for the story picture book, construct a picture book editing instruction, the picture book editing instruction carries the picture book editing content and the identification data of the story picture book, send the picture book editing instruction to the service calling module, the service calling module calls the picture book creation service based on the identification data and the picture book editing content, so that the picture book creation service adjusts the text data in the story picture book according to the identification data and the picture book editing content, determines the prompt words for guiding the visual generation large model to generate an image based on the adjusted text data, queries the visual feature data of the story picture book according to the identification data, calls the visual generation large model with the visual feature data and the prompt words as input, generates an adjusted picture book image, packages the adjusted text data and the adjusted picture book image into an edited story picture book and feeds it back to the controller, receives the edited story picture book fed back by the service calling module, and controls the display to display the edited story picture book.
9. The display device according to claim 8, characterized in that The controller is further configured to: In response to the user's editing operation on the story content, the editing state of the story content is set to an editable state, and the editing states of the story characters and the story outline are updated to a non-editable state, and the picture book editing content includes the edited story content.
10. A server, characterized in that: The server comprises: A communication module, configured to establish a communication connection with a display device; The processor is configured as: Receiving a picture book creation service call request sent by the display device, wherein the picture book creation service call request carries identification data of the story picture book and picture book editing content; Adjusting text data in the story picture book according to the identification data and the picture book editing content; Determine a prompt word based on the adjusted text data, wherein the prompt word is used to guide the visual generation model to generate an image; Querying visual feature data of the story picture book according to the identification data; Taking the visual feature data and the prompt word as input, calling the visual generation model to generate an adjusted picture book image; The adjusted text data and the adjusted picture book image are packaged into an edited story picture book, and the edited story picture book is fed back to the display device.
Citation Information
Cited By
Electronic certificate efficient generation method and system based on AI
CN120543691A
Picture book reading method and device, storage medium and program product
CN121033823A