Display device, server and story picture book generation method

By establishing a communication connection between the display device and the server, using the large language model and the storyboard prompt word extraction model to generate high-quality story picture books, the problem of low quality story picture books in the existing technology is solved and the user interaction experience is improved.

CN119991874APending Publication Date: 2025-05-13HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411982826.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The automatically generated story picture books in the prior art are of low quality, resulting in poor user interaction experience.

Method used

By establishing a communication connection between the display device and the server, a large language model is used to generate coherent and rich storyboard text, and a storyboard word extraction model extracts a storyboard word matching the character from the text, thereby generating a storyboard picture that fits the character and the storyline.

Benefits of technology

It realizes the generation of high-quality story picture books that meet users' expectations, improves user interaction experience, and saves users' actions to repeatedly adjust picture book creation needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991874A_ABST
    Figure CN119991874A_ABST
Patent Text Reader

Abstract

The invention relates to a display device, a server and a story picture book generation method. The method comprises the steps that interaction content input by a user and sent by a display device is received, a large language model is called based on the interaction content, a story text is generated, the story text comprises a plurality of split texts, a split cue word extraction model applied to the large language model is triggered, a split cue word is extracted from each split text, and the split cue word is displayed on the display device. The sub-mirror cue word extraction model is obtained by training the low-rank adaptation model based on a sub-mirror text sample containing a sub-mirror cue word label matched with the role, for each sub-mirror, a sub-mirror graph of the sub-mirror is generated based on the sub-mirror cue word of the sub-mirror, and the sub-mirror text of the plurality of sub-mirrors and the sub-mirror graph are used as story picture book data and are fed back to the display device. And displaying the story picture book data by the display equipment. By adopting the method, a high-quality story picture book can be generated, and the interaction experience of a user is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of display devices, and in particular to a display device, a server, and a story picture book generation method. Background Art

[0002] Display devices refer to terminal devices that can output specific display images. With the rapid development of display devices and the increasing diversification of user needs, people's demand for intelligent display devices such as smart TVs is also increasing, and the functions of display devices are becoming more and more abundant.

[0003] Currently, users can input some picture book creation requirements through a series of interactions with the display device. The display device sends the picture book creation requirements to the server. The server generates a story picture book based on the picture book creation requirements and feeds the story picture book back to the display device. The display device displays the story picture book, and the user can read the story picture book directly on the display device.

[0004] However, usually, the quality of story picture books generated based on picture book creation requirements input by users is often not high and difficult to meet user expectations. Users are required to repeatedly adjust the input picture book creation requirements to improve the quality of story picture books, which greatly reduces the user's interactive experience. Summary of the invention

[0005] The present application provides a display device, a server and a story picture book generation method to solve the problem that the automatically generated story picture books are of low quality, resulting in a poor user interaction experience.

[0006] In a first aspect, some embodiments provide a display device, comprising: a display and a controller. The display is configured to display a user interface of the picture book creation service in response to a start-up operation triggered by a user for the picture book creation service; the controller is configured to: identify the interactive content input by the user, and send the interactive content to the server; receive story picture book data fed back by the server based on the interactive content, and control the display to display the story picture book data;

[0007] The story book data includes storyboard texts and storyboard images of multiple storyboards, the storyboard texts are generated by calling a large language model based on the interactive content, the storyboard images are generated based on storyboard prompt words, the storyboard prompt words are extracted from the storyboard texts by triggering a storyboard prompt word extraction model applied to the large language model, and the storyboard prompt word extraction model is trained on a low-rank adaptive model based on storyboard text samples annotated with storyboard prompt words that match the characters.

[0008] The solutions of the above embodiments have the following advantages or beneficial effects:

[0009] After receiving the interactive content input by the user in the user interface of the picture book creation service, the display device can display the story picture book data generated based on the interactive content. In addition, since the generation of the story picture book is to first call the large language model based on the interactive content to generate a coherent and rich storyboard text, then, by triggering the storyboard prompt word extraction model applied to the large language model, there is no need to conduct complex training on the large language model. Only by fine-tuning the large language model, the storyboard prompt words matching the role can be efficiently extracted from the storyboard text, and then based on the extracted storyboard prompt words, a storyboard that is more suitable for the role and the storyline can be generated. In this way, a high-quality story picture book that better meets the user's expectations can be obtained. In the whole process, the user only needs to input the interactive content in the user interface of the display device, and can directly read the generated high-quality story picture book on the display device without other tedious operations. In addition, since the generated story picture book is more suitable for the role and the storyline, it can better meet the user's expectations, saving the user from repeatedly adjusting the picture book creation needs, greatly improving the user's interactive experience.

[0010] In the second aspect, some embodiments further provide a server, including: a communication module and a processor. The communication module is configured to establish a communication connection with a display device; the processor is configured to: receive interactive content input by a user sent by the display device; call a large language model based on the interactive content to generate a story text, wherein the story text includes storyboard texts of multiple storyboards; trigger a storyboard prompt word extraction model applied to the large language model to extract storyboard prompt words from each of the storyboard texts, wherein the storyboard prompt word extraction model is obtained by training a low-rank adaptive model based on storyboard text samples annotated with storyboard prompt words matching the characters; for each storyboard, generate a storyboard image of the storyboard based on the storyboard prompt words of the storyboard; and obtain story picture book data based on the storyboard texts and the storyboard images of multiple storyboards.

[0011] The solutions of the above embodiments have the following advantages or beneficial effects:

[0012] The server is deployed with a trained large language model and a storyboard prompt word extraction model. After receiving the interactive content sent by the display device, the large language model is called based on the interactive content to quickly and efficiently generate a coherent and rich storyboard text. Then, by triggering the storyboard prompt word extraction model applied to the large language model, there is no need to conduct complex training on the large language model. Only by fine-tuning the large language model, the storyboard prompt words that match the characters can be efficiently extracted from the storyboard text. Then, based on the extracted storyboard prompt words, a storyboard that is more suitable for the characters and the storyline can be generated, and finally a high-quality story picture book that meets the user's expectations can be obtained. In the whole process, the user only needs to input the interactive content in the user interface of the display device, and can directly read the generated high-quality story picture book on the display device without other tedious operations. In addition, since the story picture book generated by the server is more suitable for the characters and the storyline, it can meet the user's expectations, saving the user from repeatedly adjusting the picture book creation needs, greatly improving the user's interactive experience.

[0013] In a third aspect, some embodiments further provide a story picture book generation method, which is applied to the display device provided in the first aspect, the display device comprising: a display and a controller, the method comprising: receiving and identifying interactive content input by a user, and sending the interactive content to a server, the interactive content being input by the user on a user interface of a picture book creation service; receiving story picture book data fed back by the server, wherein the story picture book data comprises storyboard texts and storyboard images of a plurality of storyboards, the storyboard text is generated by calling a large language model based on the interactive content, the storyboard image is generated based on storyboard prompt words, the storyboard prompt words are extracted from the storyboard text by triggering a storyboard prompt word extraction model applied to the large language model, the storyboard prompt word extraction model is trained on a low-rank adaptive model based on storyboard text samples annotated with storyboard prompt words matching the characters; and displaying the storyboard data.

[0014] The solutions of the above embodiments have the following advantages or beneficial effects:

[0015] The user inputs interactive content on the user interface of the display device. After receiving the interactive content input by the user in the user interface of the picture book creation service, the display device communicates with the server, so that the server calls the large language model based on the interactive content, and quickly and efficiently generates a coherent and rich storyboard text. Then, by triggering the storyboard prompt word extraction model applied to the large language model, there is no need to conduct complex training on the large language model. Only by fine-tuning the large language model, the storyboard prompt words matching the role can be efficiently extracted from the storyboard text, and then based on the extracted storyboard prompt words, a storyboard map that is more suitable for the role and the storyline can be generated. The display device receives and displays the storyboard text and storyboard map of multiple storyboards fed back by the server, and obtains a high-quality story picture book that meets the user's expectations. In the whole process, the user only needs to input the interactive content in the user interface of the display device, and can directly read the generated high-quality story picture book on the display device without other tedious operations. In addition, since the story picture book generated by the server is more suitable for the role and the storyline, it can better meet the user's expectations, saving the user from repeatedly adjusting the picture book creation needs, greatly improving the user's interactive experience.

[0016] In a fourth aspect, some embodiments further provide a story picture book generation method, which is applied to the server provided in the second aspect, and the server includes a communication module and a processor. The method includes: obtaining interactive content input by a user; calling a large language model based on the interactive content to generate a story text, wherein the story text includes storyboard texts of multiple storyboards; triggering a storyboard prompt word extraction model applied to the large language model to extract storyboard prompt words from each of the storyboard texts, wherein the storyboard prompt word extraction model is obtained by training a low-rank adaptive model based on storyboard text samples annotated with storyboard prompt words matching the characters; for each storyboard, generating a storyboard image of the storyboard based on the storyboard prompt words of the storyboard; and obtaining storyboard data based on the storyboard texts and the storyboard images of multiple storyboards.

[0017] The solutions of the above embodiments have the following advantages or beneficial effects:

[0018] After receiving the interactive content input by the user in the user interface of the display device, the server calls the large language model based on the interactive content to quickly and efficiently generate a coherent and rich storyboard text. Then, by triggering the storyboard prompt word extraction model applied to the large language model, there is no need to conduct complex training on the large language model. Only by fine-tuning the large language model, the storyboard prompt words that match the role can be efficiently extracted from the storyboard text. Then, based on the extracted storyboard prompt words, a storyboard image that is more suitable for the role and the plot can be generated. The display device receives and displays the storyboard text and storyboard image of multiple storyboards fed back by the server, and obtains a high-quality story picture book that meets the user's expectations. In the whole process, the user only needs to input the interactive content in the user interface of the display device, and can directly read the generated high-quality story picture book on the display device without other tedious operations. In addition, since the story picture book generated by the server is more suitable for the role and the plot, it can better meet the user's expectations, saving the user from repeatedly adjusting the picture book creation needs, greatly improving the user's interactive experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 A schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application;

[0021] Figure 2 A schematic diagram of the hardware configuration of a display device provided in some embodiments of the present application;

[0022] Figure 3 A schematic diagram of the hardware configuration of a control device provided in some embodiments of the present application;

[0023] Figure 4 A schematic diagram of software configuration of a display device provided in some embodiments of the present application;

[0024] Figure 5 A schematic diagram of a user interface provided for some embodiments of the present application;

[0025] Figure 6 An interactive sequence diagram of a story picture book generation process provided by some embodiments of the present application;

[0026] Figure 7 A schematic diagram of the layout of storyboard texts and storyboard images provided for some embodiments of the present application;

[0027] Figure 8 A flowchart of a story picture book generation method executed by a server processor provided in some embodiments of the present application;

[0028] Fig. 9 A schematic diagram of a process for generating a story picture book provided in some embodiments of the present application;

[0029] Fig.10 A schematic diagram of a continuous storyboard with a consistent style provided in some embodiments of the present application;

[0030] Fig.11 A schematic diagram of a process for generating a story picture book applied to a display device provided in some embodiments of the present application;

[0031] Fig.12 A schematic diagram of a process for generating a story picture book applied to a server provided in some embodiments of the present application;

[0032] Fig.13 A schematic flow chart of steps for a server to generate a storyboard provided in some embodiments of the present application;

[0033] Fig.14 A schematic diagram of a process flow of a story picture book generation method applied to a server provided in some other embodiments of the present application;

[0034] Fig.15 A flowchart of a story picture book generation method applied to a server provided in some embodiments of the present application. DETAILED DESCRIPTION

[0035] The following embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following embodiments do not represent all implementations consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application as detailed in the claims.

[0036] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.

[0037] The terms "first", "second", "third", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar or similar objects or entities, and are not necessarily meant to limit a specific order or sequence, unless otherwise noted → It should be understood that the terms used in this way can be interchangeable under appropriate circumstances.

[0038] The terms "including" and "having" and any variations thereof are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0039] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0040] In the embodiment of the present application, the display device 200 generally refers to a device with image display and data processing capabilities. For example, the display device 200 includes but is not limited to a smart TV, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.

[0041] Figure 1 This is a schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application. Figure 1 As shown in FIG. 1 , the user can operate the display device 200 through touch operation, the mobile terminal 300 and the control device 100. For example, the control device 100 can be a remote controller, a stylus pen, a handle, etc.

[0042] The mobile terminal 300 can be used as a control device for performing human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device for establishing a communication connection with the display device 200 and performing data interaction. In some embodiments, the mobile terminal 300 can install software applications with the display device 200, and achieve connection and communication through a network communication protocol to achieve the purpose of one-to-one control operation and data communication. The audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 to achieve a synchronous display function.

[0043] like Figure 1 As also shown in FIG. 4 , the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 may be allowed to communicate and connect through a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0044] The display device 200 may provide a broadcast receiving television function, and may also additionally provide an intelligent network television function with a computer support function, including but not limited to network television, smart television, Internet Protocol television (IPTV), and the like.

[0045] Figure 2 Some embodiments of the present application provide Figure 1 2 is a block diagram of the hardware configuration of the display device 200.

[0046] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0047] In some embodiments, the detector 230 is used to collect signals of the external environment or external interaction. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of ambient light; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environment scenes, user attributes or user interaction gestures; or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.

[0048] In some embodiments, the display 260 includes a display function component for presenting a picture, a touch component for receiving a user touch operation, and a drive component for driving an image display. The display 260 is used to receive an image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, and components of a menu control interface and a user control UI interface.

[0049] In some embodiments, the communication device 220 is a component for communicating with an external device or server 400 according to various communication protocol types. The display device 200 may be provided with a plurality of communication devices 220 according to different supported communication modes. For example, when the display device 200 supports wireless network communication, the display device 200 may be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.

[0050] The communication device 220 can enable the display device 200 to communicate with the external device or server 400 by wireless or wired connection. Among them, the wired connection can connect the display device 200 with the external device through components such as data cables and interfaces. The wireless connection can connect the display device 200 with the external device through wireless signals or wireless networks. The display device 200 can establish a connection relationship with the external device directly, or indirectly establish a connection relationship through a gateway, a router, a connection device, etc.

[0051] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first interface to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to the user's operation through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.

[0052] In some embodiments, the controller 250 and the tuner-demodulator 210 may be located in different separate devices, that is, the tuner-demodulator 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0053] In some embodiments, the user may input a user command in a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).

[0054] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or may be an external audio output device of the display device 200. In particular, for the external audio output device of the display device 200, the display device 200 may also be provided with an external audio output terminal, and the audio output device may be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.

[0055] In some embodiments, the user input interface 280 may be used to receive instructions from a user.

[0056] Figure 3 Some embodiments of the present application provide Figure 1 The hardware configuration diagram of the control device in the figure. Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.

[0057] The control device 100 is configured to control the display device 200 , and can receive user input operation instructions, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, playing the role of an interactive intermediary between the user and the display device 200 .

[0058] In some embodiments, the control device 100 may be a smart device, for example, the control device 100 may be installed with various applications for controlling the display device 200 according to user needs.

[0059] In some embodiments, Figure 1 As shown, the mobile terminal 300 or other intelligent electronic devices can play a similar function as the control device 100 after installing the application for controlling the display device 200 .

[0060] The controller 110 includes a processor 112, a RAM 113, a ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between the internal components and the external and internal data processing functions.

[0061] The communication interface 130 implements communication of control signals and data signals with the display device 200 under the control of the controller 110. The communication interface 130 may include at least one of other near field communication modules such as a WiFi chip 131, a Bluetooth module 132, and an NFC module 133.

[0062] The user input / output interface 140 , wherein the input interface includes at least one of other input interfaces such as a microphone 141 , a touch panel 142 , a sensor 143 , and a button 144 .

[0063] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with a communication interface 130, such as a WiFi, Bluetooth, NFC or other module, and can encode the user input command through the WiFi protocol, Bluetooth protocol, or NFC protocol and send it to the display device 200.

[0064] The memory 190 is used to store various operating programs, data and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.

[0065] The power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller.

[0066] In order to perform user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system may provide a user interface (control the display device), allow the user to interact with the display device 200, and support the running of various application programs.

[0067] It should be noted that the operating system may be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.

[0068] The operating system can be divided into different modules or layers according to the functions implemented, such as Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom, namely, the application layer (Applications) layer (referred to as "application layer"), the application framework layer (Application Framework) layer (referred to as "framework layer"), the system library layer and the kernel layer.

[0069] In some embodiments, the application layer is used to provide services and interfaces for applications so that the display device 200 can run applications and interact with users based on the applications. At least one application can be run in the application layer, and these applications can be window programs, system settings programs, clock programs, etc. that come with the operating system; they can also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the above examples.

[0070] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications. The application framework layer includes some predefined functions. The application framework layer is equivalent to a processing center that determines the actions that applications in the application layer take. Through the API interface, applications can access system resources and obtain system services during execution.

[0071] like Figure 4 In the embodiment of the present application, the application framework layer includes a view system, managers, content providers, etc., wherein the view system can design and implement the interface and interaction of the application, and the view system includes lists, grids, text boxes, buttons, etc. The manager includes at least one of the following modules: an activity manager for interacting with all activities running in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to the application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0072] In some embodiments, the activity manager is used to manage the life cycle of each application and the usual navigation back function, such as controlling the exit, opening, and back of the application. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, capturing the screen, and controlling the display window changes, for example, reducing the display window, shaking the display, distorting the display, etc.

[0073] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction library contained in the system runtime layer, such as the C / C++ instruction library, to implement the functions to be implemented by the framework layer.

[0074] In some embodiments, the kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. Figure 4 As shown, the kernel layer may be configured with hardware drivers, and the drivers included in the kernel layer may be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0075] It should be noted that the above example is only a simple division of the operating system functions and does not constitute a limitation on the specific operating system form of the display device 200 in the embodiment of the present application. Depending on factors such as the function of the display device and the type of operating system, the number of levels and specific level types contained in the operating system may be expressed in other forms.

[0076] With the rapid development of display devices and the increasing diversification of user needs, people's demand for the intelligence of display devices such as smart TVs is also increasing, and the functions of display devices are becoming more and more abundant. At present, users can input some picture book creation requirements through a series of interactions with the display device. The display device sends the picture book creation requirements to the server, and the server generates a story picture book based on the picture book creation requirements, and feeds the story picture book back to the display device. The display device displays the story picture book, and the user can read the story picture book directly on the display device. However, under normal circumstances, the quality of the story picture book generated based on the picture book creation requirements input by the user is often not high, and it is difficult to meet the user's expectations. The user needs to repeatedly adjust the input picture book creation requirements to improve the quality of the story picture book, which greatly reduces the user's interactive experience.

[0077] In order to solve the above technical problems, an embodiment of the present application provides a display device, which includes a display and a controller.

[0078] The display is configured to display a user interface. Specifically, the display is configured to display a user interface of the picture book creation service in response to a start operation triggered by a user on the picture book creation service.

[0079] In this embodiment, the picture book creation service is a software or application based on the display device, which allows users to create and edit electronic picture books through a variety of interactive methods. This service usually includes functions such as character creation, story writing, image drawing, image uploading and style selection, aiming to help users easily create high-quality story picture books.

[0080] In actual applications, after the user opens the picture book creation service and triggers the start operation, the controller controls the display to display the user interface of the picture book creation service. The user interface diagram can be referred to Figure 5 . Users can input some data for generating story picture books on the user interface according to their own expectations and needs, including inputting or setting data such as character description data, story theme and story summary. Among them, the built-in name of the protagonist can be selected in the protagonist setting, or the user can define it by text description, or upload a character reference image. The attributes of the selected character are displayed on the protagonist description, and users are supported to modify the attributes. For example, if Tom is selected, the built-in attributes of Tom are: [Tom] a little black cat, with a big head, a chubby body, and big eyes. For the character reference image, users are supported to upload reference images or customize reference images, and add trigger words and attribute descriptions to the protagonist description display box. The story style is a built-in style, and users can choose the style they want, such as illustrations, comics, etc. The input of the story theme can be "space travel", "forest adventure", etc. After the user completes the input of the picture book creation needs according to their own needs, they can click the "Picture Book Generation" button on the user interface. Optionally, the story picture book display area displays the story title, story summary, storyboard content and storyboard illustrations, supporting users to modify the generated story content and regenerate the picture book after modification.

[0081] The controller is configured to: identify the interactive content input by the user, and control the display to display the story picture book data generated based on the interactive content. The story picture book data includes storyboard texts and storyboard images of multiple storyboards, the storyboard texts are generated by calling a large language model based on the interactive content, the storyboard images are generated based on storyboard prompt words, the storyboard prompt words are extracted from the storyboard texts by triggering a storyboard prompt word extraction model applied to the large language model, and the storyboard prompt word extraction model is based on storyboard text samples annotated with storyboard prompt words matching the characters, and is obtained by training a low-rank adaptive model.

[0082] Interactive content refers to data generated during the interaction between the user and the display device, including the user's clicks, slides, input text and images, etc. Specifically, the interactive content can be at least one of voice content, text content, picture content and video content.

[0083] Storyboards, also known as storyboards, are diagrams used to illustrate the composition of images in various image media such as movies, animations, TV series, advertisements, and music videos before they are actually shot or drawn. Storyboard text refers to the text used to describe the content of each storyboard, including dialogue, narration, and narration. These texts are used to describe the plot of the story, the dialogue and inner monologue of the characters, and the background information of the scene. Storyboards refer to graphics generated based on storyboard texts, which are used to intuitively display the content of the storyboards. Storyboards present the plot of the story, the actions of the characters, and the visual effects of the scenes. It can be understood that storyboard texts and storyboards are complementary. Storyboard texts provide the plot and dialogue of the story, and can explain the content that is not explicitly stated in the storyboard, such as the inner activities and background information of the characters, while storyboards provide visual supplements and enhancements, and intuitively display the scenes and actions described in the storyboard text. The combination of the two can make the story richer and more vivid.

[0084] Specifically, the large language model includes but is limited to GPT4 and large language models developed on the market. Storyboard prompts are keywords or phrases that guide the generation of images of specific scenes or characters. These prompts are extracted from the storyboard text and can accurately reflect the main content of the storyboard.

[0085] In this embodiment, the storyboard prompt word extraction model is a specially trained model that can extract appropriate storyboard prompt words from the storyboard text. Specifically, the storyboard prompt word extraction model can be a model obtained by training the LoRA (Low-Rank Adaptation) model based on the storyboard text sample data. LoRA is a technology for fine-tuning pre-trained models, allowing the model to be efficiently adjusted on new tasks while maintaining the original performance. The training process of the storyboard prompt word extraction model can be as follows:

[0086] First, the developer pre-sets some requirements for extracting storyboard prompts, so as to output prompts and format requirements suitable for the text map. For example, a prompt sample for extracting storyboard prompts can be as follows:

[0087] Title: [title]

[0088] Summary: [summary]

[0089] Content: [content]

[0090] Role_list: [Role_name]

[0091] Role_info (Role information): [Role_info]

[0092] Then, the controller can generate the screen description text corresponding to each PAGE (screen, which can also be regarded as storyboard) according to the provided Title, Summary, Content, Role_list and Role_info. The screen description text refers to the storyboard text. PAGE and screen description text correspond one to one. Specifically, the output of screen description text also has the following requirements:

[0093] 1) There is a Role_name in the screen, which is the role description. The screen description text format is A: [Role_name]action / bodymovements / location / background. The beginning of the sentence is represented by [Role_name].

[0094] 2) There is no Role_name in the picture, it is a description of the scenery or background, the description text format B: [NC]scenerydescription, the beginning of the sentence is indicated by [NC]. [scenery description] For example, "a large grassland, full of vitality".

[0095] 3) For the action of Role_name, use simple action words, such as "stand", "sit", "walk", "touch" and other simple body movement words to describe

[0096] 4) If there is no obvious action in Role_name, body movements are expanded according to the scene. For example, sleeping can be expanded to actions such as "lie down" and "close eyes".

[0097] 5) Each screen description text should contain the location relationship between each element and Role_name, such as "There is a green tree next to Paul", "There is a sun above Jack's head".

[0098] 6) The description text of each picture should not include any psychological or emotional actions of the elements. Delete relevant psychological descriptions such as "very happy", "feeling very fulfilled inside", "dreaming", etc., and only keep the description of body movements and element positions.

[0099] 7) Ensure that the wording is simple, including simple grammar, simple word usage, no modifiers, and adjectives, so that the image generation model can understand it.

[0100] 8) Output a simple grammatical combination of English words, and the length of each description text shall not exceed 25 words.

[0101] 9) Do not output irrelevant information and analysis information.

[0102] 10) [NC] Describes only the landscape image.

[0103] The controller determines the description text of each PAGE through the above requirements, that is, after determining the storyboard text of each storyboard, a large number of storyboard texts determined by the above method can be collected, and then, storyboard prompt words that meet the requirements are annotated on the storyboard text by manual annotation, and then the storyboard text carrying the storyboard prompt word annotation information is used as model training data to train the LoRA model to obtain the storyboard prompt word extraction model. The trained storyboard prompt word extraction model can extract more accurate storyboard prompt words from the storyboard text.

[0104] When implementing it, Figure 6 As shown, after identifying the interactive content, the controller can use the interactive content as input to call the large language model, so that the large model outputs prompt words according to the interactive content and the preset story text, and generates a story text including multiple storyboard texts.

[0105] The story text output prompt refers to a prompt used to guide the large model to generate the story text. In actual applications, developers can set some output prompts for guiding the large model to generate the story text according to actual needs, and embed the set story text output prompts into the first large language model. It is understandable that the story text output prompt can also be modified by the user, or set by the user.

[0106] For example, the sample of the story text output prompt word can be as follows:

[0107] You are a writer of children's stories.

[0108] Objective: Please create a children's story based on the story characters and themes I input. The paragraphs must be clear and the logic must be reasonable.

[0109] Requirement: Do not output my input content, directly output the results, including the title, story introduction and story content.

[0110] Output format: The output format is JSON object.

[0111] The following is an example of the output format: {\″title\″: \″\″, \″summary\″: \″\″, \″content\″: \″\″}.

[0112] The characters in the story are: [{{role}}];

[0113] The summary of the story is: [{{summary}}]

[0114] The output story text content strictly follows the following requirements:

[0115] 1. Directly output string data in standard JSON format and block other content;

[0116] 2. The content in the output JSON data must be wrapped in double quotes. If quotes are required inside the content, only single quotes can be used;

[0117] 3. The story needs to have multiple paragraphs, each paragraph ** MUST ** by ** PAGE: ** At the beginning, special symbols such as (″″\| / ) are not allowed;

[0118] 4. Directly output the story outline without providing page numbers, page segments and other useless information;

[0119] 5. The story should have 6-10 paragraphs, at least 6 and at most 10. Do not divide the story into sections.

[0120] 6. The story is narrative-based and does not require character dialogue;

[0121] 7. Don’t give random names to characters;

[0122] 8. Sensitive words related to national security and uncivilized content are strictly prohibited;

[0123] 9. Double quotes are not allowed in text content;

[0124] 10. Each story cannot exceed 2 sentences, and a character can only appear once in each story;

[0125] 11. Personal pronouns such as he, she, it, etc. are not allowed.

[0126] The controller uses the interactive content as input and calls the pre-trained large model. The large model generates a standardized story text according to the built-in story text output requirements of the model. The story text includes the storyboard texts of multiple storyboards. Then, the controller sends the generated storyboard text to the trained storyboard prompt word extraction model, which extracts the key storyboard prompt words from the storyboard text, and then inputs the extracted storyboard prompt words into the image generation model to generate a storyboard image that matches the storyboard text. Finally, the storyboard text and storyboard image of each storyboard are integrated to form a complete story picture book. Subsequently, the controller controls the display to display the complete story picture book data.

[0127] It is understandable that, if the display device has sufficient computing resources, the display device can independently generate story picture book data based on the interactive content input by the user. If the computing resources of the display device are insufficient, the display device can send the identified interactive content to the server, and the server generates corresponding story picture book data based on the interactive content and feeds it back to the display device, and the display device controls the display to display the story picture book data. The specific method can be determined according to the actual situation and is not limited here.

[0128] The above technical solution has the following specific advantages or beneficial effects: after receiving the interactive content input by the user in the user interface of the picture book creation service, the display device can display the story picture book data generated based on the interactive content, and because the generation of the story picture book first calls the large language model based on the interactive content to generate a coherent and rich storyboard text, then, by triggering the storyboard prompt word extraction model applied to the large language model, there is no need to conduct complex training on the large language model. Only by fine-tuning the large language model, the storyboard prompt words matching the role can be efficiently extracted from the storyboard text, and then based on the extracted storyboard prompt words, a storyboard that is more suitable for the role and the storyline can be generated. In this way, a high-quality story picture book that is more in line with the user's expectations can be obtained. Throughout the process, the user only needs to input the interactive content in the user interface of the display device, and can directly read the generated high-quality story picture book on the display device without other tedious operations. Moreover, since the generated story picture book is more suitable for the role and the storyline, it can better meet the user's expectations, saving the user from repeatedly adjusting the picture book creation needs, greatly improving the user's interactive experience.

[0129] In some embodiments, the controller is further configured to: layout the storyboard texts and storyboard images of multiple storyboards, and control the display to display the storyboard texts and storyboard images of each storyboard after the layout in the user interface.

[0130] In actual applications, since a storybook contains multiple storyboards, each storyboard has storyboard text and storyboard pictures. The controller needs to layout these storyboards and control the display to display the layout storyboards on the user interface in sequence to display the entire storybook.

[0131] Optionally, the layout of the storyboards includes but is not limited to a single-column layout, a double-column layout or a grid layout.

[0132] The single-column layout may be one in which each storyboard occupies a whole row, with the storyboard text on the top (or bottom) and the storyboard image on the bottom (or top), such as Figure 7As shown. The storyboard text can also be on the left (or right), and the storyboard image can be on the right (or left). A double-column layout can be that every two storyboards occupy a row, and the storyboard text and storyboard image of each storyboard are displayed side by side. A grid layout can be that multiple storyboards are arranged in a grid, and the storyboard text and storyboard image of each storyboard are in a small square. It is understandable that the grid layout is more suitable for large-screen displays. The above layout methods can automatically adjust the layout according to the screen size, or can be selected by the user, and are not limited here.

[0133] After the controller processes the storyboard text and storyboard image of each storyboard according to the set storyboard layout, it can control the display to directly display the storyboard text and storyboard image of each storyboard according to the established layout. Users can view and manage storyboards (such as turning pages, scrolling, zooming in or out, etc.) through touch operations or voice operations to enhance the interactive experience.

[0134] The above technical solution has the following advantages or beneficial effects: the controller processes the layout of the storyboard according to the preset layout template and style specifications, ensuring the consistency of the visual effect of each storyboard, improving the overall look and feel and the user's interactive experience.

[0135] In some embodiments, the controller is further configured to: when the resolution of the storyboard does not meet the preset resolution requirement, convert the storyboard into a storyboard that meets the preset resolution requirement, and control the display to display the converted storyboard.

[0136] In this embodiment, the controller is not only responsible for typeset layout of the storyboard texts and storyboard images of multiple storyboards, but is also further configured to handle the resolution of the storyboard images. Specifically, if the resolution of a storyboard image does not meet the preset resolution requirement (such as the preset resolution is 1080p), the controller will automatically convert the storyboard image to an image that meets the preset resolution requirement, and then control the display to display the converted storyboard image.

[0137] For example, if the image resolution of one of the storyboards is low, i.e., 720p, and the preset resolution requirement is 1080p, the controller may convert the storyboard image to a resolution of 1080p before performing typesetting, layout, and display.

[0138] Optionally, the conversion of low-resolution images to high-resolution images can be achieved through bilinear interpolation, bicubic interpolation, super-resolution reconstruction, wavelet transform, etc. The specific method can be determined according to the actual situation, which will not be discussed here. Super-resolution reconstruction uses deep learning technology to generate higher-resolution images by training models. Wavelet transform can decompose images into sub-bands of different frequencies, then amplify the high-frequency sub-bands, and finally synthesize high-resolution images.

[0139] The above technical solution has the following advantages or beneficial effects: by converting all storyboards to a preset resolution, the display quality of all images on the user interface is ensured to be consistent, display problems caused by inconsistent resolution are avoided, and when users browse different storyboards, they will not feel abrupt due to changes in image resolution, and the viewing process of the entire story is smoother and more natural.

[0140] In some embodiments, the display device also includes: a sound collector, configured to: obtain voice data input by the user, and send the voice data to the controller, and the interactive content includes the voice data; the controller is further configured to: perform intent recognition based on the voice data, determine the user's interaction intention, and when the interaction intention represents the start of the picture book creation service, control the display to display the user interface of the picture book creation service.

[0141] In this embodiment, the sound collector can be a microphone or other device capable of collecting sound signals. Intent recognition refers to determining the user's specific intention or need by analyzing the interactive content input by the user. Interaction intent is the result of intent analysis of the interactive content, and the interaction intent reflects the goal that the user hopes to achieve in a specific situation.

[0142] The display device supports multi-modal interaction, that is, the user can interact with the display device and input interactive content through a variety of interaction methods. Specifically, the interaction methods include touch interaction, voice interaction, gesture interaction, auditory interaction, and visual interaction.

[0143] For example, the interactive mode is touch interaction, that is, the display device has a built-in touch screen, which can support multi-touch technology, allowing multiple fingers to operate simultaneously, realizing complex gestures such as zooming and dragging, and can also support pressure sensing, and can perform different operations according to different touch strengths. In actual operation, users can touch the screen with their fingers to input handwriting or virtual keyboard input.

[0144] Taking voice interaction as an example, the display device may have a built-in or external microphone, which captures the user's voice data and sends the voice data to the controller. The controller converts the user's voice into text or commands through voice recognition, analyzes the user's intention and then performs the corresponding operation. The recognition operation of the user's voice data can refer to the relevant technology, and the embodiments of this application will not be described one by one.

[0145] Optionally, the user can control the display device to enter the voice control mode by operating a designated button of the remote controller, or can control the display device to enter the voice control mode by voice.

[0146] Optionally, when the display device is triggered to enter the voice control mode, the user can also send instructions to the display device in text form through a mobile phone, remote control or other device to prevent the display device from being unable to receive the user's voice commands when there is a problem with the microphone.

[0147] During specific implementation, the microphone collects the voice data input by the user, and sends the collected voice data to the controller, which converts the user's voice into text or commands through voice recognition. When the user is recognized to say a specific wake-up word such as "Xiaox Xiaox", the display device is controlled to turn on the voice interaction mode. The microphone collects the voice data input by the user in real time, and the controller performs intent recognition on the voice data and performs corresponding operations. For example, if the user says: "Create a picture book story" or "Open the picture book creation service", the microphone collects the user's voice data and sends it to the controller. The controller can call a trained intent recognition model (such as a model based on deep learning) to perform intent recognition on the voice data, convert the voice data into text or commands, and recognize that the user's interaction intention is "Create a picture book", so the display is controlled to display the user interface of the picture book creation service. Specifically, the user interface of the creation service provides an entry for users to upload images, as well as functions for entering character description data and other creation needs, and other interactive functions.

[0148] In this embodiment, the display device supports voice interaction, which can accurately understand the user's intention and reduce misunderstandings and erroneous operations. The user can interact with the display device through natural voice input without complicated operation steps, which simplifies the interaction process and greatly improves the interaction experience.

[0149] The embodiment of the present application further provides a server, the server comprising a communication module and a processor. The communication module is configured to establish a communication connection with a display device. Figure 8 As shown, the processor is configured to perform the following steps S602 to S610:

[0150] Step S602: receiving interactive content input by a user sent by a display device.

[0151] The interactive content sent by the display device refers to the user input data obtained by the display device in response to the user's input operation. The interactive content may include but is not limited to voice content, text content, picture content, or a combination of two or more different modal contents. The way in which the display device obtains user input data can refer to the above embodiment and will not be repeated here.

[0152] In specific implementation, the user can use voice interaction to say the voice command "create a picture book story" to make the display device control the display to open the user interface of the picture book creation service. The user can enter some data for generating a story picture book on the user interface according to his or her own expectations and needs, including entering or setting data such as character description data, story theme and story summary. Some related interactive functions provided on the user interface, including the functions of protagonist setting, style selection, theme setting, etc., can be referred to the above embodiments and will not be repeated here.

[0153] The display device receives and recognizes the interactive content input by the user in real time. The interactive content mainly includes picture book creation input data such as the protagonist attributes, story theme and picture book style set by the user, and then sends the recognized interactive content to the server.

[0154] Step S604, calling a large language model based on the interactive content to generate a story text, where the story text includes storyboard texts of multiple storyboards.

[0155] In this embodiment, the large language model may include, but is not limited to, a large language model obtained by training and fine-tuning on a large language model such as GPT, Wenxin XX or Tongyi XX, and the large language model has functions such as intelligent customer service, machine translation, content creation, and image recognition. Story text refers to text data that complements the picture book image, including storyboard texts of multiple storyboards. Storyboard text refers to text used to describe the content of each storyboard, specifically including dialogue, narration, narration, etc. These texts are used to describe the plot of the story, the dialogue and inner monologue of the characters, and the background information of the scene.

[0156] After receiving the interactive content sent by the display device, the server may first identify the interactive content, and integrate and pre-process the identified interactive content. For example, a prompt word (prompt) for guiding the large language model to generate a story picture book may be generated based on the interactive content. For example, the prompt word may be:

[0157] Objective: Please create a children's story based on the story characters and themes I input. The paragraphs must be clear and the logic must be reasonable.

[0158] Requirement: Do not output my input content, output the result directly. The output format is output as a JSON object. The following is an example of the format: {\″title\″:\″\″,\″summary\″:\″\″,\″content\″:\″\″}

[0159] The characters in this story are: [{{role}}]

[0160] The summary of the story is: [{{summary}}]

[0161] Subsequently, based on the interaction content and prompt words, the trained large language model is called. The large language model combines the above prompt words and interaction content to generate a story text that matches the user's needs. The story text includes storyboard texts of multiple storyboards, specifically including the storyline and character dialogues corresponding to each storyboard.

[0162] Step S606, triggering the storyboard prompt word extraction model applied to the large language model to extract storyboard prompt words from each storyboard text. The storyboard prompt word extraction model is trained on the low-rank adaptive model based on storyboard text samples containing storyboard prompt word annotations that match the characters.

[0163] Storyboard prompt words are keywords or phrases that guide the generation of images of specific scenes or characters. These prompt words are extracted from the storyboard text and can accurately reflect the main content of the storyboard. In this embodiment, the storyboard prompt word extraction model is a model that can extract suitable storyboard prompt words from the storyboard text by training the LoRA model. Specifically, the training process of the storyboard prompt word extraction model can refer to the training process of the storyboard prompt word extraction model in the embodiment of the above-mentioned display device, which will not be repeated here.

[0164] In some embodiments, a pre-trained BART (Bidirectional and Auto-Regressive Transformers) model can be used to extract concise storyboard prompt words from the storyboard text. Alternatively, the storyboard prompt words can be extracted using other types of pre-trained models, depending on the actual situation.

[0165] After the server generates a storyboard text including multiple storyboards by calling the large language model, it can send the generated storyboard text to the trained storyboard prompt word extraction model, and the model extracts concise storyboard prompt words from the storyboard text. For example, if the generated storyboard text is "Lily and Blue are playing in the forest, and the sun shines through the treetops onto the ground, forming mottled light and shadows", the extracted storyboard prompt words may include "Lily, Blue, forest, play, sun, treetops, light and shadows."

[0166] Step S608, for each storyboard, based on the storyboard prompt words of the storyboard, generate a storyboard image of the storyboard.

[0167] A storyboard refers to a graphic generated based on the storyboard text, which is used to intuitively display the content of the storyboard. The storyboard presents the plot of the story, the actions of the characters and the visual effects of the scene.

[0168] After the server extracts the storyboard prompt words from each storyboard text, it can generate a storyboard image that matches the storyboard prompt words of each storyboard by calling a pre-trained image generation model (text image model), thereby obtaining storyboard images of multiple storyboards. Among them, the image generation model is an intelligent model that can generate corresponding images according to a given text description. Image generation models include but are not limited to models such as generative adversarial networks, variational autoencoders, and diffusion models.

[0169] In specific implementation, the server may first convert the storyboard prompt words into a high-dimensional vector representation that captures semantic information, and then input the high-dimensional vector representation into an image generation model, which then gradually generates storyboard images corresponding to the storyboard prompt words based on the generated high-dimensional vector representation.

[0170] Step S610, obtaining story picture book data based on the storyboard texts and storyboard images of the plurality of storyboards.

[0171] Following the previous step, after the server generates the storyboard texts and storyboard pictures of the multiple storyboards, the storyboard texts and storyboard pictures of the multiple storyboards can be fed back to the display device as storybook data. After receiving the storyboard texts and storyboard pictures of the multiple storyboards, the display device can typeset the storyboard texts and storyboard pictures of each storyboard according to a preset typesetting layout to provide a complete storybook.

[0172] Optionally, the layout of the storyboard includes but is not limited to a single-column layout, a double-column layout, or a grid layout. The specific layout process can be referred to the above embodiment, which will not be described in detail here. It is understood that in other embodiments, the user can also adjust the layout of the storyboard, as well as adjust each storyboard text and storyboard image.

[0173] After the display has processed the layout of the storyboard text and storyboard image of each storyboard according to the set storyboard layout, including the position of the storyboard image, the position of the text, the margin, etc., the display can be controlled to directly display the storyboard text and storyboard image of each storyboard according to the established layout. Users can view and manage storyboards (such as turning pages, scrolling, zooming in or out, etc.) through touch operations or voice operations to enhance the interactive experience.

[0174] In some embodiments, after the server generates the storyboard texts and storyboard images of the plurality of storyboards, the server may generate the pre-set or user-set storyboard layout data, the storyboard texts and storyboard images, and the storyboard texts and storyboard images. Figure 1 The storyboards are sent to the display device at the same time, and the display device directly layouts and displays the storyboard text and storyboard images according to the received storyboard typesetting layout data.

[0175] The technical solution of the above embodiment has the following advantages and beneficial effects: a trained large model and a storyboard prompt word extraction model are deployed on the server. After receiving the interactive content sent by the display device, the large language model is called based on the interactive content, and a coherent and rich storyboard text can be quickly and efficiently generated. Then, by triggering the storyboard prompt word extraction model applied to the large language model, there is no need to conduct complex training on the large language model. Only by fine-tuning the large language model, the storyboard prompt words matching the role can be efficiently extracted from the storyboard text, and then based on the extracted storyboard prompt words, a storyboard map that is more suitable for the role and the plot can be generated, and finally a high-quality story picture book that meets the user's expectations can be obtained. In the whole process, the user only needs to input the interactive content in the user interface of the display device, and can directly read the generated high-quality story picture book on the display device without other cumbersome operations. In addition, since the story picture book generated by the server is more suitable for the role and the plot, it can better meet the user's expectations, saving the user from repeatedly adjusting the picture book creation needs, greatly improving the user's interactive experience.

[0176] In some embodiments, before executing the step of calling the large language model based on the interactive content to generate the story text, the server is further configured to: pre-process the picture book creation input data to obtain pre-processed picture book creation input data, and the pre-processing includes at least one of traditional-simplified conversion, spelling correction, and sensitive word filtering.

[0177] When the server executes the step of calling the large language model based on the interactive content to generate the story text, the server is further configured to: call the large language model based on the preprocessed picture book creation input data to generate the story text.

[0178] Traditional-Simplified conversion refers to converting traditional Chinese characters in the text into simplified Chinese characters (or vice versa) so that the model can correctly understand and process the text. Spelling correction refers to identifying typos in the text, correcting spelling errors in the input text, and improving the readability and accuracy of the text. Sensitive word filtering refers to filtering sensitive words in the text to ensure that the generated story text complies with laws, regulations and social moral standards.

[0179] Developers can pre-build a traditional-simplified character comparison table for traditional-simplified conversion processing, and can collect and maintain a word library containing sensitive words for sensitive word detection. In specific implementation, since the picture book creation input data includes some text content such as the story theme and story summary, the server can first identify the traditional Chinese characters in the picture book creation input data through regular expressions or character mapping tables after identifying the picture book creation input data input by the user, and then use the pre-built traditional-simplified character comparison table to convert the identified traditional Chinese characters into simplified Chinese characters, and finally replace the original text with the converted text. At the same time, a spelling check tool or a natural language processing library is used to identify spelling errors in the picture book creation input data, and according to the identified spelling errors, the spelling errors are automatically corrected or the user is prompted to correct them manually, and correct spelling suggestions can also be provided to improve the accuracy and readability of the text content. Regular expressions, keyword matching or sensitive word filtering algorithms can also be used to detect sensitive words (such as uncivilized words, violent words) in the picture book creation input data, and the detected sensitive words are replaced with asterisks or other symbols, or directly deleted from the text. It is understandable that in addition to the above-mentioned preprocessing, the picture book creation input data may also be processed in other aspects such as text translation, which can be specifically designed according to the needs and is not limited here.

[0180] In some embodiments, the server may also perform at least one of traditional-simplified conversion, spelling correction, and sensitive word filtering on the picture book creation input data, depending on the content of the picture book creation input data and actual conditions, and is not limited here.

[0181] The technical solution of this embodiment has the following advantages or beneficial effects: the consistency of the story picture book can be improved by converting traditional and simplified Chinese, the correctness of the story text can be improved by spelling correction, and the readability and standardization of the story text can be improved by filtering sensitive words.

[0182] In some embodiments, the large model includes a first large language model; when the server executes the step of calling the large language model based on the preprocessed picture book creation input data to generate the story text, it is further configured to: call the first large language model with the preprocessed picture book creation input data as input, so that the first large language model generates the story text according to the preprocessed picture book creation input data and the preset story text output prompt words.

[0183] In this embodiment, for ease of distinction, the large language model in this implementation is referred to as the first large language model. Specifically, the first large language model can be GPT4 as an example.

[0184] The story text output prompt word refers to a prompt word (prompt) used to guide the large model to generate the story text. For the introduction of the story text output prompt word, please refer to the description of the above embodiment, which will not be repeated here.

[0185] Following the previous embodiment, after the server performs preprocessing such as traditional-simplified conversion, spelling correction, and sensitive word filtering on the picture book creation data input by the user, the server can use the preprocessed picture book creation data as input to call the first language model, and the first language model outputs prompt words based on the built-in story text to generate a story text that meets the requirements. It can be understood that the output prompt words in the above story text are only used as examples. In actual applications, they can be set according to actual needs and are not limited here.

[0186] The technical solution of the above embodiment has the following advantages or beneficial effects: the large language model can efficiently and quickly generate high-quality story text based on the pre-processed picture book creation input data and preset story text output prompt words.

[0187] In some embodiments, after executing the step of taking the storyboard text prompt words as input, calling the first largest language model, and generating a story text, the server is further configured to: take the story text as input, call the second largest language model, so that the second largest language model scores the story text according to a preset story scoring standard to obtain a scoring result, and when the score of the story text is less than a preset scoring threshold, determine optimization suggestions based on the scoring result, and feed the optimization suggestions back to the first largest language model, so that the first largest language model adjusts the story text based on the optimization suggestions until the score of the adjusted story text is greater than or equal to the preset scoring threshold.

[0188] On the basis of the technical solution of the previous embodiment, in order to generate a higher quality story text, a story scoring mechanism can also be introduced. The story text is considered qualified only when the score of the story text generated by the first language model is higher than the preset scoring threshold, and the storyboard prompt word extraction stage is entered.

[0189] Specifically, in order to improve the credibility of the rating score, in the story rating and optimization suggestion stage, a different large language model from the first large language model used in the story text generation stage is used. For the sake of distinction, the large language model used in the story rating and optimization suggestion stage is called the second large language model. Developers can set the story rating criteria for rating the generated story text in advance according to their needs. For example, the story rating criteria can be as follows:

[0190] Story scoring criteria:

[0191] Format Correctness (20 points): Does the story have 6 - 10 paragraphs of content? Does each paragraph start with "PAGE: "? Does each paragraph of the story have no more than 2 sentences? Does each character appear only once in each paragraph of the story?

[0192] Story Structure (20 points): Does the generated story conform to the required JSON format? Does it follow all the format regulations, such as using double quotes to wrap the content of "content", and only using single quotes for quotes inside the "content" content?

[0193] Content Quality (30 points): Is the story attractive, interesting, suitable for children to read, mainly narrative, and without character dialogue?

[0194] Language Accuracy (20 points): Is the language of the story accurate? Are there no personal pronouns such as he, she, it, etc.? Are there no double quotes?

[0195] Content Compliance (10 points): Does the story content contain no sensitive words related to national security, uncivilized behavior, and violations of social morality and laws and regulations?

[0196] As Fig. 9 shown, specifically in implementation, taking the preset story scoring threshold of 90 points as an example, and taking the large language model "WenXXX" provided by a certain company as the second large language model for illustration. After the server calls the first large language model to generate the story text, the second large language model can be called to make the second large language model score the story text generated by the first large language model according to the above story scoring criteria, and obtain the scoring result. The scoring result includes the comprehensive total score of the story, the scores of the story text in multiple dimensions such as format correctness, content quality, language accuracy, and content compliance, as well as data such as the overall quality analysis opinion of the story.

[0197] After obtaining the story scoring result of the story text, compare whether the total score of the story is lower than 90 points. If the story score is lower than 90 points, corresponding optimization suggestions are generated according to the overall quality analysis opinion in the scoring result. For example, if the score of the story text in the dimension of language accuracy in the scoring result is 10 points, it indicates that the language accuracy of the story text needs to be strengthened, and the server can generate optimization suggestions such as "Improve language accuracy". Then the generated optimization suggestions are fed back to the first large language model, and the first large language model generates the story text again according to the optimization suggestions. The server calls the second large language model again to score the story text regenerated by the first large language model, obtains the scoring result. In the case where the scoring result is lower than 90 points, optimization suggestions are generated again and fed back to the first large language model, and so on, until the score of the story text generated by the first large language model is equal to or higher than 90 points, and then enter the subsequent storyboarding prompt word stage.

[0198] It is understandable that the story scoring threshold can also be 95 points, 96 points or other scores, which can be determined according to actual conditions and are not limited here. The story scoring criteria can also be set to other criteria according to actual conditions and are not limited here.

[0199] The technical solution of the above embodiment has the following advantages or beneficial effects: by adopting different large language models in the story text generation stage to score the story text, not only the credibility of the score can be improved, but also the large language model can generate higher quality story text.

[0200] In some embodiments, the interactive content includes at least one of picture book creation data and voice data. When the server executes the step of generating a storyboard of each storyboard based on the storyboard prompt words of the storyboard, the server is further configured to:

[0201] Based on the style description data and character description data in the picture book creation input data, the style prompt words and character prompt words are determined respectively. For each storyboard, the style prompt words, character prompt words and the storyboard prompt words corresponding to the storyboard are combined to obtain the combined prompt words corresponding to the storyboard. Based on the combined prompt words, a storyboard diagram of the storyboard is generated.

[0202] Picture book creation input data refers to various demand data and reference materials input by users on the user interface in order to design and produce picture books. Specifically, picture book creation input data includes character description data, style description data, story theme, story summary, and reference image data. Among them, character description data includes but is not limited to appearance characteristics, personality characteristics, and background stories. Style description data includes but is not limited to animation style, illustrations, ancient style, and realism. Reference images include but are not limited to character reference images, scene reference images, and style reference images.

[0203] In actual applications, after the server receives the picture book creation input data input by the user, the server can identify the picture book input data, determine the positive and negative character words corresponding to the character description data based on the character description data, and determine the positive and negative style prompt words corresponding to the style based on the style description data.

[0204] Among them, the positive and negative character prompts include positive character prompts and negative character prompts. Positive character prompts refer to the characteristics or elements that you hope to embody in the character, such as bravery and wisdom; negative character prompts refer to the characteristics that you do not want to appear in the character or things that need to be avoided, such as cowardice and stupidity. Style positive and negative prompts include style positive prompts and style negative prompts. Among them, style positive prompts refer to positive style descriptions provided by users, which are used to indicate the specific style or characteristics that the model should have when generating content. Negative prompts refer to negative style descriptions, which are used to instruct the model to avoid generating certain styles or characteristics that are not desired. Combined prompts refer to prompts obtained by combining positive and negative character prompts, positive and negative style prompts, and storyboard prompts.

[0205] After the server determines the positive and negative description words of the role based on the role description data and determines the positive and negative prompt words of the style based on the style description data, the positive and negative prompt words of the role, the positive and negative prompt words of the style, and the storyboard prompt words corresponding to each storyboard can be combined once to obtain the combined prompt words corresponding to the storyboard. In this way, the combined prompt words of multiple storyboards are obtained. Subsequently, the combined prompt words of the multiple storyboards can be input into the pre-trained image generation model to instruct the image generation model to generate the storyboard images of the multiple storyboards based on the input combined prompt words.

[0206] In some embodiments, the server may input the combined prompt words into the diffusion model, and the diffusion model generates the storyboard images corresponding to each storyboard based on the input combined prompt words of each storyboard. The diffusion model converts the image into random noise by simulating a process of gradually adding noise, and then learns to gradually restore the original image from the noise. Furthermore, after the storyboard images of each storyboard are generated, the storyboard images may also be post-processed, such as image enhancement, tone and contrast adjustment, and other image processing methods.

[0207] The technical solution of the above embodiment has the following advantages and beneficial effects: by combining style prompt words, role prompt words and storyboard prompt words, and then generating a storyboard based on the combined prompt words, the style and character characteristics of the generated storyboard can be accurately controlled to ensure that the generated content conforms to the overall style and character settings of the picture book.

[0208] In some embodiments, when the server executes the step of generating a storyboard of a storyboard based on the combined prompt words, the server is further configured to: for each storyboard, encode the combined prompt words corresponding to the storyboard, input the encoded combined prompt words into a preset image generation model, and obtain the storyboard of the storyboard; wherein the image generation model has a built-in consistent attention mechanism.

[0209] In this embodiment, the image generation model can be a diffusion model based on the U-Net architecture (such as SD1.5, SDXL and Kolors models). In order to better maintain the consistency of the characters, including character features and character attributes (such as clothing, hair, etc.), in this embodiment, the consistent attention mechanism of storyDiffusion can be introduced into the diffusion model, and the consistent attention mechanism can be plugged in and out of the U-Net diffusion model to replace the original self-attention of the U-Net architecture and reuse the original self-attention weights. The consistent attention mechanism can establish interactions between images in a batch to maintain character consistency.

[0210] In specific implementation, after the server obtains the combined prompt words by combining style prompt words, role prompt words and storyboard prompt words, it can input the combined prompt words of each storyboard into the text encoder, encode the combined prompt words into a vector representation, and then input the encoded vector representation into the Kolors model that introduces the consistent attention mechanism. The Kolors model first generates a series of storyboards based on the combined prompt words of each storyboard, and then establishes interactions between continuous images through the consistent attention mechanism to maintain character consistency, and outputs storyboards that ensure character consistency. Among them, the text encoder may include but is not limited to CLIP, chatGLM, etc. The processing process of the consistent attention mechanism can be found in the relevant consistent attention mechanism information, which will not be repeated here.

[0211] The technical solution of the above embodiment has the following advantages or beneficial effects: by introducing a consistent attention mechanism into the image generation model, the encoded combined prompt words are input into the image generation model introducing the consistent attention mechanism, so that a storyboard that maintains character consistency can be generated, thereby improving the quality of the picture book story.

[0212] In some embodiments, when the server executes the step of generating a storyboard of each storyboard based on the combined prompt words, the server is further configured to: when the picture book creation input data includes a reference image, for each storyboard, encode the reference image and the combined prompt words corresponding to the storyboard, input the encoded reference image and the combined prompt words corresponding to the storyboard into a preset image generation model to obtain the storyboard of the storyboard, wherein the image generation model has a built-in consistent attention mechanism.

[0213] Character reference images refer to images that depict the appearance features of the main characters, including but not limited to the characters' clothing, hairstyles, expressions, postures, etc. These images help the generative model accurately capture the visual characteristics of the characters and ensure that the appearance of the characters remains consistent in different storyboards. Scene reference images refer to images that depict the background of a specific scene, such as a forest, a river, a garden at night, etc. These images help the generative model understand the atmosphere, color, light and other elements of the scene, so as to accurately reproduce these environmental features in the generated storyboards.

[0214] In the picture book creation process, in order to ensure that the generated storyboards are consistent with the user's intentions and have a high degree of character and scene consistency, character reference images and scene reference images are usually used. These reference images can help the generation model better understand and generate images that meet expectations. In actual applications, the user interface of the picture book creation service provides an entry for uploading reference images. Users can click on the entry to upload corresponding reference images, including character reference images or scene reference images, to instruct the model to generate a story picture book that meets the user's expectations.

[0215] Taking the example of a user uploading a character reference image and a scene reference image, after the server combines multiple prompt words to obtain a combined prompt word, in order to maintain the characteristics of the character and scene, the server can input the combined prompt word of each storyboard into a pre-trained text encoder, encode the combined prompt word into a vector representation, and input the character reference image and the scene reference image into a pre-trained image encoder to encode them into a vector representation. For ease of distinction, the vector representation encoded by the text encoder can be called a text vector representation, and the vector representation encoded by the image encoder can be called an image vector representation. Subsequently, the encoded text vector representation and image vector representation are combined and input into the Kolors model that introduces a consistent attention mechanism. The Kolors model establishes interactions between images in a batch through a consistent attention mechanism to generate storyboards that are consistent in terms of characters and scenes.

[0216] In some embodiments, after the server combines multiple prompt words to obtain combined prompt words, the combined prompt words, character reference images and scene reference images of each storyboard are input into the Stacked ID (identity) model, and the Stacked ID performs text encoding and image encoding. The encoding result is then embedded into the Kolors model that introduces a consistent attention mechanism. The Kolors model establishes interaction between images in a batch through the consistent attention mechanism to generate storyboards that are consistent in terms of characters and scenes.

[0217] The Stacked ID model is a multi-layered deep learning architecture that usually stacks the encodings of multiple input ID images to form a unified ID representation. Specifically, the Stacked ID model includes multiple layers, such as layers for extracting text features and layers for extracting image features. Each layer is responsible for extracting features at different levels from the input data. After extracting text features and image features respectively through multiple layers, the image features and text features are combined into a vector as the input of the Kolors model with a consistent attention mechanism to generate a storyboard that is consistent in terms of characters and scenes.

[0218] The technical solution of the above embodiment has the following advantages or beneficial effects: by inputting the combined prompt words and the reference image into the Stacked ID model, efficient and accurate image encoding and text encoding can be achieved, and then the encoding result is embedded in the image generation model to generate high-quality storyboards with consistent characters and scenes.

[0219] In some embodiments, the style prompt words are added with style trigger words. When the server executes the step of inputting the encoded combined prompt words into a preset image generation model to obtain a storyboard of the storyboard, the server is further configured to: input the encoded combined prompt words into the preset image generation model, trigger the style generation model corresponding to the style trigger words, obtain the storyboard of the storyboard, and the storyboard matches the style represented by the style prompt words.

[0220] When the server executes the step of inputting the encoded character reference image and the combined prompt words corresponding to the storyboard into a preset image generation model to obtain a storyboard of the storyboard, the server is further configured to: input the encoded character reference image and the combined prompt words corresponding to the storyboard into the preset image generation model, trigger the style generation model corresponding to the style trigger words, and obtain the storyboard of the storyboard, and the storyboard matches the style represented by the style prompt words.

[0221] Among them, the style generation model is obtained by training the low-rank adaptation model based on the storyboards carrying tag annotation data and caption annotation data.

[0222] The low-rank adaptive model is the LoRA model, which is usually used to fine-tune a large-scale pre-trained model without retraining the entire model. In this embodiment, the style generation model is a model used to guide the image generation model to generate images with a specific style.

[0223] A style trigger is a keyword used to trigger a style generation model. When it appears in a prompt, it triggers the corresponding style generation model. These style generation models are specially trained to generate images with a specific style. For example, a style trigger may be "dream watercolor", and the corresponding style generation model will generate images with a dream watercolor style.

[0224] In practical applications, in order to improve the style effect of picture books, in this embodiment, the style generation model can be pre-trained, and corresponding style trigger words can be added to the style prompt words corresponding to each style. When the user selects a certain style, the corresponding style prompt words are first determined based on the style. Since style trigger words are added to the style prompt words, the corresponding style generation model will be triggered, and other style generation models will be disabled at the same time.

[0225] Specifically, for each style, a corresponding style generation model can be trained. The training process of each style generation model can be: select more than 200 high-quality images corresponding to the style, and then annotate the selected images. Generally speaking, image annotation methods include caption annotation and tag annotation. Since the insertion position of LoRA is the linear projection part of the QKV of the cross attention in U-Net, and the cross attention mainly affects the alignment relationship between the text token and the image Latents. If caption annotation is used, the relationship between the image style and all the words in the caption is established during training. If tag annotation is used, the relationship between the image style and a single word is established. After experimental comparison, it is also proved that in terms of picture quality and text-image matching, the training effect of using caption+tag annotation method is better.

[0226] Therefore, in this embodiment, GPT-4V is selected to perform caption annotation on the image, and a style trigger word tag is added to the annotation header. Then, the LoRA model is trained with images with caption annotations and tag annotations as training data, with a training epoch of 10 to 15 and a learning rate of 1e-5, so as to obtain a style generation model through training.

[0227] In specific implementation, for the case where the user has not input a reference image, the server inputs the encoded combined prompt words into the Kolors model. When a style trigger word in the style prompt word is detected, the server loads the style generation model corresponding to the style trigger word. Its low-rank matrix will modify the weights of the Kolors model, thereby keeping the style of the generated storyboard unified.

[0228] Similarly, when the user inputs a reference image, the server inputs the encoded combined prompt word and the reference image into the Kolors model. When a style trigger word in the style prompt word is detected, the server loads the style generation model corresponding to the style trigger word. Similarly, its low-rank matrix will modify the weights of the Kolors model, so that the generated storyboard has a specific style. Specifically, the generated storyboard with a unified style can be seen in Fig.10 .

[0229] The technical solution of the above embodiment has the following advantages or beneficial effects: by adding style trigger words to the style prompt words and training the corresponding style generation model, the image generation model can be fine-tuned to enable the image generation model to flexibly generate storyboards of different styles, thereby greatly reducing the consumption of computing resources.

[0230] In some embodiments, after the server generates a storyboard for each storyboard based on the storyboard prompt words of the storyboard, the server is further configured to: for each storyboard, if it is identified that the storyboard contains sensitive content, adjust the storyboard until the adjusted storyboard does not contain sensitive content.

[0231] Sensitive content in images includes content that violates social morality and laws and regulations, such as violence and hate speech.

[0232] In practical applications, a pre-trained deep learning model can be trained based on some images containing sensitive content in advance to obtain an image recognition model for identifying sensitive content in images. In specific implementation, after the server generates a corresponding storyboard for each storyboard, the storyboard can be subjected to image preprocessing such as cropping, scaling, and normalization to process the storyboard into a format that conforms to the input of the image recognition model. Subsequently, the pre-processed storyboard is input into the trained image recognition model, which performs sensitive content detection and outputs a confidence score. Subsequently, the server determines whether the storyboard contains sensitive content based on the confidence score.

[0233] In some embodiments, developers can also collect a large number of sensitive images containing sensitive content in advance to build a sensitive image library. The sensitive image database can be built on a server, or in a data storage system or other place accessible to the server. In specific implementations, the server can perform similarity calculations on the generated storyboards and the sensitive images in the pre-built sensitive image library, and identify whether the storyboards contain sensitive content based on the similarity calculation results. In other embodiments, it is also possible to identify whether sensitive content exists in the storyboards by distance calculation.

[0234] If the server determines that the storyboard contains sensitive content, it can blur the area where the sensitive content is located in the detected storyboard, or cover it with mosaics or other obstructions, or directly delete the storyboard and generate a new storyboard by introducing an image generation model with a consistent attention mechanism. It can also push a sensitive content correction message to notify the management staff to correct the sensitive content in the storyboard. The specific method can be determined according to the actual situation and is not limited here.

[0235] The technical solution of the above-mentioned embodiment has the following advantages and beneficial effects: by identifying whether there is sensitive content in the storyboard, and correcting the sensitive content in the storyboard when sensitive content is identified, the generated storyboard and storyboard text can be made to comply with the requirements of social morality and laws and regulations, thereby improving the quality of story picture book data.

[0236] In some embodiments, the present application further provides a picture book story generation method, which is applied to a display device as in the above embodiment. In this embodiment, Fig.11 As shown, the method comprises the following steps:

[0237] Step S702, receiving and identifying the interactive content input by the user, and sending the interactive content to the server. The interactive content is input by the user on the user interface of the picture book creation service.

[0238] Step S704, receiving story picture book data fed back by the server, wherein the story picture book data includes storyboard texts and storyboard images of multiple storyboards, the storyboard text is generated by calling a large language model based on interactive content, the storyboard image is generated based on storyboard prompt words, the storyboard prompt words are extracted from the storyboard text by triggering a storyboard prompt word extraction model applied to the large language model, and the storyboard prompt word extraction model is trained on a low-rank adaptive model based on storyboard text samples annotated with storyboard prompt words that match the characters.

[0239] Step S706, displaying story picture book data.

[0240] In some embodiments, the method further includes: typeset layout of storyboard texts and storyboard images of a plurality of storyboards, and displaying the storyboard texts and storyboard images of each storyboard after typeset layout on a user interface.

[0241] In some embodiments, the method further includes: when the resolution of the storyboard does not meet the preset resolution requirement, converting the storyboard into a storyboard that meets the preset resolution requirement, and displaying the converted storyboard.

[0242] In some embodiments, the method also includes: obtaining voice data input by the user, performing intent recognition based on the voice data, determining the user's interaction intention, and when the interaction intention represents starting the picture book creation service, displaying the user interface of the picture book creation service for the user to input interaction content.

[0243] Since the implementation solution for solving the problem provided by the above-mentioned story picture book generation method embodiment is similar to the implementation solution recorded in the above-mentioned display device embodiment, the specific limitations in one or more of the above-mentioned story picture book generation method embodiments can be referred to the above limitations on the display device and will not be repeated here.

[0244] like Fig.12 As shown, in some embodiments, the present application further provides a story picture book generation method, which is applied to the server in the above embodiment. In this embodiment, the method includes the following steps:

[0245] Step S802: receiving interactive content input by a user sent by a display device.

[0246] Step S804, calling a large language model based on the interactive content to generate a story text, where the story text includes storyboard texts of multiple storyboards.

[0247] Step S806, triggering the storyboard prompt word extraction model applied to the large language model to extract storyboard prompt words from each storyboard text. The storyboard prompt word extraction model is trained on the low-rank adaptive model based on storyboard text samples containing storyboard prompt word annotations that match the characters.

[0248] Step S808, for each storyboard, based on the storyboard prompt words of the storyboard, generate a storyboard image of the storyboard.

[0249] Step S810, feeding back the storyboard texts and storyboard images of the plurality of storyboards as storyboard data to a display device, so that the display device displays the storyboard data.

[0250] like Fig.13 As shown, in some embodiments, step S808 includes:

[0251] Step S828, determining style prompt words and role prompt words respectively based on the style description data and role description data in the picture book creation input data.

[0252] Step S848, for each storyboard, combine the style prompt words, the role prompt words and the storyboard prompt words corresponding to the storyboard to obtain the combined prompt words corresponding to the storyboard.

[0253] Step S868, when the picture book creation input data does not include a reference image, for each storyboard, the combined prompt words corresponding to the storyboard are encoded, and the encoded combined prompt words are input into a preset image generation model to obtain a storyboard image of the storyboard.

[0254] Step S888, when the picture book creation input data includes a reference image, for each storyboard, the reference image and the combined prompt words corresponding to the storyboard are encoded, and the encoded reference image and the combined prompt words corresponding to the storyboard are input into a preset image generation model to obtain a storyboard diagram of the storyboard.

[0255] In some embodiments, style trigger words are added to the style prompt words; the encoded combined prompt words are input into a preset image generation model to obtain a storyboard of the storyboard, which includes: inputting the encoded combined prompt words into a preset image generation model, triggering the style generation model corresponding to the style trigger words, and obtaining a storyboard of the storyboard, wherein the storyboard matches the style represented by the style prompt words.

[0256] Inputting the encoded reference image and the combined prompt words corresponding to the storyboard into a preset image generation model to obtain a storyboard of the storyboard includes: inputting the encoded reference image and the combined prompt words corresponding to the storyboard into the preset image generation model, triggering the style generation model corresponding to the style trigger words, and obtaining the storyboard of the storyboard, wherein the storyboard matches the style represented by the style prompt words; wherein the style generation model is obtained by training a low-rank adaptation model based on the storyboard carrying tag annotation data and caption annotation data.

[0257] like Fig.14 As shown, in some embodiments, the interactive content includes picture book creation input data. Before step S804, the method further includes: step S803, preprocessing the picture book creation input data to obtain preprocessed picture book creation input data, the preprocessing including at least one of traditional and simplified conversion, spelling correction, and sensitive word filtering;

[0258] Step S804 includes: Step S824, taking the preprocessed picture book creation input data as input, calling the first language model, so that the first language model outputs prompt words according to the preprocessed picture book creation input data and the preset story text to generate the story text.

[0259] In some embodiments, after step S824, it also includes: taking the story text as input, calling the second largest language model, so that the second largest language model scores the story text according to a preset story scoring standard, and obtaining a scoring result; when the score of the story text is less than a preset scoring threshold, determining optimization suggestions based on the scoring result, and feeding back the optimization suggestions to the first largest language model, so that the first largest language model adjusts the story text based on the optimization suggestions until the score of the adjusted story text is greater than or equal to the preset scoring threshold.

[0260] like Fig.15 As shown, in some embodiments, step S808 includes: S828, for each storyboard, based on the storyboard prompt words of the storyboard, generating a storyboard image of the storyboard, and when it is identified that the storyboard image of the storyboard contains sensitive content, adjusting the storyboard image until the adjusted storyboard image does not contain sensitive content.

[0261] Since the implementation solution for solving the problem provided by the above-mentioned story picture book generation method embodiment is similar to the implementation solution recorded in the above-mentioned server embodiment, the specific limitations in one or more of the above-mentioned story picture book generation method embodiments can be referred to the above limitations on the server and will not be repeated here.

[0262] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the method of the above embodiment is implemented when the processor executes the computer program.

[0263] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method of the above embodiment is implemented.

[0264] In one embodiment, a computer program product is provided, including a computer program, which implements the method of the above embodiment when executed by a processor.

[0265] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0266] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (M RAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited thereto.

[0267] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0268] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A display device, characterized in that: include: A display, configured to display a user interface of the picture book creation service in response to a start operation triggered by a user for the picture book creation service; The controller is configured as: Identifying the interactive content input by the user, and controlling the display to display story picture book data generated based on the interactive content; The story book data includes storyboard texts and storyboard images of multiple storyboards, the storyboard texts are generated by calling a large language model based on the interactive content, the storyboard images are generated based on storyboard prompt words, the storyboard prompt words are extracted from the storyboard texts by triggering a storyboard prompt word extraction model applied to the large language model, and the storyboard prompt word extraction model is trained on a low-rank adaptive model based on storyboard text samples annotated with storyboard prompt words that match the characters.

2. The display device according to claim 1, characterized in that The controller is further configured to: when the resolution of the storyboard does not meet the preset resolution requirement, convert the storyboard into a storyboard that meets the preset resolution requirement, and control the display to display the converted storyboard.

3. The display device according to claim 1 or 2, characterized in that: The device also includes: A sound collector is configured to: obtain voice data input by the user, and send the voice data to the controller, wherein the interactive content includes the voice data; The controller is further configured to: perform intent recognition based on the voice data to determine the user's interaction intention, and when the interaction intention represents starting a picture book creation service, control the display to display a user interface of the picture book creation service.

4. A server, characterized in that: The server comprises: A communication module, configured to establish a communication connection with a display device; The processor is configured as: Receive interactive content input by the user; Based on the interactive content, a large language model is called to generate a story text, wherein the story text includes storyboard texts of multiple storyboards; Triggering a storyboard prompt word extraction model applied to a large language model to extract storyboard prompt words from each of the storyboard texts, wherein the storyboard prompt word extraction model is obtained by training a low-rank adaptive model based on storyboard text samples containing storyboard prompt word annotations matching the characters; For each storyboard, based on the storyboard prompt words of the storyboard, generating a storyboard image of the storyboard; Based on the storyboard texts and storyboard images of the plurality of storyboards, story picture book data is obtained.

5. The server according to claim 4, characterized in that: The interactive content includes picture book creation input data; When the server executes the step of generating a storyboard image of each storyboard based on the storyboard prompt words of the storyboard, the server is further configured to: Based on the style description data and the role description data in the picture book creation input data, respectively determine the style prompt words and the role prompt words; For each storyboard, combining the style prompt word, the character prompt word and the storyboard prompt word corresponding to the storyboard to obtain a combined prompt word corresponding to the storyboard; In the case that the picture book creation input data does not include a reference image, for each storyboard, encoding the combination prompt words corresponding to the storyboard, and inputting the encoded combination prompt words into a preset image generation model to obtain a storyboard image of the storyboard; In the case where the picture book creation input data includes a reference image, for each storyboard, the reference image and a combination prompt word corresponding to the storyboard are encoded, and the encoded reference image and the combination prompt word corresponding to the storyboard are input into a preset image generation model to obtain a storyboard image of the storyboard; Among them, the image generation model has a built-in consistent attention mechanism.

6. The server according to claim 5, characterized in that: The style prompt words are added with style trigger words; When the server executes the step of inputting the encoded combined prompt words into a preset image generation model to obtain the storyboard of the storyboard, the server is further configured to: Inputting the encoded combined prompt word into a preset image generation model to trigger a style generation model corresponding to the style trigger word, and obtaining a storyboard of the storyboard, wherein the storyboard matches the style represented by the style prompt word; When the server executes the step of inputting the encoded reference image and the combined prompt words corresponding to the storyboard into a preset image generation model to obtain the storyboard image of the storyboard, the server is further configured to: Inputting the encoded reference image and the combined prompt word corresponding to the storyboard into a preset image generation model, triggering the style generation model corresponding to the style trigger word, and obtaining a storyboard image of the storyboard, wherein the storyboard image matches the style represented by the style prompt word; The style generation model is obtained by training a low-rank adaptation model based on storyboards carrying tag annotation data and caption annotation data.

7. The server according to any one of claims 4 to 6, characterized in that: The interactive content includes picture book creation input data, and the large language model includes a first large language model; when the server executes the step of calling the large language model based on the interactive content to generate a story text, the server is further configured to: Preprocessing the picture book creation input data to obtain preprocessed picture book creation input data, wherein the preprocessing includes at least one of traditional Chinese and simplified Chinese conversion, spelling correction, and sensitive word filtering; Taking the preprocessed picture book creation input data as input, calling the first large language model, so that the first large language model generates story text according to the preprocessed picture book creation input data and preset story text output prompt words.

8. The server according to claim 7, characterized in that: After executing the step of taking the pre-processed picture book creation input data as input, calling the first language model, and generating the story text, the server is further configured as follows: Taking the story text as input, calling the second largest language model, so that the second largest language model scores the story text according to a preset story scoring standard to obtain a scoring result; When the score of the story text is less than a preset score threshold, an optimization suggestion is determined based on the score result, and the optimization suggestion is fed back to the first large language model so that the first large language model adjusts the story text based on the optimization suggestion until the score of the adjusted story text is greater than or equal to the preset score threshold.

9. A method for generating a story picture book, characterized in that: Applied to the display device according to any one of claims 1 to 3, the method comprises: Receive and identify interactive content input by a user, and send the interactive content to a server, wherein the interactive content is input by the user on a user interface of a picture book creation service; Receive story picture book data fed back by the server, wherein the story picture book data includes storyboard texts and storyboard images of a plurality of storyboards, the storyboard text is generated by calling a large language model based on the interactive content, the storyboard image is generated based on storyboard prompt words, the storyboard prompt words are extracted from the storyboard text by triggering a storyboard prompt word extraction model applied to the large language model, and the storyboard prompt word extraction model is obtained by training a low-rank adaptive model based on storyboard text samples annotated with storyboard prompt words matching the characters; The story picture book data is displayed.

10. A method for generating a story picture book, characterized in that: Applied to the server according to any one of claims 4 to 8, the method comprises: Receiving interactive content input by a user sent by a display device; Based on the interactive content, a large language model is called to generate a story text, wherein the story text includes storyboard texts of a plurality of storyboards; triggering a storyboard prompt word extraction model applied to the large language model to extract storyboard prompt words from each storyboard text, wherein the storyboard prompt word extraction model is obtained by training a low-rank adaptive model based on storyboard text samples containing storyboard prompt word annotations matching the characters; For each storyboard, based on the storyboard prompt words of the storyboard, generating a storyboard image of the storyboard; The storyboard texts and storyboard images of the plurality of storyboards are fed back to the display device as storybook data, so that the display device displays the storybook data.

Citation Information

Cited By

  • Teaching document generation method and device, server and computer storage medium

    CN121071172A

  • Multi-role real scene picture book generation method for autistic children

    CN121392042A

  • Voice-driven intelligent picture book generation method and device, electronic equipment and storage medium

    CN121393444A