Picture book generation method and display equipment

By applying text processing models and automatic image generation technology on display devices, the problem of low efficiency in traditional picture book generation is solved, and high-quality picture book content is quickly and efficiently generated.

CN119991871APending Publication Date: 2025-05-13HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411982555.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

There is a problem of low generation efficiency in the generation process of traditional picture books, which requires a lot of time and energy to be invested by professional creative teams.

Method used

By applying a picture book generation method on the display device, the text processing model is used to extract picture book theme information and role information from the text information, the initial text content is automatically generated, and high-quality picture book content is generated through multiple adjustments to the text quality scoring dimension. At the same time, the image data generation process of picture book is automatically performed, and the timing synchronous synthesis is performed with background audio and aloud audio.

Benefits of technology

It has achieved the reduction of time for manual analysis and extraction of key information, quickly generated high-quality picture book content, reduced the workload of manual creation and editing, and improved the efficiency of picture book generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991871A_ABST
    Figure CN119991871A_ABST
Patent Text Reader

Abstract

The invention relates to a picture book generation method and display equipment, and relates to the technical field of display equipment. The method comprises the following steps: receiving picture book generation demand information; identifying text information corresponding to the picture book generation demand information, and inputting the text information into a text processing model to obtain picture book theme information and picture book role information; performing picture book content generation processing according to the picture book theme information and the picture book role information to obtain initial text content of the picture book; inputting the initial text content of the picture book into a text processing model to obtain target text content of the picture book; and performing picture book image data generation processing according to the target text content of the picture book to obtain picture book image data of each page in the picture book, and synthesizing the picture book corresponding to the picture book generation demand information according to the picture book image data of each page in the picture book, the data representing the background audio and the data representing the broadcast audio. By adopting the method, the picture book generation efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of display devices, and in particular to a picture book generation method and a display device. Background Art

[0002] Display devices such as smart TVs refer to devices that realize two-way human-computer interaction functions based on Internet application technology. They have multiple functions, such as playing picture books through display devices.

[0003] In traditional technology, when generating picture books, it usually requires a professional creative team to invest a lot of time and energy, including determining the story content, drawing picture book images, etc. The whole process is time-consuming, resulting in low efficiency in picture book generation.

[0004] Therefore, there is a technical problem in the traditional technology that the generation efficiency of picture books is low. Summary of the invention

[0005] The present application provides a picture book generation method and a display device to solve the technical problem of low picture book generation efficiency.

[0006] In a first aspect, some embodiments provide a picture book generation method, which is applied to a display device; the method comprises:

[0007] Receive picture book generation demand information;

[0008] Identify text information corresponding to the picture book generation requirement information, and input the text information into a text processing model to obtain picture book theme information and picture book role information; the picture book theme information is information representing the theme in the semantic information of the text information, and the picture book role information is information representing the role in the semantic information of the text information; the text processing model is used to output the picture book theme information and the picture book role information matching the semantic information according to the input semantic information of the text information;

[0009] According to the picture book theme information and the picture book character information, a picture book text content generation process is performed to obtain the initial text content of the picture book;

[0010] Inputting the initial text content of the picture book into the text processing model to obtain the target text content of the picture book; the text quality score of the target text content is higher than the text quality score of the initial text content; the text processing model is further used to perform text adjustment processing on the input initial text content in multiple text quality score dimensions, and output the adjusted target text content;

[0011] According to the target text content of the picture book, picture book image data generation processing is performed to obtain picture book image data of each page in the picture book, and according to the picture book image data of each page in the picture book, data representing background audio and data representing broadcast audio, a picture book corresponding to the picture book generation demand information is synthesized; the picture book image data of each page respectively represent the image data corresponding to the text content of each page in the target text content.

[0012] Technical effect: Automatically extracting picture book theme information and picture book character information from text information through a text processing model can help reduce the time for manual analysis and extraction of key information; automatically generating initial text content based on the extracted information and using a text processing model to adjust the text can help quickly generate high-quality picture book content; automatically generating and processing picture book images and synthesizing picture books can help reduce the workload of manual creation and editing; this fully automated processing method can help improve the efficiency of picture book generation.

[0013] In some embodiments of the present application, the step of inputting the initial text content of the picture book into the text processing model to obtain the target text content of the picture book includes:

[0014] Inputting the initial text content of the picture book into the text processing model to obtain the grammatically corrected text content of the picture book; the grammatical quality score of the grammatically corrected text content is higher than the grammatical quality score of the initial text content; the text processing model is also used to perform grammatical correction processing on the input initial text content and output the grammatically corrected text content;

[0015] Inputting the grammatically corrected text content into the text processing model to obtain the text content of the picture book after the language style is adjusted; the language style of the text content after the language style is adjusted is the same as the language style matched by the target audience of the picture book; the text processing model is also used to perform language style adjustment processing on the input grammatically corrected text content, and output the text content after the language style is adjusted;

[0016] Inputting the text content after the language style is adjusted into the text processing model to obtain the text content after the expression of the picture book is adjusted; the expression of the text content after the expression is adjusted is the same as the preset expression; the text processing model is also used to adjust the input text content after the language style is adjusted according to the preset expression, and output the text content after the expression is adjusted;

[0017] The text content after the expression mode is adjusted is input into the text processing model to obtain the target text content of the picture book; the emotional color information of the target text content of the picture book is the same as the picture book emotional color information in the picture book generation requirement information; the text processing model is also used to adjust the emotional color information of the picture book on the text content after the expression mode is adjusted, and output the target text content after the emotional color information is adjusted.

[0018] Technical effect: By using the text processing model for four consecutive processing steps, namely grammar correction, language style adjustment, expression adjustment and emotional color adjustment, it is beneficial to improve the different dimensions of text quality in each step and ensure that each dimension meets the expected evaluation standards; by conducting quality evaluation of the corresponding dimension after each processing step, it is beneficial to ensure the controllability and effectiveness of the processing results; it achieves an all-round quality improvement of the picture book text content, so that the final generated picture book text not only conforms to grammatical norms, but also adapts to the target audience, and also has standardized expressions and matching emotional colors.

[0019] In some embodiments of the present application, the method further includes:

[0020] The text information is input into the text processing model to obtain the picture book emotional color information; the picture book emotional color information is the information representing the emotional color in the semantic information of the text information; the text processing model is also used to output the picture book emotional color information matching the semantic information based on the input semantic information of the text information.

[0021] Technical effect: By inputting text information into the text processing model to automatically extract the emotional color information in the semantic information, it is helpful to accurately identify the emotional characteristics contained in user needs; by ensuring that the output picture book emotional color information matches the semantic information of the text information, it is helpful to maintain the coherence and consistency of the emotional expression of the picture book, and improve the accuracy of emotional expression in the picture book generation process.

[0022] In some embodiments of the present application, the picture book image data generation process is performed according to the target text content of the picture book to obtain the picture book image data of each page in the picture book, including:

[0023] Performing scene recognition processing on the text content of each page to obtain scene information corresponding to the text content of each page;

[0024] Performing image composition processing according to the scene information to obtain basic picture book image data corresponding to the scene information as the basic picture book image data of each page;

[0025] According to the picture book style information, style rendering processing is performed on the basic picture book image data of each page to obtain the picture book image data of each page; the picture book style information represents style information that matches the picture book theme information and the picture book character information; the style information of the picture book image data of each page is the same as the picture book style information.

[0026] Technical effect: Through the three-step progressive processing of scene recognition processing, image composition processing and style rendering processing for each page of text content, it is helpful to accurately convert the text content into an image with a reasonable layout, and ensure that the image style matches the theme of the picture book and the characteristics of the characters; by maintaining the coherence of scene information and the consistency of style information in each processing step, it is helpful to generate picture book images with unified style and coordinated content, realizing the automatic conversion from text to image, and ensuring the integrity and coordination of the picture book image in visual expression.

[0027] In some embodiments of the present application, the step of inputting the text information into a text processing model to obtain picture book theme information and picture book character information includes:

[0028] Inputting the text information into the text processing model to obtain semantic information of the text information; the text processing model is used to perform semantic analysis on the input text information and output the semantic information of the text information;

[0029] The semantic information is input into the text processing model to obtain an intent classification result of the text information; the intent classification result is used to indicate the type of picture book corresponding to the text information; the text processing model is used to perform intent classification processing on the input semantic information and output the intent classification result of the text information;

[0030] The intention classification result is input into the text processing model to obtain the picture book theme information and the picture book character information; the text processing model is used to output the picture book theme information and the picture book character information that match the semantic information according to the input intention classification result.

[0031] Technical effect: By dividing the text information into three steps and inputting it into the text processing model in sequence, the layer-by-layer conversion processing from text information to semantic information, intent classification results, and finally to picture book theme information and picture book role information is realized, which is conducive to accurately understanding user needs and converting them into specific picture book creation elements; by maintaining the consistency and matching of information in each step, it is conducive to ensuring that the final generated picture book theme and role are consistent with the user's original needs, realizing the accurate conversion from user needs to specific creation elements, and improving the pertinence and accuracy of picture book creation.

[0032] In some embodiments of the present application, the picture book text content generation process is performed according to the picture book theme information and the picture book character information to obtain the initial text content of the picture book, including:

[0033] Selecting a picture book theme that matches the picture book theme information from the preset picture book themes;

[0034] Performing role feature recognition processing on the picture book role information to obtain role features of the picture book role information;

[0035] According to the theme of the picture book and the characteristics of the characters, a storyline construction process is performed to obtain the initial text content of the picture book.

[0036] Technical effect: By selecting a picture book theme that matches the picture book theme information from the preset picture book themes, it is helpful to ensure that the generated picture book content meets the expected educational theme and creation direction; by performing role feature recognition processing on the picture book character information, it is helpful to accurately grasp the character's personality and behavioral characteristics; by constructing the picture book plot according to the picture book theme and character characteristics, it is helpful to generate the initial text content of the picture book that meets the theme and has distinct character characteristics. This processing flow from theme selection to character shaping and then to plot construction realizes the systematic generation of picture book content and ensures that the generated picture book has a clear theme orientation and rich character characteristics.

[0037] In some embodiments of the present application, the identifying text information corresponding to the picture book generation requirement information includes:

[0038] In the case where the picture book generation demand information is voice information, preprocessing the picture book generation demand information to obtain preprocessed picture book generation demand information;

[0039] Performing acoustic feature extraction processing on the preprocessed picture book generation demand information to obtain acoustic feature information of the preprocessed picture book generation demand information;

[0040] According to the acoustic feature information, speech recognition processing is performed on the preprocessed picture book generation requirement information to obtain text information corresponding to the picture book generation requirement information.

[0041] Technical effect: By preprocessing the picture book generation demand information in the form of voice, it is helpful to improve the quality and processability of the voice signal; by performing acoustic feature extraction processing on the preprocessed picture book generation demand information, it is helpful to obtain the key acoustic features in the voice information; by performing voice recognition processing based on the acoustic feature information, it is helpful to accurately convert the voice information into text information. This layered processing method from preprocessing to feature extraction to voice recognition achieves high-quality conversion from voice input to text output, providing an accurate text basis for subsequent picture book generation.

[0042] In some embodiments of the present application, the picture book generation corresponding to the demand information is synthesized according to the picture book image data of each page in the picture book, the data representing the background audio, and the data representing the broadcast audio, including a controller, and is further configured as:

[0043] According to the target text content of each picture book story page, matching background audio music is selected from the background music and background audio library to obtain data representing the background audio music of each picture book story page, and a reading sound and reading audio generation process is performed to obtain data representing the reading sound and reading audio of each picture book story page;

[0044] The picture book image of each picture book story page, the data music representing the background audio, and the data representing the reading audio by the reading sound are subjected to time-series synchronous synthesis processing to obtain the picture book story corresponding to the picture book generation demand information.

[0045] Technical effect: By selecting background audio and generating reading audio for the text content of each page of the picture book, it is helpful to configure appropriate audio effects for each page of content; by performing time-series synchronous synthesis processing on the picture book image, background audio data and reading audio data of each page, it is helpful to achieve accurate matching of image display, background music playback and text reading in the time dimension, thereby improving the quality of the picture book.

[0046] In some embodiments of the present application, the text processing model is trained in the following manner:

[0047] Acquire pre-training data and fine-tuning data as training data; the pre-training data includes text content of picture books including multiple picture book theme information, picture book style information and language structure information; the fine-tuning data includes multiple picture book writing instructions, and the picture book writing instructions represent instruction information for guiding picture book writing;

[0048] Performing data preprocessing on the training data to obtain preprocessed training data; the preprocessed training data represents data that has been cleaned, formatted, labeled, and classified;

[0049] Using the preprocessed training data, iteratively train the text processing model to be trained to obtain a trained text processing model;

[0050] According to the model performance information of the trained text processing model, the trained text processing model is subjected to model adjustment processing to obtain the trained text processing model.

[0051] Technical effect: By using pre-training data containing a variety of picture book themes, styles and language structures, as well as fine-tuning data containing specific writing instructions, it is helpful to ensure the comprehensiveness and pertinence of the training data; by pre-processing the training data through cleaning, formatting, labeling and classification, as well as iterative training and performance adjustment of the model, it is helpful to improve the quality of training data and the training effect of the model, achieve high-quality training of the text processing model on the picture book creation task, and enhance the actual application effect of the model.

[0052] In a second aspect, some embodiments further provide a display device, including: a display and a controller.

[0053] The controller is coupled to the display and is configured to:

[0054] Receive picture book generation demand information;

[0055] Identify text information corresponding to the picture book generation requirement information, and input the text information into a text processing model to obtain picture book theme information and picture book role information; the picture book theme information is information representing the theme in the semantic information of the text information, and the picture book role information is information representing the role in the semantic information of the text information; the text processing model is used to output the picture book theme information and the picture book role information matching the semantic information according to the input semantic information of the text information;

[0056] According to the picture book theme information and the picture book character information, a picture book text content generation process is performed to obtain the initial text content of the picture book;

[0057] Inputting the initial text content of the picture book into the text processing model to obtain the target text content of the picture book; the text quality score of the target text content is higher than the text quality score of the initial text content; the text processing model is further used to perform text adjustment processing on the input initial text content in multiple text quality score dimensions, and output the adjusted target text content;

[0058] According to the target text content of the picture book, picture book image data generation processing is performed to obtain picture book image data of each page in the picture book, and according to the picture book image data of each page in the picture book, data representing background audio and data representing broadcast audio, a picture book corresponding to the picture book generation demand information is synthesized; the picture book image data of each page respectively represent the image data corresponding to the text content of each page in the target text content.

[0059] Technical effect: Automatically extracting picture book theme information and picture book character information from text information through a text processing model can help reduce the time for manual analysis and extraction of key information; automatically generating initial text content based on the extracted information and using a text processing model to adjust the text can help quickly generate high-quality picture book content; automatically generating and processing picture book images and synthesizing picture books can help reduce the workload of manual creation and editing; this fully automated processing method can help improve the efficiency of picture book generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0061] Figure 1 A schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application;

[0062] Figure 2 A schematic diagram of the hardware configuration of a display device provided in some embodiments of the present application;

[0063] Figure 3 A schematic diagram of the hardware configuration of a control device provided in some embodiments of the present application;

[0064] Figure 4 A schematic diagram of software configuration of a display device provided in some embodiments of the present application;

[0065] Figure 5 A schematic diagram of a process for generating a picture book provided in some embodiments of the present application;

[0066] Figure 6 A schematic diagram of a user interface of a picture book application provided in some embodiments of the present application;

[0067] Figure 7 A flowchart of steps for determining target text content of a picture book provided in some embodiments of the present application;

[0068] Figure 8 A flowchart of steps for determining picture book image data of each page in a picture book provided in some embodiments of the present application;

[0069] Fig. 9 A flowchart of steps for determining picture book theme information and picture book character information provided in some embodiments of the present application;

[0070] Fig.10 A flowchart of steps for determining the initial text content of a picture book provided in some embodiments of the present application;

[0071] Fig.11 A flowchart of steps for determining text information corresponding to picture book generation requirement information provided in some embodiments of the present application;

[0072] Fig.12 A flowchart of steps for determining a picture book corresponding to picture book generation requirement information provided in some embodiments of the present application;

[0073] Fig.13 A flowchart of steps for training a text processing model provided in some embodiments of the present application;

[0074] Fig.14 A signaling interaction diagram of a picture book generation method provided in some embodiments of the present application;

[0075] Fig.15 A structural block diagram of a picture book generating device provided in some embodiments of the present application. DETAILED DESCRIPTION

[0076] The following embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following embodiments do not represent all implementations consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application as detailed in the claims.

[0077] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.

[0078] The terms "first", "second", "third", etc. in the specification and claims of this application and the above drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances.

[0079] The terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0080] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0081] In the embodiment of the present application, the display device 200 generally refers to a device with image display and data processing capabilities. For example, the display device 200 includes but is not limited to a smart TV, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.

[0082] Figure 1 This is a schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application. Figure 1 As shown in FIG. 1 , the user can operate the display device 200 through touch operation, the mobile terminal 300 and the control device 100. For example, the control device 100 can be a remote controller, a stylus pen, a handle, etc.

[0083] The mobile terminal 300 can be used as a control device for performing human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device for establishing a communication connection with the display device 200 and performing data interaction. In some embodiments, the mobile terminal 300 can install software applications with the display device 200, and achieve connection and communication through a network communication protocol to achieve the purpose of one-to-one control operation and data communication. The audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 to achieve a synchronous display function.

[0084] like Figure 1 As also shown in FIG. 4 , the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 may be allowed to communicate and connect through a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0085] The display device 200 may provide a broadcast receiving television function, and may also additionally provide an intelligent network television function with a computer support function, including but not limited to network television, smart television, Internet Protocol television (IPTV), and the like.

[0086] Figure 2 Some embodiments of the present application provide Figure 1 2 is a block diagram of the hardware configuration of the display device 200.

[0087] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0088] In some embodiments, the detector 230 is used to collect signals of the external environment or external interaction. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of ambient light; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environment scenes, user attributes or user interaction gestures; or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.

[0089] In some embodiments, the display 260 includes a display function component for presenting a picture, and a driving component for driving an image display. The display 260 is used to receive an image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, and components of a menu control interface and a user control UI interface.

[0090] In some embodiments, the communication device 220 is a component for communicating with an external device or server 400 according to various communication protocol types. The display device 200 may be provided with a plurality of communication devices 220 according to different supported communication modes. For example, when the display device 200 supports wireless network communication, the display device 200 may be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.

[0091] The communication device 220 can enable the display device 200 to communicate with the external device or server 400 by wireless or wired connection. Among them, the wired connection can connect the display device 200 with the external device through components such as data cables and interfaces. The wireless connection can connect the display device 200 with the external device through wireless signals or wireless networks. The display device 200 can establish a connection relationship with the external device directly, or indirectly establish a connection relationship through a gateway, a router, a connection device, etc.

[0092] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first interface to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to the user's operation through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.

[0093] In some embodiments, the controller 250 and the tuner-demodulator 210 may be located in different separate devices, that is, the tuner-demodulator 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0094] In some embodiments, the user may input a user command through a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).

[0095] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or may be an external audio output device of the display device 200. In particular, for the external audio output device of the display device 200, the display device 200 may also be provided with an external audio output terminal, and the audio output device may be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.

[0096] In some embodiments, the user input interface 280 may be used to receive instructions from a user.

[0097] Figure 3 Some embodiments of the present application provide Figure 1 The hardware configuration diagram of the control device in the figure. Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.

[0098] The control device 100 is configured to control the display device 200 , and can receive user input operation instructions, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, playing the role of an interactive intermediary between the user and the display device 200 .

[0099] In some embodiments, the control device 100 may be a smart device, for example, the control device 100 may be installed with various applications for controlling the display device 200 according to user needs.

[0100] In some embodiments, Figure 1 As shown, the mobile terminal 300 or other intelligent electronic devices can play a similar function as the control device 100 after installing the application for controlling the display device 200 .

[0101] The controller 110 includes a processor 112, a RAM 113, a ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between the internal components and the external and internal data processing functions.

[0102] The communication interface 130 implements communication of control signals and data signals with the display device 200 under the control of the controller 110. The communication interface 130 may include at least one of other near field communication modules such as a WiFi chip 131, a Bluetooth module 132, and an NFC module 133.

[0103] The user input / output interface 140 , wherein the input interface includes at least one of other input interfaces such as a microphone 141 , a touch panel 142 , a sensor 143 , and a button 144 .

[0104] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with a communication interface 130, such as a WiFi, Bluetooth, NFC or other module, and can encode the user input command through the WiFi protocol, Bluetooth protocol, or NFC protocol and send it to the display device 200.

[0105] The memory 190 is used to store various operating programs, data and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.

[0106] The power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller.

[0107] In order to perform user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system may provide a user interface (control the display device), allow the user to interact with the display device 200, and support the running of various application programs.

[0108] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.

[0109] The operating system can be divided into different modules or layers according to the functions implemented, such as Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom, namely, the application layer (Applications) layer (referred to as "application layer"), the application framework layer (Application Framework) layer (referred to as "framework layer"), the system library layer and the kernel layer.

[0110] In some embodiments, the application layer is used to provide services and interfaces for applications so that the display device 200 can run applications and interact with users based on the applications. At least one application can be run in the application layer, and these applications can be window programs, system settings programs, clock programs, etc. that come with the operating system; they can also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the above examples.

[0111] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications. The application framework layer includes some predefined functions. The application framework layer is equivalent to a processing center that determines the actions that applications in the application layer take. Through the API interface, applications can access system resources and obtain system services during execution.

[0112] like Figure 4 As shown, in the embodiment of the present application, the application framework layer includes a view system, managers, content providers, etc., wherein the view system can design and implement the interface and interaction of the application, and the view system includes lists, grids, text boxes, buttons, etc. The manager includes at least one of the following modules: an activity manager for interacting with all activities running in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to the application package currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0113] In some embodiments, the activity manager is used to manage the life cycle of each application and the usual navigation back function, such as controlling the exit, opening, and back of the application. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, capturing the screen, and controlling the display window changes, for example, reducing the display window, shaking the display, distorting the display, etc.

[0114] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction library contained in the system runtime layer, such as the C / C++ instruction library, to implement the functions to be implemented by the framework layer.

[0115] In some embodiments, the kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. Figure 4 As shown, the kernel layer may be configured with hardware drivers, and the drivers included in the kernel layer may be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0116] It should be noted that the above example is only a simple division of the operating system functions and does not constitute a limitation on the specific operating system form of the display device 200 in the embodiment of the present application. Depending on factors such as the function of the display device and the type of operating system, the number of levels and specific level types contained in the operating system may be expressed in other forms.

[0117] In some embodiments, Figure 5 As shown, a picture book generation method is provided, which is applied to a display device 200; the method may include the following steps:

[0118] Step S501, receiving picture book generation requirement information.

[0119] Step S502, identify the text information corresponding to the picture book generation requirement information, and input the text information into the text processing model to obtain the picture book theme information and picture book role information; the picture book theme information is the information representing the theme in the semantic information of the text information, and the picture book role information is the information representing the role in the semantic information of the text information; the text processing model is used to output the picture book theme information and picture book role information that match the semantic information based on the semantic information of the input text information.

[0120] Step S503, performing picture book text content generation processing according to the picture book theme information and the picture book character information to obtain the initial text content of the picture book.

[0121] Step S504, inputting the initial text content of the picture book into the text processing model to obtain the target text content of the picture book; the text quality score of the target text content is higher than the text quality score of the initial text content; the text processing model is also used to perform text adjustment processing on the input initial text content in multiple text quality score dimensions, and output the adjusted target text content.

[0122] Step S505, according to the target text content of the picture book, picture book image data generation processing is performed to obtain picture book image data of each page in the picture book, and according to the picture book image data of each page in the picture book, data representing background audio and data representing broadcast audio, a picture book corresponding to the required information is synthesized; the picture book image data of each page respectively represents the image data corresponding to the text content of each page in the target text content.

[0123] The user interface may be a human-computer interaction interface displayed on the display device 200 , for example, a graphical interface including interactive elements such as operation buttons and input boxes.

[0124] The picture book application may be an application program for creating and displaying picture books running on the display device 200, for example, a software application having functions such as story creation and image generation. For example, the user interface of the picture book application may refer to Figure 6 ,exist Figure 6 In the picture book application, the user interface prompts users to create a picture book with a sentence (use a sentence to start a fantasy picture book journey and unleash unlimited creative imagination), and prompts users to "say the outline of a story according to the voice", that is, it provides a voice input function, and users can generate a picture book by clicking the "Generate Now" button.

[0125] The picture book generation requirement information may be relevant requirements input by the user for generating a picture book, for example, it may be voice or text input such as “I want a story about a brave puppy overcoming fear”.

[0126] The text information may be text content identified from the picture book generation requirement information, for example, text converted by voice recognition or directly input text content.

[0127] Among them, the text processing model (LLM, Large Language Model) can be an artificial intelligence language model trained with picture book related data, for example, it can be a large language model fine-tuned with story picture book language pre-training data and instruction fine-tuning data.

[0128] The theme information of the picture book may be the core theme content of the story, such as educational themes such as "children must learn to share" and "brush teeth from an early age".

[0129] The picture book character information may be a feature description of a person or animal character that appears in the story, for example, it may be a story character with specific personality traits defined by a user.

[0130] Among them, semantic information can be the meaning and content contained in the text. For example, semantic information can be key concepts, story types, emotional tendencies, and other information obtained through semantic analysis.

[0131] Among them, the story text content generation process can be a process of generating story content using a text processing model, for example, it can be a processing process including steps such as theme matching, character shaping, and plot conception.

[0132] Among them, the picture book page can be a single page that constitutes a complete picture book, for example, it can be a story page containing text content and illustrations.

[0133] The initial text content may be a preliminary story text obtained after the story generation process, for example, it may be the original story content that needs to be further optimized.

[0134] Among them, the text adjustment process can be a process of optimizing the initial text content, for example, it can be a processing process including grammar correction, style matching, language polishing, emotional tone optimization and other steps.

[0135] The target text content may be the final story text after text adjustment processing, for example, it may be an optimized, more vivid and complete story content.

[0136] The text quality score may be a quantitative evaluation indicator of the quality of the text content. For example, the text quality score may include scores for aspects such as the coherence, creativity, and topic relevance of the text.

[0137] Among them, the text quality scoring dimensions can be various aspects of evaluating text quality. For example, the text quality scoring dimensions can include grammatical correctness, style adaptability, language expressiveness, and emotional coherence.

[0138] Among them, text adjustment processing can be a process of optimizing text content. For example, text adjustment processing can include processing steps such as grammar correction, style matching, language polishing and emotional tone optimization.

[0139] Among them, the picture book image data generation process can be a process of generating illustrations based on text, for example, it can be a processing process including scene understanding, image composition, style rendering and other steps.

[0140] The picture book image data may be illustrations that match the story text, for example, story illustrations in different styles such as illustrations, watercolors, and comics.

[0141] The background audio may be background music data that matches the picture book, for example, the background audio may be matching background music selected according to the storyline.

[0142] The broadcast audio may be voice data for reading the contents of the picture book aloud, for example, the broadcast audio may be story reading audio generated by speech synthesis technology.

[0143] The picture book may be a complete multimedia storytelling work including text, pictures and audio, for example, it may be personalized picture book content generated according to user needs.

[0144] Specifically, refer to Figure 2 , the controller 250 receives the picture book generation requirement information input by the user in voice or text, and performs voice recognition processing on the picture book generation requirement information to obtain text information; the controller 250 inputs the text information into the text processing model fine-tuned by the story picture book language pre-training data and the instruction fine-tuning data, and obtains the picture book theme information and the picture book role information through semantic analysis, intention analysis and key information extraction; the controller 250 generates the picture book text content through theme matching, role shaping and plot conception according to the picture book theme information and the picture book role information, and obtains the initial text content of the picture book ; The controller 250 inputs the initial text content of the picture book into the text processing model, and obtains the target text content with a higher text quality score through text adjustment processing in multiple text quality scoring dimensions such as grammar correction, style adaptation, language polishing and emotional tonality optimization; The controller 250 generates picture book image data according to the target text content of the picture book through scene understanding, image composition and style rendering, obtains picture book image data of each page in the picture book, and combines the background audio data obtained by background music selection and the broadcast audio data obtained by voice reading synthesis, and finally synthesizes the picture book to generate a picture book corresponding to the required information.

[0145] The technical solution provided in this embodiment automatically extracts picture book theme information and picture book character information from text information through a text processing model, which is conducive to reducing the time for manual analysis and extraction of key information; by automatically generating initial text content based on the extracted information and using a text processing model to perform text adjustment processing, it is conducive to quickly generating high-quality picture book content; by automatically performing picture book image generation processing and synthesizing picture books, it is conducive to reducing the workload of manual creation and editing; this full-process automated processing method is conducive to improving the efficiency of picture book generation.

[0146] In some embodiments, Figure 7 As shown, the initial text content of the picture book is input into the text processing model to obtain the target text content of the picture book, which specifically includes the following contents:

[0147] Step S701, inputting the initial text content of the picture book into the text processing model to obtain the grammatically corrected text content of the picture book; the grammatical quality score of the grammatically corrected text content is higher than the grammatical quality score of the initial text content; the text processing model is also used to perform grammatical correction processing on the input initial text content and output the grammatically corrected text content;

[0148] Step S702, inputting the grammatically corrected text content into a text processing model to obtain the text content of the picture book after the language style is adjusted; the language style of the text content after the language style adjustment is the same as the language style matched by the target audience of the picture book; the text processing model is also used to perform language style adjustment processing on the input grammatically corrected text content, and output the text content after the language style adjustment;

[0149] Step S703, inputting the text content after the language style is adjusted into the text processing model to obtain the text content after the expression of the picture book is adjusted; the expression of the text content after the expression is adjusted is the same as the preset expression; the text processing model is also used to adjust the input text content after the language style is adjusted according to the preset expression, and output the text content after the expression is adjusted;

[0150] Step S704, input the text content after the expression is adjusted into the text processing model to obtain the target text content of the picture book; the emotional color information of the target text content of the picture book is the same as the picture book emotional color information in the picture book generation requirement information; the text processing model is also used to adjust the emotional color information of the picture book for the text content after the expression is adjusted, and output the target text content after the emotional color information is adjusted.

[0151] The grammatically corrected text content may be text content after grammatical error detection and sentence structure adjustment. For example, the grammatically corrected text content may be text content after grammatical error correction and sentence readability improvement.

[0152] The grammatical quality score may be a quantitative evaluation indicator of the grammatical correctness and sentence structure rationality of the text content. For example, the grammatical quality score may include scores for compliance with grammatical rules and completeness of sentence structure.

[0153] The grammar correction process may be a process of detecting and correcting grammatical errors in text content. For example, the grammar correction process may include processing steps such as detecting and correcting grammatical errors, adjusting sentence structure to improve readability, and the like.

[0154] Among them, the text content after the language style is adjusted can be the text content that has been adapted and adjusted in language style. For example, the text content after the language style is adjusted can be the text content that has been made vivid and simplified to be more suitable for children to read.

[0155] The target audience may be the intended readers of the picture book content, for example, the target audience may be children readers of a specific age group.

[0156] The language style adjustment process may be a process of optimizing the language style of the text content according to the characteristics of the target audience. For example, the language style adjustment process may include processing steps such as adjusting the difficulty of words and simplifying sentence structures.

[0157] Among them, the text content after the expression is adjusted can be the text content that has been optimized in terms of rhetoric and word selection. For example, the text content after the expression is adjusted can be the text content after rhetoric such as metaphor and personification are added and word selection is optimized.

[0158] The preset expression method may be a predefined text expression technique and rhetoric device. For example, the preset expression method may be a combination of rhetoric devices such as metaphor, personification, and parallelism.

[0159] Among them, the emotional color information can be the emotional tendency and emotional intensity expressed in the text content. For example, the emotional color information can be emotional characteristics such as warmth, inspiration, and joy.

[0160] The picture book emotional color information may be the expected emotional features extracted from the picture book generation demand information. For example, the picture book emotional color information may be the emotional tendencies such as bravery and warmth expressed by the user in the demand.

[0161] Specifically, the controller 250 inputs the initial text content of the picture book into a text processing model that has been fine-tuned with story picture book language pre-training data and instruction fine-tuning data, and obtains grammatically corrected text content by detecting grammatical errors and adjusting sentence structure; the controller 250 inputs the grammatically corrected text content into the text processing model, and according to the characteristics of the target audience being child readers, obtains text content with a more vivid and simple language style adjustment by adjusting the difficulty of words and simplifying sentence structure; the controller 250 inputs the text content with the language style adjusted into the text processing model, and obtains text content with adjusted expression by adding rhetorical devices such as metaphors and personification and optimizing word selection; the controller 250 inputs the text content with the expression adjusted into the text processing model, and obtains target text content that matches the picture book emotional color information in the picture book generation requirement information by adjusting the emotional color of the text and ensuring the emotional coherence of the story.

[0162] The technical solution provided in this embodiment, by using the text processing model for four consecutive processing steps of grammar correction, language style adjustment, expression adjustment and emotional color adjustment, is conducive to targetedly improving different dimensions of text quality in each step and ensuring that each dimension meets the expected evaluation standards; by performing quality evaluation of the corresponding dimension after each processing step, it is conducive to ensuring the controllability and effectiveness of the processing results; it achieves an all-round quality improvement of the picture book text content, so that the final generated picture book text not only conforms to grammatical norms, but also adapts to the target audience, and also has standardized expressions and matching emotional colors.

[0163] In some embodiments, the following contents are also included: inputting text information into a text processing model to obtain picture book emotional color information; the picture book emotional color information is information representing emotional color in the semantic information of the text information; the text processing model is also used to output picture book emotional color information that matches the semantic information based on the semantic information of the input text information.

[0164] Specifically, the controller 250 performs semantic analysis on the text information to obtain semantic information including story themes, character traits, plot development, etc., and inputs the semantic information into the text processing model. By analyzing the emotional expression, emotional changes, and emotional intensity in the semantic information, the controller 250 outputs picture book emotional color information that matches the semantic information at the emotional level.

[0165] The technical solution provided in this embodiment automatically extracts emotional color information from semantic information by inputting text information into a text processing model, which is conducive to accurately identifying the emotional characteristics contained in user needs; by ensuring that the output picture book emotional color information matches the semantic information of the text information, it is conducive to maintaining the coherence and consistency of the emotional expression of the picture book, thereby improving the accuracy of emotional expression in the picture book generation process.

[0166] In some embodiments, Figure 8 As shown, according to the target text content of the picture book, the picture book image data generation process is performed to obtain the picture book image data of each page in the picture book, which specifically includes the following contents:

[0167] Step S801, performing scene recognition processing on the text content of each page to obtain scene information corresponding to the text content of each page;

[0168] Step S802, performing image composition processing according to the scene information to obtain basic picture book image data corresponding to the scene information as basic picture book image data for each page;

[0169] Step S803, according to the picture book style information, style rendering processing is performed on the basic picture book image data of each page to obtain the picture book image data of each page; the picture book style information represents the style information matching the picture book theme information and the picture book character information; the style information of the picture book image data of each page is the same as the picture book style information.

[0170] The scene recognition process may be a process of understanding and identifying the scene described in the text content, for example, it may be a process of identifying scene elements such as the environment, time, and space in which the story takes place.

[0171] The scene information may be scene-related information obtained through scene recognition processing, such as the specific place where the story takes place, the environmental atmosphere, the time background, and other information.

[0172] Among them, image composition processing can be a process of designing and planning image layout according to scene information, for example, it can be a process of determining the position, size, proportional relationship, etc. of each element in the scene.

[0173] The basic picture book image may be a preliminary image obtained after image composition processing, for example, an image with a determined scene layout but without an artistic style. The picture book style information may be descriptive information used to define the overall artistic style of the picture book, for example, the picture book style information may be descriptive information of an artistic expression form such as an illustration style, a watercolor style, or a comic style.

[0174] Among them, the style rendering process can be a process of adding a specific artistic style to the basic picture book image, for example, it can be a process of applying artistic features such as illustration style, watercolor style or comic style to the basic picture book image.

[0175] The style information may be data describing the artistic style characteristics of an image. For example, the style information may be data describing artistic expression characteristics of an image, such as color style, brushstroke style, and texture characteristics.

[0176] Specifically, the controller 250 performs scene recognition processing on the text content of each page in the target text content of the picture book, and identifies information such as scene type, scene elements and scene layout through semantic analysis to obtain scene information corresponding to the text content of each page; the controller 250 determines the perspective, composition ratio and element position of the image according to the scene type, scene elements and scene layout information in the scene information, performs image composition processing, and generates basic picture book image data containing basic scene layout and graphic elements; the controller 250 performs style rendering processing on the basic picture book image data according to the picture book style information that matches the picture book theme information and picture book character information, and adjusts the color style, brushstroke effect and texture features to make the picture book image data of each page present an artistic effect consistent with the picture book style information.

[0177] The technical solution provided in this embodiment, through the three-step progressive processing of scene recognition processing, image composition processing and style rendering processing on each page of text content, is conducive to accurately converting text content into an image with a reasonable layout and ensuring that the image style matches the picture book theme and character characteristics; by maintaining the coherence of scene information and the consistency of style information in each processing step, it is conducive to generating picture book images with unified style and coordinated content, realizing automatic conversion from text to image, and ensuring the integrity and coordination of the picture book image in visual expression.

[0178] In some embodiments, Fig. 9 As shown, the text information is input into the text processing model to obtain the picture book theme information and picture book role information, which specifically include the following contents:

[0179] Step S901, inputting text information into a text processing model to obtain semantic information of the text information; the text processing model is used to perform semantic analysis on the input text information and output the semantic information of the text information;

[0180] Step S902, inputting the semantic information into the text processing model to obtain the intent classification result of the text information; the intent classification result is used to indicate the type of picture book corresponding to the text information; the text processing model is used to perform intent classification processing on the input semantic information and output the intent classification result of the text information;

[0181] Step S903, input the intent classification result into the text processing model to obtain the picture book theme information and picture book role information; the text processing model is used to output the picture book theme information and picture book role information that match the semantic information according to the input intent classification result.

[0182] The semantic analysis process may be a process of analyzing and understanding the text information at the semantic level. The intention classification process may be a process of determining the story type of the text information, for example, it may be a process of determining whether the story described in the text information is a fairy tale type or an inspirational type.

[0183] Among them, the intention classification result can be a story type judgment result obtained through intention classification processing, for example, it can be a result of judging that the text information corresponds to a fairy tale type or an inspirational type.

[0184] The picture book type may be a specific category to which the picture book story belongs, for example, the picture book type may be a fairy tale, inspirational, educational or other story category.

[0185] Specifically, the controller 250 inputs the text information into a text processing model that has been fine-tuned using story picture book language pre-training data and instruction fine-tuning data, identifies key concepts and core appeals in the text information through semantic parsing, and obtains the semantic information of the text information; the controller 250 inputs the semantic information of the text information into the text processing model, determines the story type through intent classification, and obtains an intent classification result representing the picture book type corresponding to the text information; the controller 250 inputs the intent classification result into the text processing model, determines the story type based on the intent classification result, and outputs picture book theme information and picture book character information that matches the semantic information.

[0186] The technical solution provided in this embodiment realizes the layer-by-layer conversion processing from text information to semantic information, intention classification results, and finally to picture book theme information and picture book role information by inputting text information into the text processing model in three steps in sequence, which is conducive to accurately understanding user needs and converting them into specific picture book creation elements; by maintaining the consistency and matching of information in each step, it is conducive to ensuring that the final generated picture book theme and role are consistent with the original needs of the user, realizing the accurate conversion from user needs to specific creation elements, and improving the pertinence and accuracy of picture book creation.

[0187] In some embodiments, Fig.10 As shown, according to the picture book theme information and the picture book role information, the picture book text content generation process is performed to obtain the initial text content of the picture book, including:

[0188] Step S1001, selecting a picture book theme that matches the picture book theme information from preset picture book themes;

[0189] Step S1002, performing character feature recognition processing on the picture book character information to obtain the character features of the picture book character information;

[0190] Step S1003, constructing the storyline according to the theme of the picture book and the characteristics of the characters, and obtaining the initial text content of the picture book.

[0191] Among them, the preset picture book theme can be a picture book story theme type preset by the system, for example, it can be an educational theme such as children need to learn to share, brush their teeth from an early age, etc.

[0192] The character feature recognition process may be a process of analyzing and identifying the characteristics of the picture book character information, for example, it may be a process of identifying the character's personality traits, appearance features, behavioral habits, etc.

[0193] The character features may be character-related feature information obtained through character feature recognition processing, such as the character's personality traits (such as shyness, bravery), appearance features (such as a cute kitten), behavioral habits, and other feature information.

[0194] Among them, the storyline construction process can be a process of creating a basic framework of the story based on the story theme and character characteristics, for example, it can be a process of designing the story's beginning, development, climax and ending.

[0195] Specifically, the controller 250 first selects a picture book theme that matches the picture book theme information from preset children's education picture book themes (such as children do not like to take medicine, children must learn to share, and they must brush their teeth from an early age, etc.); then the controller 250 performs role feature recognition processing on the picture book character information, identifies the character's personality traits, appearance features, behavioral habits and other feature information, and obtains the role features of the picture book character information; finally, the controller 250 designs the story's beginning, development, climax and ending according to the selected picture book theme and the identified character features, performs story plot construction processing, and obtains the initial text content of the picture book.

[0196] The technical solution provided in this embodiment is conducive to ensuring that the generated picture book content conforms to the expected educational theme and creation direction by selecting a picture book theme that matches the picture book theme information from the preset picture book themes; by performing role feature recognition processing on the picture book character information, it is conducive to accurately grasping the character's personality characteristics and behavioral characteristics; by performing picture book plot construction processing based on the picture book theme and character characteristics, it is conducive to generating the initial text content of the picture book that conforms to the theme and has distinct character characteristics. This processing flow from theme selection to character shaping and then to plot construction realizes the systematic generation of picture book content and ensures that the generated picture book has a clear theme orientation and rich character characteristics.

[0197] In some embodiments, Fig.11 As shown, the text information corresponding to the picture book generation requirement information is identified, including the following content:

[0198] Step S1101, when the picture book generation requirement information is voice information, preprocessing the picture book generation requirement information to obtain preprocessed picture book generation requirement information;

[0199] Step S1102, performing acoustic feature extraction processing on the preprocessed picture book generation demand information to obtain acoustic feature information of the preprocessed picture book generation demand information;

[0200] Step S1103: performing speech recognition processing on the pre-processed picture book generation requirement information according to the acoustic feature information to obtain text information corresponding to the picture book generation requirement information.

[0201] The voice information may be picture book generation demand information input by the user through voice, for example, the user may say through voice “I want a story about a brave puppy overcoming fear”.

[0202] The preprocessing may be a process of performing preliminary processing on the voice information, for example, it may be a process of performing noise reduction, frame division, etc. on the voice signal.

[0203] The pre-processed picture book generation requirement information may be voice information obtained through pre-processing, for example, a voice signal processed through noise reduction, frame division, etc.

[0204] The acoustic feature extraction process may be a process of extracting acoustic features from pre-processed speech information, for example, it may be a process of extracting acoustic features of speech such as phonemes, pitch, speaking speed, etc.

[0205] The acoustic feature information may be speech feature information obtained through acoustic feature extraction processing, for example, acoustic feature information such as phonemes, pitch, speaking speed, etc. of speech.

[0206] Among them, speech recognition processing can be a processing process of converting speech information into text information, for example, it can be a processing process of converting a speech such as "I want a story about a brave puppy overcoming fear" into corresponding text based on acoustic feature information.

[0207] Specifically, when the controller 250 receives the picture book generation requirement information as voice information, it first performs preprocessing operations such as noise reduction and frame division on the picture book generation requirement information to obtain the preprocessed picture book generation requirement information; then the controller 250 performs acoustic feature extraction processing on the preprocessed picture book generation requirement information to extract acoustic feature information such as phonemes, pitch, and speaking speed; finally, the controller 250 performs voice recognition processing on the preprocessed picture book generation requirement information based on the extracted acoustic feature information, and converts the voice information into corresponding text information, thereby completing the voice-to-text conversion process.

[0208] The technical solution provided in this embodiment is conducive to improving the quality and processability of voice signals by preprocessing the picture book generation demand information in the form of voice; by performing acoustic feature extraction processing on the preprocessed picture book generation demand information, it is conducive to obtaining key acoustic features in the voice information; by performing voice recognition processing based on the acoustic feature information, it is conducive to accurately converting the voice information into text information. This hierarchical processing method from preprocessing to feature extraction to voice recognition achieves high-quality conversion from voice input to text output, providing an accurate text basis for subsequent picture book generation.

[0209] In some embodiments, Fig.12 As shown, according to the picture book image data of each page in the picture book, the data representing the background audio and the data representing the broadcast audio, the picture book corresponding to the required information is generated by synthesizing the picture book, including:

[0210] Step S1201, according to the text content of each page, select matching background audio from the background audio library to obtain data representing the background audio of each page, and perform reading audio generation processing to obtain data representing the reading audio of each page;

[0211] Step S1202 , performing time-series synchronous synthesis processing on the picture book image of each page, the data representing the background audio, and the data representing the reading audio to obtain a picture book corresponding to the picture book generation requirement information.

[0212] The data representing the background audio may be audio data used to represent the background music of the picture book. For example, the data representing the background audio may be digitized audio data of the background music that matches the storyline.

[0213] The data representing the broadcast audio may be audio data used to represent the content of the picture book reading. For example, the data representing the broadcast audio may be digitized data of the story reading audio generated by speech synthesis technology.

[0214] The background audio library may be a data set of various types of background music stored in advance. For example, the background audio library may be an audio database containing background music of different emotional types such as cheerfulness, warmth, tension, etc.

[0215] The reading audio generation process may be a process of converting text content into speech. For example, the reading audio generation process may be a process of converting a story text into reading audio through speech synthesis technology.

[0216] The data representing the reading audio may be speech data obtained through the reading audio generation process. For example, the data representing the reading audio may be digitized data of the reading audio obtained through speech synthesis of the story text.

[0217] Among them, the time-series synchronization synthesis processing can be a processing process of combining images, background audio and reading audio in chronological order. For example, the time-series synchronization synthesis processing can be a processing process of synchronizing the picture book image of each page with the corresponding background music and reading audio in time.

[0218] Specifically, the controller 250 analyzes the emotional tone based on the text content of each page of the picture book, selects background audio that matches the emotional tone from the background audio library, and obtains data representing the background audio of each page; the controller 250 performs reading audio generation processing on the text content of each page, generates reading audio with emotional expressiveness, and obtains data representing the reading audio of each page; the controller 250 performs time-series synchronous synthesis processing on the picture book image of each page, the data representing the background audio, and the data representing the reading audio, to ensure that the image display, background music playback, and text reading are accurately matched in time, and finally a picture book corresponding to the picture book generation requirement information is obtained.

[0219] The technical solution provided in this embodiment helps to configure a suitable audio effect for each page of the content by performing background audio selection and reading audio generation processing on the text content of each page of the picture book respectively; and helps to achieve accurate matching of image display, background music playback and text reading in the time dimension by performing time-series synchronous synthesis processing on the picture book image, background audio data and reading audio data of each page, thereby improving the quality of the picture book.

[0220] In some embodiments, Fig.13 As shown in Figure 1, the text processing model is trained in the following way:

[0221] Step S1301, obtaining pre-training data and fine-tuning data as training data; the pre-training data includes the text content of a picture book including multiple picture book theme information, picture book style information and language structure information; the fine-tuning data includes multiple picture book writing instructions, and the picture book writing instructions represent instruction information for guiding picture book writing;

[0222] Step S1302, preprocessing the training data to obtain preprocessed training data; the preprocessed training data means data that has been cleaned, formatted, labeled and classified;

[0223] Step S1303, using the pre-processed training data, iteratively train the text processing model to be trained to obtain a trained text processing model;

[0224] Step S1304, performing model adjustment processing on the trained text processing model according to the model performance information of the trained text processing model to obtain a trained text processing model.

[0225] The pre-training data may be a basic data set used for model training, for example, a data set containing a large amount of story picture book text.

[0226] The fine-tuning data may be a specialized data set used for fine-tuning the model, for example, a data set containing story-writing instructions such as “write a story about friendship” or “describe an animal character using anthropomorphic techniques”.

[0227] Among them, diversified story picture book texts can be story texts covering different themes, styles and language structures, such as "Once upon a time there was a little rabbit, it lived in a beautiful forest..." or "In a distant kingdom, there was a brave knight..." and other story texts.

[0228] The picture book writing instructions may be instruction information for guiding how to write a specific type of story, such as “write a story about friendship that includes at least two characters” or “describe the daily life of an animal character using anthropomorphic language”.

[0229] Among them, data preprocessing can be a process of preliminary processing of training data, for example, it can be a process of cleaning, formatting, labeling and classifying data.

[0230] The preprocessed training data may be data obtained through data preprocessing, for example, data after spelling errors are removed, format is unified, annotation information is added, and subject classification is performed.

[0231] Among them, iterative training can be a process of performing targeted training on a large language model, for example, it can be a process of performing multiple rounds of repeated training on the model using pre-training data and fine-tuning data.

[0232] The trained text processing model may be a preliminary model obtained through iterative training, for example, a model obtained through training with pre-training data.

[0233] The model performance information may be information used to evaluate the model performance, for example, it may be evaluation information on the coherence, creativity, and topic relevance of the model in generating stories.

[0234] Among them, the model adjustment process can be a process of optimizing the model according to the model performance information, for example, it can be a process of adjusting the data set, model architecture or training strategy.

[0235] Specifically, the controller 250 first obtains pre-training data including diversified story picture book texts such as "Once upon a time there was a little rabbit, it lived in a beautiful forest..." and fine-tuning data including story writing instructions such as "Write a story about friendship, the story includes at least two characters" as training data; then the controller 250 performs data preprocessing on the training data, including cleaning up spelling errors, unifying the text format, adding annotation information such as story structure and characters, and classifying according to themes and styles; then the controller 250 uses the pre-processed training data to iteratively train the large language model (LLM) to be trained, so that the model is familiar with the language style and structure of the picture book, and obtains the trained text processing model; finally, the controller 250 performs model adjustment processing on the trained text processing model according to the model performance information such as coherence, creativity and topic relevance of the trained text processing model in the story generation task, and obtains the trained text processing model.

[0236] The technical solution provided in this embodiment helps to ensure the comprehensiveness and pertinence of the training data by using pre-training data containing a variety of picture book themes, styles and language structures, as well as fine-tuning data containing specific writing instructions; by pre-processing the training data through cleaning, formatting, labeling and classification, as well as iterative training and performance adjustment of the model, it helps to improve the quality of the training data and the training effect of the model, thereby achieving high-quality training of the text processing model on the picture book creation task and improving the actual application effect of the model.

[0237] In some embodiments, Figure 1 and 2 As shown, the present application provides a display device 200 , which may include a display 260 and a controller 250 ; wherein the controller 250 is coupled to the display 260 .

[0238] like Figure 5 As shown, the controller 250 is configured as follows:

[0239] Step S501, receiving picture book generation requirement information.

[0240] Step S502, identify the text information corresponding to the picture book generation requirement information, and input the text information into the text processing model to obtain the picture book theme information and picture book role information; the picture book theme information is the information representing the theme in the semantic information of the text information, and the picture book role information is the information representing the role in the semantic information of the text information; the text processing model is used to output the picture book theme information and picture book role information that match the semantic information based on the semantic information of the input text information.

[0241] Step S503, performing picture book text content generation processing according to the picture book theme information and the picture book character information to obtain the initial text content of the picture book.

[0242] Step S504, inputting the initial text content of the picture book into the text processing model to obtain the target text content of the picture book; the text quality score of the target text content is higher than the text quality score of the initial text content; the text processing model is also used to perform text adjustment processing on the input initial text content in multiple text quality score dimensions, and output the adjusted target text content.

[0243] Step S505, according to the target text content of the picture book, picture book image data generation processing is performed to obtain picture book image data of each page in the picture book, and according to the picture book image data of each page in the picture book, data representing background audio and data representing broadcast audio, a picture book corresponding to the required information is synthesized; the picture book image data of each page respectively represents the image data corresponding to the text content of each page in the target text content.

[0244] It should be noted that, for the specific implementation process of the display device 200, reference can be made to Figure 5 The relevant embodiments of the picture book generation method shown are not repeated here.

[0245] The technical solution provided in this embodiment automatically extracts picture book theme information and picture book character information from text information through a text processing model, which is conducive to reducing the time for manual analysis and extraction of key information; by automatically generating initial text content based on the extracted information and using a text processing model to perform text adjustment processing, it is conducive to quickly generating high-quality picture book content; by automatically performing picture book image generation processing and synthesizing picture books, it is conducive to reducing the workload of manual creation and editing; this full-process automated processing method is conducive to improving the efficiency of picture book generation.

[0246] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:

[0247] The initial text content of the picture book is input into the text processing model to obtain the text content of the picture book after grammar correction; the grammar quality score of the grammar-corrected text content is higher than the grammar quality score of the initial text content; the text processing model is also used to perform grammar correction processing on the input initial text content and output the grammar-corrected text content; the grammar-corrected text content is input into the text processing model to obtain the text content of the picture book after the language style is adjusted; the language style of the text content after the language style adjustment is the same as the language style matched by the target audience of the picture book; the text processing model is also used to perform language style adjustment processing on the input grammar-corrected text content and output the text content after the language style adjustment; The text content is input into a text processing model to obtain the text content after the expression of the picture book is adjusted; the expression of the text content after the expression is adjusted is the same as the preset expression; the text processing model is also used to adjust the input text content after the language style is adjusted to the preset expression, and output the text content after the expression is adjusted; the text content after the expression is adjusted is input into the text processing model to obtain the target text content of the picture book; the emotional color information of the target text content of the picture book is the same as the picture book emotional color information in the picture book generation demand information; the text processing model is also used to adjust the picture book emotional color information of the text content after the expression is adjusted, and output the target text content after the emotional color information is adjusted.

[0248] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:

[0249] The text information is input into the text processing model to obtain the picture book emotional color information; the picture book emotional color information is the information representing the emotional color in the semantic information of the text information; the text processing model is also used to output the picture book emotional color information matching the semantic information according to the semantic information of the input text information.

[0250] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:

[0251] Perform scene recognition processing on the text content of each page to obtain scene information corresponding to the text content of each page; perform image composition processing based on the scene information to obtain basic picture book image data corresponding to the scene information as the basic picture book image data of each page; perform style rendering processing on the basic picture book image data of each page according to the picture book style information to obtain picture book image data of each page; the picture book style information represents style information that matches the picture book theme information and the picture book character information; the style information of the picture book image data of each page is the same as the picture book style information.

[0252] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:

[0253] Input text information into a text processing model to obtain semantic information of the text information; the text processing model is used to perform semantic analysis on the input text information and output the semantic information of the text information; input the semantic information into the text processing model to obtain the intent classification result of the text information; the intent classification result is used to indicate the picture book type corresponding to the text information; the text processing model is used to perform intent classification on the input semantic information and output the intent classification result of the text information; input the intent classification result into the text processing model to obtain picture book theme information and picture book role information; the text processing model is used to output picture book theme information and picture book role information that matches the semantic information according to the input intent classification result.

[0254] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:

[0255] From the preset picture book themes, a picture book theme that matches the picture book theme information is selected; the picture book character information is processed for character feature recognition to obtain the character features of the picture book character information; according to the picture book theme and the character features, a storyline is constructed to obtain the initial text content of the picture book.

[0256] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:

[0257] In the case where the picture book generation demand information is voice information, the picture book generation demand information is preprocessed to obtain the preprocessed picture book generation demand information; acoustic feature extraction is performed on the preprocessed picture book generation demand information to obtain acoustic feature information of the preprocessed picture book generation demand information; based on the acoustic feature information, voice recognition is performed on the preprocessed picture book generation demand information to obtain text information corresponding to the picture book generation demand information.

[0258] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:

[0259] According to the text content of each page, matching background audio is selected from the background audio library to obtain data representing the background audio of each page, and reading audio generation processing is performed to obtain data representing the reading audio of each page; the picture book image of each page, the data representing the background audio and the data representing the reading audio are subjected to time-series synchronous synthesis processing to obtain a picture book corresponding to the picture book generation requirement information.

[0260] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:

[0261] Pre-training data and fine-tuning data are obtained as training data; the pre-training data includes the text content of picture books including various picture book theme information, picture book style information and language structure information; the fine-tuning data includes various picture book writing instructions, and the picture book writing instructions represent the instruction information used to guide picture book writing; the training data is pre-processed to obtain pre-processed training data; the pre-processed training data represents data that has been cleaned, formatted, annotated and classified; the pre-processed training data is used to iteratively train the text processing model to be trained to obtain the trained text processing model; according to the model performance information of the trained text processing model, the trained text processing model is adjusted to obtain a trained text processing model.

[0262] In some embodiments, Fig.14 As shown, in order to more clearly describe the signaling interaction process between the modules of the display device 200, the present application also provides another picture book generation method, which may include the following steps:

[0263] Step 1: When the display 260 displays the user interface of the picture book application, the voice recognition module receives the personalized needs of the user through voice / text input.

[0264] Step 2: The speech recognition module converts the personalized needs of the user through voice / text input into text demand information, and inputs the text demand information into the large language model preprocessing module.

[0265] Step 3: The large language model preprocessing module performs intent understanding processing on the text demand information, such as semantic analysis, intent classification, and key information extraction in sequence, to obtain processed demand information, and input the processed demand information into the story generation module.

[0266] Step 4: The story generation module generates story themes and characters based on the processed demand information, such as by theme matching, character shaping, and plot conception, to generate a preliminary story text, and input the preliminary story text into the story optimization module.

[0267] Step 5: The story optimization module performs text optimization processing on the preliminary story text, such as grammar correction, style matching, language polishing, and emotional tone optimization in sequence to obtain an optimized story text, and input the optimized story text into the image generation module.

[0268] Step 6: The image generation module performs image generation processing based on the optimized story text, such as performing scene understanding, image composition, and style rendering in sequence to generate picture book images, and input the picture book images into the audio generation module.

[0269] Step 7, the audio generation module performs background music selection and voice reading synthesis to obtain corresponding background sound and reading sound, and inputs the obtained picture book picture, background sound, and reading sound into the story composition module.

[0270] Step 8: The story composition module synthesizes the picture book image, background sound, and reading sound to obtain a complete personalized picture book, and feeds the complete personalized picture book back to the user.

[0271] It should be noted that the large language model can be deployed on a display device or on a cloud server / server, and this application does not make any specific limitations.

[0272] The technical solution provided in this embodiment automatically extracts picture book theme information and picture book character information from text information through a text processing model, which is conducive to reducing the time for manual analysis and extraction of key information; by automatically generating initial text content based on the extracted information and using a text processing model to perform text adjustment processing, it is conducive to quickly generating high-quality picture book content; by automatically performing picture book image generation processing and synthesizing picture books, it is conducive to reducing the workload of manual creation and editing; this full-process automated processing method is conducive to improving the efficiency of picture book generation.

[0273] In some embodiments, in order to more clearly illustrate the picture book generation method provided in the embodiments of the present application, the picture book generation method is specifically described below with a specific embodiment. The present application also provides another picture book generation method. Based on the instruction-following ability of the large language model, a certain amount of story picture book language pre-training data and instruction fine-tuning data are used to fine-tune the large model, so that it has strong ability in picture book content writing. At the same time, according to the needs of large-screen story creation in educational scenarios, a proprietary prompt word configuration project is designed to realize user interactive personalized content creation. It mainly includes the following aspects: (1) Story theme configuration, taking into full consideration that the content of family large-screen picture books is mainly based on children's needs, focusing on creating theme customization (children do not like to take medicine, children must learn to share, and children must brush their teeth from an early age, etc.), users can select by voice, and the background large model generates story content that conforms to the theme according to the instructions; (2) Story content configuration, the story content can be personalized according to the needs of users, such as the adventure journey of the little white rabbit, and the user can define the story content he likes on the premise of embodying the theme of bravery; (3) Character personalization configuration, users can create their own personalized story characters according to language descriptions, and maintain the consistency and continuity of the user-defined characters throughout the story; (4) Story style configuration, based on the following ability of the generation model, the prompt word project is used to configure the picture book style such as illustration, watercolor, and comics.

[0274] Fine-tune the large language model to make it more capable of writing picture book content. The main steps are as follows:

[0275] Step 1: Prepare data

[0276] 1. Collect pre-training data:

[0277] 1) Collect a large number of story picture book texts, which should cover a variety of themes, styles and language structures.

[0278] 2) Ensure data diversity, including stories from target readers of different age groups and different cultural backgrounds.

[0279] 2. Collect instruction fine-tuning data:

[0280] 1) Create or collect a set of instructions, which could be about how to write a certain type of story, how to use a certain language style, etc.

[0281] 2) Instructions should be related to picture book writing, such as “write a short story about friendship” or “describe an animal character using anthropomorphism.”

[0282] Step 2: Data preprocessing

[0283] 1. Clean and format data:

[0284] 1) Clean the data to remove noise, such as spelling errors, inconsistent formatting, etc.

[0285] 2) Format the data into an input format acceptable to the model, such as dividing the text into paragraphs or sentences.

[0286] 2. Labeling and classification:

[0287] 1) Label the data to identify the story structure, characters, plot, and other information.

[0288] 2) Categorize the data to facilitate the subsequent training process, such as classification by theme or style.

[0289] Step 3: Model fine-tuning

[0290] 1. Choose a model architecture: Choose a suitable pre-trained large language model as the basis.

[0291] 2. Fine-tune the model:

[0292] 1) Use the collected pre-training data on story picture book language to perform preliminary fine-tuning on the model to make it familiar with the language style and structure of the picture books.

[0293] 2) Further fine-tune the model using instruction-based fine-tuning data so that the model can generate story content that meets the requirements based on the instructions.

[0294] 3. Adjust hyperparameters: Adjust hyperparameters such as learning rate and batch size according to the performance of the model to optimize the training effect.

[0295] Step 4: Model evaluation and optimization

[0296] 1. Evaluate model performance:

[0297] 1) Use a set of validation datasets to evaluate the performance of the model on the story generation task.

[0298] 2) Evaluation metrics can include the coherence, creativity, and topic relevance of the generated stories.

[0299] 2. Model optimization:

[0300] 1) Optimize the model based on the evaluation results, which may require adjusting the dataset, model architecture, or training strategy.

[0301] refer to Fig.14 , based on the fine-tuned large language model, creating a picture book mainly consists of the following steps:

[0302] 1. Voice / text input

[0303] Input: Users make personalized story requests through voice or text, such as "I want a story about a brave puppy overcoming his fear."

[0304] 2. Speech recognition and demand preprocessing

[0305] 1) Voice Recognition

[0306] Input: user voice;

[0307] Processing: 1. Speech signal preprocessing; 2. Acoustic feature extraction; 3. Speech to text;

[0308] Output: text requirement information.

[0309] 2) Demand preprocessing (large language model preprocessing)

[0310] Input: Text requirements

[0311] Processing steps: 1. Semantic analysis: identify key concepts in the requirements; 2. Intent classification: determine the type of story (fairy tale, inspirational, etc.); 3. Key information extraction: extract roles, themes, and emotional tendencies;

[0312] Output: Structured requirement information.

[0313] 3. Story Generation

[0314] Input: Structured requirement information

[0315] Processing steps: 1. Theme matching: select the appropriate story theme according to the needs; 2. Character creation: design character characteristics that meet the needs; 3. Plot conception: create the basic framework of the story;

[0316] Output: Preliminary story text.

[0317] 4. Text Optimization

[0318] Text optimization is a key link in the entire process and is mainly carried out in the story optimization stage.

[0319] Fine-tune the text optimization steps of the large language model:

[0320] 1) Grammar correction: detect and correct grammatical errors; adjust sentence structure to improve readability; 2) Style adaptation: adjust the language style according to the target audience; for children's picture books, use more vivid and simple language; 3) Language polishing: add rhetorical techniques (metaphor, personification, etc.); optimize word selection to improve expressiveness; 4) Emotional tonality optimization: adjust the emotional color of the text; ensure the emotional coherence of the story; highlight specific emotional dimensions according to the theme.

[0321] 5. Image Generation

[0322] Input: optimized story text;

[0323] Processing steps: 1) scene understanding; 2) image composition; 3) style rendering;

[0324] Output: picture book image.

[0325] 6. Audio Generation

[0326] Input: Picture book image

[0327] Processing steps: 1) background music selection; 2) voice reading synthesis;

[0328] Output: background sound, reading sound.

[0329] 7. Story Synthesis

[0330] Input: picture book pictures, background sound, reading sound

[0331] Output: Complete personalized picture book.

[0332] The above embodiment automatically generates a complete personalized picture book story through a fine-tuned large language model, which is conducive to improving the efficiency of picture book generation.

[0333] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0334] Based on the same inventive concept, the embodiment of the present application also provides a picture book generation device for implementing the picture book generation method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more picture book generation device embodiments provided below can refer to the limitations of the picture book generation method above, and will not be repeated here.

[0335] In some embodiments, Fig.15 As shown, a picture book generating device 1500 is provided, which can be applied to Figure 1 The display device 200 shown in FIG. Figure 2 As shown, the display device 200 may include a display 260 and a controller 250; the controller 250 is coupled to the display 260; the device may include:

[0336] The information receiving module 1501 is used to receive picture book generation requirement information.

[0337] The information identification module 1502 is used to identify the text information corresponding to the picture book generation requirement information, and input the text information into the text processing model to obtain the picture book theme information and the picture book role information; the picture book theme information is the information representing the theme in the semantic information of the text information, and the picture book role information is the information representing the role in the semantic information of the text information; the text processing model is used to output the picture book theme information and the picture book role information that match the semantic information according to the semantic information of the input text information.

[0338] The text generation module 1503 is used to generate the text content of the picture book according to the picture book theme information and the picture book role information to obtain the initial text content of the picture book.

[0339] The text adjustment module 1504 is used to input the initial text content of the picture book into the text processing model to obtain the target text content of the picture book; the text quality score of the target text content is higher than the text quality score of the initial text content; the text processing model is also used to perform text adjustment processing on the input initial text content in multiple text quality score dimensions, and output the adjusted target text content.

[0340] The picture book synthesis module 1505 is used to generate picture book image data according to the target text content of the picture book, obtain the picture book image data of each page in the picture book, and synthesize the picture book to generate a picture book corresponding to the required information according to the picture book image data of each page in the picture book, the data representing the background audio and the data representing the broadcast audio; the picture book image data of each page respectively represents the image data corresponding to the text content of each page in the target text content.

[0341] Each module in the picture book generation device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0342] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0343] In some embodiments, a computer program product is provided, including a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.

[0344] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0345] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0346] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0347] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A picture book generation method, characterized in that: Applied to display devices, including: Receive picture book generation demand information; Identify text information corresponding to the picture book generation requirement information, and input the text information into a text processing model to obtain picture book theme information and picture book role information; the picture book theme information is information representing the theme in the semantic information of the text information, and the picture book role information is information representing the role in the semantic information of the text information; the text processing model is used to output the picture book theme information and the picture book role information matching the semantic information according to the input semantic information of the text information; According to the picture book theme information and the picture book character information, a picture book text content generation process is performed to obtain the initial text content of the picture book; Inputting the initial text content of the picture book into the text processing model to obtain the target text content of the picture book; the text quality score of the target text content is higher than the text quality score of the initial text content; the text processing model is further used to perform text adjustment processing on the input initial text content in multiple text quality score dimensions, and output the adjusted target text content; According to the target text content of the picture book, picture book image data generation processing is performed to obtain picture book image data of each page in the picture book, and according to the picture book image data of each page in the picture book, data representing background audio and data representing broadcast audio, a picture book corresponding to the picture book generation demand information is synthesized; the picture book image data of each page respectively represent the image data corresponding to the text content of each page in the target text content.

2. The method according to claim 1, characterized in that The step of inputting the initial text content of the picture book into the text processing model to obtain the target text content of the picture book includes: Inputting the initial text content of the picture book into the text processing model to obtain the grammatically corrected text content of the picture book; the grammatical quality score of the grammatically corrected text content is higher than the grammatical quality score of the initial text content; the text processing model is also used to perform grammatical correction processing on the input initial text content and output the grammatically corrected text content; Inputting the grammatically corrected text content into the text processing model to obtain the text content of the picture book after the language style is adjusted; the language style of the text content after the language style is adjusted is the same as the language style matched by the target audience of the picture book; the text processing model is also used to perform language style adjustment processing on the input grammatically corrected text content, and output the text content after the language style is adjusted; Inputting the text content after the language style is adjusted into the text processing model to obtain the text content after the expression of the picture book is adjusted; the expression of the text content after the expression is adjusted is the same as the preset expression; the text processing model is also used to adjust the input text content after the language style is adjusted according to the preset expression, and output the text content after the expression is adjusted; The text content after the expression mode is adjusted is input into the text processing model to obtain the target text content of the picture book; the emotional color information of the target text content of the picture book is the same as the picture book emotional color information in the picture book generation requirement information; the text processing model is also used to adjust the emotional color information of the picture book on the text content after the expression mode is adjusted, and output the target text content after the emotional color information is adjusted.

3. The method according to claim 2, characterized in that The method further comprises: The text information is input into the text processing model to obtain the picture book emotional color information; the picture book emotional color information is the information representing the emotional color in the semantic information of the text information; the text processing model is also used to output the picture book emotional color information matching the semantic information based on the input semantic information of the text information.

4. The method according to claim 1, characterized in that: The picture book image data generation process is performed according to the target text content of the picture book to obtain the picture book image data of each page in the picture book, including: Performing scene recognition processing on the text content of each page to obtain scene information corresponding to the text content of each page; Performing image composition processing according to the scene information to obtain basic picture book image data corresponding to the scene information as the basic picture book image data of each page; According to the picture book style information, style rendering processing is performed on the basic picture book image data of each page to obtain the picture book image data of each page; the picture book style information represents style information that matches the picture book theme information and the picture book character information; the style information of the picture book image data of each page is the same as the picture book style information.

5. The method according to claim 1, characterized in that: The step of inputting the text information into a text processing model to obtain picture book theme information and picture book character information includes: Inputting the text information into the text processing model to obtain semantic information of the text information; the text processing model is used to perform semantic analysis on the input text information and output the semantic information of the text information; The semantic information is input into the text processing model to obtain an intent classification result of the text information; the intent classification result is used to indicate the type of picture book corresponding to the text information; the text processing model is used to perform intent classification processing on the input semantic information and output the intent classification result of the text information; The intention classification result is input into the text processing model to obtain the picture book theme information and the picture book character information; the text processing model is used to output the picture book theme information and the picture book character information that match the semantic information according to the input intention classification result.

6. The method according to claim 1, characterized in that The step of generating picture book text content according to the picture book theme information and the picture book character information to obtain the initial text content of the picture book includes: Selecting a picture book theme that matches the picture book theme information from the preset picture book themes; Performing role feature recognition processing on the picture book role information to obtain role features of the picture book role information; According to the theme of the picture book and the characteristics of the characters, a storyline construction process is performed to obtain the initial text content of the picture book.

7. The method according to claim 1, characterized in that The identifying of text information corresponding to the picture book generation requirement information includes: In the case where the picture book generation demand information is voice information, preprocessing the picture book generation demand information to obtain preprocessed picture book generation demand information; Performing acoustic feature extraction processing on the preprocessed picture book generation demand information to obtain acoustic feature information of the preprocessed picture book generation demand information; According to the acoustic feature information, speech recognition processing is performed on the preprocessed picture book generation requirement information to obtain text information corresponding to the picture book generation requirement information.

8. The method according to claim 1, characterized in that: The step of synthesizing the picture book corresponding to the picture book generation demand information according to the picture book image data of each page of the picture book, the data representing the background audio, and the data representing the broadcast audio includes: According to the text content of each page, matching background audio is selected from the background audio library to obtain data representing the background audio of each page, and a reading audio generation process is performed to obtain data representing the reading audio of each page; The picture book image of each page, the data representing the background audio, and the data representing the reading audio are subjected to time-series synchronous synthesis processing to obtain a picture book corresponding to the picture book generation requirement information.

9. The method according to any one of claims 1 to 8, characterized in that The text processing model is trained in the following way: Acquire pre-training data and fine-tuning data as training data; the pre-training data includes text content of picture books including multiple picture book theme information, picture book style information and language structure information; the fine-tuning data includes multiple picture book writing instructions, and the picture book writing instructions represent instruction information for guiding picture book writing; Performing data preprocessing on the training data to obtain preprocessed training data; The preprocessed training data represents data that has been cleaned, formatted, labeled, and classified; Using the preprocessed training data, iteratively train the text processing model to be trained to obtain a trained text processing model; According to the model performance information of the trained text processing model, the trained text processing model is subjected to model adjustment processing to obtain the trained text processing model.

10. A display device, characterized in that: include: monitor; A controller is coupled to the display and is configured to: Receive picture book generation demand information; Identify text information corresponding to the picture book generation requirement information, and input the text information into a text processing model to obtain picture book theme information and picture book role information; the picture book theme information is information representing the theme in the semantic information of the text information, and the picture book role information is information representing the role in the semantic information of the text information; The text processing model is used to output the picture book theme information and the picture book character information matching the semantic information according to the semantic information of the input text information; According to the picture book theme information and the picture book character information, a picture book text content generation process is performed to obtain the initial text content of the picture book; Inputting the initial text content of the picture book into the text processing model to obtain the target text content of the picture book; the text quality score of the target text content is higher than the text quality score of the initial text content; the text processing model is further used to perform text adjustment processing on the input initial text content in multiple text quality score dimensions, and output the adjusted target text content; According to the target text content of the picture book, picture book image data generation processing is performed to obtain picture book image data of each page in the picture book, and according to the picture book image data of each page in the picture book, data representing background audio and data representing broadcast audio, a picture book corresponding to the picture book generation demand information is synthesized; the picture book image data of each page respectively represent the image data corresponding to the text content of each page in the target text content.