Picture book generation method and display equipment
Patent Information
- Application Number
- CN202411982385.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
Smart Images

Figure CN119991870A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of display devices, and in particular to a picture book generation method and a display device. Background Art
[0002] Display devices such as smart TVs refer to devices that realize two-way human-computer interaction functions based on Internet application technology. They have multiple functions, such as playing picture books through display devices.
[0003] In traditional technology, when generating a picture book, users are required to manually select characters, themes, styles, etc. from a given variety of picture book elements such as characters, themes, styles, etc., or manually input characters, themes, styles, etc. The whole process is rather cumbersome, resulting in low efficiency in picture book generation.
[0004] Therefore, there is a technical problem in the traditional technology that the generation efficiency of picture books is low. Summary of the invention
[0005] The present application provides a picture book generation method and a display device method to solve the technical problem of low picture book generation efficiency.
[0006] In a first aspect, some embodiments provide a picture book generation method. The method comprises:
[0007] Receiving picture book generation requirement information, and identifying text information corresponding to the picture book generation requirement information;
[0008] Input the text information into a text processing model to obtain picture book role information and picture book theme information; the picture book role information is information representing the role in the text information, and the picture book theme information is information representing the theme in the text information; the text processing model is used to output the picture book role information and the picture book theme information matching the text information according to the input text information;
[0009] Filter out target picture book style information that matches the picture book character information and the picture book theme information from the preset picture book style information; filter out target broadcast timbre information that matches the picture book character information and the picture book theme information from the preset broadcast timbre information; filter out target background audio information that matches the picture book theme information and the target picture book style information from the preset background audio information;
[0010] Based on the picture book character information, the picture book theme information, the target picture book style information, the target broadcast timbre information and the target background audio information, a picture book generation process is performed to obtain a picture book corresponding to the picture book generation requirement information.
[0011] Technical effect: Receive the input picture book generation demand information and identify its corresponding text information; then use the text processing model to process the picture book role information and picture book theme information that match the text information; intelligently filter out the target picture book style information, target broadcast timbre information that match the picture book role information and picture book theme information, and target background music information that match the picture book theme information and target picture book style information. The precise matching of these picture book elements ensures that the generated picture book not only meets the needs of users, but also improves the fun and appeal of the picture book; finally, the picture book elements such as picture book role information, picture book theme information, target picture book style information, target broadcast timbre information and target background music information are integrated, and through picture book generation processing, a picture book that matches the picture book generation demand information is created for users to help users realize their creative ideas. This process significantly improves the user's creation convenience and the quality of picture books, making the reading experience of children and families richer and more enjoyable, and also greatly improves the efficiency and quality of picture book generation processing of display devices.
[0012] In some embodiments of the present application, the text information is input into a text processing model to obtain picture book character information and picture book theme information, including:
[0013] Input the text information into a text processing model to obtain picture book keywords; the picture book keywords represent keywords in the text information; the text processing model is used to output the picture book keywords in the text information according to the input text information;
[0014] The picture book keywords are input into the text processing model to obtain the picture book theme information; the picture book theme information represents theme information that matches the picture book keywords; the text processing model is used to output the picture book theme information that matches the picture book keywords based on the input picture book keywords.
[0015] Technical Effect: The text processing model identifies picture book keywords from text information, laying a solid foundation for the subsequent generation and construction of picture book theme information. These picture book keywords not only reflect the core ideas, but also help focus on the most important picture book themes, improving the efficiency and accuracy of picture book creation. Picture book theme generation and processing based on picture book keywords makes the generated picture book theme information more targeted and in-depth, and also ensures that it has a clear direction and connotation, further improving the quality and accuracy of picture book creation.
[0016] In some embodiments of the present application, the text information is input into a text processing model to obtain picture book character information and picture book theme information, including:
[0017] Inputting the text information into the text processing model to obtain character subject information; the character subject information represents subject information matching the text information; the text processing model is used to perform character subject information extraction processing or character subject information generation processing based on the input text information, and output the character subject information matching the text information;
[0018] The role subject information is input into the text processing model to obtain role description information; the role description information represents description information corresponding to the role subject information; the text processing model is used to perform role description information generation processing according to the input role subject information, and output the role description information corresponding to the role subject information;
[0019] The character main body information and the character description information are determined as the picture book character information.
[0020] Technical effect: By using a text processing model to extract or generate character subjects from text information, it is possible to effectively extract or create character subject information, accurately identify and generate character subject information corresponding to the theme of the picture book, and ensure that the character subject information not only conforms to the plot development, but also has a profound connection with the theme. Subsequently, by generating a character description for the character subject information, the appearance of the character subject is reasonably and comprehensively described, greatly enriching the level and depth. Finally, the character subject information and character description information are combined to determine the picture book character information, giving the picture book a vivid character image and effectively improving the generation quality of the picture book.
[0021] In some embodiments of the present application, based on the picture book character information, the picture book theme information, the target picture book style information, the target broadcast timbre information and the target background audio information, picture book generation processing is performed to obtain a picture book corresponding to the picture book generation requirement information, including:
[0022] The picture book character information and the picture book theme information are input into the text processing model to obtain the picture book text content; the picture book text content represents the text content matching the picture book character information and the picture book theme information; the picture book text content includes multiple pages of text content, and the number of words in each page of text content does not exceed a preset word count threshold; the text processing model is used to output the picture book text content matching the picture book character information and the picture book theme information according to the input picture book character information and the picture book theme information;
[0023] Input the picture book text content into the text processing model to obtain adjusted picture book text content; the adjusted picture book text content is used to represent text content adapted to the picture book image data generation process; the adjusted picture book text content includes the adjusted text content of each page of text content; the text processing model is used to adjust the input picture book text content and output the adjusted picture book text content;
[0024] Based on the picture book image information, the broadcast audio information and the target background audio information corresponding to the adjusted text content of each page of text content, a picture book corresponding to the picture book generation requirement information is synthesized.
[0025] Technical effect: Through the text processing model, multiple pages of text content are generated based on the picture book character information and picture book theme information, and the number of words in each page of text content is controlled within the preset threshold, which not only ensures the simplicity and readability of the text content, but also ensures that the display can display all the text content. Further adjustment of the picture book text content through the text processing model can enable the image processing model to more fully understand the picture content of the desired picture book image information, thereby improving the generation effect of the picture book image information. Finally, combined with the picture book image information, broadcast audio information and target background audio information corresponding to the adjusted text content of each page, a complete picture book that meets the picture book generation requirements is efficiently synthesized, improving the generation efficiency and quality of the picture book.
[0026] In some embodiments of the present application, the picture book text content is input into the text processing model to obtain the adjusted picture book text content, including:
[0027] The picture book text content is input into the text processing model to obtain a description text of an entity in each page of the text content of the picture book text content; the entity is used to represent the character subject information and / or object information in each page of the text content; the description text is used to represent at least one of the action information, position relationship information, and background information of the entity; the text processing model is used to perform text enhancement processing on the entity in each page of the text content of the input picture book text content, and output the description text of the entity in each page of the text content;
[0028] The description text of the entity in each page of text content is input into the text processing model to obtain the adjusted text content of each page of text content; the adjusted text content of each page of text content represents the text content composed of the description text of the entity in each page of text content; the text processing model is used to perform text rewriting processing on the text content of each page based on the input description text of the entity in each page of text content to obtain the adjusted text content of each page of text content.
[0029] Technical effect: Through the text processing model, text enhancement processing is performed on the entities in each page of text content, which can simply, intuitively and effectively describe the entity's action information, location information and / or background information, and can also enrich the entity's description content. Based on the entity's description text, text rewriting is performed on each page of text content, so that the adjusted text content after processing can convey the picture content more intuitively and concisely, which facilitates the image processing model to accurately understand the picture that the adjusted text content wants to express, thereby improving the quality and accuracy of the picture book image information generated by the image processing model.
[0030] In some embodiments of the present application, after inputting the picture book text content into the text processing model to obtain the adjusted picture book text content, the method further includes:
[0031] Inputting the adjusted text content of each page of text content into an image processing model to obtain initial picture book image information corresponding to the adjusted text content of each page of text content; the image processing model is used to output the initial picture book image information corresponding to the adjusted text content of each page of text content according to the input adjusted text content of each page of text content;
[0032] The target picture book style information and the initial picture book image information of the adjusted text content of each page of text content are input into the image processing model to obtain the picture book image information corresponding to the adjusted text content of each page of text content; the style information of the picture book image information corresponding to the adjusted text content of each page of text content is the same as the target picture book style information; the image content of the picture book image information corresponding to the adjusted text content of each page of text content is the same as the image content of the initial picture book image information corresponding to the adjusted text content of each page of text content; the image processing model is used to render the target picture book style information on the input initial picture book image information of the adjusted text content of each page of text content, and output the picture book image information corresponding to the adjusted text content of each page of text content.
[0033] Technical effect: Through the image processing model, image generation processing is performed based on the adjusted text content of each page, which can vividly convert the adjusted text content into initial picture book image information; then the image processing model combines the target picture book style information to perform style rendering processing on the initial picture book image information, further enhancing the artistry and personalization of the picture book, and effectively improving the image quality and picture effect of the rendered picture book image information.
[0034] In some embodiments of the present application, after inputting the picture book text content into the text processing model to obtain the adjusted picture book text content, the method further includes:
[0035] The target broadcast timbre information and the adjusted text content of each page of text content are input into a broadcast audio information synthesis model to obtain the broadcast audio information corresponding to the adjusted text content of each page of text content; the timbre information of the broadcast audio information corresponding to the adjusted text content of each page of text content is the same as the target broadcast timbre information; the broadcast audio information synthesis model is used to perform playback audio information synthesis processing on the input adjusted text content of each page of text content based on the input target broadcast timbre information to obtain the broadcast audio information corresponding to the adjusted text content of each page of text content.
[0036] Technical effect: Through the broadcast audio information synthesis model, based on the target broadcast timbre information, the adjusted text content of each page is processed by speech synthesis, which can effectively improve the expressiveness and attractiveness of the generated picture book. It can also convey the emotion and atmosphere of the picture book through the broadcast audio information, thereby improving the effect and viewing experience of the picture book.
[0037] In some embodiments of the present application, after inputting the picture book text content into the text processing model to obtain the adjusted picture book text content, the method further includes:
[0038] Input the picture book theme information and the target picture book style information into a background audio information generation model to obtain a background audio tag that matches the picture book theme information and the target picture book style information; the background audio information generation model is used to output a background audio tag that matches the picture book theme information and the target picture book style information according to the input picture book theme information and the target picture book style information;
[0039] The background audio tag is input into the background audio information generation model to obtain the target background audio information corresponding to the adjusted text content of each page of text content; the background audio information generation model is used to filter out the target background audio information matching the background audio tag from the preset background audio information according to the input background audio tag, or to generate the target background audio information corresponding to the background audio tag.
[0040] Technical effect: By broadcasting the audio information synthesis model, background audio tags that match the picture book theme information and the target picture book style information are identified, and then the target background audio information corresponding to the background audio tags is searched or generated. This can provide a richer and more fitting sound experience for the creation of picture books, ensuring that the selected target background audio information complements the content and emotion of the created picture book, greatly improving the generation effect of the picture book.
[0041] In some embodiments of the present application, before inputting the text information into a text processing model to obtain the picture book character information and the picture book theme information, the method further includes:
[0042] Upon receiving the picture book generation requirement information, if the picture book generation requirement information is voice information, performing voice recognition processing on the picture book generation requirement information to obtain text information corresponding to the picture book generation requirement information;
[0043] Input the text information into the text processing model to obtain a compliance detection result of the text information; the compliance detection result is used to indicate whether the text information is compliant; the text processing model is used to perform compliance detection on the input text information and output the compliance detection result of the text information;
[0044] The step of inputting the text information into a text processing model to obtain picture book character information and picture book theme information includes:
[0045] When the compliance detection result indicates that the text information is compliant, the text information is input into a text processing model to obtain picture book role information and picture book theme information.
[0046] Technical effect: Use the text processing model to perform compliance testing on text information to ensure that the text information corresponding to the picture book generation requirement information complies with relevant regulations and standards, which is conducive to maintaining the safety and health of the picture book creation environment.
[0047] In a second aspect, some embodiments further provide a display device, the display device comprising a display and a controller; the controller is coupled to the display and configured as follows:
[0048] Receiving picture book generation requirement information, and identifying text information corresponding to the picture book generation requirement information;
[0049] Input the text information into a text processing model to obtain picture book role information and picture book theme information; the picture book role information is information representing the role in the text information, and the picture book theme information is information representing the theme in the text information; the text processing model is used to output the picture book role information and the picture book theme information matching the text information according to the input text information;
[0050] Filter out target picture book style information that matches the picture book character information and the picture book theme information from the preset picture book style information; filter out target broadcast timbre information that matches the picture book character information and the picture book theme information from the preset broadcast timbre information; filter out target background audio information that matches the picture book theme information and the target picture book style information from the preset background audio information;
[0051] Based on the picture book character information, the picture book theme information, the target picture book style information, the target broadcast timbre information and the target background audio information, a picture book generation process is performed to obtain a picture book corresponding to the picture book generation requirement information.
[0052] Technical effect: Receive the input picture book generation demand information and identify its corresponding text information; then use the text processing model to process the picture book role information and picture book theme information that match the text information; intelligently filter out the target picture book style information, target broadcast timbre information that match the picture book role information and picture book theme information, and target background music information that match the picture book theme information and target picture book style information. The precise matching of these picture book elements ensures that the generated picture book not only meets the needs of users, but also improves the fun and appeal of the picture book; finally, the picture book elements such as picture book role information, picture book theme information, target picture book style information, target broadcast timbre information and target background music information are integrated, and through picture book generation processing, a picture book that matches the picture book generation demand information is created for users to help users realize their creative ideas. This process significantly improves the user's creation convenience and the quality of picture books, making the reading experience of children and families richer and more enjoyable, and also greatly improves the efficiency and quality of picture book generation processing of display devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0054] Figure 1 A schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application;
[0055] Figure 2 A schematic diagram of the hardware configuration of a display device provided in some embodiments of the present application;
[0056] Figure 3 A schematic diagram of the hardware configuration of a control device provided in some embodiments of the present application;
[0057] Figure 4 A schematic diagram of software configuration of a display device provided in some embodiments of the present application;
[0058] Figure 5 A schematic diagram of a process for generating a picture book provided in some embodiments of the present application;
[0059] Figure 6A schematic diagram of an interface for generating a picture book provided in some embodiments of the present application;
[0060] Figure 7 A flowchart of steps for obtaining subject information provided in some embodiments of the present application;
[0061] Figure 8 A flowchart of steps for determining role information provided in some embodiments of the present application;
[0062] Fig. 9 A flowchart of steps for synthesizing a picture book corresponding to picture book generation requirement information provided in some embodiments of the present application;
[0063] Fig.10 A schematic diagram of an image corresponding to an adjusted text provided in some embodiments of the present application;
[0064] Fig.11 A flowchart of steps for obtaining adjusted text for each page provided in some embodiments of the present application;
[0065] Fig.12 A signaling interaction diagram of a picture book generation method provided in some embodiments of the present application;
[0066] Fig.13 A schematic diagram of a process for generating a picture book provided in some embodiments of the present application;
[0067] Fig.14 A signaling interaction diagram between a task generation module, a task management module, and a task execution module provided in some embodiments of the present application;
[0068] Fig.15 A structural block diagram of a picture book generating device provided in some embodiments of the present application. DETAILED DESCRIPTION
[0069] The following embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following embodiments do not represent all implementations consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application as detailed in the claims.
[0070] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.
[0071] The terms "first", "second", "third", etc. in the specification and claims of this application and the above drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances.
[0072] The terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.
[0073] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0074] In the embodiment of the present application, the display device 200 generally refers to a device with image display and data processing capabilities. For example, the display device 200 includes but is not limited to a smart TV, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.
[0075] Figure 1 This is a schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application. Figure 1 As shown in FIG. 1 , the user can operate the display device 200 through touch operation, the mobile terminal 300 and the control device 100. For example, the control device 100 can be a remote controller, a stylus pen, a handle, etc.
[0076] The mobile terminal 300 can be used as a control device for performing human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device for establishing a communication connection with the display device 200 and performing data interaction. In some embodiments, the mobile terminal 300 can install software applications with the display device 200, and achieve connection and communication through a network communication protocol to achieve the purpose of one-to-one control operation and data communication. The audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 to achieve a synchronous display function.
[0077] like Figure 1 As also shown in FIG. 4 , the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 may be allowed to communicate and connect through a local area network (LAN), a wireless local area network (WLAN), and other networks.
[0078] The display device 200 may provide a broadcast receiving television function, and may also additionally provide an intelligent network television function with a computer support function, including but not limited to network television, smart television, Internet Protocol television (IPTV), and the like.
[0079] Figure 2 Some embodiments of the present application provide Figure 1 2 is a block diagram of the hardware configuration of the display device 200.
[0080] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.
[0081] In some embodiments, the detector 230 is used to collect signals of the external environment or external interaction. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of ambient light; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environment scenes, user attributes or user interaction gestures; or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.
[0082] In some embodiments, the display 260 includes a display function component for presenting a picture, and a driving component for driving an image display. The display 260 is used to receive an image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, and components of a menu control interface and a user control UI interface.
[0083] In some embodiments, the communication device 220 is a component for communicating with an external device or server 400 according to various communication protocol types. The display device 200 may be provided with a plurality of communication devices 220 according to different supported communication modes. For example, when the display device 200 supports wireless network communication, the display device 200 may be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.
[0084] The communication device 220 can enable the display device 200 to communicate with the external device or server 400 by wireless or wired connection. Among them, the wired connection can connect the display device 200 with the external device through components such as data cables and interfaces. The wireless connection can connect the display device 200 with the external device through wireless signals or wireless networks. The display device 200 can establish a connection relationship with the external device directly, or indirectly establish a connection relationship through a gateway, a router, a connection device, etc.
[0085] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first interface to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to the user's operation through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.
[0086] In some embodiments, the controller 250 and the tuner-demodulator 210 may be located in different separate devices, that is, the tuner-demodulator 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.
[0087] In some embodiments, the user may input a user command through a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).
[0088] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or may be an external audio output device of the display device 200. In particular, for the external audio output device of the display device 200, the display device 200 may also be provided with an external audio output terminal, and the audio output device may be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.
[0089] In some embodiments, the user input interface 280 may be used to receive instructions from a user.
[0090] Figure 3 Some embodiments of the present application provide Figure 1 The hardware configuration diagram of the control device in the figure. Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.
[0091] The control device 100 is configured to control the display device 200 , and can receive user input operation instructions, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, playing the role of an interactive intermediary between the user and the display device 200 .
[0092] In some embodiments, the control device 100 may be a smart device, for example, the control device 100 may be installed with various applications for controlling the display device 200 according to user needs.
[0093] In some embodiments, Figure 1As shown, the mobile terminal 300 or other intelligent electronic devices can play a similar function as the control device 100 after installing the application for controlling the display device 200 .
[0094] The controller 110 includes a processor 112, a RAM 113, a ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between the internal components and the external and internal data processing functions.
[0095] The communication interface 130 implements communication of control signals and data signals with the display device 200 under the control of the controller 110. The communication interface 130 may include at least one of other near field communication modules such as a WiFi chip 131, a Bluetooth module 132, and an NFC module 133.
[0096] The user input / output interface 140 , wherein the input interface includes at least one of other input interfaces such as a microphone 141 , a touch panel 142 , a sensor 143 , and a button 144 .
[0097] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with a communication interface 130, such as a WiFi, Bluetooth, NFC or other module, and can encode the user input command through the WiFi protocol, Bluetooth protocol, or NFC protocol and send it to the display device 200.
[0098] The memory 190 is used to store various operating programs, data and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.
[0099] The power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller.
[0100] In order to perform user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system may provide a user interface (control the display device), allow the user to interact with the display device 200, and support the running of various application programs.
[0101] It should be noted that the operating system may be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.
[0102] The operating system can be divided into different modules or layers according to the functions implemented, such as Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom, namely, the application layer (Applications) layer (referred to as "application layer"), the application framework layer (Application Framework) layer (referred to as "framework layer"), the system library layer and the kernel layer.
[0103] In some embodiments, the application layer is used to provide services and interfaces for applications so that the display device 200 can run applications and interact with users based on the applications. At least one application can be run in the application layer, and these applications can be window programs, system settings programs, clock programs, etc. that come with the operating system; they can also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the above examples.
[0104] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications. The application framework layer includes some predefined functions. The application framework layer is equivalent to a processing center that determines the actions that applications in the application layer take. Through the API interface, applications can access system resources and obtain system services during execution.
[0105] like Figure 4 As shown, in the embodiment of the present application, the application framework layer includes a view system, managers, content providers, etc., wherein the view system can design and implement the interface and interaction of the application, and the view system includes lists, grids, text boxes, buttons, etc. The manager includes at least one of the following modules: an activity manager for interacting with all activities running in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to the application package currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0106] In some embodiments, the activity manager is used to manage the life cycle of each application and the usual navigation back function, such as controlling the exit, opening, and back of the application. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, capturing the screen, and controlling the display window changes, for example, reducing the display window, shaking the display, distorting the display, etc.
[0107] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction library contained in the system runtime layer, such as the C / C++ instruction library, to implement the functions to be implemented by the framework layer.
[0108] In some embodiments, the kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. Figure 4 As shown, the kernel layer may be configured with hardware drivers, and the drivers included in the kernel layer may be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.
[0109] It should be noted that the above example is only a simple division of the operating system functions and does not constitute a limitation on the specific operating system form of the display device 200 in the embodiment of the present application. Depending on factors such as the function of the display device and the type of operating system, the number of levels and specific level types contained in the operating system may be expressed in other forms.
[0110] In some embodiments, Figure 1 and 2 As shown, a picture book generation method is provided, which can be applied to Figure 1 The display device 200 shown in FIG. Figure 2 As shown, the display device 200 may include a display 260 and a controller 250; the controller 250 is coupled to the display 260; the method may include the following steps:
[0111] Step S501: receiving picture book generation requirement information and identifying text information corresponding to the picture book generation requirement information.
[0112] Step S502, input the text information into the text processing model to obtain picture book role information and picture book theme information; the picture book role information is the information representing the role in the text information, and the picture book theme information is the information representing the theme in the text information; the text processing model is used to output picture book role information and picture book theme information matching the text information based on the input text information.
[0113] Step S503, filter out target picture book style information that matches the picture book character information and the picture book theme information from the preset picture book style information; filter out target broadcast timbre information that matches the picture book character information and the picture book theme information from the preset broadcast timbre information; filter out target background audio information that matches the picture book theme information and the target picture book style information from the preset background audio information.
[0114] Step S504, based on the picture book character information, the picture book theme information, the target picture book style information, the target broadcast tone information and the target background audio information, a picture book generation process is performed to obtain a picture book corresponding to the picture book generation requirement information.
[0115] The picture book generation requirement information is used to describe the content of the picture book that the user wants to generate.
[0116] The picture book generation instruction may be an instruction for controlling the display device 200 to generate a picture book.
[0117] The text processing model can be a deep learning model with a large number of parameters that is independently developed or improved from an existing model. The text processing model can learn language patterns, grammar and semantics by processing a large amount of text data, thereby understanding and generating human language.
[0118] Among them, picture book character information refers to the information describing the main characters that appear in the picture book. The characters are also the promoters and participants of the plot.
[0119] Among them, the theme information of the picture book refers to the information that describes the core idea or concept of the picture book. The theme can also be understood as a simple summary of the content, which is the essence and core of the book.
[0120] The picture book style information refers to information describing the image style (i.e., painting style) of the picture book. For example, the style information may be cartoon, realistic, hand-painted, watercolor, etc. The target style information refers to the style recommended for generating the picture book.
[0121] The announcement timbre information refers to the voice characteristics and voice style used when announcing the contents of the picture book by voice. The target announcement timbre information refers to the timbre of the announcement audio recommended for synthesizing the picture book.
[0122] Among them, background music information refers to the music used to adjust the atmosphere during the picture book playback process. Background music information is usually inserted into the dialogue, which can enhance the expression of emotions and make readers who watch the picture book feel immersive. Target background music information refers to the background music recommended for generating picture books.
[0123] Specifically, if the controller 250 receives picture book generation requirement information, the controller 250 can identify the text information corresponding to the picture book generation requirement information, and then control the display 260 to display the text information corresponding to the input picture book generation requirement information on the interface used to generate a picture book, that is, on the picture book generation interface, for the user to confirm whether the text information is correct; in actual applications, the display device 200 can be a smart TV, and the display 260 can be a display screen of the smart TV. After the user confirms that the text information is correct, it can trigger the controller 250 to send a picture book generation instruction; if the controller 250 receives the picture book generation instruction, it can input the text information into the text processing model to obtain the picture book role information and the picture book theme information; then filter out the target picture book style information that matches the picture book role information and the picture book theme information from the preset picture book style information, and filter out the target broadcast timbre information that matches the picture book role information and the picture book theme information from the preset broadcast timbre information; and then filter out the target background audio information that matches the picture book theme information and the target picture book style information from the preset background audio information; finally, generate a picture book in a video format corresponding to the picture book generation requirement information based on the picture book role information, the picture book theme information, the target picture book style information, the target broadcast timbre information and the target background music information. Finally, the controller 250 can also control the display 260 to play the picture book in video format on the interface.
[0124] For example, Figure 6 A schematic diagram of the interface for generating an interface for a picture book, such as Figure 6 As shown, the picture book generation interface in this embodiment displays a prompt word, an input box, and a generation control for triggering a picture book generation instruction. Among them, the prompt word is used to remind the user of the function of this interface and briefly remind the user how to generate a picture book. For example, the prompt word can be Figure 6 After the user inputs the picture book generation requirement information through voice, the controller 250 can control the display 260 to display the text information corresponding to the picture book generation requirement information in the input box. After the user checks the text information and confirms that it is correct, the generation control (i.e. Figure 6 ) to send a picture book generation instruction to the controller 250. After receiving the picture book generation instruction, the controller 250 may perform a picture book generation process to generate a picture book in a video format corresponding to the picture book generation requirement information.
[0125] The technical solution provided in this embodiment receives the input picture book generation demand information and identifies the corresponding text information; then uses the text processing model to process the picture book role information and picture book theme information that match the text information; intelligently screens out the target picture book style information, target broadcast timbre information that match the picture book role information and picture book theme information, and target background music information that match the picture book theme information and target picture book style information. The precise matching of these picture book elements ensures that the generated picture book not only meets the needs of users, but also improves the fun and appeal of the picture book; finally, the picture book elements such as picture book role information, picture book theme information, target picture book style information, target broadcast timbre information and target background music information are integrated, and through picture book generation processing, a picture book that matches the picture book generation demand information is created for the user to help the user realize their creative ideas. This process significantly improves the user's creation convenience and the quality of the picture book, making the reading experience of children and families richer and more enjoyable, and also greatly improves the efficiency and quality of the display device for picture book generation processing.
[0126] In some embodiments, Figure 7 As shown, in the above step S502, the text information is input into the text processing model to obtain the picture book role information and the picture book theme information, which specifically include the following contents:
[0127] Step S701, input text information into a text processing model to obtain picture book keywords; picture book keywords represent keywords in text information; the text processing model is used to output picture book keywords in the text information based on the input text information.
[0128] Step S702, input the picture book keywords into the text processing model to obtain picture book theme information; the picture book theme information represents theme information that matches the picture book keywords; the text processing model is used to output picture book theme information that matches the picture book keywords based on the input picture book keywords.
[0129] Among them, picture book keywords refer to words with substantive meanings selected from text information, which are used to summarize and describe the main content or core concepts of the text information.
[0130] Specifically, the controller 250 can input the text information corresponding to the picture book generation requirement information into the text processing model, perform keyword extraction processing on the text information through the text processing model, and then output the picture book keywords in the text information. Using the text processing model, the picture book theme is regenerated based on the keywords, and the picture book theme information matching the picture book keywords is output. The picture book theme information can be string data in standard JSON format.
[0131] In practical applications, the generated picture book theme information should be as brief as possible, but the keywords in the text information need to be retained.
[0132] (1) The input text information can be: {"qurrey":"A puppy went into the forest alone"}, the picture book keyword extracted by the text processing model is forest, and the picture book theme information output by the text processing model can be: {"story_info":"The puppy started his forest adventure with curiosity and finally gained courage"}.
[0133] (2) The input text information can be: {"qurrey":"Xiaohong and Xiaoming go to school"}, the picture book keyword extracted by the text processing model is school, and the picture book theme information output by the text processing model can be: {"story_info":"Xiaohong and Xiaoming help each other in school and finally gain friendship"}.
[0134] (3) The input text information can be: {"qurrey":"Two children go to school"}, the picture book keyword extracted by the text processing model is school, and the picture book theme information output by the text processing model can be: {"story_info":"Xiaohong and Xiaoming help each other and finally gain friendship"}.
[0135] The technical solution provided in this embodiment identifies picture book keywords from text information through a text processing model, laying a solid foundation for the subsequent generation and construction of picture book theme information. These picture book keywords not only reflect the core ideas, but also help focus on the most important picture book themes, improving the efficiency and accuracy of picture book creation. Picture book theme generation and processing based on picture book keywords makes the generated picture book theme information more targeted and in-depth, and also ensures that it has a clear direction and connotation, further improving the quality and accuracy of picture book creation.
[0136] In some embodiments, Figure 8 As shown, in the above step S502, the text information is input into the text processing model to obtain the picture book role information and the picture book theme information, which specifically include the following contents:
[0137] Step S801, input text information into a text processing model to obtain character subject information; the character subject information represents subject information that matches the text information; the text processing model is used to extract the character subject information based on the input text information, or to generate the character subject information and output the character subject information that matches the text information.
[0138] Step S802, input the role subject information into the text processing model to obtain role description information; the role description information represents the description information corresponding to the role subject information; the text processing model is used to generate the role description information according to the input role subject information, and output the role description information corresponding to the role subject information.
[0139] Step S803: determining the character main body information and the character description information as picture book character information.
[0140] The role subject information refers to the name of the central role, that is, the role subject information can be the role name. The central role can be a person or an animal.
[0141] The character description information refers to information describing the appearance of the character. For example, if the character subject information is a person, the character description information may be skin color, hairstyle, clothing, decoration, etc. If the character subject information is an animal, the character description information may be color, decoration, etc.
[0142] Specifically, the controller 250 can input the text information corresponding to the picture book generation requirement information into the text processing model; if the text processing model detects that there is subject information in the input text information, the controller 250 can use the text processing model to continue to extract the subject information in the text information, and output the character subject information that matches the text information; if the text processing model detects that there is no subject information in the input text information, the controller 250 can also use the text processing model to perform character subject generation processing based on the picture book theme information, and obtain character subject information that matches the text information. The character subject information can be string data in standard JSON format.
[0143] Furthermore, the controller 250 can also perform a suitable image description for the extracted character subject information according to the character type (e.g., person, animal) corresponding to the character subject information, for example, the skin color, hairstyle, clothing, decoration, etc. of the character subject information can be described, and then the controller 250 obtains the character description information of the character subject information. Finally, the controller combines the character subject information and the character description information to determine the picture book character information. Among them, the character subject information and the character description information can both be string data in standard JSON format.
[0144] In actual applications, the picture book character information output by the text processing model can be in the format of "[character subject information] character description information". The character name corresponds to the character subject one-to-one. When a character name contains two or more character subjects, they need to be defined separately. For example, the character name "Grandpa and Grandma" belongs to two character subjects, namely "Grandpa" and "Grandma". The controller 250 can use the python list format to store character information. In actual applications, the number of character subjects can also be limited, for example, the number of character subjects is set to be greater than 0 and less than 3. If the controller 250 cannot find the character subject information in the text information, it can also intelligently generate appropriate character subject information based on the context information of the text information. In addition, if two or more identical character subject information appears in the text information, that is, two or more character subjects have the same name, then the controller 250 can use the text processing model to reassign a non-repetitive character name to each character subject. For example:
[0145] (1) The input text information can be: {"qurrey":"A puppy wandering alone in the forest"}, the character subject information extracted by the text processing model is "puppy", and the character description information generated for the character subject information "puppy" is "a yellow Shiba Inu with a red collar around its neck", and the picture book character information output by the text processing model can be: {"role_name":["[Puppy]A yellow Shiba Inu with a red collar around its neck"]}.
[0146] (2) The input text information can be: {"qurrey":"Xiaohong and Xiaoming go to school"}. The main character information extracted by the text processing model is "Xiaohong" and "Xiaoming". The character description information generated for the main character information "Xiaohong" is "a little Chinese girl with short black hair, a white T-shirt, a black student skirt, yellow sneakers, and a red scarf". The character description information generated for the main character information "Xiaoming" is "a little Chinese boy with a black flat head, a blue T-shirt, blue jeans, white sneakers, and a red scarf". The picture book character information output by the text processing model can be: {"role_name":["[Xiaohong] A little Chinese girl with short black hair, a white T-shirt, a black student skirt, yellow sneakers, and a red scarf","[Xiaoming] A little Chinese boy with a black flat head, a blue T-shirt, blue jeans, white sneakers, and a red scarf"]}.
[0147] (3) The input text information can be: {"qurrey":"Two children go to school"}. The character subject information extracted by the text processing model is a null value, that is, the character subject information is not detected. Therefore, the text processing model can intelligently generate two character subject information "Xiaohong" and "Xiaoming" according to the context information "two children" in the text information, and then generate role description information for these two character subject information respectively. The picture book character information can be: {"role_name":["[Xiaohong] A Chinese little girl with short black hair, a white T-shirt, a black student skirt, yellow sneakers, and a red scarf","[Xiaoming] A Chinese little boy with a black flat head, a blue T-shirt, blue jeans, white sneakers, and a red scarf"]}.
[0148] The technical solution provided in this embodiment can effectively extract or create character subject information by performing character subject extraction or generation processing on text information through a text processing model, accurately identify and generate character subject information corresponding to the theme of the picture book, and ensure that the character subject information not only conforms to the plot development, but also has a profound connection with the theme. Subsequently, by generating a character description for the character subject information, the appearance of the character subject is reasonably and comprehensively described, greatly enriching the level and depth. Finally, the character subject information and the character description information are combined and determined as the picture book character information, which gives the picture book a vivid character image and effectively improves the generation quality of the picture book.
[0149] In some embodiments, Fig. 9 As shown, the above step S504 performs picture book generation processing based on picture book character information, picture book theme information, target picture book style information, target broadcast timbre information and target background audio information to obtain a picture book corresponding to the picture book generation requirement information, which specifically includes the following contents:
[0150] Step S901, input the picture book role information and the picture book theme information into the text processing model to obtain the picture book text content; the picture book text content represents the text content that matches the picture book role information and the picture book theme information; the picture book text content includes multiple pages of text content, and the number of words in each page of text content does not exceed a preset word count threshold; the text processing model is used to output the picture book text content that matches the picture book role information and the picture book theme information based on the input picture book role information and the picture book theme information.
[0151] Step S902, input the picture book text content into the text processing model to obtain the adjusted picture book text content; the adjusted picture book text content is used to represent the text content adapted to the picture book image data generation process; the adjusted picture book text content includes the adjusted text content of each page of text content; the text processing model is used to adjust the input picture book text content and output the adjusted picture book text content.
[0152] Step S903 , synthesizing a picture book corresponding to the picture book generation requirement information based on the picture book image information, broadcast audio information and target background audio information corresponding to the adjusted text content of each page of text content.
[0153] The text refers to the text information describing the content of the current page.
[0154] Specifically, the controller 250 can use the text processing model to generate multiple pages of picture book text content according to the picture book character information and picture book theme information input therein. Since the picture book text content is subsequently displayed on the display 260, the number of words in the picture book text content should not be too many, so the controller 250 can also limit the number of words in each page of the generated picture book text content to not exceed a preset word count threshold.
[0155] Furthermore, due to the limited number of characters in the picture book text content, the content length of the picture book text content description screen is limited, making it difficult for the image processing model to fully understand the picture content described in the picture book text content, affecting the quality of the picture book image information generated by the image processing model, so the controller 250 can also use the text processing model to adjust the picture book text content of each page to the adjusted picture book text content adapted to the image generation process. Then the controller 250 can synthesize the picture book corresponding to the picture book generation requirement information input by the user based on the picture book image information, broadcast audio information and target background audio information corresponding to the adjusted text content of each page of text content.
[0156] For example, the input text information can be: {"qurrey":"Friendship between the snowman and the black cat"}. The theme information of the picture book output by the text processing model can be: {"story_info":"The snowman and the black cat established a deep friendship in the winter"}. The main character information extracted by the text processing model is "snowman" and "black cat". The character description information generated for the main character information "snowman" is "a tall snowman wearing a red hat and an orange scarf", and the character description information generated for the main character information "black cat" is "a small black cat wearing a blue collar, with curious eyes". The picture book character information output by the text processing model can be: {"role_name":["[snowman] A tall snowman wearing a red hat and an orange scarf","[black cat] A small black cat wearing a blue collar, with curious eyes"]}.
[0157] The text processing model is then used to generate the text content of each page of the picture book: {"PAGE: The snowman stands alone on the vast snowy field, his red hat swaying gently in the winter wind. PAGE: ……."}.
[0158] Then, the text processing model is used to adjust the text content of each page of the picture book, and the adjusted text content of each page of the text content is obtained: {"PAGE: [Snowman] stands in the vast snow, and his red hat sways gently in the wind. PAGE: ……."}. Finally, the picture book image information is generated based on the adjusted text content of each page of the text content. Fig.10 FIG. 1 is a schematic diagram of the picture book image information corresponding to the adjusted text content of the home page in the example of this embodiment. Fig.10 , the snowman is standing on the snow with a hat on his head, and his scarf is shaking gently.
[0159] The technical solution provided in this embodiment generates multiple pages of text content based on picture book character information and picture book theme information through a text processing model, and ensures that the number of words in each page of text content is controlled within a preset threshold, which not only ensures the simplicity and readability of the text content, but also ensures that the display can display all the text content. Further adjusting the picture book text content through the text processing model can enable the image processing model to more fully understand the picture content of the desired picture book image information, thereby improving the generation effect of the picture book image information. Finally, combined with the picture book image information, broadcast audio information and target background audio information corresponding to the adjusted text content of each page, a complete picture book that meets the picture book generation requirements is efficiently synthesized, improving the generation efficiency and quality of the picture book.
[0160] In some embodiments, Fig.11 As shown, in the above step S902, the picture book text content is input into the text processing model to obtain the adjusted picture book text content, which specifically includes the following contents:
[0161] Step S1101, input the picture book text content into the text processing model to obtain the description text of the entity in each page of the picture book text content; the entity is used to represent the character subject information and / or object information in each page of the text content; the description text is used to represent at least one of the action information, position relationship information, and background information of the entity; the text processing model is used to perform text enhancement processing on the entity in each page of the text content in the input picture book text content, and output the description text of the entity in each page of the text content.
[0162] Step S1102, input the description text of the entity in each page of text content into the text processing model to obtain the adjusted text content of each page of text content; the adjusted text content of each page of text content represents the text content composed of the description text of the entity in each page of text content; the text processing model is used to perform text rewriting processing on each page of text content based on the input description text of the entity in each page of text content to obtain the adjusted text content of each page of text content.
[0163] The action information may be a behavior and / or a body movement of the character subject information. For example, the action information may be a behavior such as "walking", "standing", "stroking", etc. For another example, the action information may also be a body movement such as "lying flat", "closing eyes", etc.
[0164] The position relationship information may be information describing the position relationship between the character subject information and other objects. For example, the position relationship information may be "there is a green tree next to Paul" or "there is a sun above Jack's head".
[0165] The background information may be information describing the environment in which the character subject information is located. For example, the background information may be "snow in winter" or "beach in summer".
[0166] Specifically, the text processing model is used to perform text enhancement processing on the entities (such as character subject information and / or object information) in each page of text content, which can be to convert the psychological text and emotional text describing the entity in each page of text content into more specific description content such as action information, location information and / or background information, and then the controller 250 processes and obtains the description text of the entity in each page of text content. The controller 250 uses the description text of the entity in each page of text content to adjust the expression of each page of text content, ensuring that the adjusted text of each page of text content obtained by the final processing is simple in terms of words and grammar, and does not use modifiers and adjectives, so that the image processing model can accurately understand its meaning.
[0167] Among them, emotional texts and psychological texts can be relevant psychological descriptions and emotional descriptions such as “very happy”, “feeling very fulfilled inside”, and “dreaming”.
[0168] For example, suppose the text processing model generates the picture book text content: {"PAGE: [Jack] put on his small backpack and set off for the beach. He planned to spend an unforgettable night on the beach."}. Obviously, the "unforgettable night" in the picture book text content belongs to emotional text, and it is difficult for the image processing model to accurately understand and generate related images. Therefore, the words in the picture book text content can be simplified, and the adjusted text content can be output through the text processing model: {"PAGE: [Jack] walked on the road in front of the beach with a small backpack on his back, with blue sky and white clouds in the background."}. Obviously, the adjusted text content adds descriptive text about background information, and the descriptive text for the character's main information and actions is also more concise, such as "carrying a small backpack" and "walking", and its position relationship information is also very clear, such as "on the road in front of the beach".
[0169] The technical solution provided in this embodiment performs text enhancement processing on entities in each page of text content through a text processing model, which can simply, intuitively and effectively describe the action information, location information and / or background information of the entity, and can also enrich the description content of the entity. Based on the description text of the entity, text rewriting processing is performed on each page of text content, so that the adjusted text content obtained after processing can convey the picture content more intuitively and concisely, and facilitate the image processing model to accurately understand the picture that the adjusted text content wants to express, thereby improving the quality and accuracy of the picture book image information generated by the image processing model.
[0170] In some embodiments, after the picture book text content is input into the text processing model to obtain the adjusted picture book text content in the above step S902, the method further includes:
[0171] The adjusted text content of each page of text content is input into an image processing model to obtain initial picture book image information corresponding to the adjusted text content of each page of text content; the image processing model is used to output initial picture book image information corresponding to the adjusted text content of each page of text content according to the input adjusted text content of each page of text content; the target picture book style information and the initial picture book image information of the adjusted text content of each page of text content are input into the image processing model to obtain picture book image information corresponding to the adjusted text content of each page of text content; the style information of the picture book image information corresponding to the adjusted text content of each page of text content is the same as the target picture book style information; the image content of the picture book image information corresponding to the adjusted text content of each page of text content is the same as the image content of the initial picture book image information corresponding to the adjusted text content of each page of text content; the image processing model is used to render the target picture book style information on the initial picture book image information of the adjusted text content of each page of text content, and output the picture book image information corresponding to the adjusted text content of each page of text content.
[0172] The initial picture book image information refers to a picture book image that has not been rendered and is initially generated.
[0173] The image processing model refers to a deep learning model used to convert the content described by text data into a corresponding image. The image processing model can be obtained through independent research and development, or by improving an existing model.
[0174] Specifically, first, the image processing model is used to generate initial picture book image information corresponding to the adjusted text content of each page of text content; then, according to the target picture book style information, image style rendering processing is performed on the initial picture book image information corresponding to the adjusted text content of each page of text content. For example, the initial picture book image information is rendered into a cartoon style, realistic style, watercolor style, etc., and finally the picture book image information corresponding to the adjusted text content of each page of text content is rendered.
[0175] The technical solution provided in this embodiment, through the image processing model, performs image generation processing based on the adjusted text content of each page, and can vividly convert the adjusted text content into initial picture book image information; then the image processing model combines the target picture book style information to perform style rendering processing on the initial picture book image information, further enhancing the artistry and personalization of the picture book picture, and effectively improving the image quality and picture effect of the rendered picture book image information.
[0176] In some embodiments, after the picture book text content is input into the text processing model to obtain the adjusted picture book text content in the above step S902, the method further includes:
[0177] The target broadcast timbre information and the adjusted text content of each page of text content are input into the broadcast audio information synthesis model to obtain the broadcast audio information corresponding to the adjusted text content of each page of text content; the timbre information of the broadcast audio information corresponding to the adjusted text content of each page of text content is the same as the target broadcast timbre information; the broadcast audio information synthesis model is used to perform playback audio information synthesis processing on the adjusted text content of each page of text content based on the input target broadcast timbre information to obtain the broadcast audio information corresponding to the adjusted text content of each page of text content.
[0178] Among them, the broadcast audio information synthesis model refers to a deep learning model that can simultaneously process and generate multiple data types (such as audio data, text data). The broadcast audio information synthesis model usually combines text, images, audio and other media information, and can achieve complex and rich tasks. The broadcast audio information synthesis model can be obtained through independent research and development, or by improving existing models.
[0179] The target announcement timbre information refers to a label of the timbre used to synthesize the announcement audio. For example, the target announcement timbre information may be "sweet girl timbre".
[0180] The controller 250 can realize the dubbing of each page of text content through the broadcast audio information synthesis model. Specifically, the controller 250 can use the broadcast audio information synthesis model to accept and understand information from different sources (such as text, audio, and images), and learn the association and semantic advantages between data sources of different modes, and synthesize the broadcast audio information corresponding to the adjusted text content of each page of text content according to the target broadcast timbre information through the broadcast audio information synthesis model.
[0181] The technical solution provided in this embodiment, through the broadcast audio information synthesis model, performs speech synthesis processing on the adjusted text content of each page of text content based on the target broadcast timbre information, which can effectively improve the expressiveness and attractiveness of the generated picture book, and can also convey the emotion and atmosphere of the picture book through the broadcast audio information, thereby improving the effect and viewing experience of the picture book.
[0182] In some embodiments, after the picture book text content is input into the text processing model to obtain the adjusted picture book text content in the above step S902, the method further includes:
[0183] The picture book theme information and the target picture book style information are input into the background audio information generation model to obtain background audio tags that match the picture book theme information and the target picture book style information; the background audio information generation model is used to output background audio tags that match the picture book theme information and the target picture book style information based on the input picture book theme information and the target picture book style information; the background audio tags are input into the background audio information generation model to obtain the target background audio information corresponding to the adjusted text content of each page of text content; the background audio information generation model is used to filter out the target background audio information that matches the background audio tag from the preset background audio information based on the input background audio tag, or to generate the target background audio information corresponding to the background audio tag.
[0184] The target background audio information may be light music.
[0185] The controller 250 can also acquire the target background audio information through the broadcast audio information synthesis model. Specifically, the background audio information of different background audio tags can be pre-stored in the database; after the background audio tags that match both the picture book theme information and the target picture book style information are identified through the broadcast audio information synthesis model, the target background audio information that matches the background audio tags can be searched from the database. It is also possible to use the broadcast audio information synthesis model to intelligently generate target background audio information that meets the background audio tags by taking advantage of the ability of the broadcast audio information synthesis model to accept and understand information from different sources (such as text, audio, images), and to learn the associations and semantics between different modalities.
[0186] The technical solution provided in this embodiment, through the broadcast audio information synthesis model, identifies the background audio tags that match the picture book theme information and the target picture book style information, and then searches or generates the target background audio information corresponding to the background audio tags, which can provide a richer and more fitting sound experience for the creation of picture books, ensure that the selected target background audio information complements the content and emotion of the created picture book, and greatly improves the generation effect of the picture book.
[0187] In some embodiments, before the above step S502, inputting the text information into the text processing model to obtain the picture book character information and the picture book theme information, the method further includes:
[0188] When receiving picture book generation requirement information, and the picture book generation requirement information is voice information, voice recognition processing is performed on the picture book generation requirement information to obtain text information corresponding to the picture book generation requirement information; the text information is input into a text processing model to obtain a compliance detection result of the text information; the compliance detection result is used to indicate whether the text information is compliant; the text processing model is used to perform compliance detection on the input text information and output the compliance detection result of the text information;
[0189] In the above step S502, the text information is input into the text processing model to obtain the picture book character information and the picture book theme information, which specifically includes the following contents:
[0190] When the compliance detection result indicates that the text information is compliant, the text information is input into a text processing model to obtain picture book role information and picture book theme information.
[0191] The compliance detection result is used to indicate whether the content in the text information complies with relevant laws, regulations or rules and regulations. For example, the compliance detection result can describe whether there are any sensitive words that violate the law in the text information.
[0192] Specifically, the controller 250 first converts the picture book generation requirement information in voice format into corresponding text information. Before the controller 250 controls the display 260 to display the text information corresponding to the picture book generation requirement information on the interface, it also detects whether the text information contains sensitive words through the text processing model to obtain the compliance detection result of the text information; if the compliance detection result indicates that the text information is compliant (for example, there are no sensitive words in the text information), the control display 260 displays the text information on the picture book generation interface; otherwise, the text information is not displayed and a prompt message that the corresponding picture book generation requirement information is not compliant is provided, for example, the control display 260 displays a prompt message that sensitive words exist on the picture book generation interface.
[0193] The technical solution provided in this embodiment uses a text processing model to perform compliance detection on text information to ensure that the text information corresponding to the picture book generation requirement information complies with relevant regulations and standards, which is conducive to maintaining the safety and health of the picture book creation environment.
[0194] In some embodiments, Figure 5 As shown, the present application provides a display device 200 , which may include a display 260 and a controller 250 ; wherein the controller 250 is coupled to the display 260 .
[0195] like Figure 5As shown, the controller 250 is configured as follows:
[0196] Step S501: receiving picture book generation requirement information and identifying text information corresponding to the picture book generation requirement information.
[0197] Step S502, input the text information into the text processing model to obtain picture book role information and picture book theme information; the picture book role information is the information representing the role in the text information, and the picture book theme information is the information representing the theme in the text information; the text processing model is used to output picture book role information and picture book theme information matching the text information based on the input text information.
[0198] Step S503, filter out target picture book style information that matches the picture book character information and the picture book theme information from the preset picture book style information; filter out target broadcast timbre information that matches the picture book character information and the picture book theme information from the preset broadcast timbre information; filter out target background audio information that matches the picture book theme information and the target picture book style information from the preset background audio information.
[0199] Step S504, based on the picture book character information, the picture book theme information, the target picture book style information, the target broadcast tone information and the target background audio information, a picture book generation process is performed to obtain a picture book corresponding to the picture book generation requirement information.
[0200] It should be noted that for the specific implementation process of the controller 250, reference can be made to Figure 5 The relevant embodiments of the picture book generation method shown are not repeated here.
[0201] The technical solution provided in this embodiment receives the input picture book generation demand information and identifies the corresponding text information; then uses the text processing model to process the picture book role information and picture book theme information that match the text information; intelligently screens out the target picture book style information, target broadcast timbre information that match the picture book role information and picture book theme information, and target background music information that match the picture book theme information and target picture book style information. The precise matching of these picture book elements ensures that the generated picture book not only meets the needs of users, but also improves the fun and appeal of the picture book; finally, the picture book elements such as picture book role information, picture book theme information, target picture book style information, target broadcast timbre information and target background music information are integrated, and through picture book generation processing, a picture book that matches the picture book generation demand information is created for the user to help the user realize their creative ideas. This process significantly improves the user's creation convenience and the quality of the picture book, making the reading experience of children and families richer and more enjoyable, and also greatly improves the efficiency and quality of the display device for picture book generation processing.
[0202] In some embodiments, the controller 250 is further configured to:
[0203] Input text information into the text processing model to obtain picture book keywords; picture book keywords represent keywords in the text information; the text processing model is used to output picture book keywords in the text information based on the input text information; input picture book keywords into the text processing model to obtain picture book theme information; picture book theme information represents theme information that matches the picture book keywords; the text processing model is used to output picture book theme information that matches the picture book keywords based on the input picture book keywords.
[0204] In some embodiments, the controller 250 is further configured to:
[0205] Input text information into a text processing model to obtain character subject information; the character subject information represents subject information that matches the text information; the text processing model is used to extract character subject information based on the input text information, or to generate character subject information and output character subject information that matches the text information; input character subject information into a text processing model to obtain character description information; the character description information represents description information corresponding to the character subject information; the text processing model is used to generate character description information based on the input character subject information and output character description information corresponding to the character subject information; the character subject information and the character description information are determined as picture book character information.
[0206] In some embodiments, the controller 250 is further configured to:
[0207] Input the picture book role information and picture book theme information into the text processing model to obtain the picture book text content; the picture book text content represents the text content matching the picture book role information and picture book theme information; the picture book text content includes multiple pages of text content, and the number of words in each page of text content does not exceed the preset word count threshold; the text processing model is used to output the picture book text content matching the picture book role information and picture book theme information according to the input picture book role information and picture book theme information; input the picture book text content into the text processing model to obtain the adjusted picture book text content; the adjusted picture book text content is used to represent the text content adapted to the picture book image data generation and processing; the adjusted picture book text content includes the adjusted text content of each page of text content; the text processing model is used to adjust the input picture book text content and output the adjusted picture book text content; based on the picture book image information, broadcast audio information and target background audio information corresponding to the adjusted text content of each page of text content, synthesize the picture book corresponding to the picture book generation requirement information.
[0208] In some embodiments, the controller 250 is further configured to:
[0209] The picture book text content is input into a text processing model to obtain description text of the entities in each page of the text content in the picture book text content; the entities are used to represent the character subject information and / or object information in each page of the text content; the description text is used to represent at least one of the action information, position relationship information, and background information of the entity; the text processing model is used to perform text enhancement processing on the entities in each page of the input picture book text content, and output the description text of the entities in each page of the text content; the description text of the entities in each page of the text content is input into the text processing model to obtain the adjusted text content of each page of the text content; the adjusted text content of each page of the text content represents the text content composed of the description text of the entities in each page of the text content; the text processing model is used to perform text rewriting processing on each page of the text content based on the description text of the entities in each page of the text content input to obtain the adjusted text content of each page of the text content.
[0210] In some embodiments, the controller 250 is further configured to:
[0211] The adjusted text content of each page of text content is input into an image processing model to obtain initial picture book image information corresponding to the adjusted text content of each page of text content; the image processing model is used to output initial picture book image information corresponding to the adjusted text content of each page of text content according to the input adjusted text content of each page of text content; the target picture book style information and the initial picture book image information of the adjusted text content of each page of text content are input into the image processing model to obtain picture book image information corresponding to the adjusted text content of each page of text content; the style information of the picture book image information corresponding to the adjusted text content of each page of text content is the same as the target picture book style information; the image content of the picture book image information corresponding to the adjusted text content of each page of text content is the same as the image content of the initial picture book image information corresponding to the adjusted text content of each page of text content; the image processing model is used to render the target picture book style information on the initial picture book image information of the adjusted text content of each page of text content, and output the picture book image information corresponding to the adjusted text content of each page of text content.
[0212] In some embodiments, the controller 250 is further configured to:
[0213] The target broadcast timbre information and the adjusted text content of each page of text content are input into the broadcast audio information synthesis model to obtain the broadcast audio information corresponding to the adjusted text content of each page of text content; the timbre information of the broadcast audio information corresponding to the adjusted text content of each page of text content is the same as the target broadcast timbre information; the broadcast audio information synthesis model is used to perform playback audio information synthesis processing on the adjusted text content of each page of text content based on the input target broadcast timbre information to obtain the broadcast audio information corresponding to the adjusted text content of each page of text content.
[0214] In some embodiments, the controller 250 is further configured to:
[0215] The picture book theme information and the target picture book style information are input into the background audio information generation model to obtain background audio tags that match the picture book theme information and the target picture book style information; the background audio information generation model is used to output background audio tags that match the picture book theme information and the target picture book style information based on the input picture book theme information and the target picture book style information; the background audio tags are input into the background audio information generation model to obtain the target background audio information corresponding to the adjusted text content of each page of text content; the background audio information generation model is used to filter out the target background audio information that matches the background audio tag from the preset background audio information based on the input background audio tag, or to generate the target background audio information corresponding to the background audio tag.
[0216] In some embodiments, the controller 250 is further configured to:
[0217] When receiving picture book generation requirement information, and the picture book generation requirement information is voice information, voice recognition processing is performed on the picture book generation requirement information to obtain text information corresponding to the picture book generation requirement information; the text information is input into a text processing model to obtain a compliance detection result of the text information; the compliance detection result is used to indicate whether the text information is compliant; the text processing model is used to perform compliance detection on the input text information and output the compliance detection result of the text information;
[0218] The controller 250 is further configured to:
[0219] When the compliance detection result indicates that the text information is compliant, the text information is input into a text processing model to obtain picture book role information and picture book theme information.
[0220] In some embodiments, Fig.12 As shown, in order to more clearly describe the signaling interaction process between the modules of the display device, the present application also provides another picture book generation method, which may include the following steps:
[0221] Step 1, when the user opens the APP (such as a picture book application), the APP control module responds to the opening operation and displays the user interaction interface to the user. The interaction interface can also display a radio button, a text input window and an expression component (audio and video alignment, display broadcast).
[0222] Step 2, when the user wants to generate a picture book, the APP control module can call the voice receiving module to monitor the user's far-field voice / near-field voice through the receiving microphone; the voice receiving module sends the acquired far-field voice / near-field voice to the voice recognition module; the voice recognition module converts the received far-field voice / near-field voice into text, and then sends the converted text information to the APP control module.
[0223] Step 3: When the user wants to generate a picture book, the APP control module can also call the text entry module to monitor the keys on the keyboard; the text entry module sends the text information corresponding to the keys to the APP control module.
[0224] Step 4: The APP control module sends the text information to the compliance review module so that the compliance review module can identify whether the text information corresponding to the user input data (such as voice, key text) contains sensitive words. If no sensitive words are detected, the compliance review module sends the text information to the role extraction module, theme extraction module, style extraction module and broadcast tone recommendation module respectively.
[0225] Step 5: The role extraction module uses the text processing model to give the appropriate role and role attributes according to the text information. If the user does not provide a role, the corresponding role name and role attribute information are generated according to the theme, and the role name and role attribute information are sent to the generation module. The theme extraction module uses the text processing model to recommend the picture book theme information according to the text information, and sends the picture book theme information to the generation module and the music generation module respectively. The style extraction module recommends the picture book style according to the text information, and sends the picture book style to the text and picture module and the music generation module respectively. The broadcast tone recommendation module recommends the broadcast tone according to the text information and the picture book style, and sends the broadcast tone to the speech synthesis module.
[0226] Step 6: The generation module generates chapters and the content of each chapter based on the character name and character attribute information, as well as the picture book theme information, and sends the generated content to the text and picture rewriting module.
[0227] Step 7: The text image rewriting module rewrites the prompt words based on the content to obtain prompt words suitable for the text image, and sends the prompt words suitable for the text image to the text image module. The text image module uses the image processing model to generate an image based on the prompt words suitable for the text image, and sends the image to the APP control module.
[0228] Step 8: The speech synthesis module synthesizes the broadcast audio based on the broadcast timbre and sends the broadcast audio to the APP control module.
[0229] Step 9: The music generation module synthesizes background music based on the picture book theme information and picture book style, and sends the background music to the APP control module.
[0230] Step 10, the APP control module performs audio and video alignment processing on the background music, broadcast audio and image to obtain a picture book image, and can also perform picture book broadcast processing on the picture book image.
[0231] It should be noted that models such as text processing models can be deployed on display devices or on cloud servers / servers, and this application does not limit this.
[0232] The technical solution provided in this embodiment utilizes a text processing model to extract key character information and theme information from text information; intelligently screens out target style information, target broadcast timbre information, and target background music information that match the character information and theme information. The precise matching of these picture book elements ensures that the generated picture book not only meets the needs of users, but also enhances the fun and appeal of the picture book; finally, the picture book elements such as character information, theme information, target style information, target broadcast timbre information, and target background music information are integrated, and through picture book generation processing, a picture book that matches the picture book generation demand information is created for users to help users realize their creative ideas.
[0233] In some embodiments, in order to more clearly illustrate the picture book generation method provided in the embodiments of the present application, as shown in FIG. Fig.13 As shown, the picture book generation method is specifically described below with a specific embodiment.
[0234] (1) Recognize a user request in voice format input into a user interface of a picture book application and obtain corresponding text information.
[0235] (2) Inputting text information into a text processing model, extracting and generating role themes, and rewriting text images through the text processing model to obtain role themes, text, and rewritten text; and searching and selecting roles from a role library.
[0236] (3) The rewritten text is input into the multimodal large model and image generation is performed to obtain an image. The multimodal large model is used to perform reading sound synthesis and background music synthesis to obtain the reading audio and background music of the text.
[0237] (4) Use character themes, text, images, reading audio, and background music to synthesize picture book images.
[0238] In practical applications, it can also be achieved through task generation module, task management module and task execution module Fig.12 The signaling interaction diagram between the task generation module, task management module and task execution module is shown in Fig.13As shown in the figure. After the homepage submits the request, the task generation module first performs a pre-judgment, then generates the task, and submits the generated task to the task management module. After a period of time, the task generation module queries the task management module for the queue status, and the task management module returns the queue status to the task generation module, so that the task generation module returns the task submission result to the homepage. After that, the task management module distributes the task to the task execution module, and the task execution module selects an executor to execute the task.
[0239] The above embodiment achieves the effect of automatically recommending themes, characters, styles, and background sounds, without the need for users to manually select characters, themes, styles, etc., nor does it require users to manually input characters, themes, styles, etc., thereby simplifying the picture book generation process and improving the picture book generation efficiency.
[0240] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0241] Based on the same inventive concept, the embodiment of the present application also provides a picture book generation device for implementing the picture book generation method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more picture book generation device embodiments provided below can refer to the limitations of the picture book generation method above, and will not be repeated here.
[0242] In some embodiments, Fig.15 As shown, a picture book generating device 1500 is provided, which can be applied to Figure 1 The display device 200 shown in FIG. Figure 2 As shown, the display device 200 includes a display 260 and a controller 250; the controller 250 is coupled to the display 260; the device may include:
[0243] The demand acquisition module 1501 is used to receive the picture book generation demand information and identify the text information corresponding to the picture book generation demand information.
[0244] The information extraction module 1502 is used to input text information into the text processing model to obtain picture book role information and picture book theme information; the picture book role information is the information representing the role in the text information, and the picture book theme information is the information representing the theme in the text information; the text processing model is used to output picture book role information and picture book theme information matching the text information based on the input text information.
[0245] The information filtering module 1503 is used to filter out target picture book style information that matches the picture book character information and the picture book theme information from the preset picture book style information; filter out target broadcast timbre information that matches the picture book character information and the picture book theme information from the preset broadcast timbre information; and filter out target background audio information that matches the picture book theme information and the target picture book style information from the preset background audio information.
[0246] The picture book generation module 1504 is used to perform picture book generation processing based on picture book role information, picture book theme information, target picture book style information, target broadcast tone information and target background audio information to obtain a picture book corresponding to the picture book generation requirement information.
[0247] Each module in the picture book generation device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.
[0248] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0249] In some embodiments, a computer program product is provided, including a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.
[0250] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0251] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0252] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0253] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A picture book generation method, characterized in that: include: Receiving picture book generation requirement information, and identifying text information corresponding to the picture book generation requirement information; Input the text information into a text processing model to obtain picture book role information and picture book theme information; the picture book role information is information representing the role in the text information, and the picture book theme information is information representing the theme in the text information; the text processing model is used to output the picture book role information and the picture book theme information matching the text information according to the input text information; Filter out target picture book style information that matches the picture book character information and the picture book theme information from the preset picture book style information; filter out target broadcast timbre information that matches the picture book character information and the picture book theme information from the preset broadcast timbre information; Filtering out target background audio information that matches the picture book theme information and the target picture book style information from the preset background audio information; Based on the picture book character information, the picture book theme information, the target picture book style information, the target broadcast timbre information and the target background audio information, a picture book generation process is performed to obtain a picture book corresponding to the picture book generation requirement information.
2. The method according to claim 1, characterized in that: The step of inputting the text information into a text processing model to obtain picture book character information and picture book theme information includes: Input the text information into a text processing model to obtain picture book keywords; the picture book keywords represent keywords in the text information; the text processing model is used to output the picture book keywords in the text information according to the input text information; The picture book keywords are input into the text processing model to obtain the picture book theme information; the picture book theme information represents theme information that matches the picture book keywords; the text processing model is used to output the picture book theme information that matches the picture book keywords based on the input picture book keywords.
3. The method according to claim 2, characterized in that The step of inputting the text information into a text processing model to obtain picture book character information and picture book theme information includes: Inputting the text information into the text processing model to obtain character subject information; the character subject information represents subject information matching the text information; the text processing model is used to perform character subject information extraction processing or character subject information generation processing based on the input text information, and output the character subject information matching the text information; The role subject information is input into the text processing model to obtain role description information; the role description information represents description information corresponding to the role subject information; the text processing model is used to perform role description information generation processing according to the input role subject information, and output the role description information corresponding to the role subject information; The character main body information and the character description information are determined as the picture book character information.
4. The method according to claim 1, characterized in that: The picture book generation process is performed based on the picture book character information, the picture book theme information, the target picture book style information, the target broadcast timbre information and the target background audio information to obtain a picture book corresponding to the picture book generation requirement information, including: The picture book character information and the picture book theme information are input into the text processing model to obtain the picture book text content; the picture book text content represents the text content matching the picture book character information and the picture book theme information; the picture book text content includes multiple pages of text content, and the number of words in each page of text content does not exceed a preset word count threshold; the text processing model is used to output the picture book text content matching the picture book character information and the picture book theme information according to the input picture book character information and the picture book theme information; Input the picture book text content into the text processing model to obtain adjusted picture book text content; the adjusted picture book text content is used to represent text content adapted to the picture book image data generation process; the adjusted picture book text content includes the adjusted text content of each page of text content; the text processing model is used to adjust the input picture book text content and output the adjusted picture book text content; Based on the picture book image information, the broadcast audio information and the target background audio information corresponding to the adjusted text content of each page of text content, a picture book corresponding to the picture book generation requirement information is synthesized.
5. The method according to claim 4, characterized in that The step of inputting the picture book text content into the text processing model to obtain adjusted picture book text content includes: The picture book text content is input into the text processing model to obtain a description text of an entity in each page of the text content of the picture book text content; the entity is used to represent the character subject information and / or object information in each page of the text content; the description text is used to represent at least one of the action information, position relationship information, and background information of the entity; the text processing model is used to perform text enhancement processing on the entity in each page of the text content of the input picture book text content, and output the description text of the entity in each page of the text content; The description text of the entity in each page of text content is input into the text processing model to obtain the adjusted text content of each page of text content; the adjusted text content of each page of text content represents the text content composed of the description text of the entity in each page of text content; the text processing model is used to perform text rewriting processing on the text content of each page based on the input description text of the entity in each page of text content to obtain the adjusted text content of each page of text content.
6. The method according to claim 4, characterized in that After inputting the picture book text content into the text processing model to obtain the adjusted picture book text content, the method further includes: Inputting the adjusted text content of each page of text content into an image processing model to obtain initial picture book image information corresponding to the adjusted text content of each page of text content; the image processing model is used to output the initial picture book image information corresponding to the adjusted text content of each page of text content according to the input adjusted text content of each page of text content; The target picture book style information and the initial picture book image information of the adjusted text content of each page of text content are input into the image processing model to obtain the picture book image information corresponding to the adjusted text content of each page of text content; the style information of the picture book image information corresponding to the adjusted text content of each page of text content is the same as the target picture book style information; the image content of the picture book image information corresponding to the adjusted text content of each page of text content is the same as the image content of the initial picture book image information corresponding to the adjusted text content of each page of text content; the image processing model is used to render the target picture book style information on the input initial picture book image information of the adjusted text content of each page of text content, and output the picture book image information corresponding to the adjusted text content of each page of text content.
7. The method according to claim 4, characterized in that After inputting the picture book text content into the text processing model to obtain the adjusted picture book text content, the method further includes: The target broadcast timbre information and the adjusted text content of each page of text content are input into a broadcast audio information synthesis model to obtain the broadcast audio information corresponding to the adjusted text content of each page of text content; the timbre information of the broadcast audio information corresponding to the adjusted text content of each page of text content is the same as the target broadcast timbre information; the broadcast audio information synthesis model is used to perform playback audio information synthesis processing on the input adjusted text content of each page of text content based on the input target broadcast timbre information to obtain the broadcast audio information corresponding to the adjusted text content of each page of text content.
8. The method according to claim 4, characterized in that After inputting the picture book text content into the text processing model to obtain the adjusted picture book text content, the method further includes: Input the picture book theme information and the target picture book style information into a background audio information generation model to obtain a background audio tag that matches the picture book theme information and the target picture book style information; the background audio information generation model is used to output a background audio tag that matches the picture book theme information and the target picture book style information according to the input picture book theme information and the target picture book style information; The background audio tag is input into the background audio information generation model to obtain the target background audio information corresponding to the adjusted text content of each page of text content; the background audio information generation model is used to filter out the target background audio information matching the background audio tag from the preset background audio information according to the input background audio tag, or to generate the target background audio information corresponding to the background audio tag.
9. The method according to any one of claims 1 to 8, characterized in that: Before inputting the text information into a text processing model to obtain the picture book character information and the picture book theme information, the method further includes: Upon receiving the picture book generation requirement information, if the picture book generation requirement information is voice information, performing voice recognition processing on the picture book generation requirement information to obtain text information corresponding to the picture book generation requirement information; The text information is input into the text processing model to obtain a compliance detection result of the text information; the compliance detection result is used to indicate whether the text information is compliant; the text processing model is used to perform compliance detection on the input text information and output the compliance detection result of the text information; The step of inputting the text information into a text processing model to obtain picture book character information and picture book theme information includes: When the compliance detection result indicates that the text information is compliant, the text information is input into a text processing model to obtain picture book role information and picture book theme information.
10. A display device, characterized in that: include: monitor; A controller is coupled to the display and configured to: Receiving picture book generation requirement information, and identifying text information corresponding to the picture book generation requirement information; Input the text information into a text processing model to obtain picture book role information and picture book theme information; the picture book role information is information representing the role in the text information, and the picture book theme information is information representing the theme in the text information; the text processing model is used to output the picture book role information and the picture book theme information matching the text information according to the input text information; Filter out target picture book style information that matches the picture book character information and the picture book theme information from the preset picture book style information; filter out target broadcast timbre information that matches the picture book character information and the picture book theme information from the preset broadcast timbre information; Filtering out target background audio information that matches the picture book theme information and the target picture book style information from the preset background audio information; Based on the picture book character information, the picture book theme information, the target picture book style information, the target broadcast timbre information and the target background audio information, a picture book generation process is performed to obtain a picture book corresponding to the picture book generation requirement information.
Citation Information
Cited By
Voice-driven intelligent picture book generation method and device, electronic equipment and storage medium
CN121393444A