Picture book rewriting method, picture book continuous writing method and display equipment
By receiving the rewriting requirement information on the display device, using the multi-dimensional information and processing model in the database to generate rewriting content, the problem of low rewriting efficiency in picture books in traditional technology is solved, and automated rewriting is achieved and content consistency is maintained.
Patent Information
- Application Number
- CN202411982780.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-31
AI Technical Summary
In traditional technology, picture books are less efficient in rewriting, and the complete picture book generation process needs to be re-executed, resulting in a cumbersome rewriting process.
When displaying picture book content on the user interface, it receives the rewrite demand information, recognizes the relevant text information, and obtains multi-dimensional information from the database based on the picture book identification. Enter text and multi-dimensional information into the text processing model and image processing model, generate overwritten text and image data, and perform audio data synthesis processing to replace the original picture book content.
The automatic rewriting of picture books is realized, manual operation steps are reduced, rewriting efficiency is improved, and the consistency between the rewriting content and the original picture books in terms of roles, styles, themes, etc.
Smart Images

Figure CN119991873A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of display devices, and in particular to a picture book rewriting method, a picture book continuation method and a display device. Background Art
[0002] Display devices such as smart TVs refer to devices that realize two-way human-computer interaction functions based on Internet application technology.
[0003] In traditional technology, when the generated picture book does not meet the user's expectations and the user wants to rewrite the picture book, a complete picture book is usually regenerated according to the user's new needs; however, each time the picture book is rewritten, the picture book generation process needs to be re-executed, which makes the picture book rewriting process more cumbersome and results in low efficiency in rewriting the picture book.
[0004] Therefore, there is a technical problem in the traditional technology that the rewriting efficiency of the picture book is low. Summary of the invention
[0005] The present application provides a picture book rewriting method, a picture book continuation method and a display device to solve the technical problem of low picture book rewriting efficiency.
[0006] In a first aspect, some embodiments provide a picture book rewriting method, applied to a display device, the method comprising:
[0007] In the case where the user interface displays picture book content, receiving rewriting requirement information for the picture book content, and using the identifier of the current picture book as the picture book identifier corresponding to the rewriting requirement information;
[0008] Identify the text information corresponding to the rewriting requirement information, and acquire the picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the rewriting requirement information; the picture book information represents multi-dimensional information of the picture book corresponding to the picture book identifier;
[0009] Inputting the text information and the information representing the characters in the picture book information into a text processing model to obtain rewritten text content; the rewritten text content represents text content composed of the text information and the information representing the characters; the text processing model is used to perform picture book text content generation processing based on the input text information and the information representing the characters, and output the rewritten text content;
[0010] Inputting the rewritten text content, the information representing the style in the picture book information, and the information representing the character characteristics in the picture book information into an image processing model to obtain rewritten image data; the rewritten image data represents image data that matches the rewritten text content, the information representing the style, and the information representing the character characteristics; the image processing model is used to perform picture book image data generation processing based on the input rewritten text content, the information representing the style, and the information representing the character characteristics, and output the rewritten image data;
[0011] Based on the rewritten text content and the information representing the broadcast tone in the picture book information, a playback audio data synthesis process is performed to obtain the broadcast audio data corresponding to the rewritten image data; and based on the information representing the theme and the information representing the style in the picture book information, a background audio data generation process is performed to obtain the background audio data corresponding to the rewritten image data;
[0012] The picture book image data in the picture book content in the picture book corresponding to the picture book identifier is replaced by the rewritten image data, and the broadcast audio data and background audio data corresponding to the picture book image data are replaced by the broadcast audio data and background audio data corresponding to the rewritten image data, to obtain a rewritten picture book; the number of pages of the rewritten picture book and the number of pages of the picture book corresponding to the picture book identifier remain unchanged.
[0013] Technical effect: By associating picture book identification with multi-dimensional information and storing it in a database, and uniformly obtaining information such as role, style, and theme based on the picture book identification when rewriting, it is helpful to avoid redundant operations of repeated acquisition and manual input of information; by inputting the obtained information into a text processing model, an image processing model, and an audio processing model for automatic processing, and accurately replacing them while keeping the number of pages unchanged, it is helpful to maintain the consistency of the rewritten content with the original picture book in terms of role, style, theme, etc.; this is conducive to realizing the automated rewriting of picture books, reducing manual operation steps, and thus improving the efficiency of picture book rewriting.
[0014] In some embodiments of the present application, when the picture book content is displayed on the user interface, receiving rewriting requirement information for the picture book content includes:
[0015] When the user interface displays picture book content, receiving an input picture book content selection operation;
[0016] When the selected picture book content is identified according to the picture book content selection operation, rewriting requirement information for the picture book content is received.
[0017] Technical effect: By first receiving the user's selection operation on the displayed picture book content, then identifying the specific picture book content based on the selection operation, and then receiving the rewriting demand information for the selected content, it is helpful to accurately locate the specific content location that the user wants to rewrite, and avoid the problem of unclear rewriting range; it is helpful to ensure that there is a clear correspondence between the rewriting demand information and the picture book content selected by the user, thereby improving the accuracy of picture book rewriting.
[0018] In some embodiments of the present application, after receiving the input picture book content selection operation, the method further includes:
[0019] In response to the picture book content selection operation, identifying the picture book content corresponding to the current focus;
[0020] The picture book content corresponding to the current focus is identified as the selected picture book content.
[0021] Technical effect: By responding to the picture book content selection operation and identifying the picture book content corresponding to the current focus, it is helpful to accurately locate the specific content position that the user wants to rewrite and avoid misidentifying the user's rewriting intention; by identifying the picture book content corresponding to the current focus as the selected picture book content, it is helpful to establish an accurate correspondence between the focus position and the specific content; thereby facilitating accurate understanding of the user's rewriting needs and improving the accuracy of picture book rewriting.
[0022] In some embodiments of the present application, before obtaining the picture book information corresponding to the picture book identifier from the database according to the picture book identifier corresponding to the rewriting requirement information, the method further includes:
[0023] Performing intent recognition processing on the text information corresponding to the rewriting requirement information to obtain an intent recognition result corresponding to the rewriting requirement information; the intent recognition result is used to indicate whether to obtain data from the database;
[0024] Based on the intention recognition result, determining whether to obtain data from the database;
[0025] The acquiring, from a database, picture book information corresponding to the picture book identifier according to the rewriting requirement information, includes:
[0026] When it is determined that the data is to be obtained from the database, the picture book information corresponding to the picture book identifier is obtained from the database according to the picture book identifier corresponding to the rewriting requirement information.
[0027] Technical effect: By performing intent recognition processing on the text information corresponding to the rewriting requirement information and judging whether it is necessary to obtain data from the database based on the intent recognition result, it is helpful to avoid unnecessary database access operations; by obtaining picture book information from the database only when the data is really needed, it is helpful to realize the on-demand execution of data acquisition operations; thereby, it is helpful to reduce resource consumption and improve overall operating efficiency while ensuring the normal implementation of the rewriting function.
[0028] In some embodiments of the present application, the information characterizing the characteristics of the role is also obtained in the following manner:
[0029] Selecting picture book image data containing the information representing the character from the picture book corresponding to the picture book identifier as reference picture book image data;
[0030] The information representing the characteristics of the character is extracted from the reference picture book image data.
[0031] Technical effect: By selecting picture book image data containing information representing the character from the picture book as a reference, it is helpful to obtain the standardized visual features of the character; by extracting information representing the character characteristics from the reference picture book image data, it is helpful to establish baseline data for the character characteristics; this is helpful to maintain the consistency of the character image in the subsequent picture book generation process and improve the quality of picture book generation.
[0032] In some embodiments of the present application, inputting the information representing the characters in the text information and the picture book information into a text processing model to obtain rewritten text content includes:
[0033] Inputting the text information into a text processing model to obtain key text information; the key text information represents key information in the text information; the text processing model is used to output the key text information in the text information according to the input text information;
[0034] The key text information and the information representing the character are input into the text processing model to obtain the rewritten text content; the rewritten text content represents the text content composed of the key text information and the information representing the character; the text processing model is used to generate picture book text content based on the input key text information and the information representing the character, and output the rewritten text content.
[0035] Technical effect: By first inputting the text information into the text processing model to extract the key text information, it is helpful to accurately grasp the core content of the user's rewriting needs; by inputting the extracted key text information together with the information representing the role into the text processing model, it is helpful to generate new text content while maintaining the role characteristics; thereby ensuring that the rewritten content meets the user's needs while maintaining the consistency of the picture book character image and improving the generation quality of the rewritten text content.
[0036] In some embodiments of the present application, inputting the rewritten text content, the information representing the style in the picture book information, and the information representing the character features in the picture book information into an image processing model to obtain the rewritten image data includes:
[0037] The rewritten text content is input into the text processing model to obtain image generation prompt information corresponding to the rewritten text content; the image generation prompt information is used to represent the text content adapted to the picture book image data generation process; the text processing model is used to adjust the input rewritten text content and output the adjusted text content; the adjusted text content is the image generation prompt information;
[0038] The image generation prompt information, the information representing the style and the information representing the character characteristics are input into an image processing model to obtain rewritten image data; the rewritten image data represents image data that matches the image generation prompt information, the information representing the style and the information representing the character characteristics; the image processing model is used to perform picture book image data generation processing based on the input image generation prompt information, the information representing the style and the information representing the character characteristics, and output the rewritten image data.
[0039] Technical effect: By converting the rewritten text content into prompt information that is more suitable for image generation, it is helpful to improve the accuracy of text-to-image conversion; by inputting the image generation prompt information together with the information representing the style and the information representing the character characteristics into the image processing model, it is helpful to maintain the consistency of style and character characteristics during the image generation process; thus, it is helpful to generate high-quality image data that meets the user's rewriting needs while maintaining the original picture book style and character characteristics.
[0040] In some embodiments of the present application, the step of inputting the image generation prompt information, the information representing the style, and the information representing the character features into an image processing model to obtain the rewritten image data includes:
[0041] Inputting the image generation prompt information and the information representing the character characteristics into an image processing model to obtain initial rewritten image data; the initial rewritten image data represents image data that matches the image generation prompt information and the information representing the character characteristics; the image processing model is used to perform picture book image data generation processing based on the input image generation prompt information and the information representing the character characteristics, and output the initial rewritten image data;
[0042] The initial rewritten image data and the information representing the style are input into the image processing model to obtain the rewritten image data; the style information of the rewritten image data is the same as the information representing the style, and the image content of the rewritten image data is the same as the image content of the initial rewritten image data; the image processing model is used to render the information representing the style on the input initial rewritten image data, and output the rewritten image data.
[0043] Technical effect: By dividing the generation process of rewritten image data into two steps, the initial rewritten image data is first generated according to the image generation prompt information and the information representing the character characteristics, which is conducive to ensuring the accuracy of the content of the generated image and the consistency of the character characteristics; then the initial rewritten image data is subjected to style rendering processing to obtain the rewritten image data, which is conducive to achieving style unification while keeping the image content unchanged; thereby, it is conducive to improving the generation quality of the rewritten image data, so that the generated image can accurately express the content and maintain the consistency of style.
[0044] In some embodiments of the present application, the playing audio data synthesis processing is performed based on the rewritten text content and the information representing the broadcast tone in the picture book information to obtain the broadcast audio data corresponding to the rewritten image data, and the background audio data generation processing is performed based on the information representing the theme and the information representing the style in the picture book information to obtain the background audio data corresponding to the rewritten image data, including:
[0045] The rewritten text content and the information representing the broadcast timbre are input into a playback audio data processing model to obtain the broadcast audio data corresponding to the rewritten image data; the timbre information of the broadcast audio data is the same as the information representing the broadcast timbre, and the text content corresponding to the broadcast audio data is the same as the rewritten text content; the playback audio data processing model is used to convert the input rewritten text content into corresponding initial broadcast audio data, adjust the information representing the broadcast timbre on the initial broadcast audio data, and output the broadcast audio data corresponding to the rewritten image data;
[0046] The information representing the theme and the information representing the style are input into a background audio data processing model to obtain background audio data corresponding to the rewritten image data; the background audio tag of the background audio data is the background audio tag corresponding to the information representing the theme and the information representing the style; the background audio data processing model is used to generate background audio data based on the background audio tag corresponding to the input information representing the theme and the information representing the style, and output the background audio data corresponding to the rewritten image data.
[0047] Technical effect: By separating the generation process of broadcast audio data and background audio data, it is helpful to control the timbre characteristics of the broadcast content and the style characteristics of the background music respectively; by adjusting the timbre of the broadcast audio data and matching the theme and style of the background audio data, it is helpful to ensure the consistency of the generated audio with the rewritten content, theme and style.
[0048] In a second aspect, some embodiments further provide a picture book continuation method, applied to a display device, the method comprising:
[0049] In the case where the user interface displays a picture book, receiving the continuation requirement information for the picture book, and using the identifier of the current picture book as the picture book identifier corresponding to the continuation requirement information;
[0050] Identify the text information corresponding to the continuation demand information, and acquire the picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the continuation demand information; the picture book information represents multi-dimensional information of the picture book corresponding to the picture book identifier;
[0051] Inputting the text information and the information representing the characters in the picture book information into a text processing model to obtain a continued text content; the continued text content represents the text content composed of the text information and the information representing the characters; the text processing model is used to perform picture book text content generation processing based on the input text information and the information representing the characters, and output the continued text content;
[0052] Inputting the continued text content, the information representing the style in the picture book information, and the information representing the character characteristics in the picture book information into an image processing model to obtain continued image data; the continued image data represents image data that matches the continued text content, the information representing the style, and the information representing the character characteristics; the image processing model is used to perform picture book image data generation processing based on the input continued text content, the information representing the style, and the information representing the character characteristics, and output the continued image data;
[0053] Based on the content of the continued text and the information representing the tone of the announcement in the picture book information, a playback audio data synthesis process is performed to obtain the announcement audio data corresponding to the continued image data; and based on the information representing the theme and the information representing the style in the picture book information, a background audio data generation process is performed to obtain the background audio data corresponding to the continued image data;
[0054] Perform semantic recognition processing on the text information corresponding to the continuation demand information to obtain the continuation position of the continuation image data in the picture book corresponding to the picture book identifier; add the continuation image data to the continuation position in the picture book corresponding to the picture book identifier, and configure the broadcast audio data and background audio data corresponding to the continuation image data to obtain a continued picture book; the number of pages of the continued picture book is greater than the number of pages of the picture book corresponding to the picture book identifier.
[0055] Technical effect: By obtaining multi-dimensional information of the original picture book and using it as a constraint for continuation, it is helpful to ensure that the sequel content is consistent with the original picture book in terms of character image, artistic style and story theme; by collaboratively generating text content, image data, broadcast audio and background audio, and determining the appropriate continuation position based on semantic recognition, it is helpful to achieve the natural connection of multimodal content, while expanding the content of the picture book while maintaining the coherence and integrity of the story; this is conducive to the automated continuation of picture books, reducing manual operation steps, and thus improving the efficiency of picture book continuation.
[0056] In some embodiments of the present application, when the picture book is displayed on the user interface, receiving the continuation requirement information for the picture book includes:
[0057] When the user interface displays a picture book, receiving an input picture book selection operation;
[0058] In response to the picture book selection operation, identifying the picture book corresponding to the current focus;
[0059] Identify the picture book corresponding to the current focus as the selected picture book;
[0060] When the selected picture book is identified according to the picture book selection operation, continuation requirement information for the picture book is received.
[0061] Technical effect: By responding to the user's picture book selection operation and identifying the picture book corresponding to the current focus position, it is helpful to establish a clear correspondence between the user operation and the specific picture book; by identifying the picture book corresponding to the current focus as the picture book selected by the user, it is helpful to accurately locate the specific picture book that the user wants to continue writing; thereby helping to improve the accuracy of the user's picture book selection operation; by receiving continuation demand information after identifying the picture book selected by the user, it is helpful to ensure the accurate correspondence between the continuation demand and the specific picture book; thereby helping to establish an orderly picture book continuation operation process, avoid mismatching between continuation demand and picture books, and improve the accuracy of picture book continuation.
[0062] In some embodiments of the present application, the method of adding the continued writing image data to the continued writing position in the picture book corresponding to the picture book identifier, and configuring the broadcast audio data and background audio data corresponding to the continued writing image data to obtain the continued writing picture book includes:
[0063] When the continuing position is the end of the picture book corresponding to the picture book identifier, the continuing image data is added after the last page of the picture book corresponding to the picture book identifier, and the broadcasting audio data and background audio data corresponding to the continuing image data are configured to obtain a continuing picture book;
[0064] When the continuation position is the beginning of the picture book corresponding to the picture book identifier, the continuation image data is added before the first page of the picture book corresponding to the picture book identifier, and the broadcast audio data and background audio data corresponding to the continuation image data are configured to obtain the continued picture book.
[0065] Technical effect: By flexibly adding continuation image data at the beginning or end of the picture book, it is helpful for users to choose the appropriate continuation position according to actual needs; by configuring corresponding broadcast audio data and background audio data for the continuation image data, it is helpful to maintain the consistency of the continuation content with the original picture book in visual and auditory experience.
[0066] In a third aspect, some embodiments further provide a display device, including:
[0067] monitor;
[0068] A controller is coupled to the display and is configured to:
[0069] In the case where the user interface displays the picture book content, receiving rewriting requirement information for the picture book content, and using the identifier of the current picture book as the picture book identifier corresponding to the rewriting requirement information;
[0070] Identify the text information corresponding to the rewriting requirement information, and acquire the picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the rewriting requirement information; the picture book information represents multi-dimensional information of the picture book corresponding to the picture book identifier;
[0071] Inputting the text information and the information representing the characters in the picture book information into a text processing model to obtain rewritten text content; the rewritten text content represents text content composed of the text information and the information representing the characters; the text processing model is used to perform picture book text content generation processing based on the input text information and the information representing the characters, and output the rewritten text content;
[0072] Inputting the rewritten text content, the information representing the style in the picture book information, and the information representing the character characteristics in the picture book information into an image processing model to obtain rewritten image data; the rewritten image data represents image data that matches the rewritten text content, the information representing the style, and the information representing the character characteristics; the image processing model is used to perform picture book image data generation processing based on the input rewritten text content, the information representing the style, and the information representing the character characteristics, and output the rewritten image data;
[0073] Based on the rewritten text content and the information representing the broadcast tone in the picture book information, a playback audio data synthesis process is performed to obtain the broadcast audio data corresponding to the rewritten image data; and based on the information representing the theme and the information representing the style in the picture book information, a background audio data generation process is performed to obtain the background audio data corresponding to the rewritten image data;
[0074] The picture book image data in the picture book content in the picture book corresponding to the picture book identifier is replaced by the rewritten image data, and the broadcast audio data and background audio data corresponding to the picture book image data are replaced by the broadcast audio data and background audio data corresponding to the rewritten image data, to obtain a rewritten picture book; the number of pages of the rewritten picture book and the number of pages of the picture book corresponding to the picture book identifier remain unchanged.
[0075] Technical effect: By associating picture book identification with multi-dimensional information and storing it in a database, and uniformly obtaining information such as role, style, and theme based on the picture book identification when rewriting, it is helpful to avoid redundant operations of repeated acquisition and manual input of information; by inputting the obtained information into a text processing model, an image processing model, and an audio processing model for automatic processing, and accurately replacing them while keeping the number of pages unchanged, it is helpful to maintain the consistency of the rewritten content with the original picture book in terms of role, style, theme, etc.; this is conducive to realizing the automated rewriting of picture books, reducing manual operation steps, and thus improving the efficiency of picture book rewriting.
[0076] In a fourth aspect, some embodiments further provide a display device, including:
[0077] monitor;
[0078] A controller is coupled to the display and is configured to:
[0079] In the case where the user interface displays a picture book, receiving the continuation requirement information for the picture book, and using the identifier of the current picture book as the picture book identifier corresponding to the continuation requirement information;
[0080] Identify the text information corresponding to the continuation demand information, and acquire the picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the continuation demand information; the picture book information represents multi-dimensional information of the picture book corresponding to the picture book identifier;
[0081] Inputting the text information and the information representing the characters in the picture book information into a text processing model to obtain a continued text content; the continued text content represents the text content composed of the text information and the information representing the characters; the text processing model is used to perform picture book text content generation processing based on the input text information and the information representing the characters, and output the continued text content;
[0082] Inputting the continued text content, the information representing the style in the picture book information, and the information representing the character characteristics in the picture book information into an image processing model to obtain continued image data; the continued image data represents image data that matches the continued text content, the information representing the style, and the information representing the character characteristics; the image processing model is used to perform picture book image data generation processing based on the input continued text content, the information representing the style, and the information representing the character characteristics, and output the continued image data;
[0083] Based on the content of the continued text and the information representing the tone of the announcement in the picture book information, a playback audio data synthesis process is performed to obtain the announcement audio data corresponding to the continued image data; and based on the information representing the theme and the information representing the style in the picture book information, a background audio data generation process is performed to obtain the background audio data corresponding to the continued image data;
[0084] Perform semantic recognition processing on the text information corresponding to the continuation demand information to obtain the continuation position of the continuation image data in the picture book corresponding to the picture book identifier; add the continuation image data to the continuation position in the picture book corresponding to the picture book identifier, and configure the broadcast audio data and background audio data corresponding to the continuation image data to obtain a continued picture book; the number of pages of the continued picture book is greater than the number of pages of the picture book corresponding to the picture book identifier.
[0085] Technical effect: By obtaining multi-dimensional information of the original picture book and using it as a constraint for continuation, it is helpful to ensure that the sequel content is consistent with the original picture book in terms of character image, artistic style and story theme; by collaboratively generating text content, image data, broadcast audio and background audio, and determining the appropriate continuation position based on semantic recognition, it is helpful to achieve the natural connection of multimodal content, while expanding the content of the picture book while maintaining the coherence and integrity of the story; this is conducive to the automated continuation of picture books, reducing manual operation steps, and thus improving the efficiency of picture book continuation. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0087] Figure 1 A schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application;
[0088] Figure 2 A schematic diagram of the hardware configuration of a display device provided in some embodiments of the present application;
[0089] Figure 3 A schematic diagram of the hardware configuration of a control device provided in some embodiments of the present application;
[0090] Figure 4 A schematic diagram of software configuration of a display device provided in some embodiments of the present application;
[0091] Figure 5 A schematic diagram of a process of rewriting a picture book provided in some embodiments of the present application;
[0092] Figure 6 A flowchart of steps for receiving rewriting requirement information provided in some embodiments of the present application;
[0093] Figure 7 A schematic diagram of the contents of a picture book provided in some embodiments of the present application;
[0094] Figure 8 A flowchart of steps for identifying picture book content provided in some embodiments of the present application;
[0095] Fig. 9 A schematic diagram of a process for continuing a picture book provided in some embodiments of the present application;
[0096] Fig.10 A schematic diagram of a picture book provided in some embodiments of the present application;
[0097] Fig.11 A signaling interaction diagram of a picture book rewriting method provided in some embodiments of the present application;
[0098] Fig.12 An overall architecture diagram of a picture book rewriting method provided in some embodiments of the present application;
[0099] Fig.13 A main branch architecture diagram of a picture book rewriting method provided in some embodiments of the present application;
[0100] Fig.14 A structural block diagram of a picture book rewriting device provided in some embodiments of the present application;
[0101] Fig.15 A structural block diagram of a picture book continuation device provided in some embodiments of the present application. DETAILED DESCRIPTION
[0102] The following embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following embodiments do not represent all implementations consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application as detailed in the claims.
[0103] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.
[0104] The terms "first", "second", "third", etc. in the specification and claims of this application and the above drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances.
[0105] The terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.
[0106] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0107] In the embodiment of the present application, the display device 200 generally refers to a device with image display and data processing capabilities. For example, the display device 200 includes but is not limited to a smart TV, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.
[0108] Figure 1 This is a schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application. Figure 1 As shown in FIG. 1 , the user can operate the display device 200 through touch operation, the mobile terminal 300 and the control device 100. For example, the control device 100 can be a remote controller, a stylus pen, a handle, etc.
[0109] The mobile terminal 300 can be used as a control device for performing human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device for establishing a communication connection with the display device 200 and performing data interaction. In some embodiments, the mobile terminal 300 can install software applications with the display device 200, and achieve connection and communication through a network communication protocol to achieve the purpose of one-to-one control operation and data communication. The audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 to achieve a synchronous display function.
[0110] like Figure 1 As also shown in FIG. 4 , the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 may be allowed to communicate and connect through a local area network (LAN), a wireless local area network (WLAN), and other networks.
[0111] The display device 200 may provide a broadcast receiving television function, and may also additionally provide an intelligent network television function with a computer support function, including but not limited to network television, smart television, Internet Protocol television (IPTV), and the like.
[0112] Figure 2 Some embodiments of the present application provide Figure 1 2 is a block diagram of the hardware configuration of the display device 200.
[0113] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.
[0114] In some embodiments, the detector 230 is used to collect signals of the external environment or external interaction. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of ambient light; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environment scenes, user attributes or user interaction gestures; or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.
[0115] In some embodiments, the display 260 includes a display function component for presenting a picture, and a driving component for driving an image display. The display 260 is used to receive an image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, and components of a menu control interface and a user control UI interface.
[0116] In some embodiments, the communication device 220 is a component for communicating with an external device or server 400 according to various communication protocol types. The display device 200 may be provided with a plurality of communication devices 220 according to different supported communication modes. For example, when the display device 200 supports wireless network communication, the display device 200 may be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.
[0117] The communication device 220 can enable the display device 200 to communicate with the external device or server 400 by wireless or wired connection. Among them, the wired connection can connect the display device 200 with the external device through components such as data cables and interfaces. The wireless connection can connect the display device 200 with the external device through wireless signals or wireless networks. The display device 200 can establish a connection relationship with the external device directly, or indirectly establish a connection relationship through a gateway, a router, a connection device, etc.
[0118] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first interface to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to the user's operation through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.
[0119] In some embodiments, the controller 250 and the tuner-demodulator 210 may be located in different separate devices, that is, the tuner-demodulator 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.
[0120] In some embodiments, the user may input a user command through a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).
[0121] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or may be an external audio output device of the display device 200. In particular, for the external audio output device of the display device 200, the display device 200 may also be provided with an external audio output terminal, and the audio output device may be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.
[0122] In some embodiments, the user input interface 280 may be used to receive instructions from a user.
[0123] Figure 3 Some embodiments of the present application provide Figure 1 The hardware configuration diagram of the control device in the figure. Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.
[0124] The control device 100 is configured to control the display device 200 , and can receive user input operation instructions, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, playing the role of an interactive intermediary between the user and the display device 200 .
[0125] In some embodiments, the control device 100 may be a smart device, for example, the control device 100 may be installed with various applications for controlling the display device 200 according to user needs.
[0126] In some embodiments, Figure 1 As shown, the mobile terminal 300 or other intelligent electronic devices can play a similar function as the control device 100 after installing the application for controlling the display device 200 .
[0127] The controller 110 includes a processor 112, a RAM 113, a ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between the internal components and the external and internal data processing functions.
[0128] The communication interface 130 implements communication of control signals and data signals with the display device 200 under the control of the controller 110. The communication interface 130 may include at least one of other near field communication modules such as a WiFi chip 131, a Bluetooth module 132, and an NFC module 133.
[0129] The user input / output interface 140 , wherein the input interface includes at least one of other input interfaces such as a microphone 141 , a touch panel 142 , a sensor 143 , and a button 144 .
[0130] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with a communication interface 130, such as a WiFi, Bluetooth, NFC or other module, and can encode the user input command through the WiFi protocol, Bluetooth protocol, or NFC protocol and send it to the display device 200.
[0131] The memory 190 is used to store various operating programs, data and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.
[0132] The power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller.
[0133] In order to perform user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system may provide a user interface (control the display device), allow the user to interact with the display device 200, and support the running of various application programs.
[0134] It should be noted that the operating system may be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.
[0135] The operating system can be divided into different modules or layers according to the functions implemented, such as Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom, namely, the application layer (Applications) layer (referred to as "application layer"), the application framework layer (Application Framework) layer (referred to as "framework layer"), the system library layer and the kernel layer.
[0136] In some embodiments, the application layer is used to provide services and interfaces for applications so that the display device 200 can run applications and interact with users based on the applications. At least one application can be run in the application layer, and these applications can be window programs, system settings programs, clock programs, etc. that come with the operating system; they can also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the above examples.
[0137] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications. The application framework layer includes some predefined functions. The application framework layer is equivalent to a processing center that determines the actions that applications in the application layer take. Through the API interface, applications can access system resources and obtain system services during execution.
[0138] like Figure 4 As shown, in the embodiment of the present application, the application framework layer includes a view system, managers, content providers, etc., wherein the view system can design and implement the interface and interaction of the application, and the view system includes lists, grids, text boxes, buttons, etc. The manager includes at least one of the following modules: an activity manager for interacting with all activities running in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to the application package currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0139] In some embodiments, the activity manager is used to manage the life cycle of each application and the usual navigation back function, such as controlling the exit, opening, and back of the application. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, capturing the screen, and controlling the display window changes, for example, reducing the display window, shaking the display, distorting the display, etc.
[0140] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction library contained in the system runtime layer, such as the C / C++ instruction library, to implement the functions to be implemented by the framework layer.
[0141] In some embodiments, the kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. Figure 4 As shown, the kernel layer may be configured with hardware drivers, and the drivers included in the kernel layer may be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.
[0142] It should be noted that the above example is only a simple division of the operating system functions and does not constitute a limitation on the specific operating system form of the display device 200 in the embodiment of the present application. Depending on factors such as the function of the display device and the type of operating system, the number of levels and specific level types contained in the operating system may be expressed in other forms.
[0143] In some embodiments, Figure 5 As shown, the present application provides a picture book rewriting method, which is applied to a display device 200. The method may include the following steps:
[0144] Step S501, when the user interface displays the picture book content, receiving rewriting requirement information for the picture book content, and using the identifier of the current picture book as the picture book identifier corresponding to the rewriting requirement information;
[0145] Step S502, identifying text information corresponding to the rewriting requirement information, and acquiring picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the rewriting requirement information; the picture book information represents multi-dimensional information of the picture book corresponding to the picture book identifier;
[0146] Step S503, inputting the information representing the characters in the text information and the picture book information into the text processing model to obtain rewritten text content; the rewritten text content represents the text content composed of the text information and the information representing the characters; the text processing model is used to generate the picture book text content according to the input text information and the information representing the characters, and output the rewritten text content;
[0147] Step S504, inputting the rewritten text content, the information representing the style in the picture book information, and the information representing the character characteristics in the picture book information into the image processing model to obtain rewritten image data; the rewritten image data represents image data that matches the rewritten text content, the information representing the style, and the information representing the character characteristics; the image processing model is used to generate picture book image data based on the input rewritten text content, the information representing the style, and the information representing the character characteristics, and output the rewritten image data;
[0148] Step S505, based on the rewritten text content and the information representing the broadcast tone in the picture book information, performing a playback audio data synthesis process to obtain the broadcast audio data corresponding to the rewritten image data, and based on the information representing the theme and the information representing the style in the picture book information, performing a background audio data generation process to obtain the background audio data corresponding to the rewritten image data;
[0149] Step S506, replacing the picture book image data in the picture book content in the picture book corresponding to the picture book identifier with the rewritten image data, and replacing the broadcast audio data and background audio data corresponding to the picture book image data with the broadcast audio data and background audio data corresponding to the rewritten image data, to obtain a rewritten picture book; the number of pages of the rewritten picture book and the number of pages of the picture book corresponding to the picture book identifier remain unchanged.
[0150] The user interface may be an interactive interface for displaying picture book content and receiving user operations, for example, it may be a display interface of a picture book.
[0151] The picture book content may be a combination of images and texts, for example, a graphic content including a storyline.
[0152] The rewriting requirement information may be a modification request made by a user for the content of a picture book, for example, it may be modification content input by a user through voice, such as "rewrite the content of this story to a child crying loudly and knocking over the medicine bottle on the ground."
[0153] The picture book identifier may be identification information used to uniquely identify a picture book, for example, may be a unique identifier generated for each story.
[0154] The database may be a data storage system for storing information related to picture books, for example, a database system for storing information such as character information, style information, story theme, story content, broadcast tone ID, character reference features, etc.
[0155] The text information may be specific text content identified from the rewriting requirement information, for example, text content converted from user voice input.
[0156] The picture book information may be all information related to the picture book stored in the database, such as character information, style information, story theme, story content, broadcast tone ID, character reference features and other information.
[0157] The multi-dimensional information may be a collection of information describing various aspects of the picture book, for example, it may be information including multiple dimensions such as characters, styles, themes, contents, and timbre.
[0158] The text processing model may be a model for processing text generation tasks, for example, a large language model.
[0159] The rewritten text content may be new text content generated by the model based on user input and character information, for example, new story content that meets the user's rewriting requirements and maintains character consistency.
[0160] The information representing the role may be information describing the characteristics of the story role, such as the role name, role attributes, and the like.
[0161] The image processing model may be a model for generating images, for example, a large Wenshengtu model.
[0162] The rewritten image data may be a new image generated according to the rewritten text content, for example, an image generated by a text graph model and matching the new story content.
[0163] The information representing the style may be information describing the style of the picture book, for example, style information such as cartoon, realism, hand-painted, watercolor, etc.
[0164] The information representing the role characteristics may be feature information used to maintain the consistency of the role, for example, it may be role attribute related features extracted from the text graph model level.
[0165] The information representing the announcement timbre may be timbre information used for speech synthesis, for example, it may be system integrated timbre such as a little girl, a little boy, a mature man, a mature woman, etc.
[0166] The broadcast audio data may be speech data converted from text content, for example, story broadcast audio generated by a speech synthesis model.
[0167] The information representing the theme may be information describing the theme of the story, for example, the overall theme style information of the story.
[0168] The background audio data may be background music data matching the story, for example, light music generated according to the theme and style of the story.
[0169] Specifically, when the controller 250 displays the picture book content on the user interface, it receives the rewriting requirement information input by the user through the voice recognition module, and generates a unique picture book identifier for the current picture book; the controller 250 inputs the rewriting requirement information into the large language model for voice-to-text processing to obtain text information, and at the same time obtains the picture book information corresponding to the picture book identifier from the database, and the picture book information includes multi-dimensional information such as role information, style information, theme information, and broadcast tone information; the controller 250 inputs the information representing the role in the text information and the picture book information into the large language model to generate rewritten text content that meets the role setting; The controller 250 inputs the rewritten text content and the information representing the style and the information representing the character characteristics in the picture book information into the Wenshengtu large model, maintains the consistency of the character image through the consistent attention mechanism, and generates rewritten image data; the controller 250 uses the multimodal generation large model to synthesize the broadcast audio data based on the rewritten text content and the information representing the broadcast timbre, and generates matching background audio data based on the information representing the theme and the information representing the style; finally, the controller 250 replaces the image, broadcast audio and background audio in the original picture book content with the newly generated content to obtain the rewritten picture book, while keeping the number of pages unchanged.
[0170] The technical solution provided in this embodiment helps avoid redundant operations of repeatedly acquiring and manually inputting information by associating picture book identification with multi-dimensional information and storing them in a database, and uniformly acquiring information such as characters, styles, and themes based on the picture book identification when rewriting. It helps maintain consistency in characters, styles, themes, and other aspects between the rewritten content and the original picture book by inputting the acquired information into a text processing model, an image processing model, and an audio processing model for automated processing, and accurately replacing the information while keeping the number of pages unchanged. This helps realize automated rewriting of picture books, reduces manual operation steps, and thus improves the efficiency of picture book rewriting.
[0171] In some embodiments, Figure 6 As shown, when the user interface displays the picture book content, receiving the rewriting requirement information for the picture book content includes:
[0172] Step S601, when the user interface displays picture book content, receiving an input picture book content selection operation;
[0173] Step S602: When the selected picture book content is identified according to the picture book content selection operation, rewriting requirement information for the picture book content is received.
[0174] The picture book content selection operation may be an operation in which a user selects picture book content through an input device, for example, it may be an operation in which a user focuses on a certain picture book image through a remote control.
[0175] For example, refer to Figure 7The display 260 displays the picture book image and the picture book text content of the picture book story page of the picture book story in the user interface of the picture book application. The picture book image may refer to Figure 7 The child holding medicine in the picture book can refer to Figure 7 For example, "The child smiled, stretched out his hand, gently picked up the medicine, and prepared to swallow it." Figure 7 The picture book story page can be the 8th picture book story page in the picture book story with the theme of "Yaoyao's Adventure Journey". The user can also return to the homepage by clicking the back button.
[0176] Specifically, when the controller 250 displays the picture book content on the user interface, it receives the picture book content selection operation input by the user through the remote control. When the user moves the focus to a page of the picture book content, the controller 250 identifies the specific picture book content selected by the user and activates the receiving microphone of the voice receiving module. After the controller 250 identifies the specific picture book content selected by the user according to the picture book content selection operation, the controller 250 receives the rewriting requirement information for the selected picture book content input by the user through the voice receiving module.
[0177] The technical solution provided in this embodiment helps to accurately locate the specific content that the user wants to rewrite, and avoid the problem of unclear rewriting range, by first receiving the user's selection operation on the displayed picture book content, and then identifying the specific picture book content based on the selection operation, and then receiving the rewriting requirement information for the selected content. This helps to ensure that there is a clear correspondence between the rewriting requirement information and the picture book content selected by the user, thereby improving the accuracy of picture book rewriting.
[0178] In some embodiments, Figure 8 As shown, after receiving the input picture book content selection operation, it also includes:
[0179] Step S801, responding to the picture book content selection operation, identifying the picture book content corresponding to the current focus;
[0180] Step S802: Identify the picture book content corresponding to the current focus as the selected picture book content.
[0181] The current focus may be the position currently selected by the remote controller on the user interface.
[0182] Among them, the picture book content corresponding to the current focus can be the picture book image displayed at the current focus position of the remote control, for example, it can be the picture book image displayed on page 8 of the picture book story page when the user focuses on the picture book image displayed on page 8 through the remote control.
[0183] Specifically, the controller 250 receives the picture book content selection operation input by the user through the remote control through the application control module, and monitors the focus position of the remote control on the user interface of the picture book application in real time. When it is detected that the focus position of the remote control has changed, the controller 250 obtains the coordinate information of the current focus position, and locates the corresponding picture book content in the user interface of the picture book application according to the coordinate information; after locating the picture book content corresponding to the current focus, the controller 250 uses the picture book content as the picture book content selected by the user, so as to subsequently receive the rewriting requirement information for the picture book content.
[0184] The technical solution provided in this embodiment, by responding to the picture book content selection operation and identifying the picture book content corresponding to the current focus, is conducive to accurately locating the specific content position that the user wants to rewrite and avoiding misidentification of the user's rewriting intention; by identifying the picture book content corresponding to the current focus as the selected picture book content, it is conducive to establishing an accurate correspondence between the focus position and the specific content; thereby facilitating accurate understanding of the user's rewriting needs and improving the accuracy of picture book rewriting.
[0185] In some embodiments, before obtaining the picture book information corresponding to the picture book identifier from the database according to the picture book identifier corresponding to the rewriting requirement information, it also includes: performing intent recognition processing on the text information corresponding to the rewriting requirement information to obtain the intention recognition result corresponding to the rewriting requirement information; the intention recognition result is used to characterize whether to obtain data from the database; based on the intention recognition result, judging whether to obtain data from the database; obtaining the picture book information corresponding to the picture book identifier from the database according to the picture book identifier corresponding to the rewriting requirement information, including: when it is judged that the data is to be obtained from the database, obtaining the picture book information corresponding to the picture book identifier from the database according to the picture book identifier corresponding to the rewriting requirement information.
[0186] Among them, the intention recognition processing can be a process of analyzing and understanding the rewriting requirement information input by the user.
[0187] Among them, the intention recognition result can be a classification result obtained after intention recognition processing of the rewriting requirement information input by the user. For example, it can be a result of determining that the user wants to rewrite the content of the picture book and needs to obtain relevant data from the database.
[0188] Specifically, after receiving the rewriting requirement information, the controller 250 sends the text information corresponding to the rewriting requirement information to the domain intention recognition module for intent recognition processing; the controller 250 receives the intention recognition result returned by the domain intention recognition module, and the intention recognition result is used to characterize whether it is necessary to obtain data from the database; the controller 250 determines whether it is necessary to obtain data from the database based on the intention recognition result. When it is determined that data needs to be obtained from the database, the controller 250 obtains the character information, style information, story theme, story content, broadcast tone ID, character reference characteristics and other picture book information corresponding to the picture book identifier corresponding to the rewriting requirement information from the database.
[0189] The technical solution provided in this embodiment helps to avoid unnecessary database access operations by performing intent recognition processing on the text information corresponding to the rewriting requirement information and judging whether it is necessary to obtain data from the database based on the intent recognition result; it helps to realize on-demand execution of data acquisition operations by obtaining picture book information from the database only when it is really necessary to obtain data; thereby helping to reduce resource consumption and improve overall operating efficiency while ensuring the normal implementation of the rewriting function.
[0190] In some embodiments, the information representing the characteristics of the character is also obtained by: selecting picture book image data containing information representing the character from the picture book corresponding to the picture book identifier as reference picture book image data; and extracting information representing the characteristics of the character from the reference picture book image data.
[0191] The picture book image data may be a generated image frame containing a specific character in the picture book, for example, it may be image data corresponding to a page of the picture book containing the character "Yaoyao".
[0192] Among them, the reference picture book image data can be an image frame selected from the picture book for extracting character features, for example, it can be an image frame selected from multiple image frames containing the "Yaoyao" character, which is used to maintain the consistency of character features in the subsequent generation process.
[0193] Specifically, the controller 250 selects a frame of picture book image data containing a character representation from the picture book corresponding to the picture book identifier as reference picture book image data through a random sampling function; the controller 250 inputs the reference picture book image data into the U-Net (U network) diffusion model, and extracts tokens features through the attention layer; the controller 250 uses the extracted tokens features as information representing the character features, and stores them in a database for maintaining the consistency of the character features in the subsequent generation process.
[0194] The technical solution provided in this embodiment is conducive to obtaining standardized visual features of the character by selecting picture book image data containing information representing the character from the picture book as a reference; by extracting information representing the character characteristics from the reference picture book image data, it is conducive to establishing baseline data of the character characteristics; thereby facilitating maintaining the consistency of the character image in the subsequent picture book generation process and improving the quality of picture book generation.
[0195] In some embodiments, information representing characters in text information and picture book information is input into a text processing model to obtain rewritten text content, including: inputting text information into a text processing model to obtain key text information; the key text information represents key information in the text information; the text processing model is used to output the key text information in the text information based on the input text information; the key text information and information representing characters are input into the text processing model to obtain rewritten text content; the rewritten text content represents text content composed of the key text information and information representing characters; the text processing model is used to perform picture book text content generation processing based on the input key text information and information representing characters, and output the rewritten text content.
[0196] Among them, the key text information can be the core information content in the user input text obtained after being processed by the text processing model. For example, it can be the key information such as "the child cried" and "the medicine bottle knocked over" extracted from the content "rewrite this story content as the child cried loudly and knocked over the medicine bottle on the ground."
[0197] Specifically, the controller 250 inputs the text information input by the user into the large language model for processing, and extracts the core content of the text information through the key information extraction module to obtain the key text information; the controller 250 inputs the extracted key text information together with the information representing the role obtained from the database into the large language model, and the large language model generates new picture book text content based on the input information, and converts the generated picture book text content into content suitable for image generation, thereby obtaining rewritten text content.
[0198] The technical solution provided in this embodiment helps to accurately grasp the core content of the user's rewriting needs by first inputting text information into a text processing model to extract key text information; by inputting the extracted key text information together with the information representing the role into the text processing model, it helps to generate new text content while maintaining the character characteristics; thereby ensuring that the rewritten content meets user needs while maintaining the consistency of the image of the picture book character and improving the generation quality of the rewritten text content.
[0199] In some embodiments, rewritten text content, information representing style in picture book information, and information representing character characteristics in picture book information are input into an image processing model to obtain rewritten image data, including: inputting the rewritten text content into a text processing model to obtain image generation prompt information corresponding to the rewritten text content; the image generation prompt information is used to represent text content that is compatible with picture book image data generation processing; the text processing model is used to adjust the input rewritten text content and output the adjusted text content; the adjusted text content is image generation prompt information; the image generation prompt information, information representing style, and information representing character characteristics are input into an image processing model to obtain rewritten image data; the rewritten image data represents image data that matches the image generation prompt information, information representing style, and information representing character characteristics; the image processing model is used to perform picture book image data generation processing based on the input image generation prompt information, information representing style, and information representing character characteristics, and output the rewritten image data.
[0200] The image generation prompt information may be a text description that has been adjusted and optimized to be more suitable for image generation. For example, "the child is very happy" may be rewritten into a more specific and detailed descriptive text such as "the child raised his hands high with a smile on his face".
[0201] Specifically, the controller 250 inputs the rewritten text content into the large language model for processing, and converts a simple description such as "the child is very happy" into a detailed prompt word that is more suitable for image generation; the controller 250 obtains information representing the style and information representing the character features from the database, and inputs the converted image generation prompt information, information representing the style, and information representing the character features into the Wenshengtu large model; the controller 250 processes the input information through the consistency attention module CAB (consistency attention module) in the U-Net (U network) diffusion model, uses the cross-attention mechanism in the downsampling layer and the upsampling layer to maintain the consistency of the character features, and finally outputs the rewritten image data.
[0202] The technical solution provided in this embodiment is conducive to improving the accuracy of text-to-image conversion by converting the rewritten text content into prompt information that is more suitable for image generation; by inputting the image generation prompt information together with the information representing the style and the information representing the character characteristics into the image processing model, it is conducive to maintaining the consistency of style and character characteristics during the image generation process; thereby, it is conducive to generating high-quality image data that meets the user's rewriting needs while maintaining the original picture book style and character characteristics.
[0203] In some embodiments, image generation prompt information, information representing style, and information representing character characteristics are input into an image processing model to obtain rewritten image data, including: inputting image generation prompt information and information representing character characteristics into an image processing model to obtain initial rewritten image data; the initial rewritten image data represents image data that matches the image generation prompt information and information representing character characteristics; the image processing model is used to perform picture book image data generation processing based on the input image generation prompt information and information representing character characteristics, and output the initial rewritten image data; the initial rewritten image data and information representing style are input into the image processing model to obtain rewritten image data; the style information of the rewritten image data is the same as the information representing style, and the image content of the rewritten image data is the same as the image content of the initial rewritten image data; the image processing model is used to render the information representing style on the input initial rewritten image data, and output the rewritten image data.
[0204] The initial rewritten image data may be original image data that is generated by an image processing model according to image generation prompt information and information representing character features and has not yet been style rendered.
[0205] Specifically, the controller 250 first inputs the image generation prompt information and the information representing the character features into the Vincent graph large model, and processes them through the consistency attention module CAB (consistency attention module) in the downsampling layer and upsampling layer of the U-Net (U network) diffusion model; the controller 250 selects reference image features containing the same character from the database through the RandSample (random sampling) function, and calculates the cross-attention between the current generated image features and the reference image features in the U-Net (U network) diffusion model to generate initial rewritten image data; the controller 250 inputs the generated initial rewritten image data and the information representing the style into the Vincent graph large model, performs style rendering processing on the initial rewritten image data, and obtains rewritten image data with a specified style.
[0206] The technical solution provided in this embodiment divides the generation process of rewritten image data into two steps. First, initial rewritten image data is generated according to image generation prompt information and information representing character characteristics, which is conducive to ensuring the accuracy of the content of the generated image and the consistency of the character characteristics; then, the initial rewritten image data is subjected to style rendering processing to obtain rewritten image data, which is conducive to achieving style unification while keeping the image content unchanged; thereby, it is conducive to improving the generation quality of the rewritten image data, so that the generated image can accurately express the content and maintain consistency in style.
[0207] In some embodiments, based on the rewritten text content and the information representing the broadcast timbre in the picture book information, a playback audio data synthesis process is performed to obtain the broadcast audio data corresponding to the rewritten image data, and based on the information representing the theme and the information representing the style in the picture book information, a background audio data generation process is performed to obtain the background audio data corresponding to the rewritten image data, including: inputting the rewritten text content and the information representing the broadcast timbre into a playback audio data processing model to obtain the broadcast audio data corresponding to the rewritten image data; the timbre information of the broadcast audio data is the same as the information representing the broadcast timbre, and the text content corresponding to the broadcast audio data is the same as the rewritten text content; the playback audio data processing model is used to The input rewritten text content is converted into corresponding initial broadcast audio data, the initial broadcast audio data is adjusted to represent the broadcast timbre, and the broadcast audio data corresponding to the rewritten image data is output; the information representing the theme and the information representing the style are input into the background audio data processing model to obtain the background audio data corresponding to the rewritten image data; the background audio tag of the background audio data is the background audio tag corresponding to the information representing the theme and the information representing the style; the background audio data processing model is used to generate background audio data based on the background audio tag corresponding to the input information representing the theme and the information representing the style, and output the background audio data corresponding to the rewritten image data.
[0208] Among them, the playback audio data processing model can be an artificial intelligence model for converting text content into speech, for example, it can be a multimodal generation large model.
[0209] The initial broadcast audio data may be voice data after the audio data processing model converts text into voice but has not yet adjusted the timbre, for example, it may be voice data after the text content is converted into a default timbre.
[0210] Among them, the background audio data processing model can be an artificial intelligence model for generating background music, for example, it can be a multimodal generation large model.
[0211] The background audio tag may be identification information for specifying the type of background music, for example, a music style tag such as cheerful or relaxing.
[0212] Specifically, the controller 250 inputs the rewritten text content and the information representing the broadcast timbre into the multimodal generation model, first converts the rewritten text content into initial broadcast audio data, and then performs timbre adjustment processing on the initial broadcast audio data according to the information representing the broadcast timbre, to generate broadcast audio data corresponding to the rewritten image data; the controller 250 simultaneously inputs the information representing the theme and the information representing the style into the multimodal generation model, determines the corresponding background audio tags according to the story theme and style characteristics, generates background audio data matching the rewritten image data through the background audio data processing model, and finally performs audio mixing processing on the broadcast audio data and the background audio data to form a complete audio output.
[0213] The technical solution provided in this embodiment, by separating the generation process of the broadcast audio data and the background audio data, is conducive to separately controlling the timbre characteristics of the broadcast content and the style characteristics of the background music; by performing timbre adjustment processing on the broadcast audio data and performing theme and style label matching processing on the background audio data, it is conducive to ensuring the consistency of the generated audio with the rewritten content, theme and style.
[0214] In some embodiments, reference Fig. 9 The present application also provides a picture book continuation method, which is applied to the display device 200. The method may include the following steps:
[0215] Step S901, when the user interface displays a picture book, receiving the continuation requirement information for the picture book, and using the identifier of the current picture book as the picture book identifier corresponding to the continuation requirement information.
[0216] Step S902, identifying text information corresponding to the continuation demand information, and acquiring picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the continuation demand information; the picture book information represents multi-dimensional information of the picture book corresponding to the picture book identifier.
[0217] Step S903, input the information representing the characters in the text information and the picture book information into the text processing model to obtain the continued text content; the continued text content represents the text content composed of the text information and the information representing the characters; the text processing model is used to generate the picture book text content according to the input text information and the information representing the characters, and output the continued text content.
[0218] Step S904, input the continued text content, the information representing the style in the picture book information, and the information representing the character characteristics in the picture book information into the image processing model to obtain continued image data; the continued image data represents image data that matches the continued text content, the information representing the style, and the information representing the character characteristics; the image processing model is used to generate picture book image data based on the input continued text content, the information representing the style, and the information representing the character characteristics, and output the continued image data.
[0219] Step S905, based on the information representing the broadcast tone in the continued text content and the picture book information, perform playback audio data synthesis processing to obtain the broadcast audio data corresponding to the continued image data, and based on the information representing the theme and the information representing the style in the picture book information, perform background audio data generation processing to obtain the background audio data corresponding to the continued image data.
[0220] Step S906, perform semantic recognition processing on the text information corresponding to the continuation demand information, and obtain the continuation position of the continuation image data in the picture book corresponding to the picture book identifier; add the continuation image data to the continuation position in the picture book corresponding to the picture book identifier, and configure the broadcast audio data and background audio data corresponding to the continuation image data to obtain the continued picture book; the number of pages of the continued picture book is greater than the number of pages of the picture book corresponding to the picture book identifier.
[0221] The continuation demand information may be content expansion request information proposed by a user for an existing picture book. For example, the user may input a continuation request through voice input such as "At the end of the story, Wanwan invented a magical medicine. After taking it, children will never get sick again."
[0222] The continued text content may be new story content generated by a text processing model based on text information input by a user and original character information. For example, it may be a new storyline description generated on the basis of maintaining the original character characteristics.
[0223] Among them, the continued image data can be new image data generated by the image processing model according to the continued text content, style information and character feature information. For example, it can be a picture book image corresponding to the newly added storyline that maintains the original character image and artistic style.
[0224] The continuation position may be the insertion position of the new content in the original picture book determined by semantic recognition, for example, the continuation position may be at the beginning or the end of the picture book.
[0225] The continued picture book may be a new picture book formed by adding continued content to the original picture book. For example, it may be an 11-page picture book formed by adding 1 page of new content at the end of the original 10-page picture book.
[0226] Specifically, the controller 250 receives the user's continuation demand information when the picture book is displayed on the user interface, and assigns a picture book identifier to the continuation demand information; the controller 250 recognizes the text information in the continuation demand information through the large language model, and obtains the picture book information containing the role information, style information and theme information from the database; the controller 250 inputs the text information and the information representing the role into the large language model to generate the continuation text content, and then inputs the continuation text content, the information representing the style and the information representing the role characteristics into the Wenshengtu large model, and the consistency in the downsampling layer and the upsampling layer of the U-Net (U network) diffusion model is obtained. The attention module CAB (consistent attention module) performs processing, uses a random sampling function to select reference image features containing the same character from the database, and generates continued image data by calculating cross-attention; the controller 250 inputs the continued text content and the information representing the broadcast timbre into the multimodal generation model to generate broadcast audio data, and generates background audio data based on the information representing the theme and the information representing the style; finally, the controller 250 determines the insertion position of the continued content through semantic recognition, and adds the continued image data, broadcast audio data and background audio data to the original picture book to form a continued picture book with more pages. Exemplarily, the contents of steps S902, S903, S904 and S905 can refer to the contents of steps S502, S503, S504 and S505 in the above-mentioned picture book rewriting method.
[0227] For example, refer to Fig.10 The display 260 displays multiple picture book stories in the user interface of the picture book application. The multiple picture book stories may include "The Wonderful Journey of Pills", "The Wonderful Life of Kittens", "The Cute World of So-and-so", etc. On this interface, the user can see the introduction of the picture book story, click on AI Picture Book Generation to automatically generate a picture book, and delete the customized picture book by pressing the menu key.
[0228] The technical solution provided in this embodiment, by obtaining multi-dimensional information of the original picture book and using it as a constraint condition for continuation, is conducive to ensuring that the content of the continuation is consistent with the original picture book in terms of character image, artistic style and story theme; by collaboratively generating text content, image data, broadcast audio and background audio, and determining the appropriate continuation position based on semantic recognition, it is conducive to achieving the natural connection of multimodal content, maintaining the coherence and integrity of the story while expanding the content of the picture book; thereby facilitating the automatic continuation of the picture book, reducing manual operation steps, and thus improving the efficiency of picture book continuation.
[0229] In some embodiments, when a picture book is displayed on a user interface, receiving information on a need to continue writing a picture book includes: when a picture book is displayed on the user interface, receiving an input picture book selection operation; when a selected picture book is identified according to the picture book selection operation, receiving information on a need to continue writing a picture book.
[0230] The picture book selection operation may be an operation in which a user selects a specific picture book through an input device, for example, an operation in which a user controls a remote control to place a focus on a certain picture book.
[0231] Specifically, the controller 250 displays a picture book list on the user interface. When receiving a picture book selection operation input by the user through the remote control, the controller 250 focuses on the picture book selected by the user and activates the voice input function; the controller 250 calls the voice receiving module to turn on the receiving microphone and receives the continuation requirement information for the picture book.
[0232] The technical solution provided in this embodiment helps users to accurately locate the specific picture book that needs to be continued by displaying picture books on the user interface and receiving the user's picture book selection operation; and helps to ensure the accurate correspondence between the continuation demand and the specific picture book by receiving the continuation demand information after identifying the picture book selected by the user; thereby helping to establish an orderly picture book continuation operation process, avoid mismatching between the continuation demand and the picture book, and improve the accuracy of picture book continuation.
[0233] In some embodiments, after receiving the input picture book selection operation, it also includes: responding to the picture book selection operation, identifying the picture book corresponding to the current focus; identifying the picture book corresponding to the current focus as the selected picture book.
[0234] The selected picture book may be a specific picture book selected by the user as determined by the current focus position. For example, when the current focus falls on the picture book "Yaoyao's Adventure Journey", it is recognized that the picture book selected by the user is "Yaoyao's Adventure Journey".
[0235] Specifically, the controller 250 displays a picture book list on the user interface. When receiving a picture book selection operation input by the user through the remote control, the controller 250 responds to the picture book selection operation, identifies the specific position of the current focus in the user interface through the focus positioning module, and identifies the picture book corresponding to the current focus from the displayed picture book list based on the position information of the current focus, and identifies it as the selected picture book.
[0236] For example, when a picture book is displayed on the user interface, receiving the information on the need to continue writing a picture book specifically includes the following contents: when a picture book is displayed on the user interface, the controller 250 receives an input picture book selection operation; responds to the picture book selection operation, identifies the picture book corresponding to the current focus; identifies the picture book corresponding to the current focus as the selected picture book; and receives the information on the need to continue writing a picture book when the selected picture book is identified according to the picture book selection operation. By responding to the user's picture book selection operation and identifying the picture book corresponding to the current focus position, it is helpful to establish a clear correspondence between the user operation and the specific picture book; by identifying the picture book corresponding to the current focus as the picture book selected by the user, it is helpful to accurately locate the specific picture book that the user wants to continue writing; thereby, it is helpful to improve the accuracy of the user's picture book selection operation; by receiving the information on the need to continue writing after identifying the picture book selected by the user, it is helpful to ensure the accurate correspondence between the need to continue writing and the specific picture book; thereby, it is helpful to establish an orderly picture book continuation operation process, avoid mismatching between the need to continue writing and the picture book, and improve the accuracy of picture book continuation.
[0237] The technical solution provided in this embodiment is conducive to establishing a clear correspondence between user operations and specific picture books by responding to the user's picture book selection operation and identifying the picture book corresponding to the current focus position; by identifying the picture book corresponding to the current focus as the picture book selected by the user, it is conducive to accurately locating the specific picture book that the user wants to continue writing; thereby helping to improve the accuracy of the user's picture book selection operation.
[0238] In some embodiments, in the continuation position in the picture book corresponding to the picture book identifier, continuation image data is added, and the broadcast audio data and background audio data corresponding to the continuation image data are configured to obtain the continued picture book, including: when the continuation position is the end of the picture book corresponding to the picture book identifier, after the last page in the picture book corresponding to the picture book identifier, the continuation image data is added, and the broadcast audio data and background audio data corresponding to the continuation image data are configured to obtain the continued picture book; when the continuation position is the beginning of the picture book corresponding to the picture book identifier, before the first page in the picture book corresponding to the picture book identifier, the continuation image data is added, and the broadcast audio data and background audio data corresponding to the continuation image data are configured to obtain the continued picture book.
[0239] Specifically, the controller 250 obtains the character information, style information and story theme of the picture book from the database according to the picture book identifier, and determines the content adding method according to the continuation position; when the continuation position is the end of the picture book, the controller 250 adds the continuation image data after the last page of the picture book, and when the continuation position is the beginning of the picture book, the controller 250 adds the continuation image data before the first page of the picture book; when generating the continuation image data, the controller 250 maintains the consistency of the character image through the consistency attention mechanism, and calls the speech synthesis module to generate the broadcast audio data according to the story theme, and calls the music generation module to generate the background audio data according to the picture book style, and finally integrates the continuation image data, the broadcast audio data and the background audio data to form the continued picture book.
[0240] The technical solution provided in this embodiment helps users to select a suitable continuation position according to actual needs by flexibly adding continuation image data at the beginning or end of the picture book; and helps to maintain the consistency of the continuation content with the original picture book in visual and auditory experience by configuring corresponding broadcast audio data and background audio data for the continuation image data.
[0241] In some embodiments, Figure 1 and 2 As shown, the present application also provides a display device 200, which may include a display 260 and a controller 250; wherein the controller 250 is coupled to the display 260. Figure 5 As shown, the controller 250 is configured as follows:
[0242] Step S501, when the user interface displays the picture book content, receiving rewriting requirement information for the picture book content, and using the identifier of the current picture book as the picture book identifier corresponding to the rewriting requirement information;
[0243] Step S502, identifying text information corresponding to the rewriting requirement information, and acquiring picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the rewriting requirement information; the picture book information represents multi-dimensional information of the picture book corresponding to the picture book identifier;
[0244] Step S503, inputting the information representing the characters in the text information and the picture book information into the text processing model to obtain rewritten text content; the rewritten text content represents the text content composed of the text information and the information representing the characters; the text processing model is used to generate the picture book text content according to the input text information and the information representing the characters, and output the rewritten text content;
[0245] Step S504, inputting the rewritten text content, the information representing the style in the picture book information, and the information representing the character characteristics in the picture book information into the image processing model to obtain rewritten image data; the rewritten image data represents image data that matches the rewritten text content, the information representing the style, and the information representing the character characteristics; the image processing model is used to generate picture book image data based on the input rewritten text content, the information representing the style, and the information representing the character characteristics, and output the rewritten image data;
[0246] Step S505, based on the rewritten text content and the information representing the broadcast tone in the picture book information, performing a playback audio data synthesis process to obtain the broadcast audio data corresponding to the rewritten image data, and based on the information representing the theme and the information representing the style in the picture book information, performing a background audio data generation process to obtain the background audio data corresponding to the rewritten image data;
[0247] Step S506, replacing the picture book image data in the picture book content in the picture book corresponding to the picture book identifier with the rewritten image data, and replacing the broadcast audio data and background audio data corresponding to the picture book image data with the broadcast audio data and background audio data corresponding to the rewritten image data, to obtain a rewritten picture book; the number of pages of the rewritten picture book and the number of pages of the picture book corresponding to the picture book identifier remain unchanged.
[0248] It should be noted that, for the specific implementation process of the display device 200, reference can be made to Figure 5 The relevant embodiments of the picture book rewriting method shown will not be repeated here.
[0249] The technical solution provided in this embodiment helps avoid redundant operations of repeatedly acquiring and manually inputting information by associating picture book identification with multi-dimensional information and storing them in a database, and uniformly acquiring information such as characters, styles, and themes based on the picture book identification when rewriting. It helps maintain consistency in characters, styles, themes, and other aspects between the rewritten content and the original picture book by inputting the acquired information into a text processing model, an image processing model, and an audio processing model for automated processing, and accurately replacing the information while keeping the number of pages unchanged. This helps realize automated rewriting of picture books, reduces manual operation steps, and thus improves the efficiency of picture book rewriting.
[0250] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:
[0251] When the user interface displays picture book content, an input picture book content selection operation is received; when the selected picture book content is identified according to the picture book content selection operation, rewriting requirement information for the picture book content is received.
[0252] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:
[0253] In response to a picture book content selection operation, the picture book content corresponding to the current focus is identified; and the picture book content corresponding to the current focus is identified as the selected picture book content.
[0254] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:
[0255] Perform intent recognition processing on the text information corresponding to the rewriting requirement information to obtain the intent recognition result corresponding to the rewriting requirement information; the intent recognition result is used to characterize whether to obtain data from the database; based on the intent recognition result, determine whether to obtain data from the database; when it is determined that the data is to be obtained from the database, obtain the picture book information corresponding to the picture book identifier from the database according to the picture book identifier corresponding to the rewriting requirement information.
[0256] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:
[0257] Picture book image data containing information representing the character is selected from the picture book corresponding to the picture book identifier as reference picture book image data; and information representing the character characteristics is extracted from the reference picture book image data.
[0258] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:
[0259] Input text information into a text processing model to obtain key text information; the key text information represents key information in the text information; the text processing model is used to output the key text information in the text information based on the input text information; input the key text information and information representing the role into the text processing model to obtain rewritten text content; the rewritten text content represents text content composed of key text information and information representing the role; the text processing model is used to generate picture book text content based on the input key text information and information representing the role, and output the rewritten text content.
[0260] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:
[0261] The rewritten text content is input into the text processing model to obtain image generation prompt information corresponding to the rewritten text content; the image generation prompt information is used to represent the text content adapted to the picture book image data generation process; the text processing model is used to adjust the input rewritten text content and output the adjusted text content; the adjusted text content is the image generation prompt information; the image generation prompt information, the information representing the style and the information representing the character characteristics are input into the image processing model to obtain the rewritten image data; the rewritten image data represents the image data matching the image generation prompt information, the information representing the style and the information representing the character characteristics; the image processing model is used to perform picture book image data generation process according to the input image generation prompt information, the information representing the style and the information representing the character characteristics, and output the rewritten image data.
[0262] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:
[0263] Inputting image generation prompt information and information representing character characteristics into an image processing model to obtain initial rewritten image data; the initial rewritten image data represents image data that matches the image generation prompt information and information representing character characteristics; the image processing model is used to generate picture book image data based on the input image generation prompt information and information representing character characteristics, and output the initial rewritten image data; inputting the initial rewritten image data and information representing style into the image processing model to obtain rewritten image data; the style information of the rewritten image data is the same as the information representing the style, and the image content of the rewritten image data is the same as the image content of the initial rewritten image data; the image processing model is used to render the information representing the style on the input initial rewritten image data, and output the rewritten image data.
[0264] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:
[0265] The rewritten text content and the information representing the broadcast timbre are input into the playback audio data processing model to obtain the broadcast audio data corresponding to the rewritten image data; the timbre information of the broadcast audio data is the same as the information representing the broadcast timbre, and the text content corresponding to the broadcast audio data is the same as the rewritten text content; the playback audio data processing model is used to convert the input rewritten text content into the corresponding initial broadcast audio data, adjust the initial broadcast audio data with the information representing the broadcast timbre, and output the broadcast audio data corresponding to the rewritten image data; the information representing the theme and the information representing the style are input into the background audio data processing model to obtain the background audio data corresponding to the rewritten image data; the background audio tag of the background audio data is the background audio tag corresponding to the information representing the theme and the information representing the style; the background audio data processing model is used to generate background audio data according to the background audio tag corresponding to the input information representing the theme and the information representing the style, and output the background audio data corresponding to the rewritten image data.
[0266] In some embodiments, Figure 1 and 2 As shown, the present application also provides a display device 200, which may include a display 260 and a controller 250; wherein the controller 250 is coupled to the display 260. Fig. 9 As shown, the controller 250 is configured as follows:
[0267] Step S901, when the user interface displays a picture book, receiving the continuation requirement information for the picture book, and using the identifier of the current picture book as the picture book identifier corresponding to the continuation requirement information.
[0268] Step S902, identifying text information corresponding to the continuation demand information, and acquiring picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the continuation demand information; the picture book information represents multi-dimensional information of the picture book corresponding to the picture book identifier.
[0269] Step S903, input the information representing the characters in the text information and the picture book information into the text processing model to obtain the continued text content; the continued text content represents the text content composed of the text information and the information representing the characters; the text processing model is used to generate the picture book text content according to the input text information and the information representing the characters, and output the continued text content.
[0270] Step S904, input the continued text content, the information representing the style in the picture book information, and the information representing the character characteristics in the picture book information into the image processing model to obtain continued image data; the continued image data represents image data that matches the continued text content, the information representing the style, and the information representing the character characteristics; the image processing model is used to generate picture book image data based on the input continued text content, the information representing the style, and the information representing the character characteristics, and output the continued image data.
[0271] Step S905, based on the information representing the broadcast tone in the continued text content and the picture book information, perform playback audio data synthesis processing to obtain the broadcast audio data corresponding to the continued image data, and based on the information representing the theme and the information representing the style in the picture book information, perform background audio data generation processing to obtain the background audio data corresponding to the continued image data.
[0272] Step S906, perform semantic recognition processing on the text information corresponding to the continuation demand information, and obtain the continuation position of the continuation image data in the picture book corresponding to the picture book identifier; add the continuation image data to the continuation position in the picture book corresponding to the picture book identifier, and configure the broadcast audio data and background audio data corresponding to the continuation image data to obtain the continued picture book; the number of pages of the continued picture book is greater than the number of pages of the picture book corresponding to the picture book identifier.
[0273] It should be noted that, for the specific implementation process of the display device 200, reference can be made to Fig. 9 The relevant embodiments of the picture book continuation method shown will not be repeated here.
[0274] The technical solution provided in this embodiment, by obtaining multi-dimensional information of the original picture book and using it as a constraint condition for continuation, is conducive to ensuring that the content of the continuation is consistent with the original picture book in terms of character image, artistic style and story theme; by collaboratively generating text content, image data, broadcast audio and background audio, and determining the appropriate continuation position based on semantic recognition, it is conducive to achieving the natural connection of multimodal content, maintaining the coherence and integrity of the story while expanding the content of the picture book; thereby facilitating the automatic continuation of the picture book, reducing manual operation steps, and thus improving the efficiency of picture book continuation.
[0275] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:
[0276] When a picture book is displayed on a user interface, an input picture book selection operation is received; in response to the picture book selection operation, a picture book corresponding to a current focus is identified; the picture book corresponding to the current focus is identified as the selected picture book; and when the selected picture book is identified according to the picture book selection operation, continuation requirement information for the picture book is received.
[0277] In some embodiments, the controller 250 of the display device 200 may further perform the following steps:
[0278] When the continuation position is the end of the picture book corresponding to the picture book identifier, the continuation image data is added after the last page of the picture book corresponding to the picture book identifier, and the broadcast audio data and background audio data corresponding to the continuation image data are configured to obtain the continued picture book; when the continuation position is the beginning of the picture book corresponding to the picture book identifier, the continuation image data is added before the first page of the picture book corresponding to the picture book identifier, and the broadcast audio data and background audio data corresponding to the continuation image data are configured to obtain the continued picture book.
[0279] In some embodiments, Fig.11 As shown, in order to more clearly describe the signaling interaction process between the modules of the display device 200, the present application also provides another picture book rewriting method, which may include the following steps:
[0280] Step 1: The picture book control in the control center receives the content input by the user.
[0281] Step 2: The picture book central control initiates a data pull request to the database (including a character information pull request, a story theme pull request, a style information pull request, and a broadcast tone id pull request).
[0282] Step 3: The database returns character information, story theme, style information, character reference features and broadcast tone ID to the picture book central control.
[0283] Step 4: The picture book central control sends the user input content and the character information obtained from the database to the story rewriting / continuation module in the multimodal generation model.
[0284] Step 5: The story rewriting / continuation module processes the received information and generates corresponding story content.
[0285] Step 6: The story rewriting / continuation module sends the generated story content to the text and picture rewriting module, and can also send the story content to the picture book central control.
[0286] Step 7: The text-to-raw picture rewriting module converts the received story content into prompt words suitable for raw pictures.
[0287] Step 8: The text-image rewriting module sends the converted prompt words to the text-image module.
[0288] Step 9: The text image module generates a composite image according to the prompt word (in which the style model can be switched), and returns the composite image to the picture book control.
[0289] Step 10: The speech synthesis module receives the announcement tone ID sent by the picture book central control, generates a speech announcement and returns it to the picture book central control.
[0290] Step 11, the music generation module receives the story theme information and style information sent by the picture book central control, generates background music based on the story theme information and style information, and returns it to the picture book central control.
[0291] Step 12: The picture book central control stores the key information into the database (updates the story content) and returns the image / broadcast audio / background audio to the user.
[0292] The technical solution provided in this embodiment helps avoid redundant operations of repeatedly acquiring and manually inputting information by associating picture book identification with multi-dimensional information and storing them in a database, and uniformly acquiring information such as characters, styles, and themes based on the picture book identification when rewriting. It helps maintain consistency in characters, styles, themes, and other aspects between the rewritten content and the original picture book by inputting the acquired information into a text processing model, an image processing model, and an audio processing model for automated processing, and accurately replacing the information while keeping the number of pages unchanged. This helps realize automated rewriting of picture books, reduces manual operation steps, and thus improves the efficiency of picture book rewriting.
[0293] In some embodiments, in order to more clearly illustrate the picture book rewriting method provided in the embodiments of the present application, the picture book rewriting method is specifically described below with a specific embodiment. The present application also provides another picture book rewriting method.
[0294] In order to meet the personalized needs of users, it is necessary to grant users the permission to edit and rewrite picture books. For example, if the user feels that the content of the 9th sentence of the picture book does not meet his expectations, he can rewrite the story according to his needs. In addition, the user is allowed to continue the story according to the plot.
[0295] In terms of the interactive link, in the TV scenario, implementing this function requires voice recognition and a large language model to understand the user's intention, and then send the rewrite or continuation request to the generation service.
[0296] For the generation service, it is necessary to ensure that the current rewrite or continued generated image maintains the same characters and style as the original story. The algorithm can use a conditional control solution or a Train Free (train-free) inter-frame Cross Attention solution. This task requires the requester to provide story context information or conditional information to control generation in addition to the current continued or rewritten content. One approach is to save the story generation dependencies in a unified database system and give the story a unique identifier, so that users can edit and rewrite the story after a period of time.
[0297] The difficulty of the continuation and rewriting tasks for the text graph model lies in maintaining the consistency of the generated character attributes. This part involves Train Free or conditional control technology:
[0298] (1) Train Free: If you choose a link like Story Diffusion, you need to save the feature input of different roles at different steps in the early story generation through the Attention layer in the U-shaped network. ij ∈F, where i traverses 1-N, j traverses 1-M, N is the number of generation steps, and M is the number of model layers. The subsequent continuation is the current step, and the output of the current layer will be the same as the reference f ij Use Cross Attention to ensure the consistency between the output image and the overall story.
[0299] (2) Conditional control technology, such as Instance ID, randomly selects pictures in the story that contain the image of the current rewritten character as a condition to control the generation of the rewritten picture book.
[0300] The user-operable positions of the picture book rewrite can be referenced Figure 7 , the user clicks the remote control to focus on a certain picture, and can now input voice input such as: "Rewrite the content of this story as the child bursts into tears and knocks the medicine bottle to the ground."
[0301] The picture book application (APK, Android Package) will package the unique identifier of the story and the corresponding modified content and send them to the picture book central control.
[0302] The trigger position for continuing writing picture books can be referenced Fig.10 The user controls the story through the remote control, focuses on a certain story, and then inputs voice input such as "At the end of the story, Yaowanwan invented a magical medicine. After taking it, children will never get sick again."
[0303] The picture book application will package the unique identifier of the story and the user request and send them to the picture book central control.
[0304] The following are the module components in the timing diagram without continuation and rewriting:
[0305] 1. APP (application) control module: user interaction interface, including perception (receiver button, text input window) and expression components (audio and video alignment, display broadcast).
[0306] 2. Voice reception: call up the microphone to receive user audio.
[0307] 3. Speech recognition: Speech recognition module converts speech into text.
[0308] 4. Compliance detection: A large language model is used to identify whether the user input contains sensitive words. If so, the input is rejected and the user is prompted.
[0309] 5. Role extraction: A large language model is used to give appropriate roles and role attributes based on user input. If the user does not provide a role, the corresponding role and attributes are generated based on the story theme.
[0310] 6. Theme extraction: The large language model gives the appropriate story theme based on user requests, and the overall theme style is positive.
[0311] 7. Style extraction: The large language model specifies the corresponding picture book style, such as "cartoon", "realistic", "hand-painted", "watercolor", etc., based on user input and story theme, to make it easier to interpret the story in the picture-to-text process.
[0312] 8. Recommended voice for announcements: For the large language model, the default is "little girl". If the theme of the story is more courageous and positive, "little boy" can be recommended. Other voice recommendations (system integrated voice) can also be given based on the theme of the story, such as "mature man" and "mature woman".
[0313] 9. Story generation: A large language model generates story content and story chapters based on user input (usually 10-15 pages, with each page containing about 20 words).
[0314] 10. Text-to-image rewriting: The large language model rewrites each page of the story into prompt words suitable for image generation, such as "the child is very happy" is rewritten as "the child raised his hands high with a smile on his face."
[0315] 11. Wenshengtu: A large model of Wenshengtu, which rewrites prompt words and recommended styles according to Wenshengtu and generates pictures.
[0316] 12. Speech synthesis: Generate a large multimodal model to convert story generation content into speech broadcast.
[0317] 13. Music generation: A multimodal large model is used to generate appropriate background music based on the recommended style and story theme.
[0318] When including continuation and rewriting, the following modules need to be added:
[0319] 1. Picture book central control: Receive user needs from the application, send them to the domain intent recognition module, decide whether to retrieve data from the database based on the intent classification results, collect the results returned by each subtask, send the results to the application, and store character information, story theme, story content and other information in the database system.
[0320] 2. Continuation and rewriting of story picture books: Since user input is generally random, key information needs to be extracted. This module is used to extract user key information, and then integrate the extracted information with the pulled database information to output the overall content of the story.
[0321] 3. Database: used to store character information, style information, story theme, story content, broadcast tone identification, character reference features (used to maintain the consistency of generated characters, which are character attribute-related features extracted from the text graph model level).
[0322] For the specific process, please refer to Fig.11 , Fig.12 and Fig.13 .
[0323] Fig.11 The information flow of rewriting and rewriting branch graphs is shown. Please refer to the above Fig.11 The description will not be repeated here.
[0324] Fig.12 The information flow of the overall architecture diagram is shown, and the information flow is as follows: the user opens the APP, the APP control module calls the voice monitoring, the user sends the far-field / near-field voice to the voice receiving module in the information perception module, the voice receiving module converts the voice into text, and sends it to the voice recognition module, the voice recognition module updates the text input area of the APP control module, the APP control module sends the content entered by the user to the picture book central control in the control center, if the picture book central control detects non-compliant content, it will feedback the non-compliant status to the APP control module, if the picture book central control detects compliance and is judged to be generated normally, it will be processed by the main branch in the story generation branch, if it is judged to be rewritten / continued, it will be processed by the rewrite and continuation branch, the subsequent main branch or rewrite and continuation branch will return the image / broadcast audio / background audio to the picture book central control, the picture book central control returns the image / broadcast audio / background audio to the APP control module, and the APP control module performs audio and video alignment and picture book broadcast synthesis processing.
[0325] Fig.13The information flow of the main branch graph is shown: the content input by the user is first sent to the picture book control center in the control center, and the picture book control center stores the key information in the database. The language model key information extraction / completion module includes modules such as role extraction, theme extraction, style extraction, and broadcast tone recommendation. The role extraction module generates role name and role attribute information; the theme extraction module generates story theme information; the style extraction module generates recommended style; the broadcast tone recommendation module generates recommended tone according to the content input by the user, and returns the broadcast tone information to the picture book control center. This information is sent to the multimodal generation model. The story generation module sends the story content to the Wenshengtu rewriting module, and the Wenshengtu rewriting module converts the story content into prompt words suitable for the raw picture and sends it to the Wenshengtu module. The Wenshengtu module (which can switch the style model) generates a synthetic image based on the prompt words and recommended style and returns it to the picture book control center, and can also return the role reference features to the picture book control center. The speech synthesis module receives the recommended tone, generates a synthetic speech broadcast and returns it to the picture book control center. The music generation module generates background music based on the recommended style and story theme information and returns it to the picture book control center. The database is also responsible for unique identification generation, character information storage, theme information storage, style information storage, timbre information storage, story content storage, and character reference feature storage. Finally, the picture book control center returns the image, broadcast audio, and background audio to the user.
[0326] Among them, at the model level: the main solution to maintain character consistency is that for a certain character, a frame is randomly selected as a reference frame in the text generation process, and the generation of all other characters of the same character refers to the generation process characteristics of this frame, and the reference method is the cross-attention mechanism.
[0327] The following is a systematic introduction to the model's cross-attention mechanism:
[0328] The consistent attention mechanism is used to maintain the consistency of roles within an image batch. This attention mechanism does not require training and can be plugged in and out of the U-Net diffusion model. When used, the consistent attention is directly inserted into the original attention position of the U-Net architecture in the diffusion model, and the original self-attention weight is reused, so no training is required.
[0329] Given a batch of image features I∈R B×N×C , where B, N, and C are the batch size, the number of tokens in each image, and the number of channels, respectively. Define a function Attention(X q ,X k ,X v ) to calculate self-attention. Xq, X k , X v Denote the query Q, key K, and value V used in the attention calculation respectively. The original self-attention is calculated on each image feature I iThe features I are independently performed, and there is no mutual reference between batch images. i Projection to Q i , K i , V i , and feed it into the attention function to get:
[0330]
[0331] To establish interactions between images within a batch to maintain topic consistency, consistent attention randomly samples feature tokens of some reference frames from other images in the batch (saving all previously generated image features).
[0332] S i =RandSample(I1,I2,…I i-1 ,I i+1 ,I B-1 ,I B )
[0333] RandSample represents a random sampling function. After sampling, the sampled tokens (reference frame features ref) are compared with the image features I i Pair to form a new tokens set P i Then for P i Perform linear projection to generate new keys and consistent attention value Here, the original Q i The query does not change. Finally, the self-attention is calculated as follows:
[0334]
[0335]
[0336]
[0337] in, is the weight of the reference image feature and the image feature, K ref 、V ref The key and value of the reference feature.
[0338] Given paired tokens, self-attention is performed on image batches, promoting interactions between different image features. This type of interaction promotes the convergence of the model to characters, faces, and clothing during generation. Although it is simple and training-free, consistent self-attention is effective in maintaining the consistency of the main characters.
[0339] For example, at the technical implementation level, the solution adopts a text-to-image diffusion model architecture with a consistent self-attention mechanism. Consistent attention modules are added to the downsampling and upsampling layers of the diffusion U-network. The structure of the consistent self-attention module contains a linear layer and corresponding dimensional information. The entire processing flow includes a connection module, a labeled sampling module, a multiplication operation module, a soft maximization processing module, a cross attention module, a normalization processing module, a consistent self-attention module, a residual network module, and a consistent layer attention module. These modules work together to ensure that the consistency of the generated content can be maintained throughout the entire processing process from input to output. In the specific implementation, the consistent attention mechanism is used to solve the problem of maintaining role consistency within an image batch. This mechanism does not require training and can be plug-and-play in the diffusion model. It is implemented by inserting consistent attention at the original attention position and reusing the original self-attention weights. The system maintains thematic consistency by establishing interactions between images within the batch, randomly sampling features of the reference frame from other images in the batch, and pairing these features with the current image features to form a new feature set. This method promotes the interaction between different image features and effectively maintains the consistency of people, faces, and clothing during the generation process.
[0340] The above-mentioned embodiments are conducive to realizing the automatic updating of picture book stories and reducing the manual operation steps, thereby facilitating improving the updating efficiency of picture book stories.
[0341] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0342] Based on the same inventive concept, the embodiment of the present application also provides a picture book rewriting device and a picture book continuation device for implementing the picture book rewriting method and the picture book continuation method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more picture book rewriting device embodiments and picture book continuation device embodiments provided below can refer to the limitations of the picture book rewriting method and the picture book continuation method above, and will not be repeated here.
[0343] In an exemplary embodiment, Fig.14 As shown, a picture book rewriting device 1400 is provided, which can be applied to Figure 1 The display device 200 shown in FIG. Figure 2 As shown, the display device 200 may include a display 260 and a controller 250; the controller 250 is coupled to the display 260; the device may include:
[0344] The first receiving module 1401 is used to receive rewriting requirement information for picture book content when the user interface displays the picture book content, and use the identifier of the current picture book as the picture book identifier corresponding to the rewriting requirement information;
[0345] The first acquisition module 1402 is used to identify the text information corresponding to the rewriting requirement information, and acquire the picture book information corresponding to the picture book identifier from the database according to the picture book identifier corresponding to the rewriting requirement information; the picture book information represents the multi-dimensional information of the picture book corresponding to the picture book identifier;
[0346] The first input module 1403 is used to input the information representing the characters in the text information and the picture book information into the text processing model to obtain the rewritten text content; the rewritten text content represents the text content composed of the text information and the information representing the characters; the text processing model is used to generate the picture book text content according to the input text information and the information representing the characters, and output the rewritten text content;
[0347] The second input module 1404 is used to input the rewritten text content, the information representing the style in the picture book information, and the information representing the character characteristics in the picture book information into the image processing model to obtain rewritten image data; the rewritten image data represents image data that matches the rewritten text content, the information representing the style, and the information representing the character characteristics; the image processing model is used to generate the picture book image data according to the input rewritten text content, the information representing the style, and the information representing the character characteristics, and output the rewritten image data;
[0348] The first synthesis module 1405 is used to perform a synthesis process on the playback audio data based on the rewritten text content and the information representing the broadcast tone in the picture book information to obtain the broadcast audio data corresponding to the rewritten image data, and to perform a background audio data generation process based on the information representing the theme and the information representing the style in the picture book information to obtain the background audio data corresponding to the rewritten image data;
[0349] The data replacement module 1406 is used to replace the picture book image data in the picture book content in the picture book corresponding to the picture book identifier with the rewritten image data, and to replace the broadcast audio data and background audio data corresponding to the picture book image data with the broadcast audio data and background audio data corresponding to the rewritten image data, so as to obtain a rewritten picture book; the number of pages of the rewritten picture book and the number of pages of the picture book corresponding to the picture book identifier remain unchanged.
[0350] In an exemplary embodiment, Fig.15 As shown, a picture book continuation device 1500 is provided, which can be applied to Figure 1 The display device 200 shown in FIG. Figure 2 As shown, the display device 200 may include a display 260 and a controller 250; the controller 250 is coupled to the display 260; the device may include:
[0351] The second receiving module 1501 is used to receive the continuation requirement information for the picture book when the picture book is displayed on the user interface, and use the identifier of the current picture book as the picture book identifier corresponding to the continuation requirement information;
[0352] The second acquisition module 1502 is used to identify the text information corresponding to the continuation demand information, and acquire the picture book information corresponding to the picture book identifier from the database according to the picture book identifier corresponding to the continuation demand information; the picture book information represents the multi-dimensional information of the picture book corresponding to the picture book identifier;
[0353] The third input module 1503 is used to input the information representing the characters in the text information and the picture book information into the text processing model to obtain the continued text content; the continued text content means the text content composed of the text information and the information representing the characters; the text processing model is used to generate the picture book text content according to the input text information and the information representing the characters, and output the continued text content;
[0354] The fourth input module 1504 is used to input the continued text content, the information representing the style in the picture book information, and the information representing the character characteristics in the picture book information into the image processing model to obtain the continued image data; the continued image data represents the image data matching the continued text content, the information representing the style, and the information representing the character characteristics; the image processing model is used to generate the picture book image data according to the input continued text content, the information representing the style, and the information representing the character characteristics, and output the continued image data;
[0355] The second synthesis module 1505 is used to perform a synthesis process on the playback audio data based on the information representing the broadcast tone in the continued text content and the picture book information to obtain the broadcast audio data corresponding to the continued image data, and to perform a background audio data generation process based on the information representing the theme and the information representing the style in the picture book information to obtain the background audio data corresponding to the continued image data;
[0356] The data adding module 1506 is used to perform semantic recognition processing on the text information corresponding to the continuation demand information to obtain the continuation position of the continuation image data in the picture book corresponding to the picture book identifier; add the continuation image data to the continuation position in the picture book corresponding to the picture book identifier, and configure the broadcast audio data and background audio data corresponding to the continuation image data to obtain the continued picture book; the number of pages of the continued picture book is greater than the number of pages of the picture book corresponding to the picture book identifier.
[0357] Each module in the above picture book rewriting device and picture book continuation device can be implemented in whole or in part by software, hardware and their combination. Each module can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each module above.
[0358] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0359] In some embodiments, a computer program product is provided, including a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.
[0360] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0361] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited thereto.
[0362] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0363] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A picture book rewriting method, characterized in that: Applied to display devices, including: In the case where the user interface displays picture book content, receiving rewriting requirement information for the picture book content, and using the identifier of the current picture book as the picture book identifier corresponding to the rewriting requirement information; Identify the text information corresponding to the rewriting requirement information, and acquire the picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the rewriting requirement information; the picture book information represents multi-dimensional information of the picture book corresponding to the picture book identifier; Inputting the text information and the information representing the characters in the picture book information into a text processing model to obtain rewritten text content; the rewritten text content represents text content composed of the text information and the information representing the characters; the text processing model is used to perform picture book text content generation processing based on the input text information and the information representing the characters, and output the rewritten text content; Inputting the rewritten text content, the information representing the style in the picture book information, and the information representing the character characteristics in the picture book information into an image processing model to obtain rewritten image data; the rewritten image data represents image data that matches the rewritten text content, the information representing the style, and the information representing the character characteristics; the image processing model is used to perform picture book image data generation processing based on the input rewritten text content, the information representing the style, and the information representing the character characteristics, and output the rewritten image data; Based on the rewritten text content and the information representing the broadcast tone in the picture book information, a playback audio data synthesis process is performed to obtain the broadcast audio data corresponding to the rewritten image data; and based on the information representing the theme and the information representing the style in the picture book information, a background audio data generation process is performed to obtain the background audio data corresponding to the rewritten image data; The picture book image data in the picture book content in the picture book corresponding to the picture book identifier is replaced by the rewritten image data, and the broadcast audio data and background audio data corresponding to the picture book image data are replaced by the broadcast audio data and background audio data corresponding to the rewritten image data, to obtain a rewritten picture book; the number of pages of the rewritten picture book and the number of pages of the picture book corresponding to the picture book identifier remain unchanged.
2. The method according to claim 1, characterized in that In the case where the picture book content is displayed on the user interface, receiving rewriting requirement information for the picture book content includes: When the user interface displays picture book content, receiving an input picture book content selection operation; When the selected picture book content is identified according to the picture book content selection operation, rewriting requirement information for the picture book content is received.
3. The method according to claim 2, characterized in that After receiving the input picture book content selection operation, it also includes: In response to the picture book content selection operation, identifying the picture book content corresponding to the current focus; The picture book content corresponding to the current focus is identified as the selected picture book content.
4. The method according to claim 1, characterized in that Before acquiring picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the rewriting requirement information, the method further includes: Performing intent recognition processing on the text information corresponding to the rewriting requirement information to obtain an intent recognition result corresponding to the rewriting requirement information; the intent recognition result is used to indicate whether to obtain data from the database; Based on the intention recognition result, determining whether to obtain data from the database; The acquiring, from a database, picture book information corresponding to the picture book identifier according to the rewriting requirement information, includes: When it is determined that the data is to be obtained from the database, the picture book information corresponding to the picture book identifier is obtained from the database according to the picture book identifier corresponding to the rewriting requirement information.
5. The method according to claim 1, characterized in that The information characterizing the character characteristics is also obtained in the following manner: Selecting picture book image data containing the information representing the character from the picture book corresponding to the picture book identifier as reference picture book image data; The information representing the characteristics of the character is extracted from the reference picture book image data.
6. The method according to claim 1, characterized in that The step of inputting the text information and the information representing the characters in the picture book information into a text processing model to obtain rewritten text content includes: Inputting the text information into a text processing model to obtain key text information; the key text information represents key information in the text information; the text processing model is used to output the key text information in the text information according to the input text information; The key text information and the information representing the character are input into the text processing model to obtain the rewritten text content; the rewritten text content represents the text content composed of the key text information and the information representing the character; the text processing model is used to generate picture book text content based on the input key text information and the information representing the character, and output the rewritten text content.
7. The method according to claim 1, characterized in that The step of inputting the rewritten text content, the information representing the style in the picture book information, and the information representing the character features in the picture book information into an image processing model to obtain the rewritten image data includes: The rewritten text content is input into the text processing model to obtain image generation prompt information corresponding to the rewritten text content; the image generation prompt information is used to represent the text content adapted to the picture book image data generation process; the text processing model is used to adjust the input rewritten text content and output the adjusted text content; the adjusted text content is the image generation prompt information; The image generation prompt information, the information representing the style and the information representing the character characteristics are input into an image processing model to obtain rewritten image data; the rewritten image data represents image data that matches the image generation prompt information, the information representing the style and the information representing the character characteristics; the image processing model is used to perform picture book image data generation processing based on the input image generation prompt information, the information representing the style and the information representing the character characteristics, and output the rewritten image data.
8. The method according to claim 7, characterized in that The step of inputting the image generation prompt information, the information representing the style, and the information representing the character characteristics into an image processing model to obtain rewritten image data includes: Inputting the image generation prompt information and the information representing the character characteristics into an image processing model to obtain initial rewritten image data; the initial rewritten image data represents image data that matches the image generation prompt information and the information representing the character characteristics; the image processing model is used to perform picture book image data generation processing based on the input image generation prompt information and the information representing the character characteristics, and output the initial rewritten image data; The initial rewritten image data and the information representing the style are input into the image processing model to obtain the rewritten image data; the style information of the rewritten image data is the same as the information representing the style, and the image content of the rewritten image data is the same as the image content of the initial rewritten image data; the image processing model is used to render the information representing the style on the input initial rewritten image data, and output the rewritten image data.
9. The method according to claim 1, characterized in that: The synthesizing process of playing audio data based on the rewritten text content and the information representing the broadcast tone in the picture book information to obtain the broadcast audio data corresponding to the rewritten image data, and the generating process of background audio data based on the information representing the theme and the information representing the style in the picture book information to obtain the background audio data corresponding to the rewritten image data include: The rewritten text content and the information representing the broadcast timbre are input into a playback audio data processing model to obtain the broadcast audio data corresponding to the rewritten image data; the timbre information of the broadcast audio data is the same as the information representing the broadcast timbre, and the text content corresponding to the broadcast audio data is the same as the rewritten text content; the playback audio data processing model is used to convert the input rewritten text content into corresponding initial broadcast audio data, adjust the information representing the broadcast timbre on the initial broadcast audio data, and output the broadcast audio data corresponding to the rewritten image data; The information representing the theme and the information representing the style are input into a background audio data processing model to obtain background audio data corresponding to the rewritten image data; the background audio tag of the background audio data is the background audio tag corresponding to the information representing the theme and the information representing the style; the background audio data processing model is used to generate background audio data based on the background audio tag corresponding to the input information representing the theme and the information representing the style, and output the background audio data corresponding to the rewritten image data.
10. A method for continuing a picture book, characterized in that: Applied to display devices, including: In the case where the user interface displays a picture book, receiving the continuation requirement information for the picture book, and using the identifier of the current picture book as the picture book identifier corresponding to the continuation requirement information; Identify the text information corresponding to the continuation demand information, and acquire the picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the continuation demand information; the picture book information represents multi-dimensional information of the picture book corresponding to the picture book identifier; Inputting the text information and the information representing the characters in the picture book information into a text processing model to obtain a continued text content; the continued text content represents the text content composed of the text information and the information representing the characters; the text processing model is used to perform picture book text content generation processing based on the input text information and the information representing the characters, and output the continued text content; Inputting the continued text content, the information representing the style in the picture book information, and the information representing the character characteristics in the picture book information into an image processing model to obtain continued image data; the continued image data represents image data that matches the continued text content, the information representing the style, and the information representing the character characteristics; the image processing model is used to perform picture book image data generation processing based on the input continued text content, the information representing the style, and the information representing the character characteristics, and output the continued image data; Based on the content of the continued text and the information representing the tone of the announcement in the picture book information, a playback audio data synthesis process is performed to obtain the announcement audio data corresponding to the continued image data; and based on the information representing the theme and the information representing the style in the picture book information, a background audio data generation process is performed to obtain the background audio data corresponding to the continued image data; Perform semantic recognition processing on the text information corresponding to the continuation demand information to obtain the continuation position of the continuation image data in the picture book corresponding to the picture book identifier; add the continuation image data to the continuation position in the picture book corresponding to the picture book identifier, and configure the broadcast audio data and background audio data corresponding to the continuation image data to obtain a continued picture book; the number of pages of the continued picture book is greater than the number of pages of the picture book corresponding to the picture book identifier.
11. The method according to claim 10, characterized in that In the case where the picture book is displayed on the user interface, receiving the continuation requirement information for the picture book includes: When the user interface displays a picture book, receiving an input picture book selection operation; In response to the picture book selection operation, identifying the picture book corresponding to the current focus; Identify the picture book corresponding to the current focus as the selected picture book; When the selected picture book is identified according to the picture book selection operation, continuation requirement information for the picture book is received.
12. The method according to claim 10, characterized in that The method of adding the continued writing image data to the continued writing position in the picture book corresponding to the picture book identifier, and configuring the broadcast audio data and background audio data corresponding to the continued writing image data to obtain the continued writing picture book includes: When the continuing position is the end of the picture book corresponding to the picture book identifier, the continuing image data is added after the last page of the picture book corresponding to the picture book identifier, and the broadcasting audio data and background audio data corresponding to the continuing image data are configured to obtain a continuing picture book; When the continuation position is the beginning of the picture book corresponding to the picture book identifier, the continuation image data is added before the first page of the picture book corresponding to the picture book identifier, and the broadcast audio data and background audio data corresponding to the continuation image data are configured to obtain the continued picture book.
13. A display device, characterized in that: include: monitor; A controller is coupled to the display and is configured to: In the case where the user interface displays picture book content, receiving rewriting requirement information for the picture book content, and using the identifier of the current picture book as the picture book identifier corresponding to the rewriting requirement information; Identify the text information corresponding to the rewriting requirement information, and acquire the picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the rewriting requirement information; the picture book information represents multi-dimensional information of the picture book corresponding to the picture book identifier; Inputting the text information and the information representing the characters in the picture book information into a text processing model to obtain rewritten text content; the rewritten text content represents text content composed of the text information and the information representing the characters; the text processing model is used to perform picture book text content generation processing based on the input text information and the information representing the characters, and output the rewritten text content; Inputting the rewritten text content, the information representing the style in the picture book information, and the information representing the character characteristics in the picture book information into an image processing model to obtain rewritten image data; the rewritten image data represents image data that matches the rewritten text content, the information representing the style, and the information representing the character characteristics; the image processing model is used to perform picture book image data generation processing based on the input rewritten text content, the information representing the style, and the information representing the character characteristics, and output the rewritten image data; Based on the rewritten text content and the information representing the broadcast tone in the picture book information, a playback audio data synthesis process is performed to obtain the broadcast audio data corresponding to the rewritten image data; and based on the information representing the theme and the information representing the style in the picture book information, a background audio data generation process is performed to obtain the background audio data corresponding to the rewritten image data; The picture book image data in the picture book content in the picture book corresponding to the picture book identifier is replaced by the rewritten image data, and the broadcast audio data and background audio data corresponding to the picture book image data are replaced by the broadcast audio data and background audio data corresponding to the rewritten image data, to obtain a rewritten picture book; the number of pages of the rewritten picture book and the number of pages of the picture book corresponding to the picture book identifier remain unchanged.
14. A display device, characterized in that: include: monitor; A controller is coupled to the display and is configured to: In the case where the user interface displays a picture book, receiving the continuation requirement information for the picture book, and using the identifier of the current picture book as the picture book identifier corresponding to the continuation requirement information; Identify the text information corresponding to the continuation demand information, and acquire the picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the continuation demand information; the picture book information represents multi-dimensional information of the picture book corresponding to the picture book identifier; Inputting the text information and the information representing the characters in the picture book information into a text processing model to obtain a continued text content; the continued text content represents the text content composed of the text information and the information representing the characters; the text processing model is used to perform picture book text content generation processing based on the input text information and the information representing the characters, and output the continued text content; Inputting the continued text content, the information representing the style in the picture book information, and the information representing the character characteristics in the picture book information into an image processing model to obtain continued image data; the continued image data represents image data that matches the continued text content, the information representing the style, and the information representing the character characteristics; the image processing model is used to perform picture book image data generation processing based on the input continued text content, the information representing the style, and the information representing the character characteristics, and output the continued image data; Based on the content of the continued text and the information representing the tone of the announcement in the picture book information, a playback audio data synthesis process is performed to obtain the announcement audio data corresponding to the continued image data; and based on the information representing the theme and the information representing the style in the picture book information, a background audio data generation process is performed to obtain the background audio data corresponding to the continued image data; Perform semantic recognition processing on the text information corresponding to the continuation demand information to obtain the continuation position of the continuation image data in the picture book corresponding to the picture book identifier; add the continuation image data to the continuation position in the picture book corresponding to the picture book identifier, and configure the broadcast audio data and background audio data corresponding to the continuation image data to obtain a continued picture book; the number of pages of the continued picture book is greater than the number of pages of the picture book corresponding to the picture book identifier.
Citation Information
Patent Citations
Electronic picture book generation method and device and electronic equipment
CN114693844A
Interactive story picture book generation method and device, electronic equipment and storage medium
CN117877052A
Digital picture book identification method and system, electronic equipment and storage medium
CN118885969A
User Customized Animated Video and Method For Making the Same
US20110064388A1