Electronic reading data processing method and device, electronic equipment and medium

By using a large language model to convert e-reading data into audio and images, high-quality promotional videos are generated, solving the problem that traditional novel promotion methods struggle to attract attention, improving video generation speed and quality, and providing a better user experience.

CN122269095APending Publication Date: 2026-06-23BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2024-12-20
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Traditional methods of promoting novels struggle to attract and retain the attention of potential readers. With the rise of social media and video platforms, traditional methods of promoting printed novels face challenges.

Method used

The process involves generating promotional videos from e-reading data using a large language model, including converting text into audio and script text, generating corresponding images, and finally synthesizing high-quality promotional videos.

Benefits of technology

It improved the speed and quality of promotional video generation, provided a more intuitive and engaging reading experience, and enhanced user appeal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122269095A_ABST
    Figure CN122269095A_ABST
Patent Text Reader

Abstract

The present disclosure provides a processing method and device of electronic reading data, electronic equipment, computer readable storage medium and computer program product, relates to the field of artificial intelligence, and particularly relates to the field of video generation and large model technology. The implementation scheme is: obtaining a first text of to-be-processed electronic reading data; generating audio data based on the first text; inputting the first text into a first language model to obtain a set of script texts, and the script texts in the set of script texts are texts obtained by performing content summarization on the text content in the corresponding text block in the first text; for each script text in the set of script texts, inputting the script text into a second language model to obtain a picture corresponding to the script text; and generating a promotion video of the electronic reading data based on the pictures corresponding to the set of script texts respectively and the audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence, and in particular to the fields of video generation and large model technology, specifically to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing electronic reading data. Background Technology

[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0003] With the advent of the digital age, the novel, as a traditional literary form, is facing unprecedented challenges. The widespread use of social media, video platforms, and mobile devices has led to increasingly fragmented attention spans. These platforms, with their rich multimedia content and high interactivity, attract a large number of users, making it difficult for traditional text-based novel promotion methods to attract and retain the attention of potential readers.

[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0005] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing electronic reading data.

[0006] According to one aspect of this disclosure, a method for processing electronic reading data is provided, comprising: acquiring a first text of electronic reading data to be processed; generating audio data based on the first text; inputting the first text into a first language model to obtain a set of script texts, wherein the script texts in the set of script texts are texts obtained by summarizing the text content in corresponding text blocks of the first text; for each script text in the set of script texts, inputting the script text into a second language model to obtain an image corresponding to the script text; and generating a promotional video for the electronic reading data based on the images corresponding to the set of script texts and the audio data.

[0007] According to another aspect of this disclosure, an apparatus for processing electronic reading data is provided, comprising: an acquisition unit configured to acquire a first text of electronic reading data to be processed; an audio generation unit configured to generate audio data based on the first text; a first invocation unit configured to input the first text into a first language model to obtain a set of script texts, wherein the script texts in the set of script texts are texts obtained by summarizing the text content in corresponding text blocks of the first text; a second invocation unit configured to input each script text in the set of script texts into a second language model to obtain an image corresponding to the script text; and a video generation unit configured to generate a promotional video of the electronic reading data based on the images and audio data corresponding to the set of script texts respectively.

[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor to enable the at least one processor to perform the methods described in this disclosure.

[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods described in this disclosure.

[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods described in this disclosure.

[0011] According to one or more embodiments of this disclosure, electronic reading data promotion videos are automatically generated based on a large language model. The large language model enables rapid understanding and content summarization of the corresponding text, thereby generating corresponding script text and further generating images corresponding to the script text, thus improving the generation speed and quality of subsequent promotion videos.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0014] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown;

[0015] Figure 2 A flowchart illustrating a method for processing electronic reading data according to an embodiment of the present disclosure is shown;

[0016] Figure 3 A flowchart of a method for obtaining first text according to an embodiment of the present disclosure is shown;

[0017] Figure 4 A flowchart illustrating a method for obtaining an image corresponding to a corresponding script text according to an embodiment of the present disclosure is shown;

[0018] Figure 5 A structural block diagram of an electronic reading data processing apparatus according to embodiments of the present disclosure is shown; and

[0019] Figure 6 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0021] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0022] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0023] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0024] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.

[0025] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of methods for processing electronic reading data.

[0026] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105, and / or 106 under a Software as a Service (SaaS) model.

[0027] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.

[0028] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to obtain e-reading data, videos, and other information, and to set corresponding configuration information. The client devices can provide an interface that allows users to interact with the client devices. The client devices can also output information to users through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.

[0029] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0030] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0031] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0032] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0033] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.

[0034] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0035] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.

[0036] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.

[0037] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.

[0038] In the current e-reading field, users' demand for audiobooks and reading experiences is growing. To attract people's reading interest, promotional videos, as an emerging marketing tool, can leverage the visual and auditory effects of short video platforms to bring viewers a more intuitive and engaging experience.

[0039] Therefore, according to embodiments of this disclosure, a method for processing electronic reading data is provided. Figure 2 A flowchart illustrating a method for processing electronic reading data according to an embodiment of the present disclosure is shown, such as... Figure 2 As shown, method 200 includes: acquiring a first text of the e-reading data to be processed (step 210); generating audio data based on the first text (step 220); inputting the first text into a first language model to obtain a set of script texts, wherein the script texts in the set of script texts are texts obtained by summarizing the text content in the corresponding text blocks of the first text (step 230); for each script text in the set of script texts, inputting the script text into a second language model to obtain an image corresponding to the script text (step 240); and generating a promotional video for the e-reading data based on the images corresponding to the set of script texts and the audio data (step 250).

[0040] According to embodiments of this disclosure, electronic reading data promotion videos are automatically generated based on a large language model. The large language model enables rapid understanding and content summarization of the corresponding text, thereby generating corresponding script text and further generating images corresponding to the script text, thus improving the generation speed and quality of subsequent promotion videos.

[0041] In step 210, the first text of the electronic reading data to be processed is obtained.

[0042] The electronic reading data in this embodiment can be electronic text content from the Internet, such as web pages, documents, pictures, videos, etc. Any multimedia data from which readable electronic text can be extracted can fall into the category of electronic reading data in this embodiment.

[0043] According to some embodiments, obtaining the first text of the electronic reading data to be processed includes: obtaining the source text data of the electronic reading data to be processed; and inputting the source text data into a third language model to extract preset text content to obtain the first text.

[0044] In some examples, the source text data can be text data in any format, including but not limited to URL links, TXT, DOCX, or PDF documents, so as to extract the corresponding first text.

[0045] For example, the e-reading data can be novel reading data, and the source text data can be an electronic novel document, such as the novel webpage text linked by a URL. In the embodiments of this disclosure, multiple different file formats are compatible and supported to ensure seamless integration of novel text from various sources; structured processing facilitates subsequent processing and analysis, improving the efficiency of promotional video generation.

[0046] In some examples, the extracted preset text content can be key plot points including at least one of the following: turning points, climaxes, and conflict points. Additionally or alternatively, the preset text content can also be preset text content from the first few chapters, for example, the first three chapters of text content within 2000 words.

[0047] In some examples, corresponding preset instructions (i.e., prompts) are constructed, and then the storyline in the source text data is analyzed through a third language model (i.e., a large language model) to extract preset text content in order to obtain the first text.

[0048] For example, when the e-reading data is novel reading data, given that novel texts are usually quite long, it is necessary to extract coherent and key plot points from the novel text in order to attract readers within a limited time. Automatic extraction of novel text using a large language model ensures the coherence of the storyline and enhances the content's appeal to the audience.

[0049] According to some embodiments, before inputting the source text data into the third language model, the method further includes: performing data cleaning on the source text data to obtain cleaned source text data.

[0050] In some examples, after obtaining the source text data of the e-reading data to be processed, the source text data can be cleaned first. For example, data cleaning can remove some data that is not related to the generation of audio and video, such as text data other than the novel's plot, including but not limited to special characters, chapter information, etc.

[0051] Figure 3 A flowchart illustrating a method for obtaining first text according to an embodiment of this disclosure is shown. Figure 3As shown, obtaining the first text of the electronic reading data to be processed (step 210) includes: obtaining the source text data of the electronic reading data to be processed (step 310); cleaning the source text data to obtain cleaned source text data (step 320); and inputting the cleaned source text data into a third language model to extract preset text content to obtain the first text (step 330).

[0052] In some examples, data cleaning of the source text data can be achieved through a corresponding language model to obtain cleaned source text data. For instance, a corresponding instruction (i.e., a prompt) can be constructed to enable the language model to perform data cleaning on the source text data.

[0053] It is understood that data cleaning of the source text data can be achieved by any suitable method in the embodiments of this disclosure, and no limitation is made herein.

[0054] In step 220, audio data is generated based on the first text.

[0055] According to some embodiments, generating audio data based on the first text includes: obtaining audio configuration information, wherein the audio configuration information includes at least one of the following: timbre, speech rate; and generating audio data based on the first text and the audio configuration information.

[0056] In some examples, timbre information can be selected for generating audio data. For instance, this timbre information could be a male voice, a female voice, or the voice of an authorized broadcaster. By configuring the appropriate timbre information, the initial text can be read aloud in a specific voice, thereby improving the quality of subsequent video generation and enhancing the user experience.

[0057] In some examples, speech rate information can also be selected for generating audio data. For instance, it may be possible to slide to select the corresponding speech rate information within a range of speech rates, thereby achieving the effect of reading the first text aloud based on that speech rate information, further improving the user experience.

[0058] It is understood that the above audio data can be generated by any suitable speech generator to perform the text-to-speech operation in this disclosure, and no limitation is made herein.

[0059] In some examples, speech data derived from text can include stress information. For instance, when implementing text-to-speech operations using a language model, a corresponding prompt can be constructed to allow the language model to analyze the text content and emphasize key words or expressions during the reading, thereby enabling listeners to more clearly understand the main idea and focus of the e-reading data.

[0060] In some examples, the audio data may include background music information in addition to the speech data derived from the text. For instance, background music can be matched to a material library based on preset type tags of the e-reading data and edited to fit the speech data, thus forming the audio data from the background music and the speech data derived from the text. In some examples, the audio data may also include subtitle information.

[0061] In step 230, the first text is input into the first language model to obtain a set of script text.

[0062] For example, when the e-reading data is novel reading data, and novel texts are usually quite long, it usually takes a long time to directly generate a set of pictures representing the development of the plot from the first text, thus affecting the user experience.

[0063] In the embodiments of this disclosure, the script text in the set of script texts is text obtained by summarizing the text content of corresponding text blocks in the first text. That is, this set of script texts serves as individual storyboard information. By summarizing the content of corresponding text blocks in the first text, the first language model can quickly understand the storyline, thereby improving the generation speed of images corresponding to subsequent script texts.

[0064] In some examples, the script text may include, but is not limited to: character information, location information, and plot information obtained by highly summarizing the plot of the corresponding text block. In some examples, the script text may also be a single sentence from the corresponding text block that can be used to characterize the development of the plot of that text block.

[0065] In some examples, corresponding preset instructions (i.e., prompts) are constructed to guide the first language model to understand the storyline in the first text and sequentially break it down into multiple text blocks, from which its corresponding script text is extracted. That is, the order of this set of script texts is consistent with the development order of the storyline in the first text.

[0066] In step 240, for each script text in the set of script texts, the script text is input into the second language model to obtain the image corresponding to the script text.

[0067] According to some embodiments, the method according to this disclosure may further include: acquiring a second text of the electronic reading data to be processed; and inputting the second text into a fourth language model to obtain story background information corresponding to the electronic reading data.

[0068] Therefore, inputting the script text into a second language model to obtain an image corresponding to the script text may include: inputting the story background information and the script text into a second language model to obtain an image corresponding to the script text.

[0069] In the above embodiments, the story background information can be, for example, the historical context of the story. The historical context of the story is obtained through a large language model to assist in image generation.

[0070] For example, when the e-reading data is novel reading data, the historical background of the novel story can be ancient or modern. If the current first text or the extracted set of script texts does not include descriptions to represent its corresponding historical background information, it may cause the generated image background or character image to be inconsistent with the novel's plot, thereby affecting the promotion effect of subsequent novel videos.

[0071] Through the above embodiments, the story background information corresponding to the e-reading data is determined in advance, thereby accurately guiding the subsequent image generation based on the second language model.

[0072] In some examples, the second text includes background information about the story. For example, it can be the first text mentioned above. Alternatively, the second text may not be the first text mentioned above; for example, it could be text obtained after data cleaning of the source text data of the e-reading data, without limitation.

[0073] In some examples, when using a fourth language model to generate story background information, one or more preset prompts can be simultaneously input into the model. These prompts guide the model to organize the textual storyline to obtain the relevant background information. Therefore, by introducing at least one preset prompt, the model is guided to generate story background information that better meets user expectations, improving the quality of the generated information and enhancing the user experience.

[0074] According to some embodiments, the method of this disclosure may further include: inputting the second text into a fifth language model to obtain first role information; matching the first role information with second role information in a preset role library to obtain second role information that matches the first role information. The preset role library includes multiple sets of second role information, wherein each set of second role information is descriptive information of one or more feature points of a corresponding role.

[0075] Therefore, inputting the script text into a second language model to obtain an image corresponding to the script text may further include: inputting the matched second character information and the script text into a second language model to obtain an image corresponding to the script text.

[0076] In some examples, the preset role library includes multiple second role information, each of which is a description of one or more feature points of the corresponding role. For example, these feature points could be: hairstyle, skin color, gender, identity, age, etc. Identity is used to indicate the role's background, social status, etc., such as student, professional, scientist, etc.

[0077] By pre-setting multiple second-character information in a character library, the first-character information obtained based on the fifth language model is matched with these multiple second-character information in the character library to obtain the second-character information most similar to the first-character information. This second-character information then guides the generation of subsequent images. By pre-setting the character library, the consistency of character features in each image in the set of images is ensured, thereby preventing frame skipping effects and impacting the user experience during subsequent video playback.

[0078] Figure 4 A flowchart illustrating a method for obtaining an image corresponding to corresponding script text according to an embodiment of this disclosure is shown. Figure 4 As shown, the process involves: obtaining the second text of the electronic reading data to be processed (step 410); inputting the second text into a fourth language model to obtain the story background information corresponding to the electronic reading data (step 420); inputting the second text into a fifth language model to obtain the first character information (step 430); matching the first character information with the second character information in a preset character library to obtain the second character information that matches the first character information (step 440); and inputting the story background information, the matched second character information, and the script text into a second language model to obtain the image corresponding to the script text (step 450).

[0079] This embodiment helps second language models generate storyboards with characters having a unified story background and similar style.

[0080] In some examples, when using a fifth language model to generate story background information, one or more preset prompts can be simultaneously input into the model. These prompts guide the model to organize the textual storyline and obtain character information. Therefore, by introducing at least one preset prompt, the model is guided to generate information that better meets user expectations, improving the quality of the generated information and enhancing the user experience.

[0081] In some examples, the first role information may include information about multiple roles, such as: the corresponding information for each of the multiple roles and the relationship between that role and other roles. The corresponding information for each role may be, for example, the role's name, identity, physical characteristics, clothing characteristics, etc. Therefore, the matching second role information may also be information about multiple roles.

[0082] In step 250, a promotional video for the e-reading data is generated based on the images and audio data corresponding to the set of script texts.

[0083] According to some embodiments, generating a promotional video for the e-reading data based on the images and audio data corresponding to the set of script texts includes: obtaining video configuration information, wherein the video configuration information includes at least one of the following: speed information, resolution, and aspect ratio; and generating a promotional video for the e-reading data based on the video configuration information, the images and audio data corresponding to the set of script texts.

[0084] In some examples, configuring information such as resolution and aspect ratio ensures a good viewing experience for promotional videos across different playback platforms and devices. Configuring appropriate speed information synchronizes it with the corresponding audio data, thus guaranteeing the quality of the generated promotional video.

[0085] In some examples, the images in the image group corresponding to the set of script text can be automatically resized and frame-interpolated. Additionally or alternatively, the audio data is automatically mixed, and finally, the image group and audio data are combined into a promotional video to automatically generate a promotional video for e-reading data. Frame interpolation is used to smoothly connect the various frames, increasing visual fluidity and making the entire promotional video look more natural and attractive; audio mixing is used to balance multiple audio tracks, preventing any one sound from overpowering others and affecting the overall auditory experience, such as background music and dialogue.

[0086] In some examples, promotional videos for the e-reading data can be generated using a corresponding language model. For instance, by constructing corresponding instructions (i.e., prompts), the language model can automatically adjust or interpolate the images in a group of pictures.

[0087] It is understood that the above-mentioned promotional video can be generated by any suitable method in the embodiments of this disclosure, and no limitation is made herein.

[0088] In some embodiments, any of the above language models can be knowledge-enhanced large language models for dialogue (such as ERNIE bot), which are trained on massive knowledge resources and dialogue data (such as trillions of web pages, billions of search data, hundreds of millions of image data, billions of voice requests per day, more than 50 billion text requests per day, and more than 550 billion factual knowledge).

[0089] Therefore, applying this type of model as a language model can not only directly process chat-type dialogue information, but also directly generate response information for dialogue information of logical reasoning, common sense, and image generation, thereby improving generation efficiency while generating higher quality response information.

[0090] It is understood that the above multiple language models can be the same language model or different language models, and no restrictions are imposed here.

[0091] According to some embodiments, the method according to this disclosure may further include: obtaining address information of the electronic reading data, so as to publish the address information and the promotional video together to a social media platform.

[0092] According to this disclosure, a social platform refers to a tool and platform used by people to share opinions, insights, experiences, and perspectives. With the support of a robust technological platform, this social platform allows internet users to smoothly publish, browse, and share video works online, including any video playback platform and live streaming platform that can utilize this disclosure.

[0093] In some examples, by publishing the address information and the promotional video together on a social media platform, users on that platform can click or copy the address information to jump to the reading interface of the e-reading data after watching the promotional video, thereby achieving the promotional effect of the e-reading data.

[0094] In some examples, when the promotional video is published, the address information can be published simultaneously with the promotional video in any suitable manner, including but not limited to, in the form of a pop-up window at a designated location in the promotional video, in the form of copying the address link in the comment section of the promotional video, in the form of forwarding it as two separate pieces of information alongside the promotional video, etc.

[0095] According to embodiments of this disclosure, such as Figure 5As shown, an electronic reading data processing device 500 is also provided, comprising: an acquisition unit 510 configured to acquire a first text of electronic reading data to be processed; an audio generation unit 520 configured to generate audio data based on the first text; a first invocation unit 530 configured to input the first text into a first language model to obtain a set of script texts, wherein the script texts in the set of script texts are texts obtained by summarizing the text content in corresponding text blocks of the first text; a second invocation unit 540 configured to input each script text in the set of script texts into a second language model to obtain an image corresponding to the script text; and a video generation unit 550 configured to generate a promotional video of the electronic reading data based on the images and audio data corresponding to the set of script texts.

[0096] Here, the operation of each of the above-mentioned units 510 to 550 of the electronic reading data processing device 500 is similar to the operation of steps 210 to 250 described above, and will not be repeated here.

[0097] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0098] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.

[0099] refer to Figure 6 The present invention describes a structural block diagram of an electronic device 600 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0100] like Figure 6As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0101] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, output unit 607, storage unit 608, and communication unit 609. Input unit 606 can be any type of device capable of inputting information to electronic device 600. Input unit 606 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and can include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 607 can be any type of device capable of presenting information, and can include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 608 can include, but is not limited to, disk and optical disk. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and can include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0102] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of method 200 described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform method 200 by any other suitable means (e.g., by means of firmware).

[0103] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0104] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0105] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0106] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0107] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0108] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0109] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0110] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A method for processing electronic reading data, comprising: Obtain the first text of the e-reading data to be processed; Audio data is generated based on the first text; The first text is input into the first language model to obtain a set of script texts, wherein the script texts in the set of script texts are texts obtained by summarizing the text content in the corresponding text blocks of the first text; For each script text in the set of script texts, the script text is input into the second language model to obtain the image corresponding to the script text; Based on the images and audio data corresponding to the set of script texts, a promotional video for the e-reading data is generated.

2. The method of claim 1, wherein, The first text of the e-reading data to be processed includes: Obtain the source text data of the e-reading data to be processed; and The source text data is input into a third language model to extract preset text content, thereby obtaining the first text.

3. The method of claim 2, wherein, Before inputting the source text data into the third language model, the method further includes: performing data cleaning on the source text data to obtain cleaned source text data.

4. The method of claim 1, wherein, The process of generating audio data based on the first text includes: Obtain audio configuration information, wherein the audio configuration information includes at least one of the following: timbre, speech rate; and Audio data is generated based on the first text and the audio configuration information.

5. The method of claim 1, further comprising: Obtain the second text of the electronic reading data to be processed; as well as The second text is input into the fourth language model to obtain the story background information corresponding to the e-reading data. Furthermore, inputting the script text into a second language model to obtain an image corresponding to the script text includes: inputting the story background information and the script text into a second language model to obtain an image corresponding to the script text.

6. The method of claim 1 or 5, further comprising: The second text is input into the fifth language model to obtain the first role information; The first character information is matched with second character information in a preset character database to obtain second character information that matches the first character information. The preset character database includes multiple sets of second character information, where each set of second character information is a description of one or more feature points of the corresponding character. Furthermore, inputting the script text into a second language model to obtain an image corresponding to the script text includes: inputting the matched second character information and the script text into a second language model to obtain an image corresponding to the script text.

7. The method of claim 1, further comprising: Obtain the address information of the electronic reading data, and publish the address information and the promotional video together to a social media platform.

8. The method of claim 1, wherein, The promotional video for generating the e-reading data based on the images and audio data corresponding to the set of script texts includes: Obtain video configuration information, wherein the video configuration information includes at least one of the following: speed information, resolution, aspect ratio; and Based on the video configuration information, the images corresponding to the set of script texts, and the audio data, a promotional video for the e-reading data is generated.

9. An electronic reading data processing device, comprising: The acquisition unit is configured to acquire the first text of the electronic reading data to be processed. An audio generation unit is configured to generate audio data based on the first text; The first calling unit is configured to input the first text into the first language model to obtain a set of script texts, wherein the script texts in the set of script texts are texts obtained by summarizing the text content in the corresponding text blocks of the first text; The second calling unit is configured to input each script text in the set of script texts into the second language model to obtain an image corresponding to the script text. The video generation unit is configured to generate a promotional video for the e-reading data based on the images and audio data corresponding to the set of script texts.

10. The apparatus of claim 9, wherein, The acquisition unit includes: A unit for acquiring source text data of e-reading data to be processed; and A unit used to input the source text data into a third language model to extract preset text content and obtain the first text.

11. The apparatus of claim 10, wherein, The acquisition unit includes a unit for cleaning the source text data before inputting it into a third language model to obtain cleaned source text data.

12. The apparatus of claim 9, wherein, The audio generation unit includes: A unit for acquiring audio configuration information, wherein the audio configuration information includes at least one of the following: timbre, speech rate; and A unit for generating audio data based on the first text and the audio configuration information.

13. The apparatus of claim 9, further comprising: A unit for acquiring the second text of the electronic reading data to be processed; as well as A unit for inputting the second text into a fourth language model to obtain the story background information corresponding to the electronic reading data. Furthermore, the second calling unit includes a unit for inputting the story background information and the script text into a second language model to obtain an image corresponding to the script text.

14. The apparatus of claim 9 or 13, further comprising: A unit used to input the second text into a fifth language model to obtain the first role information; A unit for matching the first character information with second character information in a preset character database to obtain second character information that matches the first character information, wherein the preset character database includes multiple pieces of second character information, and each piece of second character information is descriptive information of one or more feature points of the corresponding character. Furthermore, the second calling unit includes a unit for inputting the matched second role information and the script text into a second language model to obtain an image corresponding to the script text.

15. The apparatus of claim 9, further comprising: A unit for obtaining the address information of the electronic reading data and publishing the address information and the promotional video together on a social media platform.

16. The apparatus of claim 9, wherein, The video generation unit includes: A unit for acquiring video configuration information, wherein the video configuration information includes at least one of the following: speed information, resolution, aspect ratio; and A unit for generating a promotional video for the e-reading data based on the video configuration information, the images corresponding to the set of script texts, and the audio data.

17. An electronic device comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.

18. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.

19. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-8.