Method and apparatus for presenting text works, device, and storage medium

By displaying character images based on text fragments in e-book readers, the problem of users having difficulty obtaining details of text works from voice data is solved, richer visual information presentation is achieved, and the reading experience is improved.

WO2025199827A1PCT designated stage Publication Date: 2025-10-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/084218
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

It is difficult for users to understand the details of text works, such as relevant scene information and character information, based only on voice data.

Method used

In an electronic book reader, in response to a playback request for audio data, an image generated based on a character in a text segment corresponding to a playback position of the audio data is displayed.

Benefits of technology

It provides richer visual information to help users better understand text works, improve the fun of reading experience and the efficiency of information acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024084218_02102025_PF_FP_ABST
    Figure CN2024084218_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method and apparatus for presenting text works, a device, and a storage medium. The method comprises: in an electronic book reader for reading text works, in response to receiving a playing request used for playing audio data corresponding to a text work, playing the audio data; and displaying images corresponding to the text work, wherein the images are generated on the basis of characters in text sections in the text work, and the text sections correspond to a playing position of the audio data. In this way, the displayed images vary with the playing position of the audio data, thereby providing users with richer visual information about the text work while providing the audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device and storage medium for presenting text works Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for presenting text works. Background Art

[0002] With the development of digital technology, more and more applications and websites are able to be used to present electronic publications, also known as e-books. Users can read the text content in e-books (especially text works such as novels, essays, prose, poems, and plays). At present, a conversion solution based on text to speech (TTS) has been proposed, and the text content in the e-book can be converted into voice data, and then the voice data can be played. However, it is difficult for users to understand the details of the text work based on voice data alone, and it is desired to provide users with richer information.

[0003] Summary of the Invention

[0004] In a first aspect of the present disclosure, a method for presenting a text work is provided. The method includes: in an electronic book reader for reading a text work, in response to receiving a playback request for audio data corresponding to the text work, playing the audio data; and displaying an image corresponding to the text work, the image being generated based on a character in a text segment in the text work, the text segment corresponding to a playback position of the audio data.

[0005] In a second aspect of the present disclosure, a device for presenting a text work is provided. The device includes: a playback module configured to, in an electronic book reader for reading a text work, play audio data in response to receiving a playback request for audio data corresponding to the text work; and a display module configured to display an image corresponding to the text work, the image being generated based on a character in a text segment in the text work, the text segment corresponding to a playback position of the audio data.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.

[0008] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0010] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0011] FIG2 shows a schematic diagram of a page for presenting a text work according to some embodiments of the present disclosure;

[0012] 3A and 3B are schematic diagrams respectively illustrating example architectures for generating images for textual works according to some embodiments of the present disclosure;

[0013] FIG4 shows a schematic diagram of an example character image according to some embodiments of the present disclosure;

[0014] FIG5 shows a schematic diagram for storing an image according to some embodiments of the present disclosure;

[0015] FIG6 shows a flowchart of a process for text work presentation according to some embodiments of the present disclosure;

[0016] FIG7 shows a schematic structural block diagram of an apparatus for presenting a text work according to some embodiments of the present disclosure; and

[0017] FIG8 illustrates a block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION

[0018] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0019] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below.

[0020] Herein, unless explicitly stated otherwise, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.

[0021] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0022] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0023] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to the electronic device, application, server or storage medium and other software or hardware that performs the operation of the technical solution of the present disclosure based on the prompt message.

[0024] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0025] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0026] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0027] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs. It typically includes an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence so that the output of the previous layer is provided as input to the next layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes the input from the previous layer.

[0028] Generally speaking, machine learning can be roughly divided into three stages, namely the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values ​​are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association between input and output (also called input-to-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. The testing stage can sometimes be integrated into the training stage. In the application or inference stage, the trained model can be used to process the actual model input based on the parameter values ​​obtained through training to determine the corresponding model output.

[0029] As briefly mentioned above, in order to facilitate improving the reading efficiency of users, a text-to-sound conversion process has been proposed to convert text works into voice data, and then play the voice data to the user. Referring to Figure 1, the application environment of the present disclosure is described. Figure 1 shows a schematic diagram 100 of an example environment in which an embodiment of the present disclosure can be implemented. As shown in Figure 1, an e-book reader can provide a page 110. The user can select the text work that he wants to play; alternatively and / or additionally, the user can select the chapter in the text work that he wants to play. Further, the user can use the control 140 to play the corresponding audio data. At this time, the display area 120 can present general information of the text work, such as the name, author, the title of the chapter being played, and so on.

[0030] However, audio data can only provide limited information. It is difficult for users to understand the details of the text work (for example, relevant scene information and character information, etc.) based on voice data alone. At this time, it is expected that more abundant information can be provided to users.

[0031] To at least partially address the aforementioned issues, embodiments of the present disclosure provide a method for presenting a text work. In this method, in an e-book reader for reading a text work, upon receiving a request to play audio data corresponding to the text work, the audio data is played. Furthermore, an image corresponding to the text work is displayed. The image is generated based on a character in a text segment within the text work, the text segment corresponding to the playback position of the audio data.

[0032] An overview of embodiments of the present disclosure is provided with reference to FIG2 , which illustrates a schematic diagram 200 of a page for presenting a textual work according to some embodiments of the present disclosure. For ease of description, the following textual work is hereinafter exemplified by a novel. Alternatively and / or additionally, the textual work may include, but is not limited to, novels, essays, prose, poetry, and scripts.

[0033] The method of the present disclosure can be performed by an e-book reader. The user can select the text work that he or she wishes to play; alternatively and / or additionally, when the text work includes multiple chapters, the user can select the chapter that he or she wishes to play. As shown in Figure 2, the e-book reader can provide page 210. In page 210, the user can use control 240 to request to play audio data. In response to receiving a play request for playing audio data corresponding to the text work, the audio data can be played. Further, page 210 can include a display area 220, and an image 222 corresponding to the text work can be displayed in the display area 220. Here, the image 222 is generated based on a character in a text segment in the text work, and the text segment corresponds to the playback position of the audio data. In this way, the image presented in the display area 220 will change with the playback position of the audio data. Thus, while providing audio data, richer visual information about the text work can be provided to the user.

[0034] In the context of the present disclosure, an e-book reader can be installed at an electronic device. The electronic device may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, server devices, etc. The terminal device may be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a personal communication system (Personal Communication System, PCS) device, a personal navigation device, a personal digital assistant (Personal Digital Assistant, PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. The server device may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. The server-side device may include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like.

[0035] In some embodiments, a pre-processing process may be performed on the text work to generate corresponding images for each text segment in the text work. These images may be stored in the electronic device or obtained by the electronic device from a remote server.

[0036] In some embodiments, the initial playback position in the audio data specified by the play request can be obtained, and then the audio segment corresponding to the initial playback position in the audio data (that is, the portion after the initial playback position in the audio data) can be played. The user can press the control 240 to start playing. At this time, the display area 220 will present the corresponding images of each text segment in the text work one by one according to the length of time of the playback. It should be understood that the image may include an image of a character in the corresponding text segment of the audio data currently being played. In this way, as a supplement to the audio data currently being played, the e-book reader can present more visual information about the content being played to the user, thereby facilitating the user's understanding of the text work.

[0037] In some embodiments, in response to receiving a request to specify another initial playback position for audio data, an audio segment in the audio data corresponding to the other initial playback position is played; based on the other initial playback position and the playback time, a current playback position of the audio data is determined; a current text segment in the text work corresponding to the current playback position is determined; and a current image corresponding to the current text segment is presented. In this way, the user is allowed to adjust the playback position as desired. For example, the user can use a fast forward control, a fast rewind control, or a progress adjustment control to adjust the playback position, and then an image corresponding to the playback position can be accurately presented.

[0038] 2 , the user can set the initial play position via control 250, and play the audio segment corresponding to the set initial play position in the audio data via control 240. In this way, the audio data can be played from the desired play position specified by the user, thereby facilitating user operation.

[0039] In some embodiments, the text in the text segment is displayed in the e-book reader. Continuing with FIG2 , page 210 may provide a display area 230, and the text in the text segment currently being played may be displayed in the display area. For example, the text may be displayed in a scrolling manner, or the text may be displayed in a text block manner. In this way, as a supplement to the audio data, the relevant text of the currently playing audio may be presented to the user, thereby facilitating the user to obtain more information. For example, the text may be used to assist in understanding the audio data, particularly unclear portions of the audio data.

[0040] Images of each text segment can be generated at the electronic device, or pre-generated images can be obtained by the electronic device from a remote server. In the context of the present disclosure, images can be referred to as illustrations of text works. Specifically, images that match each text segment in a text work can be generated with the help of a trained machine learning model. The machine learning model can be, for example, an image generation model. The machine learning model can include, for example, but is not limited to, any appropriate model such as a Transformer model, a convolutional neural network (CNN), a recurrent neural network (RNN), and a deep neural network (DNN). The machine learning model can be a model local to the electronic device, or a model installed in other electronic devices (for example, installed in a remote device). It should be noted that the machine learning model can include multiple models, and the present disclosure does not limit the number and type of models specifically included in the machine learning model.

[0041] In this way, multiple images can be generated quickly and easily, which can improve the efficiency of generating images. In addition, inserting relevant images of characters into novels based on the content of the novels can increase the interest of readers when reading novels.

[0042] See Figure 3A for more details about generating images, which shows a schematic diagram of an example architecture 300A for generating images for text works according to some embodiments of the present disclosure. Architecture 300A can be implemented at an electronic device, and Figure 3A shows an overview of the image generation process. The text work 301 may include multiple text segments 302-1, ..., 302-N (individually and / or collectively referred to as text segments 302). Descriptive information 304 of a text segment (e.g., text segment 302-1) can be determined based on a machine learning model. Role information 303 of at least one character of the text work 301 can be determined, and the role information 303 includes attribute information of the character in at least one dimension. Here, the dimensions may include, for example, the character's name, gender, age, occupation, appearance, expression, and clothing, etc. For a character in at least one character, a character image diagram 305 of the character can be generated based on the character information 303 of the character.

[0043] Furthermore, an image 222 of the text segment 302 can be generated based on the description information 304 of the text segment 302 and the character image 305. In this way, corresponding images can be generated for each text segment in the text work, thereby increasing the interest of readers when reading the text work.

[0044] In some embodiments, the text work 301 may include multiple text fragments, and the text fragments may be divided based on predetermined rules. For example, the structural information of the text work may be obtained, and the text work may be divided into multiple text fragments based on the structural information. Here, the structural information may include the directory structure of the text work, for example, the multiple text fragments may be determined according to the hierarchy of multi-level titles defined in the directory structure. Rules for dividing text fragments may be pre-specified, for example, a text fragment may include a "chapter", a "section", or one or more paragraphs, etc. In this way, the text fragments may be divided according to different accuracies, and an image that better matches the content of the text fragment may be generated.

[0045] For another example, the predetermined rule may indicate the maximum number of text units (e.g., chapters, paragraphs, sentences, or words, etc.) that each text segment may include. For example, the predetermined rule may indicate that each text segment may include at most one chapter, one section, one paragraph, or 50 text units, etc. Based on such predetermined rules, the entire text of the novel may be segmented (e.g., with one chapter, one section, one paragraph, or 50 text units as one text segment) to obtain multiple text segments.

[0046] In some embodiments, based on the multiple text segments that have been divided, predetermined rules can adjust the text segments according to the frequency or time length of the specified switching images. For example, the user can specify the frequency of switching images, for example, switching once per minute, switching once every two minutes, and so on. At this time, the number of words included in the text segment can be determined based on the speed information of the audio data (for example, 150 words / minute, etc.) and the switching frequency, and then the scope of the text segment can be determined. Alternatively and / or additionally, it can be specified that the text segment cannot span the natural paragraphs in the text work, and so on. At this time, the text segments will be divided according to the natural paragraphs. In some embodiments, it can be considered whether the user starts "double-speed playback" during playback, and the scope of the text segment is adjusted accordingly based on the "double-speed" of playback. In this way, the switching frequency of the image can be adjusted accordingly, thereby preventing the image from being switched too frequently.

[0047] See FIG3B for more details about image generation, which shows a schematic diagram of an example architecture 300B for generating images for a text work according to some embodiments of the present disclosure. As shown in FIG3B , the architecture 300B includes a description information extraction unit 310 and a character information acquisition unit 320. The description information extraction unit 310 can be used, for example, to extract description information 315 (e.g., a summary, etc.) of a text fragment from a text fragment of a novel. In some embodiments, the electronic device can obtain each text fragment in the novel and provide it to the description information extraction unit 310. The description information extraction unit 310 can perform processing on each text fragment in the novel based on a predetermined machine learning model to obtain description information of multiple text fragments.

[0048] In some embodiments, the description information may include environmental information about the character's environment and the character's action information. The description information extraction unit 310 may, for example, determine description information 315 for a text segment by summarizing at least one character in the text segment and the environmental information and action information associated with the at least one character. For example, for text segment A, "Character A woke up early, made breakfast, put away the mess of toys in the living room, mopped the floor, and then took two steamed buns and went out."

[0049] The description information extraction unit 310 can determine that text segment A only includes character A. The environmental information in the description information generated by the description information extraction unit 310 can be, for example, "home", and the action information can be, for example, "character A gets up early to do housework." It should be noted that not every text segment includes a character. For example, the text segment B "From now on, they will go on an AA basis, and each will pay half of any expenses" does not include any characters or actions associated with the characters. Therefore, the description information extraction unit 310 may not generate description information corresponding to text segment B. At this time, the text segment can continue to use the image of the previous text segment.

[0050] The character information acquisition unit 320 can acquire character information 325 of at least one character. The at least one character here can be determined by the character information acquisition unit 320 based on the full text of the text work or the current text fragment, or it can be determined by the electronic device and provided to the character information acquisition unit 320. Specifically, the electronic device / character information acquisition unit 320 can determine at least one character in the novel. For example, for description information including "Character A gets up early to do housework," "Character A hears Mom and Dad arguing," and "Character A decides to go out for a walk," the electronic device / character information acquisition unit 320 can determine that the three characters include "Character A," "Dad," and "Mom."

[0051] In some embodiments, the role information acquisition unit 320 can also determine the number of occurrences of the target role in the text segment for the target role among at least one role, and thus determine the main role in the text segment. For example, the role information of the target role can be acquired in response to determining that the number of occurrences meets a predetermined condition. The predetermined condition here can, for example, indicate a predetermined number of times (for example, 3 times, 5 times, or any other number), and the role information acquisition unit 320 can, for example, acquire the role information of the target role in response to determining that the number of occurrences reaches a predetermined number. Thus, only the role information of the role with a larger number of occurrences can be acquired, which can reduce the final image generation cost. Alternatively and / or additionally, assuming that the current text segment only includes one role and the number of occurrences of the role is lower than the predetermined number, the role information of the role can still be acquired.

[0052] In some embodiments, the role information 225 of at least one role may include multiple attributes of the at least one role, such as any one or more of the role's name, gender, age, occupation, appearance, demeanor, and clothing.

[0053] In some embodiments, to determine the character information of at least one character, the character information of the target character can be determined based on the portion of the text work associated with the target character. Furthermore, the character information of the target character can be updated based on the portion of the text segment associated with the target character. For example, the basic attributes of the character, such as name, gender, age, and occupation, can be determined from the entire text work. Furthermore, the special attributes of the character in the current segment of the text being processed, such as the current appearance, expression, and clothing, etc., can be determined.

[0054] For example, if text segment 1 depicts a winter scene, then based on this text segment 1, the character's attire can be determined to be a "coat." If text segment 2 depicts a summer scene, then based on this text segment 2, the character's attire can be determined to be a "dress." A mapping relationship can exist between character information and each text segment in the text work. This ensures that the character information matches the character's basic characteristics and reflects the character's current state as the story progresses within the text work.

[0055] In some embodiments, according to a specific determination method, multiple attributes can be divided into two parts, wherein the first part can be directly determined from the novel, and the second part can be indirectly determined from the novel, or can be manually set. For example, the name, gender, etc. of the character in the multiple attributes can be the attributes of the first part. The character information acquisition unit 320 can, for example, determine the attributes of the first part from the novel. It should be noted that the character information acquisition unit 320 can obtain the first part of the attributes of character A from the full text of the novel, and is not limited to the current text segment.

[0056] For example, if the novel clearly records the attributes of the character's appearance, demeanor, clothing, etc., then the above attributes can be used as the attributes of the first part. If the novel does not include text associated with the attributes of the second part, the electronic device can receive user input from a user (e.g., a relevant staff member) and determine the attributes of the second part based on the user input. For example, the electronic device can provide a setting control for setting the second part of the multiple attributes of the character in an electronic book reader, and in response to receiving a setting operation for the setting control, set the second part of the multiple attributes based on the setting operation.

[0057] Such a setting control may be, for example, an input box. The electronic device may, for example, receive user input via the input box and determine the second portion of the multiple attributes based on the user input. The electronic device may, for example, provide the determined second portion of attributes to the character information acquisition unit 320 so that the character information acquisition unit 320 acquires the second portion of attributes. In this way, the user is allowed to specify the attributes of the character (e.g., clothing style and color, etc.) according to their needs during the reading process, thereby generating an image that meets their needs.

[0058] The electronic device can generate an image 222 of the novel based on the description information 315 of the text segment and the role information 325 of at least one character. The electronic device can generate the image 222 based on any appropriate method and using the description information 315 of the text segment and the role information 325 of at least one character. The present disclosure does not limit the specific method of generating the image. For example, the electronic device and / or other devices can generate the image based on pre-acquired rules or algorithms. In some embodiments, the electronic device can generate the image with the help of a trained machine learning model. In this case, the architecture 300B can also include a prompt word determination unit 330 and a machine learning model 370.

[0059] The prompt word determination unit 330 can, for example, be configured to generate prompt words 335 for the machine learning model 370 based on the description information and the role information. The prompt word determination unit 330 can, for example, obtain a predetermined prompt word template and fill the prompt word template with the description information and the role information to generate the prompt word 335. For example, the prompt word template can include: environmental information, role information, and action information. The obtained various information can be filled into the corresponding positions of the template to generate the prompt word.

[0060] In some embodiments, in order to ensure the uniformity of the characters in subsequently generated images, the prompt word determination unit 330 may also call the character image determination unit 250 and the character model generation unit 360. Here, the character image may, for example, represent the character image of the character from multiple angles, and the character image determination unit 350 may, for example, determine a character image 355 for at least one character based on the character information of at least one character in the novel. The character image determination unit 350 may determine the character image 355 in any appropriate manner. For example, the character image determination unit 350 may generate a character image 355 for at least one character based on the character information of at least one character using a trained image generation model. Alternatively or additionally, in some embodiments, the character image determination unit 350 may also directly obtain a character image input by a user (e.g., a character image drawn by an illustrator for at least one character).

[0061] FIG4 illustrates a schematic diagram of an example character image 400 according to some embodiments of the present disclosure. The character image determination unit 350 may, for example, generate character image 400 for character A based on character information for character A, such as "character A, 25 years old, female, curly hair, wearing a dress." Character image 400 may include multiple images of character A at various angles (e.g., image 401 tilted 45 degrees to the side, a side view image 402, a front view image 403, and a back view image 404).

[0062] For each character, the character model generation unit 360 can generate a character model 365 describing the character based on the character image 355 corresponding to the character. Character model 365 can be, for example, a LoRA model. The LoRA model can be understood as a plug-in to the Stable Diffusion (SD) model (a generative model), which can be used to meet a specific style or specified character attributes.

[0063] The process of generating a character model based on a character image can be understood as storing the character image in the form of a character model. The prompt word determination unit 330 can subsequently flexibly call different character models to call different character images. The prompt word determination unit 330 can obtain a character model 365 for at least one character and determine a prompt word 335 based on the character model 365 and the description information 315 of the text segment. For example, the prompt word determination unit 330 can generate the prompt word "Character A <Model A> makes breakfast at home" based on the description information "Character A gets up early to do housework" and the character model A corresponding to character A.

[0064] In some embodiments, the prompt word determination unit 330 can also update the prompt word 335 based on the weight index of the character model. The weight index can be used to indicate the similarity between the character in the image and the character image corresponding to the character. For example, if the prompt word is "Character A <Model A, 0.5> making breakfast at home", then the prompt word indicates that the similarity between the character A in the subsequently generated image and the character image corresponding to character A is 50%. It can be understood that the higher the weight index, the higher the similarity between the character in the image and the character image corresponding to the character, and the more similar the two are. Using the embodiments of the present disclosure, the character details in each image can be adjusted while ensuring the consistency of the appearance of the novel character.

[0065] In some embodiments, the prompt word determination unit 330 can also determine the style of the image based on the background environment of the novel, and update the prompt word 335 based on the style. For example, if the background of the novel is a modern urban background, the style of the image can be determined to be "comic style", and the prompt word can be, for example, "Character A <Model A, 0.5> making breakfast at home comic style". If the background of the novel is an ancient martial arts background, the style of the image can be determined to be "ink style", and the prompt word can be, for example, "Character A <Model A, 0.5> making breakfast at home ink style". Utilizing the embodiments of the present disclosure, images with richer visual effects can be generated in a more flexible manner.

[0066] In some embodiments, during the prompt word generation process, the image style can be determined based on the sound properties of the audio data, and the prompt word can then be updated based on the style. For example, if the text is written by a deep male voice, a rugged image style can be generated; if the voice is a sweet female voice, a refined image style can be generated, and so on. This ensures that the style of the image data matches the style of the audio data, thereby providing a more harmonious visual and auditory experience.

[0067] The prompt word determination unit 330 may provide the determined prompt word 335 to the machine learning model 370. The machine learning model 370 may then generate the image 222 based on the acquired prompt word 335. If the prompt word is "Character A <Model A, 0.5> making breakfast at home," the machine learning model 120 may call Model A and generate an image showing Character A making breakfast at home based on the weight index 0.5.

[0068] In some embodiments, an association relationship can be established between the generated image and the text fragment, and the image and the text fragment can be stored in association. At least one image corresponding to at least one character can be obtained (each character can correspond to one or more images), and the image can be inserted into the position associated with the text fragment in the novel. At this time, the generated image can be stored in the text work as an illustration of the text work. Exemplarily, image A can be inserted into text fragment A (for example, inside text fragment A, before / after text fragment A, etc.). In this way, during the playback process, the image in the text fragment corresponding to the audio fragment being played can be searched and then presented.

[0069] Alternatively and / or additionally, the individual images may be stored in an image database. See FIG5 for further details, which shows a schematic diagram 500 for storing images according to some embodiments of the present disclosure. As shown in FIG5 , an image database 520 may be established, and images corresponding to individual text segments in the text work 210 may be stored in the image database 520. The audio data 510 of the text work 201 may include a plurality of audio segments 512, each of which may correspond to a text segment. In the context of the present disclosure, the audio data 510 may be pre-recorded audio data or may be audio data generated based on text-to-audio conversion.

[0070] Furthermore, an index can be established between the text segment 201-1, the image 222 generated based on the character in the text segment, and the audio segment 510 of the text segment 202-1. During playback, the audio segment can be played based on the index and the corresponding image can be displayed. In this way, image data can be stored in a centralized manner, thereby facilitating unified management of multiple image data.

[0071] In summary, according to the embodiments of the present disclosure, multiple images can be generated based on the text content of a novel quickly and easily while ensuring image quality, which can improve the efficiency of image generation. In addition, while providing audio data, users can be provided with richer visual information about the text work, which can increase their interest in listening to the novel.

[0072] The above describes the specific details of each step of generating an image for a text work, providing a method for presenting a text work. Figure 6 shows a flowchart of a process 600 for generating an image for a text work according to some embodiments of the present disclosure. Process 600 can be implemented at an electronic device.

[0073] At block 610 , in an electronic book reader for reading a text work, in response to receiving a play request for playing audio data corresponding to the text work, audio data is played.

[0074] At block 620 , an image corresponding to the text work is displayed, the image being generated based on a character in a text segment in the text work, the text segment corresponding to a playback position of the audio data.

[0075] In some embodiments, playing the audio data includes: obtaining an initial playback position in the audio data specified by the playback request; and playing an audio segment in the audio data corresponding to the initial playback position.

[0076] In some embodiments, the process 600 further includes: in response to receiving a specified request for specifying another initial playback position of the audio data, playing an audio segment in the audio data corresponding to the other initial playback position; determining a current playback position of the audio data based on the other initial playback position and the playback time; determining a current text segment in the text work corresponding to the current playback position; and presenting an image corresponding to the current text segment.

[0077] In some embodiments, the process 600 further includes displaying the text in the text snippet in an electronic book reader.

[0078] In some embodiments, an image is generated based on: determining descriptive information of a text segment and role information of at least one character based on a text work; wherein the text work includes multiple text segments, and the role information includes attribute information of the character in at least one dimension; for a role in at least one character, generating a role image diagram of the role based on the role information of the role; and generating an image of the text segment based on the descriptive information of the text segment and the role image diagram of the character.

[0079] In some embodiments, the plurality of text segments is determined based on: obtaining structural information of the text work; and dividing the text work into the plurality of text segments based on the structural information.

[0080] In some embodiments, at least one role is determined based on: determining the number of occurrences of a target role among multiple roles in a text work in a text segment; and in response to determining that the number of occurrences meets a predetermined condition, using the target role as a role among at least one role.

[0081] In some embodiments, determining role information of at least one role includes: determining the role information of the target role based on a portion of the text work associated with the target role; and updating the role information of the target role based on a portion of the text fragment associated with the target role.

[0082] In some embodiments, the description information includes environmental information of the character's environment and action information of the character, and generating an image of the text fragment includes: generating a character model for describing the character based on the character image; generating prompt words for the machine learning model using the environmental information, action information and the character model; and generating an image based on the prompt words.

[0083] In some embodiments, generating the prompt word further includes: determining a style of the image based on sound attributes of the audio data; and updating the prompt word based on the style.

[0084] In some embodiments, the attribute information of at least one dimension includes at least any one of the following: the character's name, gender, age, occupation, appearance, expression, and clothing, and the first part of the attribute information of at least one dimension is determined from the text work.

[0085] In some embodiments, the process 600 further includes: providing a setting control in the electronic book reader for setting a second portion of the attribute information of at least one dimension; and in response to receiving a setting operation for the setting control, setting the second portion of the attribute information of at least one dimension based on the setting operation.

[0086] According to some embodiments of the present disclosure, a device for presenting a text work is also provided. Figure 7 shows a schematic block diagram of a device 700 for generating an image for a text work according to some embodiments of the present disclosure. Device 700 can be implemented as or included in an electronic device. The various modules / components in device 700 can be implemented using hardware, software, firmware, or any combination thereof.

[0087] As shown in Figure 7, the device 700 includes: a playback module 710, which is configured to play audio data in an electronic book reader for reading a text work in response to receiving a playback request for playing audio data corresponding to the text work; and a display module 720, which is configured to display an image corresponding to the text work, where the image is generated based on a character in a text segment in the text work, and the text segment corresponds to the playback position of the audio data.

[0088] In some embodiments, the playback module 710 is further configured to: obtain an initial playback position in the audio data specified by the playback request; and play an audio segment in the audio data corresponding to the initial playback position.

[0089] In some embodiments, the device 700 further includes: a designated playback module, configured to play an audio segment in the audio data corresponding to another initial playback position in response to receiving a designated request for specifying another initial playback position of the audio data; a position determination module, configured to determine the current playback position of the audio data based on another initial playback position and the playback time; a text segment determination module, configured to determine the current text segment in the text work corresponding to the current playback position; and an image presentation module, configured to present an image corresponding to the current text segment.

[0090] In some embodiments, the apparatus 700 further includes: a text display module configured to display the text in the text segment in the electronic book reader.

[0091] In some embodiments, the image is generated based on: an information determination module, configured to determine the description information of the text fragment and the role information of at least one character based on the text work; wherein the text work includes multiple text fragments, and the role information includes attribute information of the character in at least one dimension; an image generation module, configured to generate a character image diagram of the character based on the role information of the character in at least one character; and an image generation module, configured to generate an image of the text fragment based on the description information of the text fragment and the role image diagram of the character.

[0092] In some embodiments, the plurality of text segments is determined based on: a receiving module configured to obtain structural information of the text work; and a dividing module configured to divide the text work into the plurality of text segments based on the structural information.

[0093] In some embodiments, at least one role is determined based on: a number determination module configured to determine the number of occurrences of a target role among multiple roles in a text work in a text segment; and a role determination module configured to select the target role as a role among at least one role in response to determining that the number of occurrences meets a predetermined condition.

[0094] In some embodiments, the information determination module includes: a basic information determination module, configured to determine the role information of the target character based on the part of the text work associated with the target character; and an update module, configured to update the role information of the target character based on the part of the text fragment associated with the target character.

[0095] In some embodiments, the description information includes environmental information of the character's environment and action information of the character, and the image generation module includes: a model generation module, configured to generate a character model for describing the character based on the character image; a prompt word generation module, configured to use the environmental information, action information and character model to generate prompt words for the machine learning model; and a prompt word-based generation module, configured to generate an image based on the prompt words.

[0096] In some embodiments, the prompt word generation module further includes: a style determination module configured to determine the style of the image based on the sound attributes of the audio data; and a prompt word update module configured to update the prompt word based on the style.

[0097] In some embodiments, the attribute information of at least one dimension includes at least any one of the following: the character's name, gender, age, occupation, appearance, expression, and clothing, and the first part of the attribute information of at least one dimension is determined from the text work.

[0098] In some embodiments, the device 700 further includes: a providing module configured to provide a setting control for setting the second part of the attribute information of at least one dimension in the electronic book reader; and a property setting module configured to set the second part of the attribute information of at least one dimension based on the setting operation in response to receiving a setting operation for the setting control.

[0099] The units and / or modules included in the device 700 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units and / or modules in the device 700 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0100] FIG8 shows a block diagram of an electronic device 800 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 800 shown in FIG8 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 800 shown in FIG8 may be used to implement the electronic device of FIG1 and / or the apparatus 600 of FIG6 .

[0101] As shown in FIG8 , electronic device 800 is in the form of a general-purpose computing device. Components of electronic device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. Processing unit 810 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 800.

[0102] The electronic device 800 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 800.

[0103] The electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG8 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 820 may include a computer program product 825 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.

[0104] The communication unit 840 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the electronic device 800 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 800 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0105] The input device 850 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 860 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 800 may also communicate with one or more external devices (not shown) via the communication unit 840 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 800, or with any device that allows the electronic device 800 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0106] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0107] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0108] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0109] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0110] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0111] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for presenting a text work, comprising: In an electronic book reader for reading a text work, in response to receiving a playback request for playing audio data corresponding to the text work, playing the audio data; as well as An image corresponding to the text work is displayed, the image being generated based on a character in a text segment in the text work, the text segment corresponding to a playback position of the audio data.

2. The method according to claim 1, wherein playing the audio data comprises: Obtaining an initial playback position in the audio data specified by the playback request; as well as Play the audio segment corresponding to the initial playback position in the audio data.

3. The method according to claim 2, further comprising: In response to receiving a designation request for designating another initial playback position of the audio data, playing an audio segment in the audio data corresponding to the another initial playback position; determining a current playback position of the audio data based on the other initial playback position and the playback time; Determining a current text segment in the text work corresponding to the current playback position; and An image corresponding to the current text segment is presented.

4. The method according to claim 1, further comprising: The text in the text segment is displayed in the electronic book reader.

5. The method of claim 1 , wherein the image is generated based on: Determine the description information of the text segment and the role information of at least one role according to the text work; wherein, The text work includes a plurality of text segments, and the character information includes attribute information of the character in at least one dimension; For a role among the at least one role, generating a role image of the role based on the role information of the role; as well as An image of the text segment is generated according to the description information of the text segment and the character image of the character.

6. The method of claim 5, wherein the plurality of text segments are determined based on: Obtaining structural information of the textual work; and The text work is divided into the plurality of text segments based on the structural information.

7. The method of claim 5, wherein the at least one role is determined based on: For a target character among the multiple characters in the text work, determining the number of occurrences of the target character in the text segment; and In response to determining that the number of occurrences satisfies a predetermined condition, the target character is selected as a character in the at least one character.

8. The method according to claim 7, wherein determining the role information of the at least one role comprises: For the target character, determining the character information of the target character based on a portion of the text work associated with the target character; as well as Based on the portion of the text segment associated with the target role, the role information of the target role is updated.

9. The method according to claim 5, wherein the description information includes environmental information of the environment in which the character is located and action information of the character, and generating the image of the text segment comprises: generating a role model for describing the role based on the role image; generating prompt words for a machine learning model using the environmental information, the action information, and the role model; as well as The image is generated based on the prompt word.

10. The method according to claim 9, wherein generating the prompt word further comprises: determining a style of the image based on sound attributes of the audio data; as well as The prompt word is updated based on the style.

11. The method according to claim 5, wherein the attribute information of at least one dimension includes at least any one of the following: the name, gender, age, occupation, appearance, expression, and clothing of the character, and the first part of the attribute information of at least one dimension is determined from the text work.

12. The method according to claim 11, further comprising: Providing a setting control in the electronic book reader for setting a second part of the attribute information of the at least one dimension; as well as In response to receiving a setting operation for the setting control, a second portion of the attribute information of the at least one dimension is set based on the setting operation.

13. A text work presentation device, comprising: a playing module configured to play the audio data corresponding to the text work in response to receiving a play request for playing the audio data in an electronic book reader for reading the text work; as well as The display module is configured to display an image corresponding to the text work, wherein the image is generated based on a character in a text segment in the text work, and the text segment corresponds to a playback position of the audio data.

14. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 12 when executed by the at least one processing unit.

15. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Image presentation method and device

    CN112328088A

  • Article voice playing method, apparatus and device, and computer readable storage medium

    CN113010138A

  • Image generation method and device, electronic equipment and storage medium

    CN116894881A

  • Synchronizing the playing and displaying of digital content

    US8290777B1