Text processing method, system and apparatus, device, and storage medium

By processing text in segments and obtaining storage location information in advance, the problem of time spent generating media content by multi-paragraph text content is solved, achieving smooth display of media content and improving user experience.

WO2025113021A1PCT designated stage expired Publication Date: 2025-06-05BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/127509
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-10-25
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

When processing multi-paragraph text content, the overall time of generating media content is long, resulting in a longer loading time after the media content is generated, and a longer time for users to wait for the first frame of media content to be displayed.

Method used

By processing the target text in segments, the media content corresponding to the first text segment is obtained, and the media content storage location information corresponding to the second text segment is obtained in advance. After displaying the first text clip and the first media content, the media content of the second text clip is obtained from the cache based on the storage location information and displayed.

Benefits of technology

It reduces the loading time after media content generation, reduces the waiting time for the first frame media content display, and ensures the smooth display and user experience of media content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024127509_05062025_PF_FP_ABST
    Figure CN2024127509_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a text processing method, system and apparatus, a device, and a storage medium. The method comprises: firstly, in response to a content generation triggering operation for target text, obtaining first media content corresponding to a first text segment in the target text and storage location information corresponding to a second text segment in the target text; then displaying the first text segment and the first media content; and then, on the basis of the storage location information corresponding to the second text segment, obtaining, from a preset cache, second media content corresponding to the second text segment, and displaying the second text segment and the second media content.
Need to check novelty before this filing date? Find Prior Art

Description

Text processing method, system, device, equipment and storage medium

[0001] This application claims priority to the Chinese invention patent application entitled “A text processing method, system, device, equipment and storage medium” and application number 202311632619.X filed on November 30, 2023. The entire contents of that application are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to the field of data processing, and in particular to a text processing method, system, device, equipment, and storage medium. Background Art

[0003] With the continuous development of text processing technology, more and more users use text processing technology to generate media content such as pictures and audio based on text content to enrich the presentation of text content.

[0004] Currently, for multi-paragraph text content, the media content corresponding to each paragraph is usually generated separately before being displayed sequentially. Because the overall time required to generate media content for multi-paragraph text content is relatively long, the loading time after triggering the media content generation operation for multi-paragraph text content is relatively long, which in turn causes users to wait longer for the first frame of media content to be displayed.

[0005] Summary of the Invention

[0006] In order to solve the above technical problems, an embodiment of the present disclosure provides a text processing method.

[0007] In a first aspect, the present disclosure provides a text processing method, the method comprising:

[0008] In response to a content generation trigger operation for a target text, first media content corresponding to a first text segment in the target text and storage location information corresponding to a second text segment are obtained; wherein the first text segment is the first text segment among a plurality of text segments obtained after segmenting the target text, the second text segment is a non-first text segment among the plurality of text segments, the first media content includes media content of a preset content carrier generated based on the first text segment, and the storage location information is used to identify a storage location in a preset cache of second media content pre-generated based on the second text segment;

[0009] displaying the first text segment and the first media content;

[0010] Based on the storage location information corresponding to the second text segment, the second media content corresponding to the second text segment is obtained from the preset cache, and the second text segment and the second media content are displayed; wherein, the second media content includes the media content of the preset content carrier generated based on the second text segment.

[0011] In an optional implementation, the first media content includes first image content and a first audio clip, and the displaying of the first text clip and the first media content includes:

[0012] Display the first text segment and the first image content, and synchronously play the first audio segment; wherein the first audio segment includes a human voice reading voice segment generated based on the first text segment, and the first image content includes a picture, animation or video segment generated based on the first text segment.

[0013] In an optional implementation, the first audio segment further includes a background music segment generated based on the first text segment.

[0014] In an optional embodiment, before obtaining the first media content corresponding to the first text segment in the target text and the storage location information corresponding to the second text segment in response to the triggering operation for the content of the target text, the method further includes:

[0015] Receive the target text entered by the user.

[0016] In an optional embodiment, before obtaining the first media content corresponding to the first text segment in the target text and the storage location information corresponding to the second text segment in response to the triggering operation for the content of the target text, the method further includes:

[0017] Receive content generation parameters set for the target text; wherein the content generation parameters are used to determine the display attributes of the media content of the preset content carrier.

[0018] In an optional embodiment, in response to a trigger operation for generating content of a target text, obtaining first media content corresponding to a first text segment in the target text and storage information corresponding to a second text segment includes:

[0019] In response to a content generation trigger operation for a target text, a content generation request carrying the target text is sent to a target server; wherein the target server is configured to segment the target text to obtain a first text segment and a second text segment, and generate first media content based on the first text segment; wherein the first media content includes media content of a preset content carrier;

[0020] In response to the content generation request, receiving storage location information corresponding to the first media content and the second text segment; wherein the storage location information is used to identify a storage location in a preset cache of the second media content pre-generated based on the second text segment;

[0021] The first text segment and the first media content are displayed.

[0022] In an optional embodiment, the acquiring, based on the storage location information corresponding to the second text segment, the second media content corresponding to the second text segment from the preset cache, and displaying the second text segment and the second media content includes:

[0023] Sending a media content request carrying storage location information corresponding to a target second text segment to the target server; wherein the target second text segment is a text segment adjacent to the currently displayed text segment, and the media content request is used to instruct the target server to obtain second media content corresponding to the target second text segment from the preset cache based on the storage location information;

[0024] The second media content is received, and in response to the completion of presentation of the currently presented text segment, the target second text segment and the second media content are presented.

[0025] In a second aspect, the present disclosure further provides a text processing method, the method comprising:

[0026] In response to a content generation request for a target text, generating first media content based on a first text segment in the target text; wherein the first text segment is a first text segment among a plurality of text segments obtained after segmenting the target text, and the first media content includes media content of a preset content carrier;

[0027] Determining storage location information corresponding to a second text segment in the target text; wherein the storage location information is used to identify a storage location in a preset cache of second media content generated based on the second text segment, the second text segment being a non-first text segment among the multiple text segments;

[0028] Return the storage location information corresponding to the first media content and the second text segment in response to the content generation request:

[0029] A second media content is asynchronously generated based on a second text segment in the target text, and the second media content is stored in a preset cache based on storage location information corresponding to the second text segment.

[0030] In an optional implementation, the method further includes:

[0031] In response to a media content request carrying storage location information corresponding to a target second text segment, obtaining second media content from the preset cache based on the storage location information;

[0032] The second media content is returned in response to the media content request.

[0033] In a third aspect, the present disclosure further provides a text processing system, the system comprising a target client and a target server;

[0034] The target client is configured to send a content generation request carrying the target text to the target server in response to a content generation trigger operation for the target text;

[0035] The target server is configured to generate first media content based on a first text segment in the target text, determine storage location information corresponding to a second text segment in the target text, and return the storage location information corresponding to the first media content and the second text segment to the target client; wherein the first text segment is the first text segment among a plurality of text segments obtained after segmenting the target text, the second text segment is a non-first text segment among the plurality of text segments, the first media content includes media content of a preset content carrier, and the storage location information is used to identify a storage location in a preset cache of the second media content pre-generated based on the second text segment;

[0036] The target client is further configured to display the first text segment and the first media content.

[0037] In an optional implementation, the target server is further configured to asynchronously generate second media content based on a second text segment in the target text, and store the second media content in a preset cache based on storage location information corresponding to the second text segment.

[0038] In an optional embodiment, the target client is further configured to send a media content request carrying storage location information corresponding to a target second text segment to the target server; wherein the target second text segment is the next text segment adjacent to the currently displayed text segment;

[0039] The target server is further configured to, in response to the media content request, obtain second media content from the preset cache based on the storage location information, and return the second media content to the target client;

[0040] The target client is further configured to display the target second text segment and the second media content in response to completion of display of the previously displayed text segment.

[0041] In a fourth aspect, the present disclosure provides a text processing device, comprising:

[0042] A first acquisition module is configured to, in response to a content generation trigger operation for a target text, acquire first media content corresponding to a first text segment in the target text and storage location information corresponding to a second text segment; wherein the first text segment is the first text segment among a plurality of text segments obtained after segmenting the target text, the second text segment is a non-first text segment among the plurality of text segments, the first media content includes media content of a preset content carrier generated based on the first text segment, and the storage location information is used to identify a storage location in a preset cache of second media content pre-generated based on the second text segment;

[0043] A first display module, configured to display the first text segment and the first media content;

[0044] The second display module is used to obtain the second media content corresponding to the second text segment from the preset cache based on the storage location information corresponding to the second text segment, and display the second text segment and the second media content; wherein the second media content includes the media content of the preset content carrier generated based on the second text segment.

[0045] In a fifth aspect, the present disclosure further provides a text processing device, the device comprising:

[0046] a generation module configured to generate first media content based on a first text segment in the target text in response to a content generation request for the target text; wherein the first text segment is a first text segment among a plurality of text segments obtained after segmenting the target text, and the first media content includes media content of a preset content carrier;

[0047] a determination module, configured to determine storage location information corresponding to a second text segment in the target text; wherein the storage location information is used to identify a storage location in a preset cache of second media content generated based on the second text segment, the second text segment being a non-first text segment among the plurality of text segments;

[0048] A first returning module, configured to return storage location information corresponding to the first media content and the second text segment in response to the content generation request;

[0049] The storage module is configured to asynchronously generate second media content based on a second text segment in the target text, and store the second media content in a preset cache based on storage location information corresponding to the second text segment.

[0050] In a sixth aspect, the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device implements the above-mentioned method.

[0051] In a seventh aspect, the present disclosure provides a text processing device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method when executing the computer program.

[0052] In an eighth aspect, the present disclosure provides a computer program product, which includes a computer program / instructions, and the computer program / instructions implement the above method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0054] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0055] FIG1 is a flowchart of a text processing method provided by an embodiment of the present disclosure;

[0056] FIG2 is a schematic diagram of a text processing page provided by an embodiment of the present disclosure;

[0057] FIG3 is a schematic diagram of a media content display page provided by an embodiment of the present disclosure;

[0058] FIG4 is a flowchart of another text processing method provided by an embodiment of the present disclosure;

[0059] FIG5 is a schematic diagram of the structure of a text processing system provided by an embodiment of the present disclosure;

[0060] FIG6 is an interactive diagram of a text processing process provided by an embodiment of the present disclosure;

[0061] FIG7 is a schematic structural diagram of a text processing device provided by an embodiment of the present disclosure;

[0062] FIG8 is a schematic structural diagram of another text processing device provided by an embodiment of the present disclosure;

[0063] FIG9 is a schematic structural diagram of a text processing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0064] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.

[0065] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0066] With the continuous development of text processing technology, more and more users use text processing technology to generate media content such as pictures and audio based on text paragraphs to enrich the presentation of text content.

[0067] Currently, for multi-paragraph text content, the media content corresponding to each paragraph is usually generated separately before being displayed sequentially. Because the overall time required to generate media content for multi-paragraph text content is relatively long, the loading time after triggering the media content generation operation for multi-paragraph text content is relatively long, which in turn causes users to wait longer for the first frame of media content to be displayed.

[0068] To this end, the present disclosure provides a text processing method, which firstly, in response to a content generation trigger operation for a target text, obtains the first media content corresponding to the first text segment in the target text, and the storage location information corresponding to the second text segment; then displays the first text segment and the first media content; then, based on the storage location information corresponding to the second text segment, obtains the second media content corresponding to the second text segment from a preset cache, and displays the second text segment and the second media content. It can be seen that the embodiment of the present disclosure can synchronously obtain the media content corresponding to the first text paragraph of the target text and display it after receiving the content generation trigger operation for the target text, thereby reducing the media content loading time after the content generation operation is triggered, and reducing the user's waiting time for the first frame of media content to be displayed.

[0069] In addition, by pre-acquiring the storage location information of the media content corresponding to the non-first text paragraph, the embodiment of the present disclosure can promptly acquire and display the pre-generated media content of the next text paragraph after the media content of the adjacent previous text paragraph is displayed, thereby ensuring the smooth display of the media content generated based on the target text, thereby ensuring the overall user experience.

[0070] Based on this, an embodiment of the present disclosure provides a text processing method that can be applied to a target client. Referring to FIG1 , which is a flowchart of a text processing method provided by an embodiment of the present disclosure, the method specifically includes:

[0071] S101: In response to a content generation triggering operation for a target text, first media content corresponding to a first text segment in the target text and storage location information corresponding to a second text segment are acquired.

[0072] Among them, the first text segment is the first text paragraph among multiple text paragraphs obtained after segmenting the target text, the second text segment is a non-first text paragraph among the multiple text paragraphs, the first media content includes media content of a preset content carrier generated based on the first text segment, and the storage location information is used to identify the storage location of the second media content pre-generated based on the second text segment in the preset cache.

[0073] The text processing method provided by the embodiments of the present disclosure may be applied to a target client. For example, the target client may include a client deployed on a smart phone, a client deployed on a tablet computer, and the like.

[0074] In the embodiments of the present disclosure, the target text can be used to generate different types of media content, wherein the target text can include text types such as novels, essays, poems, news, etc. In addition, the target text can also include language types such as English and Chinese, which are not limited in the embodiments of the present disclosure.

[0075] The content generation trigger operation for the target text is used to trigger the acquisition of the first media content corresponding to the first text segment in the target text and the storage location information corresponding to the second text segment. The content generation trigger operation for the target text may include a trigger operation for a content generation control corresponding to the target text.

[0076] As shown in Figure 2, a schematic diagram of a text processing page provided in an embodiment of the present disclosure is shown. A content generation control 201 is displayed on the text processing page. When a user clicks on the control, the first media content corresponding to the first text segment and the storage location information corresponding to the second text segment are obtained.

[0077] In the embodiment of the present disclosure, the first text segment may be the first text paragraph among multiple text paragraphs obtained after segmenting the target text.

[0078] In an optional implementation, in order to dynamically adjust the length of each text segment, the target text can also be sent to the target server, which then processes the target text in segments. The specific implementation will be described in detail in the subsequent embodiments and will not be repeated here.

[0079] In the embodiment of the present disclosure, the first media content corresponding to the first text segment may include media content of a preset content carrier generated based on the first text segment, wherein the preset content carrier may include a picture, audio, animation, or video type of content carrier. After the first media content corresponding to the first text segment is generated, the media content may be displayed on a user page so that the user can better read the first text segment based on the first media content.

[0080] In the embodiment of the present disclosure, the second text segment may be a non-first text paragraph among multiple text segments. For example, assuming that three text paragraphs A, B, and C are obtained after segmenting the target text, text paragraph A is the first text segment, and text paragraphs B and C are the second text segments.

[0081] In an optional embodiment, after obtaining multiple text segments corresponding to the target text, the first media content corresponding to the first text segment and the storage location information corresponding to the second text segment in the multiple text segments may also be obtained. The storage location information corresponding to the second text segment may be used to identify the storage location of the second media content pre-generated by the second text segment in a preset cache.

[0082] In actual applications, after the first text segment is displayed, since the storage location information corresponding to the second text segment has been obtained, the second media content corresponding to the second text segment can be directly obtained from the preset cache based on the storage location information and displayed, ensuring the smooth display of the media content generated based on the target text, thereby ensuring the overall user experience.

[0083] S102: Display the first text segment and the first media content.

[0084] In an optional embodiment, the first media content corresponding to the first text segment may include a first image content and a first audio segment. Therefore, after acquiring the first media content, the first text segment and the first image content may be displayed to the user, and the first audio segment may be played synchronously.

[0085] The first audio segment may include a human voice reading segment generated based on the first text segment, and the first image content may include a picture, animation or video segment generated based on the first text segment.

[0086] In an optional embodiment, the first audio segment may also include a background music segment generated based on the first text segment. Therefore, after obtaining the first media content, the first text segment and the first image content may also be displayed to the user, and the human voice reading voice segment and the background audio segment may be played synchronously.

[0087] Figure 3 shows a schematic diagram of a media content display page provided by an embodiment of the present disclosure. A first text segment can be displayed in the text display area, and first image content corresponding to the first text segment, such as an image or animation, can be displayed in the image content display area. Furthermore, a playback control bar can be displayed on the media content display page, allowing the user to adjust the playback progress of the first audio segment.

[0088] Since the embodiment of the present disclosure can synchronously obtain and display the media content corresponding to the first text paragraph of the target text after receiving the content generation trigger operation for the target text, the media content loading time after the content generation operation is triggered is reduced, and the user's waiting time for the display of the first frame of media content is reduced.

[0089] S103: Based on the storage location information corresponding to the second text segment, obtain the second media content corresponding to the second text segment from the preset cache, and display the second text segment and the second media content.

[0090] The second media content includes media content of the preset content carrier generated based on the second text segment.

[0091] Since the second media content corresponding to the second text segment is stored in the preset cache based on the storage location information in advance, after the first text segment and the first media content are displayed, the second media content corresponding to the second text segment can be directly obtained from the preset cache based on the storage location information and displayed, and the user does not need to wait for a long time to watch the second media content corresponding to the first text segment.

[0092] In the text processing method provided by the embodiment of the present disclosure, first, in response to a content generation trigger operation for a target text, the first media content corresponding to the first text segment in the target text and the storage location information corresponding to the second text segment are obtained; then the first text segment and the first media content are displayed; then, based on the storage location information corresponding to the second text segment, the second media content corresponding to the second text segment is obtained from a preset cache, and the second text segment and the second media content are displayed.

[0093] It can be seen that the embodiment of the present disclosure can synchronously obtain and display the media content corresponding to the first text paragraph of the target text after receiving the content generation trigger operation for the target text, thereby reducing the media content loading time after the content generation operation is triggered and reducing the user's waiting time for the display of the first frame of media content.

[0094] In addition, by pre-acquiring the storage location information of the media content corresponding to the non-first text paragraph, the embodiment of the present disclosure can promptly acquire and display the pre-generated media content of the next text paragraph after the media content of the adjacent previous text paragraph is displayed, thereby ensuring the smooth display of the media content generated based on the target text, thereby ensuring the overall user experience.

[0095] In an optional implementation, in response to a triggering operation for generating content of a target text, before obtaining the first media content corresponding to the first text segment in the target text, the method may further include: receiving the target text input by a user.

[0096] In actual applications, a text input box may be displayed on the text processing page, and the target text input by the user may be received based on the text input box.

[0097] As shown in Figure 2, the text processing page may also display a text input box 202, in which the user may enter target text. After receiving the target text entered by the user, in response to a trigger operation generated for the content of the target text, the first media content corresponding to the first text segment is obtained.

[0098] In an optional implementation, before obtaining the first media content corresponding to the first text segment in the target text in response to the content generation triggering operation for the target text, the method may further include: receiving content generation parameters set for the target text.

[0099] The content generation parameters can be used to determine the presentation attributes of the media content of the preset content carrier. For example, if the media content is a human voice reading speech, the content generation parameters set for the target text may include the speed or intonation of the human voice reading speech; if the media content is image content, the content generation parameters set for the target text may include the presentation style of the image content, such as oil painting or sketch style.

[0100] In actual applications, a parameter adjustment control may be displayed on the text processing page, and further, content generation parameters set by the user for the target text may be received based on the parameter adjustment control.

[0101] As shown in FIG2 , the text processing page may also display content adjustment controls 203, such as a speech speed adjustment control and a tone adjustment control. The user can use the content adjustment controls to adjust the speech speed and tone of the generated human voice reading speech, for example, to control the human voice reading speech to play at a faster speed.

[0102] In an optional embodiment, the target client can also send a content generation request carrying the target text to the target server in response to a content generation trigger operation for the target text, so that the target server can segment the target text to obtain a first text segment and a second text segment, and generate the first media content based on the first text segment.

[0103] The first media content may include media content in a preset content carrier, such as pictures, animations, or videos.

[0104] After receiving the content generation request carrying the target text, the target server segments the target text to obtain a first text segment and a second text segment, and determines the first media content corresponding to the first text segment and the storage location information corresponding to the second text segment. Then, the storage location information corresponding to the first media content and the second text segment is sent to the client corresponding to the content generation request, that is, the target client.

[0105] Next, the target client receives the storage location information corresponding to the first media content and the second text segment in response to the content generation request, wherein the storage location information can be used to identify the storage location of the second media content pre-generated based on the second text segment in a preset cache.

[0106] Furthermore, the target client displays the first text segment and the second media content.

[0107] In an optional implementation, based on the storage location information corresponding to the second text segment, obtaining the second media content corresponding to the second text segment from the preset cache, and displaying the second text segment and the second media content may also include:

[0108] A media content request carrying storage location information corresponding to the target second text segment is sent to the target server.

[0109] In the embodiment of the present disclosure, the target second text segment may be a text segment adjacent to the currently displayed text segment, wherein the currently displayed text paragraph may include any one of the first text segment and the second text segment.

[0110] For example, assume that the target text corresponds to multiple text segments, namely, text segment A, text segment B, text segment C, and text segment D, where text segment A is the first text segment. Assuming that the currently displayed text segment is C, the target second text segment is the next text segment adjacent to text segment C, namely, text segment D.

[0111] In the embodiment of the present disclosure, the media content request may be used to instruct the target server to obtain the second media content corresponding to the target second text segment from a preset cache based on the storage location information carried in the request.

[0112] In an embodiment of the present disclosure, when the target server receives a media content request carrying storage location information corresponding to the target second text segment, it obtains the second media content corresponding to the target second text segment from a preset cache based on the storage location information, and sends the second media content to the client corresponding to the media content request, that is, the target client.

[0113] In another optional implementation, the target client receives the second media content from the target server, and displays the target second text segment and the second media content when the currently displayed text segment is finished being displayed.

[0114] It can be seen that the embodiment of the present disclosure can obtain and display the pre-generated media content of the next text segment in a timely manner after the media content corresponding to the currently displayed text segment is displayed, by pre-acquiring the storage location information of the media content corresponding to the non-means text segment.

[0115] To facilitate further understanding of the text processing method provided by the present disclosure, the present disclosure also provides a text processing method. Referring to FIG4 , there is shown a flowchart of another text processing method provided by the present disclosure. Specifically, the text processing method is applied to the target server and specifically includes:

[0116] S401: In response to a content generation request for a target text, generate first media content based on a first text segment in the target text.

[0117] The first text segment is the first text segment among multiple text segments obtained after segmenting the target text, and the first media content includes media content of a preset content carrier.

[0118] In actual applications, upon receiving a trigger operation for generating content for a target text, the target client may further send a content generation request for the target text to the target server to instruct the target server to generate the first media content based on the target text.

[0119] The content generation request may carry a target text. When the target server receives the content generation request for the target text, it may generate the first media content based on the first-end text segment in the target text.

[0120] In an optional manner, before generating the first media content based on the first text segment in the target text, the method may further include: segmenting the target text to obtain multiple text segments; wherein the first text segment may be the first text segment among the multiple text segments.

[0121] In practical applications, a long text classification model based on machine learning or a long text window segmentation algorithm can be used to segment the target text to obtain multiple text paragraphs.

[0122] In order to generate the second media content corresponding to the second text segment as soon as possible before the first text segment and the first media content are displayed, in actual applications, the length of the first text segment can also be controlled during segmentation processing.

[0123] In an optional embodiment, upon receiving a content generation trigger operation for a target text, the target text may be segmented based on a preset text segmentation strategy so that the number of words in the first text segment is no less than a preset threshold. The preset text segmentation strategy is used to control the number of words in the first text paragraph to be no less than a preset threshold.

[0124] In actual applications, in the process of displaying the first text segment and the first media content, the second media content corresponding to the second text segment can be generated and stored in a preset cache based on the storage location information, so that when a media content request carrying the storage location information is subsequently received, the second media content can be obtained from the preset cache in a timely manner based on the storage location information, and the second media content can be returned in response to the media content request, ensuring the smooth display of the media content generated based on the target text, thereby ensuring the overall user experience.

[0125] In an optional implementation, the first media content may include first image content and a first audio segment, wherein the first image content may include a picture, animation, or video segment generated based on the first text segment.

[0126] In practical applications, a deep learning algorithm or a text-to-image model can be used to generate an image corresponding to the first text segment, and an animation or video segment corresponding to the first text segment can be generated using a text-to-video model.

[0127] In an optional implementation, in the process of generating the media content corresponding to the first text segment of the target text, the media content corresponding to the next or multiple adjacent text segments can also be generated synchronously, so that after the media content of the current text segment is displayed, the pre-generated media content of the next text segment can be obtained and displayed in a timely manner, ensuring the smooth display of the media content generated based on the target text, thereby ensuring the overall user experience.

[0128] S402: Determine storage location information corresponding to a second text segment in the target text.

[0129] The storage location information is used to identify a storage location in a preset cache of the second media content generated based on the second text segment, and the second text segment is a non-first text segment among the multiple text segments.

[0130] In practical applications, after the target text is segmented to obtain multiple text paragraphs, corresponding storage location information may be determined for each of the multiple text paragraphs other than the first text paragraph (ie, the second text segment).

[0131] In an optional implementation, a random string can be generated using random number generation technology as the storage location information corresponding to the second text segment to identify the storage location of the second media content in the preset cache, so that the second media content corresponding to the second text segment can be obtained from the preset cache in a timely manner based on the storage location information and displayed, thereby ensuring the smooth display of the media content generated based on the target text, thereby ensuring the overall user experience.

[0132] S403: Return storage location information corresponding to the first media content and the second text segment in response to the content generation request.

[0133] In an embodiment of the present disclosure, after determining the storage location information corresponding to the second text segment, the storage location information corresponding to the second text segment can also be sent to the target client corresponding to the content generation request, so that after receiving the storage location information corresponding to the second text segment, the target client can obtain the second media content from the preset cache based on the storage location information corresponding to the second text segment and display it.

[0134] S404: Asynchronously generate second media content based on a second text segment in the target text, and store the second media content in a preset cache based on storage location information corresponding to the second text segment.

[0135] In actual applications, asynchronous and synchronous are two different message communication mechanisms. Synchronous means responding to a call request and returning the corresponding result synchronously, while asynchronous means receiving a call request and returning the call request at different times.

[0136] In the embodiment of the present disclosure, asynchronously generating the second media content means that before receiving the media content request, the second media content is generated in advance based on the second text segment and stored in a preset cache, and then when the media content request is received later, the second media content is returned in response to the media content request.

[0137] In an optional embodiment, after storing the second media content in the preset cache, the second media content can be retrieved from the preset cache based on the storage location information in response to a media content request carrying the storage location information corresponding to the target second text segment. Furthermore, the second media content can be returned in response to the media content request.

[0138] In the embodiment of the present disclosure, when the target server receives a content generation request for the target text, it can send the first media content corresponding to the first text paragraph to the target client, without having to wait for all paragraphs to generate corresponding media content before returning the media content corresponding to the first paragraph, thereby reducing the media content loading time after the content generation operation is triggered, and reducing the user's waiting time for the first frame of media content to be displayed.

[0139] In addition, by pre-acquiring the storage location information of the media content corresponding to the non-first text paragraph, the embodiment of the present disclosure can promptly acquire and display the pre-generated media content of the next text paragraph after the media content of the adjacent previous text paragraph is displayed, thereby ensuring the smooth display of the media content generated based on the target text, thereby ensuring the overall user experience.

[0140] To facilitate understanding of the overall solution, the present disclosure provides a text processing system. Specifically, the text processing system includes a target client and a target server. Referring to FIG5 , there is shown a schematic diagram of the structure of a text processing system provided in an embodiment of the present disclosure.

[0141] The text processing system 500 includes a target client 501 and a target server 502 .

[0142] Specifically, the target client 501 is configured to send a content generation request carrying the target text to the target server in response to a content generation trigger operation for the target text.

[0143] The target server 502 is used to generate first media content based on the first text segment in the target text, determine the storage location information corresponding to the second text segment in the target text, and return the storage location information corresponding to the first media content and the second text segment to the target client; wherein, the first text segment is the first text paragraph among multiple text paragraphs obtained after segmenting the target text, the second text segment is a non-first text paragraph among the multiple text paragraphs, the first media content includes media content of a preset content carrier, and the storage location information is used to identify the storage location of the second media content pre-generated based on the second text segment in a preset cache.

[0144] In addition, the target client 501 is also used to display the first text segment and the first media content.

[0145] In an embodiment of the present disclosure, when a target client receives a content generation trigger operation for a text, it sends a content generation request carrying the target text to a target server.

[0146] After receiving the content generation request sent by the target client, the target server first segments the target text based on the target text carried in the content generation request to obtain multiple text segments, generates the first media content based on the first text segment among the multiple text segments, and determines the storage location information corresponding to the second text segment among the multiple text segments.

[0147] After generating the first media content based on the first text segment and determining the storage location information corresponding to the second text segment, the target server returns the first media content and the storage location information corresponding to the second text segment to the target client.

[0148] After receiving the storage location information corresponding to the first media content and the second text segment returned by the target server, the target client displays the first text segment and the first media content.

[0149] The target client is further configured to send a media content request carrying storage location information corresponding to a target second text segment to the target server, wherein the target second text segment is the next text segment adjacent to the currently displayed text segment.

[0150] When receiving the media content request sent by the target client, the target server obtains the second media content from the preset cache based on the storage location information and returns the second media content to the target client.

[0151] When determining that the display of the currently displayed text segment is finished, the target client displays the target second text segment and the second media content.

[0152] In the text processing system provided by the embodiments of the present disclosure, the target server can send the first media content corresponding to the first text paragraph to the target client upon receiving a content generation request for the target text, without having to wait for all paragraphs to generate corresponding media content before returning the media content corresponding to the first paragraph. This reduces the media content loading time after the content generation operation is triggered, and reduces the user's waiting time for the first frame of media content to be displayed.

[0153] In addition, by pre-acquiring the storage location information of the media content corresponding to the non-first text paragraph, the embodiment of the present disclosure can promptly acquire and display the pre-generated media content of the next text paragraph after the media content of the adjacent previous text paragraph is displayed, thereby ensuring the smooth display of the media content generated based on the target text, thereby ensuring the overall user experience.

[0154] To facilitate understanding of the above embodiments, the present disclosure will be described using a specific interaction scenario. As shown in Figure 6, the present disclosure provides an interaction diagram of the text processing process. Specifically, the text processing process is described using the example of generating a human voice reading, background music, and images corresponding to the target text.

[0155] The target client receives the target text input by the user and the content generation parameters set for the target text, such as speaking speed and intonation, and sends a content generation request carrying the target text and content generation parameters to the target server.

[0156] The target server receives a content generation request carrying a target text and content generation parameters, segments the target text to obtain multiple text paragraphs, and storage location information corresponding to the second text segment; and sends a speech generation request carrying the first text segment (i.e., the first text segment among the multiple text paragraphs) to the text-to-speech service.

[0157] The text-to-speech service receives a speech generation request carrying a first text segment, generates a human voice reading speech segment corresponding to the first text segment, and returns the human voice reading speech segment in response to the speech generation request.

[0158] The target server may also generate corresponding background music based on the first text segment, or select relevant music from a preset music library as the background music corresponding to the first text segment.

[0159] In actual applications, the target server may further send a picture generation request carrying a keyword corresponding to the first text segment to the text-picture service to instruct the text-picture service to generate a picture corresponding to the first text segment.

[0160] Before sending the keywords corresponding to the first text segment to the text-image service, the target text can also be segmented to obtain the keywords corresponding to the first text segment. In practical applications, for English text segments, spaces can be used for word segmentation, while for Chinese text segments, third-party libraries can be used for word segmentation. After word segmentation, the text segments can be more accurately extracted from the keywords.

[0161] The text-image service receives an image generation request carrying a first text segment, generates an image corresponding to the first text segment, and returns the image corresponding to the first text segment in response to the image generation request.

[0162] The target server returns the segmentation information, the first media content (ie, the human voice reading speech, background music, and pictures corresponding to the first text segment), and the storage location information corresponding to the second text segment to the target client.

[0163] The target client receives the segmentation information, the first media content, and the storage location information corresponding to the second text segment, and displays the first text segment and the first media content.

[0164] While the target client is displaying the first text segment and the first media content, the target server continues to execute the step of obtaining the human voice reading speech, background music, and pictures corresponding to the second text segment.

[0165] When receiving the media content request from the target client, the target server obtains the second media content from the preset cache based on the storage location information; and returns the second media content to the target client.

[0166] The target client receives the second media content corresponding to the target second text segment, and displays the target second text segment and the second media content.

[0167] Corresponding to the above method embodiment, the present disclosure further provides a text processing device. Referring to FIG7 , there is shown a schematic diagram of the structure of a text processing device provided in an embodiment of the present disclosure. Specifically, the device includes:

[0168] A first acquisition module 701 is configured to, in response to a content generation trigger operation for a target text, acquire first media content corresponding to a first text segment in the target text and storage location information corresponding to a second text segment; wherein the first text segment is the first text segment among a plurality of text segments obtained after segmenting the target text, the second text segment is a non-first text segment among the plurality of text segments, the first media content includes media content of a preset content carrier generated based on the first text segment, and the storage location information is used to identify a storage location in a preset cache of second media content pre-generated based on the second text segment;

[0169] A first display module 702, configured to display the first text segment and the first media content;

[0170] The second display module 703 is used to obtain the second media content corresponding to the second text segment from the preset cache based on the storage location information corresponding to the second text segment, and display the second text segment and the second media content; wherein the second media content includes the media content of the preset content carrier generated based on the second text segment.

[0171] In an optional embodiment, the display module includes:

[0172] The first display submodule is used to display the first text segment and the first image content, and to synchronously play the first audio segment; wherein the first audio segment includes a human voice reading voice segment generated based on the first text segment, and the first image content includes a picture, animation or video segment generated based on the first text segment.

[0173] In an optional implementation, the first audio segment further includes a background music segment generated based on the first text segment.

[0174] In an optional embodiment, the device further includes:

[0175] The first receiving module is configured to receive a target text input by a user.

[0176] In an optional embodiment, the device further includes:

[0177] The second receiving module is configured to receive content generation parameters set for the target text; wherein the content generation parameters are used to determine display attributes of the media content of the preset content carrier.

[0178] In an optional implementation, the acquisition module includes:

[0179] A first sending submodule is configured to, in response to a content generation trigger operation for a target text, send a content generation request carrying the target text to a target server; wherein the target server is configured to segment the target text to obtain a first text segment and a second text segment, and generate first media content based on the first text segment; wherein the first media content includes media content of a preset content carrier;

[0180] A first receiving submodule is configured to receive, in response to the content generation request, storage location information corresponding to the first media content and the second text segment; wherein the storage location information is used to identify a storage location in a preset cache of the second media content pre-generated based on the second text segment;

[0181] The second display submodule is configured to display the first text segment and the first media content.

[0182] In an optional embodiment, the second display module includes:

[0183] a second sending submodule for sending a media content request carrying storage location information corresponding to a target second text segment to the target server; wherein the target second text segment is a text segment adjacent to the currently displayed text segment, and the media content request is used to instruct the target server to obtain second media content corresponding to the target second text segment from the preset cache based on the storage location information;

[0184] The second receiving submodule is configured to receive the second media content and, in response to completion of display of the currently displayed text segment, display the target second text segment and the second media content.

[0185] In addition, an embodiment of the present disclosure further provides a text processing device. Referring to FIG8 , which is a schematic structural diagram of another text processing device provided by an embodiment of the present disclosure, the device specifically includes:

[0186] A generation module 801 is configured to generate first media content based on a first text segment in the target text in response to a content generation request for the target text; wherein the first text segment is a first text segment among a plurality of text segments obtained after segmenting the target text, and the first media content includes media content of a preset content carrier;

[0187] Determining module 802, configured to determine storage location information corresponding to a second text segment in the target text; wherein the storage location information is used to identify a storage location in a preset cache of second media content generated based on the second text segment, the second text segment being a non-first text segment among the plurality of text segments;

[0188] A first returning module 803 is configured to return storage location information corresponding to the first media content and the second text segment in response to the content generation request;

[0189] The storage module 804 is configured to asynchronously generate second media content based on a second text segment in the target text, and store the second media content in a preset cache based on storage location information corresponding to the second text segment.

[0190] In an optional embodiment, the device further includes:

[0191] a second acquisition module, configured to respond to a media content request carrying storage location information corresponding to a target second text segment and acquire second media content from the preset cache based on the storage location information;

[0192] The second returning module is configured to return the second media content in response to the media content request.

[0193] The embodiment of the present disclosure provides a text processing device, which firstly, in response to a content generation trigger operation for a target text, obtains the first media content corresponding to the first text segment in the target text, and the storage location information corresponding to the second text segment; then displays the first text segment and the first media content; then, based on the storage location information corresponding to the second text segment, obtains the second media content corresponding to the second text segment from a preset cache, and displays the second text segment and the second media content. It can be seen that the embodiment of the present disclosure can synchronously obtain the media content corresponding to the first text paragraph of the target text and display it after receiving the content generation trigger operation for the target text, thereby reducing the media content loading time after the content generation operation is triggered, and reducing the user's waiting time for the first frame of media content to be displayed.

[0194] In addition, by pre-acquiring the storage location information of the media content corresponding to the non-first text paragraph, the embodiment of the present disclosure can promptly acquire and display the pre-generated media content of the next text paragraph after the media content of the adjacent previous text paragraph is displayed, thereby ensuring the smooth display of the media content generated based on the target text, thereby ensuring the overall user experience.

[0195] In addition to the above-mentioned method and apparatus, the embodiments of the present disclosure further provide a computer-readable storage medium, which stores instructions. When the instructions are executed on a terminal device, the terminal device implements the text processing method described in the embodiments of the present disclosure.

[0196] The embodiments of the present disclosure further provide a computer program product, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, the text processing method described in the embodiments of the present disclosure is implemented.

[0197] In addition, the present disclosure also provides a text processing device, as shown in FIG9 , which may include:

[0198] Processor 901, memory 902, input device 903, and output device 909. The text processing device may include one or more processors 901, with one processor being used as an example in FIG9 . In some embodiments of the present disclosure, processor 901, memory 902, input device 903, and output device 909 may be connected via a bus or other means, with FIG9 using a bus as an example.

[0199] The memory 902 can be used to store software programs and modules. The processor 901 executes the various functional applications and data processing of the text processing device by running the software programs and modules stored in the memory 902. The memory 902 may primarily include a program storage area and a data storage area. The program storage area may store an operating system, at least one application required for a function, and the like. Furthermore, the memory 902 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. The input device 903 may be used to receive input digital or character information and generate signal input related to user settings and function control of the text processing device.

[0200] Specifically in this embodiment, the processor 901 will load the executable files corresponding to the processes of one or more applications into the memory 902 according to the following instructions, and the processor 901 will run the applications stored in the memory 902, thereby realizing the various functions of the above-mentioned text processing device.

[0201] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0202] The foregoing description is intended only to provide specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but rather to be construed in the broadest manner consistent with the principles and novel features disclosed herein.

Claims

1. A text processing method, the method comprising: In response to a content generation trigger operation for a target text, first media content corresponding to a first text segment in the target text and storage location information corresponding to a second text segment are obtained; wherein the first text segment is the first text segment among multiple text segments obtained after segmenting the target text, the second text segment is a non-first text segment among the multiple text segments, the first media content includes media content of a preset content carrier generated based on the first text segment, and the storage location information is used to identify a storage location in a preset cache of second media content pre-generated based on the second text segment; Displaying the first text segment and the first media content; Based on the storage location information corresponding to the second text segment, the second media content corresponding to the second text segment is obtained from the preset cache, and the second text segment and the second media content are displayed; wherein the second media content includes the media content of the preset content carrier generated based on the second text segment.

2. The method according to claim 1, wherein: The first media content includes a first image content and a first audio segment, and the displaying of the first text segment and the first media content includes: The first text segment and the first image content are displayed, and the first audio segment is played synchronously; wherein the first audio segment includes a voice segment generated based on the first text segment, and the first image content includes a picture, animation or video segment generated based on the first text segment.

3. The method according to claim 2, wherein: The first audio segment also includes a background music segment generated based on the first text segment.

4. The method according to claim 1, wherein: Before obtaining the first media content corresponding to the first text segment in the target text and the storage location information corresponding to the second text segment in response to the trigger operation for the content generation of the target text, the method further includes: Receive the target text entered by the user.

5. The method according to claim 4, wherein: Before obtaining the first media content corresponding to the first text segment in the target text and the storage location information corresponding to the second text segment in response to the trigger operation for the content generation of the target text, the method further includes: Receive content generation parameters set for the target text; wherein the content generation parameters are used to determine the display attributes of the media content of the preset content carrier.

6. The method according to claim 1, wherein: The step of obtaining, in response to a trigger operation for generating content of a target text, first media content corresponding to a first text segment in the target text and storage information corresponding to a second text segment includes: In response to a content generation trigger operation for a target text, a content generation request carrying the target text is sent to a target server; wherein the target server is used to segment the target text to obtain a first text segment and a second text segment, and generating first media content based on the first text segment; wherein the first media content includes media content of a preset content carrier; In response to the content generation request, receiving storage location information corresponding to the first media content and the second text segment; wherein the storage location information is used to identify a storage location in a preset cache of the second media content pre-generated based on the second text segment; The first text segment and the first media content are displayed.

7. The method according to claim 6, wherein: The acquiring, based on the storage location information corresponding to the second text segment, the second media content corresponding to the second text segment from the preset cache, and displaying the second text segment and the second media content, includes: Sending a media content request carrying storage location information corresponding to a target second text segment to the target server; wherein the target second text segment is the next adjacent text segment of the currently displayed text segment, and the media content request is used to instruct the target server to obtain the second media content corresponding to the target second text segment from the preset cache based on the storage location information; The second media content is received, and in response to the completion of display of the currently displayed text segment, the target second text segment and the second media content are displayed.

8. A text processing method, the method comprising: In response to a content generation request for a target text, generating first media content based on a first text segment in the target text; wherein the first text segment is a first text segment among a plurality of text segments obtained after segmenting the target text, and the first media content includes media content of a preset content carrier; Determine storage location information corresponding to a second text segment in the target text; wherein the storage location information is used to identify a storage location in a preset cache of second media content generated based on the second text segment, and the second text segment is a non-first text segment among the multiple text segments; Return the storage location information corresponding to the first media content and the second text segment in response to the content generation request: A second media content is asynchronously generated based on a second text segment in the target text, and the second media content is stored in a preset cache based on storage location information corresponding to the second text segment.

9. The method according to claim 8, wherein: The method further comprises: In response to a media content request carrying storage location information corresponding to a target second text segment, obtaining second media content from the preset cache based on the storage location information; The second media content is returned in response to the media content request.

10. A text processing system, the system comprising a target client and a target server; The target client is used to send a content generation request carrying the target text to the target server in response to a content generation trigger operation for the target text; The target server is configured to generate first media content based on a first text segment in the target text, and Determine the storage location information corresponding to the second text segment in the target text, and return the storage location information corresponding to the first media content and the second text segment to the target client; wherein, The first text segment is the first text segment among multiple text segments obtained after segmenting the target text, the second text segment is a non-first text segment among the multiple text segments, the first media content includes media content of a preset content carrier, and the storage location information is used to identify a storage location of second media content pre-generated based on the second text segment in a preset cache; The target client is also used to display the first text segment and the first media content.

11. The system according to claim 10, wherein: The target server is further configured to asynchronously generate second media content based on a second text segment in the target text, and store the second media content in a preset cache based on storage location information corresponding to the second text segment.

12. The system according to claim 10, wherein: The target client is further configured to send a media content request carrying storage location information corresponding to a target second text segment to the target server; wherein the target second text segment is a next text segment adjacent to the currently displayed text segment; The target server is further configured to, in response to the media content request, obtain second media content from the preset cache based on the storage location information, and return the second media content to the target client; The target client is further configured to display the target second text segment and the second media content in response to the completion of display of the currently displayed text segment.

13. A text processing device, comprising: A first acquisition module is used to obtain, in response to a content generation trigger operation for a target text, first media content corresponding to a first text segment in the target text and storage location information corresponding to a second text segment; wherein the first text segment is the first text segment among multiple text segments obtained after segmenting the target text, the second text segment is a non-first text segment among the multiple text segments, the first media content includes media content of a preset content carrier generated based on the first text segment, and the storage location information is used to identify a storage location in a preset cache of second media content pre-generated based on the second text segment; A first display module, used to display the first text segment and the first media content; A second display module is used to obtain the second media content corresponding to the second text segment from the preset cache based on the storage location information corresponding to the second text segment, and to display the second text segment and the second media content; wherein the second media content includes the media content of the preset content carrier generated based on the second text segment.

14. A text processing device, comprising: A generation module, configured to generate first media content based on a first text segment in the target text in response to a content generation request for the target text; wherein the first text segment is a first text segment among a plurality of text segments obtained after segmenting the target text, and the first media content includes media content of a preset content carrier; A determination module, used to determine storage location information corresponding to a second text segment in the target text; wherein the storage location information is used to identify a storage location in a preset cache of second media content generated based on the second text segment, and the second text segment is a non-first text segment among the multiple text segments; A first returning module, configured to return storage location information corresponding to the first media content and the second text segment in response to the content generation request; The storage module is used to asynchronously generate second media content based on a second text segment in the target text, and store the second media content in a preset cache based on storage location information corresponding to the second text segment.

15. A computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, and when the instructions are executed on a terminal device, the terminal device implements the method according to any one of claims 1 to 9.

16. A text processing device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Video playing method, terminal and storage medium

    CN108260014A

  • Media content display method and device, electronic equipment and storage medium

    CN112861041A

  • Video generation method, video generation device, neural network training method and neural network training device

    CN114254158A

  • Program Based Caching In Live Media Distribution

    US20140189140A1

  • Method and apparatus for generating multi-media presentations

    US6081262A