Long text voice broadcast method and device and medium
By splitting long text files into text segments and synthesizing speech using short text synthesis voice interfaces, the problem of delay in online synthesis of long text is solved, and high timeliness and smooth voice broadcasting is achieved.
Patent Information
- Application Number
- CN202510384377.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-27
AI Technical Summary
When processing long text files, the existing long text online synthesis voice interface has minute-level delays, which is difficult to meet the high-time requirements of real-time scenarios.
By obtaining a long text file, dividing it into several text segments based on preset segmentation rules, generating corresponding text voice segments, and calling the short text synthesis voice interface to obtain and synthesize voice files in sequence according to the order of text voice segments to realize long text voice broadcast.
It effectively solves the minute-level delay problem of long text online synthesized voice interface, meets the high-timedness needs of real-time scenarios for voice broadcasts, and improves user experience.
Smart Images

Figure CN120220644A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a long text voice broadcast method, device, and medium. Background Art
[0002] Online speech synthesis (Text to Speech, TTS) is a technology that converts text into natural speech and is widely used in fields such as audiobooks, voice assistants, navigation systems, and barrier-free services. When a terminal uses a text-to-speech interface, it generally uses a short text online speech synthesis interface, a long text online speech synthesis interface, or a streaming text online synthesis interface. The short text online speech synthesis interface has a fast return speed, the long text online speech synthesis interface can process a large amount of text at one time, and the streaming text online synthesis interface is suitable for real-time scenarios.
[0003] However, when processing a long text file, if it is processed through the long text online speech synthesis interface, it cannot be immediately returned and broadcast, and its minute-level delay is difficult to meet the requirements of real-time scenarios such as navigation and news. Therefore, there is an urgent need for a long text voice broadcast method that can intelligently parse the text structure and optimize the synthesis process. Summary of the Invention
[0004] To solve the above problems, this application proposes a long text voice broadcast method, including: Obtain a long text file, and based on a preset segmentation rule, divide the long text file into several text segments according to the string information included in the long text file; Generate corresponding text voice segments for the text segments, and determine the paragraph identifier and voice file address corresponding to the text voice segments; Call a preset short text speech synthesis interface, and sequentially obtain the corresponding text voice segments from the voice file addresses in the order of the text voice segments, and synthesize the text voice segments into the long text voice file for voice broadcast.
[0005] In an implementation manner of this application, sequentially obtaining the corresponding text voice segments from the voice file addresses in the order of the text voice segments specifically includes: Determine the order of the text voice segments according to the paragraph identifier corresponding to the text voice segment; Determine the first text voice segment in the text voice segments, and generate a first broadcast request for the first text voice segment; In response to the first broadcast request, obtain the first text voice segment according to the first voice file address corresponding to the first text voice segment, and perform voice broadcast on the first text voice segment; While the first text voice segment is being voice broadcasted, other text voice segments located after the first text voice segment are pre-loaded in the above-mentioned order, so that after the voice broadcast of the first text voice segment is completed, the other text voice segments are automatically broadcasted, realizing the synthesized voice broadcast of the text voice segments.
[0006] In an implementation manner of the present application, other text voice segments located after the first text voice segment are pre-loaded in the above-mentioned order, so that after the voice broadcast of the first text voice segment is completed, the other text voice segments are automatically broadcasted. Specifically, it includes: Generate a next broadcast request for the next text voice segment located after the first text voice segment in the above-mentioned order; In response to the next broadcast request, obtain the next text voice segment according to the second voice file address corresponding to the next text voice segment, and broadcast the next text voice segment; While the next text voice segment is being voice broadcasted, repeat the above process until the pre-loading and broadcast of the other text voice segments are completed.
[0007] In an implementation manner of the present application, generating a next broadcast request for the next text voice segment located after the first text voice segment specifically includes: Determine the broadcast device for broadcasting the long text voice file; According to the device performance and network load information of the broadcast device, determine the synthesis strategy corresponding to the next broadcast request; According to the synthesis strategy, determine whether the next broadcast request is a single broadcast request or a combined broadcast request.
[0008] In an implementation manner of the present application, based on a preset segmentation rule, according to the string information included in the long text file, the long text file is segmented into several text segments. Specifically, it includes: Based on a preset segmentation rule, determine the string length required for each text segment; According to the string length, segment the long text file into several text segments; among them, the string length of the first text segment in the text segments is the smallest.
[0009] In an implementation manner of the present application, segmenting the long text file into several text segments according to the string length specifically includes: Segment the long text file into several text segments according to the string length, For other text segments except the first text segment, perform semantic analysis on other string information included in the other text segments to extract the core text in the other string information; If the core text corresponds to multiple text segments, merge the multiple text segments into the same text segment.
[0010] In one implementation manner of the present application, merging the multiple text segments into the same text segment specifically includes: In the case where the length of the merged text segment exceeds a preset string length threshold or the string length corresponding to the subsequent text segment, perform combined broadcast on the text voice segment corresponding to the core text.
[0011] In one implementation manner of the present application, the method further includes: Determine the usage scenario corresponding to the long text file; In the case where the usage scenario is a specified usage scenario, generate a text summary corresponding to the long text file according to the core text, so as to broadcast the text summary when broadcasting the long text file.
[0012] An embodiment of the present application provides a long text voice broadcast device, and the device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a long text voice broadcast method as described in any one of the above.
[0013] An embodiment of the present application provides a non-volatile computer storage medium, storing computer-executable instructions, and the computer-executable instructions are set as: A long text voice broadcast method as described in any one of the above.
[0014] The long text voice broadcast method proposed by the present application can bring the following beneficial effects: Generate corresponding text voice segments for each text segment, determine their paragraph identifiers and voice file addresses, and then use a preset short text synthesis voice interface to sequentially obtain and synthesize a long text voice file from the voice file addresses in the order of the text voice segments for broadcast, effectively solving the problem of minute-level delay existing in the long text online synthesis voice interface, meeting the high timeliness requirements of voice broadcast in real-time scenarios, and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings: Figure 1 It is a schematic flowchart of a long text voice broadcast method provided by an embodiment of the present application; Figure 2 It is a schematic structural diagram of a long text voice broadcast device provided by an embodiment of the present application. Detailed implementation manners
[0016] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0017] The following will detail the technical solutions provided by each embodiment of the present application in conjunction with the drawings.
[0018] As Figure 1 shown, a long text voice broadcast method provided by an embodiment of the present application includes: S101: Obtain a long text file, and based on a preset segmentation rule, divide the long text file into several text segments according to the string information included in the long text file.
[0019] A long text file refers to text content with a relatively long length, such as novels, news reports, academic papers, etc., and usually contains multiple paragraphs or chapters. When broadcasting a long text file, if the long text online synthesis voice interface is directly used to perform text-to-speech broadcasting on the long text file, response delays may occur due to the excessive length of the text content. Therefore, the independent text segments are segmented from the long text file according to the preset segmentation rule in the embodiments of the present application. The core logic is to analyze the string information of the voice content and combine the preset segmentation rule to achieve automated and intelligent text segmentation.
[0020] S102: Generate corresponding text voice segments for the text segments, and determine the paragraph identifiers and voice file addresses corresponding to the text voice segments.
[0021] For each text segment obtained after segmentation, a text-to-speech segment is generated by converting it into corresponding speech data using text-to-speech technology. Each text-to-speech segment corresponds to a corresponding paragraph identifier, which is used to mark the segmentation position of the text-to-speech segment and can help determine the synthesis order of the text-to-speech segments. The voice file address is the specific location identifier for storing the voice file on the network or locally. It is usually a string containing information such as the server address and file path. Through this address, the corresponding voice file can be accurately found and obtained.
[0022] When the server performs paragraph segmentation on a long text file, it generates a unique paragraph identifier for each text-to-speech segment corresponding to a text paragraph. Then, each generated text-to-speech segment is stored in the form of a file on the server or a local storage device. When storing, a unique voice file address is generated for each voice file according to certain rules. For example, in server storage, voice files may be stored according to a certain directory structure, and the voice file address contains information such as the server's IP address, port number, and file path; in local storage, the voice file address may be the file path on the local hard disk. Finally, a correspondence relationship is established between the text-to-speech segment, the paragraph identifier, and the voice file address. Each paragraph identifier is associated with the corresponding voice file address, which facilitates quickly and accurately finding the corresponding voice file for playback in subsequent voice broadcast and other applications.
[0023] In one embodiment, when segmenting a long text file, the preset segmentation rule specifies the character length of each text segment. According to the preset segmentation rule, the number of strings that each text segment should contain is determined. Starting from the beginning of the long text file, several text segments are sequentially segmented according to the determined string length. The string length of each text segment needs to be no greater than the string length threshold. Starting from the first segmented text segment, the string length gradually increases. For text segments with a relatively early order, the requested text is shorter, which can ensure the quick start of voice broadcast and reduce the response delay. At the same time, as resources are gradually released after the device completes the playback of the first segment, the single-segment processing load is gradually increased, matching the dynamic changes in device performance.
[0024] In one embodiment, after splitting a long text file into several text segments according to the string length, it is also necessary to consider whether the splitting of each text segment meets the core requirements for obtaining user information. Therefore, for other text segments except the first one, natural language processing technology is used to perform semantic analysis on these other text segments. During the semantic analysis process, it is necessary to include analyzing the theme, keywords, sentiment tendency, etc. of the text segment, with the aim of extracting the core text information in each text segment. Here, the core text information can reflect the main idea information in a long text file and is applicable to online speech synthesis scenarios such as intelligent navigation and intelligent customer service. After extracting the core text, it is determined whether the core text is split into multiple text segments. If the text voice segments are directly generated according to the split text segments and then synthesized for voice broadcast, it may cause the loss of semantic coherence. Therefore, for such cases where the core text corresponds to multiple text segments, it is necessary to merge the multiple text segments into the same text segment. By merging the paragraphs with the same core text, the narration around the same theme or information is ensured to be coherent. The merged text segment makes the voice broadcast more natural and fluent, enabling the listener to better follow the content without being confused or interrupted in thinking due to frequent paragraph switching, thus improving the user's listening experience.
[0025] It should be noted that after merging the text segments, it is necessary to determine whether the string length of the merged text segment exceeds the preset string length threshold or whether there is a mismatch with the string length of the subsequent text segment. For example, if the preset threshold is 60 characters and the merged text segment reaches 120 characters, or the string length of the subsequent text segment is significantly smaller than that of the merged text segment, subsequent combined broadcast processing is triggered. When the above judgment conditions are met, the server will perform combined broadcast on the text voice segments corresponding to the core text. Specifically, when requesting the text voice segments, multiple text voice segments corresponding to the core text are requested simultaneously, so as to achieve combined broadcast of these multiple text voice segments.
[0026] In one embodiment, for some special scenarios that require voice announcements, such as navigation scenarios, meeting recording scenarios, etc., users only focus on specific key points of content and do not have to repeat the navigation information and the overall meeting content verbatim. Therefore, when performing long-text voice announcements, it is necessary to first determine the usage scenario corresponding to the long-text file. When the usage scenario is the specified usage scenario, according to the core text, generate a text summary corresponding to the long-text file. In this way, when announcing the long-text file, directly announce the text summary instead of the complete long-text content. Through the text summary, when users face long information, they can understand the most important content in a short time without spending a lot of time listening to unnecessary detailed information, effectively improving the efficiency of users' information acquisition. At the same time, it also effectively reduces the amount of content for text-to-speech synthesis and saves computing power. For example, when performing long-text news announcements, viewers usually hope to understand the most important news points within a limited time rather than being overwhelmed by excessive background information or details. Especially for reports on some major events, what viewers care most about are the core content of the event, the latest progress, etc. Generate a text summary based on its content to highlight the core information of the event, such as the time, place, main characters, key events, etc. of the major event. When announcing, first announce this summary part to allow viewers to quickly understand the key points of the news and meet the viewers' need to quickly obtain key information.
[0027] S103: Call the preset short-text synthesis voice interface, and sequentially obtain the corresponding text voice segments from the voice file address according to the order of the text voice segments, and synthesize the text voice segments into a long-text voice file for voice announcement.
[0028] The above process splits the long-text file into multiple text paragraphs and generates corresponding short-text voice segments for the text paragraphs. When performing voice announcements on the long-text voice file, for these segmented text voice segments, call the short-text synthesis voice interface, sequentially obtain the corresponding text voice segments from the voice file address according to the order of the text voice segments, and then combine the multiple text voice segments together in order to form a complete and longer voice file, and finally present the voice content of a long text completely for voice announcement for users to listen to.
[0029] In one embodiment, the paragraph identifier corresponding to each text voice segment is identified, and then, according to the text order represented by the paragraph identifier, the playback order of all text voice segments from the first to the last is determined. From all the text voice segments, the first text voice segment ranked at the forefront is found, and a corresponding first playback request is generated for the first text voice segment. The playback request is a request signal or instruction for obtaining and playing a specific text voice segment. The server will, in response to the playback request, obtain the voice data of the first text voice segment, that is, the first text voice segment, according to the first voice file address corresponding to the first text voice segment, and then perform voice playback of the first text voice segment for the user to listen to.
[0030] While playing the first text voice segment, the server will, according to the previously determined order, pre-obtain the voice data of other text voice segments located after the first text voice segment, that is, perform a preloading operation. In this way, when the first text voice segment is finished playing, the subsequent text voice segments that have been pre-loaded can be automatically played without manual intervention, so as to realize the coherent synthesis of all text voice segments for voice playback, bringing a smooth auditory experience to the user and avoiding stuttering or delay caused by the connection between text voice segments.
[0031] Preloading essentially means that during the playback of the previous text voice segment, the subsequent text voice segments are obtained simultaneously, so that at the moment when the playback of the previous text voice segment is completed, the subsequent text voice segments can be directly played, making the entire text voice sequence coherent and smooth.
[0032] Specifically, while completing the voice playback of the first text voice segment, according to the order of the text voice segments, a next playback request for the next text voice segment located after the first text voice segment is generated. The server will, in response to the generated next playback request, obtain the next text voice segment corresponding to the next text voice segment according to the second voice file address corresponding to the next text voice segment. After obtaining the next text voice segment, the next text voice segment is automatically played. While playing the next text voice segment, the system continues to perform a preloading operation on other text voice segments located after this text voice segment according to the order. That is to say, the server will repeat the previous steps, generate a playback request for the next text voice segment, obtain the voice data and prepare for playback. This process will be repeated continuously until all text voice segments are pre-loaded and played.
[0033] For example, for a 300-word string, after cutting it into multiple text segments according to the character counts of 10, 20, 40, 60... corresponding text-to-speech segments will be generated for the text segments. When playing the text-to-speech segments, first request text-to-speech segment 1. After the voice file address of text-to-speech segment 1 is returned, play it and request the voice file of text-to-speech segment 2. After the voice file address corresponding to text-to-speech segment 2 is returned and saved locally, request the voice file of text-to-speech segment 3, and so on. After text-to-speech segment 1 is played, play the voice file of text-to-speech segment 2, and so on.
[0034] It should be noted that to improve the playback efficiency of long text files, according to the terminal performance, the corresponding synthesis strategy of text-to-speech segments can be selected, that is, when playing other text-to-speech segments except the first one, whether to play them one by one or in groups.
[0035] Specifically, determine the playback device currently used for playing the long text voice file, and real-time monitor the performance status of the playback device, such as the usage rate of the processor, the remaining space of the memory, etc. At the same time, the load information of the current network will also be obtained, such as parameters such as network bandwidth and latency, to judge the stability and transmission ability of the network. According to the evaluation results of device performance and network load, generate the corresponding synthesis strategy. If the device performance is good and the network load is light, a more complex speech synthesis algorithm can be selected, that is, request multiple groups of text-to-speech segments for combined synthesis at the same time; if the device performance is limited or the network load is heavy, for example, when using an older model terminal device with limited processing performance for text playback, in order to ensure the smoothness of playback, just play them one by one in the order of text-to-speech segments.
[0036] Based on the synthesis strategy, decide whether the next playback request is a single playback request or a combined playback request. A single playback request is a playback request issued separately for one text-to-speech segment. When the server performs voice playback, it will process and play each individual playback request in sequence. A combined playback request means combining multiple text-to-speech segments into one request for processing and playback, which can reduce the number of requests and improve the playback efficiency. For example, when the processing performance of the terminal device is good, request text-to-speech segment 2 and text-to-speech segment 3 at the same time, and only one playback request needs to be made to complete the combined playback of these two text-to-speech segments.
[0037] The above technical solution can be applied to multiple speech synthesis fields. For example, in the field of audiobook production, the server can automatically identify chapter titles, scene transitions, and character dialogues in novels, divide long works into sub-paragraphs that combine semantic integrity and synthesis efficiency, and ensure emotional coherence during speech synthesis. In the enterprise training scenario, complex technical documents can be segmented by knowledge modules, and voice scripts with index markers can be generated in combination with PPT content to realize the production of intelligent courseware with audio-video synchronization. In the intelligent customer service system, this technology can decompose long text consultations input by users into short sentences for real-time broadcast, and cooperate with the streaming synthesis interface to achieve millisecond-level response.
[0038] The above is the method embodiment proposed in this application. Based on the same idea, some embodiments of this application also provide the devices and non-volatile computer storage media corresponding to the above method.
[0039] Figure 2 It is a schematic structural diagram of a long text voice broadcast device provided by an embodiment of this application. As Figure 2 shown, it includes: At least one processor; and, A memory communicatively connected to at least one processor; wherein, The memory stores instructions executable by at least one processor. The instructions are executed by at least one processor so that at least one processor can: Obtain a long text voice file, and based on a preset segmentation rule, according to the string information included in the long text voice file, divide the long text voice file into several text voice segments; Determine the paragraph identifier and voice file address corresponding to the text voice segment; Call a preset short text synthesis voice interface, and sequentially obtain the corresponding text voice segments from the voice file address according to the order of the text voice segments, and synthesize the text voice segments into a long text voice file for voice broadcast.
[0040] An embodiment of this application provides a non-volatile computer storage medium, storing computer-executable instructions, and the computer-executable instructions are set as: A long text voice broadcast method as described in any one of the above.
[0041] Each embodiment in this application is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0042] The devices, media, and methods provided by the embodiments of this application correspond one by one. Therefore, the devices and media also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.
[0043] Those skilled in the art should understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0044] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more flows and / or Figure 1 blocks.
[0045] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more flows and / or Figure 1 blocks.
[0046] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more flows and / or Figure 1 blocks.
[0047] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0048] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0049] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0050] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0051] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A long text voice broadcast method, characterized in that: The method comprises: Acquire a long text file, and divide the long text file into a plurality of text segments based on a preset segmentation rule and according to character string information contained in the long text file; Generate a corresponding text voice segment for the text segment, and determine a paragraph identifier and a voice file address corresponding to the text voice segment; The preset short text-to-speech synthesis interface is called, and the corresponding text-to-speech segments are obtained from the voice file address in sequence according to the sequence of the text-to-speech segments, and the text-to-speech segments are synthesized into the long text-to-speech file for voice broadcast.
2. A long text voice broadcasting method according to claim 1, characterized in that: According to the sequence of the text and voice segments, the corresponding text and voice segments are obtained from the voice file addresses in sequence, specifically including: Determining the sequence of the text and speech segments according to the paragraph identifiers corresponding to the text and speech segments; Determine a first text-speech segment in the text-speech segments, and generate a first broadcast request for the first text-speech segment; In response to the first broadcast request, the first text voice segment is acquired according to the first voice file address corresponding to the first text voice segment, and the first text voice segment is voice broadcasted; While the first text speech segment is being voiced, other text speech segments following the first text speech segment are preloaded in the order of precedence, so that after the voice broadcast of the first text speech segment is completed, the other text speech segments are automatically broadcast, thereby realizing the synthesized voice broadcast of the text speech segments.
3. A long text voice broadcasting method according to claim 2, characterized in that: Preloading other text and speech segments after the first text and speech segment in the order described above, so as to automatically broadcast the other text and speech segments after the voice broadcast of the first text and speech segment is completed, specifically includes: Generating a next broadcast request for a next text voice segment after the first text voice segment according to the sequence; In response to the next broadcast request, the next text voice segment is acquired according to the second voice file address corresponding to the next text voice segment, and the next text voice segment is broadcasted; While the next text voice segment is being voiced, the above process is repeated until the preloading and the voice broadcasting of the other text voice segments are completed.
4. A long text voice broadcasting method according to claim 3, characterized in that: Generating a next broadcast request for a next text speech segment after the first text speech segment specifically includes: Determining a broadcasting device for broadcasting the long text voice file; Determining a synthesis strategy corresponding to the next broadcast request according to the device performance and network load information of the broadcast device; According to the synthesis strategy, it is determined that the next broadcast request is a single broadcast request or a combined broadcast request.
5. A long text voice broadcasting method according to claim 1, characterized in that: Based on the preset segmentation rules, the long text file is segmented into a plurality of text segments according to the character string information contained in the long text file, specifically including: Based on the preset segmentation rules, determine the string length that each text segment needs to contain; According to the length of the character string, the long text file is divided into a plurality of text segments; wherein the character string length of the first text segment among the text segments is the smallest.
6. A long text voice broadcasting method according to claim 5, characterized in that: According to the length of the string, the long text file is divided into several text segments, specifically including: According to the length of the string, the long text file is divided into several text segments, For other text segments except the first text segment, semantic analysis is performed on other character string information contained in the other text segments to extract core text from the other character string information; If the core text corresponds to multiple text segments, the multiple text segments are merged into the same text segment.
7. A long text voice broadcasting method according to claim 6, characterized in that: Merging the multiple text segments into a single text segment specifically includes: When the merged text segment exceeds a preset character string length threshold or the character string length corresponding to the subsequent text segment, the text and voice segments corresponding to the core text are combined and broadcasted.
8. A long text voice broadcasting method according to claim 7, characterized in that: The method further comprises: Determine a usage scenario corresponding to the long text file; In the case where the usage scenario is a designated usage scenario, a text summary corresponding to the long text file is generated according to the core text, so that the text summary is broadcasted when the long text file is broadcasted.
9. A long text voice broadcasting device, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a long text voice broadcast method as described in any one of claims 1-8.
10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to: A method for voice broadcasting a long text as described in any one of claims 1 to 8.