Method and device for voice broadcasting based on streaming text, electronic device and medium
By performing punctuation mark recognition and special text filtering on streaming text blocks, the problems of delay and incoherence in traditional speech synthesis methods are solved, achieving real-time and smooth speech broadcasting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2024-11-12
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional speech synthesis methods suffer from latency and inconsistency when processing long texts or real-time speech playback, especially when processing special texts that are not yet spoken, which can easily result in garbled text.
By performing character recognition on the streamed text blocks, identifying preset punctuation marks and filtering special text, the text to be read is determined for voice broadcast, ensuring real-time performance and smoothness.
It achieves real-time and smooth voice broadcasting, avoids garbled text broadcasting, and improves user experience.
Smart Images

Figure CN119474461B_ABST
Abstract
Description
Streaming text-based voice broadcasting methods, devices, electronic equipment, and media Technical Field
[0001] This disclosure relates to the field of artificial intelligence, and in particular to the fields of natural language processing, digital humans, voice broadcasting, intelligent agents, and generative search technology. Specifically, it relates to a voice broadcasting method, device, electronic device, computer-readable storage medium, and computer program product based on streaming text. Background Technology
[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0003] With the development of artificial intelligence technology, speech synthesis technology has been widely applied in various scenarios, such as digital humans, intelligent assistants, and voice navigation. Traditional speech synthesis methods generate audio data from text, which is then played back by a player. However, when processing long texts or real-time voice broadcasts, problems such as latency and incoherence often arise. Summary of the Invention
[0004] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for voice broadcasting based on streaming text.
[0005] According to one aspect of this disclosure, a method for voice broadcasting based on streaming text is provided, comprising: acquiring one or more text blocks that have not yet been broadcast and are being streamed, wherein the one or more text blocks are at least a portion of a text segment; sequentially performing character recognition on the one or more text blocks to identify preset punctuation marks in the one or more text blocks; in response to identifying the preset punctuation marks and determining that the one or more text blocks contain special text preceding the preset punctuation marks, acquiring first text preceding the preset punctuation marks in the one or more text blocks, excluding the special text, and using the first text as text to be broadcast, wherein the special text is preset text that is not suitable for voice broadcasting; and inputting the text to be broadcast into a voice broadcaster for voice broadcasting.
[0006] According to another aspect of this disclosure, a voice broadcasting device based on streaming text is provided, comprising: an acquisition unit configured to acquire one or more text blocks that have been streamed but not yet broadcast, wherein the one or more text blocks are at least a portion of a text segment; a recognition unit configured to sequentially perform character recognition on the one or more text blocks to recognize preset punctuation marks in the one or more text blocks; a first determination unit configured to, in response to recognizing the preset punctuation marks and determining that the one or more text blocks contain special text preceding the preset punctuation marks, acquire first text preceding the preset punctuation marks in the one or more text blocks, excluding the special text, and use the first text as text to be broadcast, wherein the special text is preset text that is not suitable for voice broadcasting; and a broadcasting unit configured to input the text to be broadcast into a voice broadcaster for voice broadcasting.
[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor to enable the at least one processor to perform the methods described in this disclosure.
[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods described in this disclosure.
[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods described in this disclosure.
[0010] According to one or more embodiments of this disclosure, preset punctuation marks are recognized based on the currently received text to broadcast complete sentences or half sentences, thereby ensuring the real-time performance and fluency of subsequent voice broadcasts; and special text in the text to be broadcast is filtered to avoid the occurrence of "garbled" broadcasts.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0012] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0013] Figure 1 illustrates a schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure;
[0014] Figure 2 shows a flowchart of a streaming text-based voice broadcasting method according to an embodiment of the present disclosure;
[0015] Figure 3 shows a structural block diagram of a streaming text-based voice broadcasting device according to an embodiment of the present disclosure; and
[0016] Figure 4 shows a structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure. Detailed Implementation
[0017] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0018] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0019] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0020] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0021] Figure 1 illustrates a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein may be implemented according to embodiments of the present disclosure. Referring to Figure 1, the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 may be configured to execute one or more applications.
[0022] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of methods for streaming text-based voice broadcasting.
[0023] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105, and / or 106 under a Software as a Service (SaaS) model.
[0024] In the configuration shown in Figure 1, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0025] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to perform voice broadcasts. The client devices can provide an interface that allows users to interact with the client devices. The client devices can also output information to the user via this interface. Although Figure 1 depicts only six client devices, those skilled in the art will understand that this disclosure can support any number of client devices.
[0026] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0027] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0028] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0029] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0030] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0031] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0032] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.
[0033] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.
[0034] The system 100 of Figure 1 can be configured and operated in various ways to enable the application of various methods and apparatuses described in this disclosure.
[0035] In generative AI (artificial intelligence) applications, large models return text in a streaming manner. Therefore, in the scenario of automatic speech broadcasting (automatic broadcasting) of digital humans, that is, real-time broadcasting, in order to have a good broadcasting experience, it means that the streaming text needs to be analyzed and segmented in real time to ensure that the broadcasting effect is smooth and natural. Usually, the entire text is fed to the TTS (Text To Speech) broadcaster for full broadcasting, which has limited real-time performance. (1) Broadcasting according to the currently received text will cause pauses in sentences; (2) There is no filtering capability for special characters, which leads to the broadcasting of some special rules of markdown syntax.
[0036] Therefore, an embodiment of the present disclosure provides a method for voice broadcasting based on streaming text. Figure 2 shows a flowchart of the method for voice broadcasting based on streaming text according to an embodiment of the present disclosure. As shown in Figure 2, method 200 includes: acquiring one or more text blocks that have not yet been broadcast in streaming output, wherein the one or more text blocks are at least a portion of a text segment (step 210); sequentially performing character recognition on the one or more text blocks to identify preset punctuation marks in the one or more text blocks (step 220); in response to identifying the preset punctuation marks and determining that the one or more text blocks contain special text before the preset punctuation marks, acquiring first text other than the special text before the preset punctuation marks in the one or more text blocks, and using the first text as the text to be broadcast, wherein the special text is preset text that is not suitable for voice broadcasting (step 230); inputting the text to be broadcast into a voice broadcaster for voice broadcasting (step 240).
[0037] According to embodiments of this disclosure, preset punctuation marks are recognized based on the currently received text to broadcast complete sentences or half sentences, thereby ensuring the real-time performance and fluency of subsequent voice broadcasts; furthermore, special text in the text to be broadcast is filtered to avoid "garbled" broadcasts.
[0038] In this disclosure, streaming text generally refers to text generated based on streaming output. Streaming output is a data transmission method that allows data to be displayed progressively during transmission, rather than waiting for all data to load completely before displaying it. This method is particularly suitable for scenarios involving big data processing and real-time data transmission, such as live web streaming, video streaming, and real-time data analysis. Streaming output can significantly improve the user experience because it allows users to begin using data without waiting for all data to load completely.
[0039] In this disclosure, the lost text is generated sequentially in the reading order. Therefore, character recognition is performed on the one or more text blocks sequentially, that is, character recognition is performed on the one or more text blocks sequentially based on their reading order.
[0040] In streaming text processing, the output text is typically in the form of text chunks (i.e., blocks of text data). This means that a whole segment of text is divided into blocks of text data according to certain rules or sizes and output sequentially. In data processing and storage, a chunk usually refers to a smaller portion or block within a larger dataset. For example, when processing large amounts of text or data, it might be divided into smaller chunks for easier processing.
[0041] It is understood that text blocks can be divided and streamed in any suitable manner in this disclosure, and no restrictions are imposed herein.
[0042] In some embodiments, the text returned by a large model can be read aloud using the method described in the embodiments of this disclosure.
[0043] Understandably, when a large model returns a lot of text, returning the complete result can take a considerable amount of time. Waiting for the large model to generate a complete answer before displaying it to the user would obviously provide a poor user experience. Therefore, most AI applications present results to users in a streaming manner. This way, the model doesn't output all the results at once, but rather gradually and continuously generates the output. This approach is similar to the human thought process in spoken or written communication—thinking and expressing simultaneously. Streaming output offers greater flexibility and interactivity, allowing for dynamic adjustments to content during long text generation, while also allowing users to provide real-time feedback or guidance during the generation process.
[0044] In this disclosure, "special text" refers to text that is pre-defined as unsuitable for voice playback. For example, text such as formulas and codes, when played as voice, can easily produce garbled text for the user, thus affecting the user experience. According to embodiments of this disclosure, by identifying and filtering special text in the text block to be played as voice, the occurrence of "garbled text" playback is effectively avoided, improving the smoothness of voice playback.
[0045] Therefore, according to some embodiments, the special text includes at least one of the following: markup language text, formulas. In some examples, the special text in the one or more text blocks can be identified by recognizing corresponding identifiers.
[0046] Markup languages are a special type of computer text encoding that uses specific tags or marks to identify different parts of text, thereby defining the purpose and relationships of these parts. Markup languages provide semantic information to text content, enabling computers to understand the structure and meaning of the text. Typical examples of markup languages include HTML, XML, and Markdown.
[0047] Markdown is a lightweight markup language typically used for formatting text. Markdown uses simple symbols to represent text formatting, such as bold, italics, lists, and links. By recognizing Markdown syntax (e.g., headings, lists, and code blocks), it intelligently divides content according to its structure and hierarchy, resulting in more semantically coherent blocks.
[0048] HTML elements supported by Markdown, and tags not covered by Markdown, can be written directly in HTML within the document.
[0049] In some examples, Markdown can also be used with LaTeX and other markup languages to represent mathematical formulas. In this case, the special text in formula form is also markup language text. For instance, when you need to insert a mathematical formula, you can wrap the TeX or LaTeX formatted mathematical formula with two dollar signs "$$". The formula is then identified, for example, by recognizing the dollar sign "$$".
[0050] HTML (Hypertext Markup Language) is the basic markup language used to create web pages. It uses tags to describe various elements on a webpage, such as text, images, links, and tables. HTML tags provide semantic information for each element, enabling the browser to correctly render the webpage content. XML (Extensible Markup Language) is a markup language used for storing and exchanging data. Similar to HTML, XML uses tags to define data structures and is commonly used for data exchange, configuration files, and metadata management. For example, text in HTML markup language... <img src=’http: xxx.jpg’> This can be marked as an image. For example, this can be done by identifying the identifier " This is used to identify the HTML markup language text.
[0051] In an embodiment according to this disclosure, after the text to be played is determined in one or more text blocks, the text following the text to be played in the one or more text blocks is taken as a text block that has not yet been voice-played, and together with the subsequently newly received text block, it is taken as a new text block that has not yet been voice-played, so as to determine the new text to be played in the new text block that has not yet been voice-played.
[0052] For example, when multiple text blocks corresponding to the phrase "From the snowy plateau of the north to the tropical rainforest of the south, the mountains and rivers are magnificent, and the rivers are rushing, nurturing the five thousand years of brilliant Chinese civilization" are obtained, for instance, if the preset punctuation mark is a comma, and the text before the comma is recognized as "From the snowy plateau of the north to the tropical rainforest of the south, the mountains and rivers are magnificent, and the rivers are rushing," then this text can be used as the text to be broadcast. Furthermore, the unbroadcasted text "nurturing" from these multiple text blocks, along with subsequently received text blocks, are considered as the next round of text blocks that have not yet been broadcast, and the preset punctuation mark recognition is performed sequentially.
[0053] It is understandable that, in the above examples, in some embodiments, when performing preset punctuation mark recognition, the first "comma" can be recognized, that is, "from the snowy plateau of the north to the tropical rainforest of the south" is taken as the text to be broadcast in this round, while "the mountains and rivers are magnificent and the rivers are surging, which have nurtured" and the subsequently received text blocks are taken as the text blocks that have not yet been broadcast in the next round, and preset punctuation mark recognition is performed in sequence.
[0054] Therefore, in some embodiments, when multiple preset punctuation marks are identified in one or more text blocks, the text to be broadcast can be determined based on any text before the preset punctuation mark is identified, as long as the determined text to be broadcast does not exceed the maximum number of input characters of the voice broadcaster.
[0055] Preferably, in some examples, when multiple preset punctuation marks are detected in one or more text blocks, if the number of preset punctuation marks does not exceed the maximum number of input characters of the voice broadcaster, the text before the last preset punctuation mark can be determined as the text to be broadcast, so as to broadcast enough received text and thus ensure the continuity of the voice broadcast.
[0056] According to some embodiments, character recognition of the one or more text blocks to identify preset punctuation marks in the one or more text blocks includes: in response to determining that the first sentence in the text to be spoken is the first sentence, identifying preset punctuation marks representing the whole sentence in the one or more text blocks.
[0057] This embodiment ensures the integrity of the voice broadcast and the consistency of the tone by broadcasting a complete sentence when the product starts broadcasting (when the first sentence is broadcast), thus improving the user experience.
[0058] According to some embodiments, the preset punctuation marks for representing complete sentences include at least one of the following: period, exclamation mark, question mark.
[0059] For example, in an embodiment of speech playback of text returned by a large model, the text block can be cached upon receiving the streaming text block from the large model. When the first sentence of the text to be spoken is determined, punctuation marks representing the entire sentence, such as periods, exclamation marks, and question marks, are identified in the cached text block.
[0060] Therefore, in some examples, in Chinese broadcast scenarios, the punctuation mark representing a complete sentence can be "。!?"; in English broadcast scenarios, the punctuation mark representing a complete sentence can be ".!?". Specifically, in scenarios involving alphabetical text such as English, to prevent misinterpretation of decimal points, such as the number 3.14, the punctuation mark "." representing a complete sentence is determined by recognizing "." followed by a space.
[0061] In some examples, to ensure rapid start-up and to begin voice playback for the user as quickly as possible, the first sentence of a text segment is sent to the voice playback device immediately. Therefore, when determining the first sentence of the text segment to be played, the first punctuation mark representing a complete sentence in one or more text blocks can be identified, and the text preceding this first punctuation mark is taken as the text to be played. This minimizes the user's waiting time.
[0062] According to some embodiments, the method according to this disclosure further includes: when determining that the first sentence in the text to be spoken is not recognized, in response to the fact that the preset punctuation mark is not recognized in the one or more text blocks and the special text is recognized in the one or more text blocks, obtaining the second text in the one or more text blocks before the special text, so as to use the second text as the text to be spoken.
[0063] Specifically, when determining the first sentence of the text to be spoken, if the special text is recognized before the preset punctuation marks are recognized, appropriate degradation can be performed, that is, the first sentence is not guaranteed to be a complete sentence.
[0064] For example, suppose the large model's response text is: "The following is a quicksort algorithm written in C++:\n\n```\nc++\n......\n```\n, you can also try other ways of asking questions, which will provide you with more detailed and comprehensive answers." And the text block output by the large model in streaming is shown below:
[0065]
[0066] During the first round of voice broadcasting, after accumulating multiple text blocks corresponding to the text "The following is a quicksort algorithm written in C++:\n\n```", if the identifier "\n\n```" detects that the following text is a piece of special code text, then the text "The following is a quicksort algorithm written in C++:" before the identifier "\n\n```" (i.e. before the special text) can be directly used as the content of the first round of voice broadcasting. After accumulating multiple text blocks corresponding to the text "The following is a quicksort algorithm written in C++:\n\n```\nc++\n......\n```\n, you can also try other questioning methods, for", the text that has not yet been broadcasted is "You can also try other questioning methods, for" (i.e., the special text has been filtered out), and then subsequent preset punctuation mark recognition and determination of the text to be broadcast in the next round can be performed. If the preset punctuation mark to be recognized in the next round includes a comma ",", then the text "You can also try other ways of asking questions, for" which has not been broadcast yet can be used as the text to be broadcast in the second round.
[0067] According to some embodiments, character recognition of the one or more text blocks to identify preset punctuation marks in the one or more text blocks includes: in response to determining that a non-first sentence in the text to be spoken, identifying preset punctuation marks representing half sentences in the one or more text blocks.
[0068] In this embodiment, starting from the second sentence of the text segment, it is acceptable to broadcast in units of half sentences (half sentences formed by punctuation marks) to ensure the smoothness of subsequent broadcasts.
[0069] In some embodiments, to capture as much text as possible, when recognizing punctuation marks representing half-sentences in a preset manner, the last punctuation mark representing a half-sentence in the acquired one or more text blocks can be identified. If the length of the text before the last punctuation mark representing a half-sentence exceeds the maximum input character length of the voice player, then the system can try to find the previous (i.e., the second to last) punctuation mark representing a half-sentence... until the maximum input character length requirement of the voice player is met.
[0070] According to some embodiments, the preset punctuation marks for representing half sentences include at least one of the following: period, exclamation mark, question mark, semicolon, and comma.
[0071] In some examples, in Chinese broadcast scenarios, the punctuation marks representing a half-sentence can be: ",;。!?"; while in English broadcast scenarios, the punctuation marks representing a half-sentence can be: ",.;!?". It can be seen that a half-sentence can include a whole sentence, and a half-sentence provides a more granular division of the text.
[0072] For example, suppose the large model's response text is: "First, Sun Wukong is a fearless person who dares to wreak havoc in Heaven and storm Hell; second, Sun Wukong is a righteous sword of wisdom and courage, sweeping away all demons. Fearlessness and wisdom make this character shine, which is also the value and essence of this book. This fearless spirit of daring to challenge all authority and despise all sacredness is the fundamental characteristic of the artistic image of Sun Wukong, and also the individual spirit that the author strongly praises. Because Sun Wukong is wise and courageous, he can see through all pretense and perceive all truths." Furthermore, the large model's streaming output text block is shown below:
[0073]
[0074] During the first round of voice broadcasting, after acquiring multiple text blocks corresponding to the text "First, Sun Wukong is a fearless person who dares to cause trouble in Heaven and the Underworld; second, Sun Wukong is a righteous sword with wisdom and courage who sweeps away all demons.", the first punctuation mark "." representing a complete sentence in that text is identified, and that sentence is then sent to the voice broadcaster as the text to be broadcast.
[0075] In the subsequent second round of voice broadcasting, assuming that enough text blocks have been received during the broadcasting of the first sentence, in order to ensure the continuity of the tone (i.e., to extract as much text as possible), the last punctuation mark representing half a sentence is found, and the text to be broadcast in the next round is determined based on the text before the last punctuation mark representing half a sentence. For example, when determining the second round of audio text to be broadcast, assuming the text "First, Sun Wukong is a fearless person who dares to cause trouble in Heaven and Hell; second, Sun Wukong is a righteous sword with wisdom and courage, sweeping away all demons. Fearlessness and wisdom make this character shine, which is also the value and essence of this book. This fearless spirit of daring to challenge all authority and daring to despise all sacredness is the basic characteristic of the artistic image of Sun Wukong, and also" has been received, the text that has not yet been broadcast is "Fearlessness and wisdom make this character shine, which is also the value and essence of this book. This fearless spirit of daring to challenge all authority and daring to despise all sacredness is the basic characteristic of the artistic image of Sun Wukong, and also" (i.e., special text has been filtered out), then subsequent preset punctuation mark recognition and determination of the next round of audio text can be carried out. If the preset punctuation mark to be identified in the next round includes a comma, then the following text can be used as the second round of text to be broadcast: "Fearlessness and wisdom make this character shine, which is also the value and essence of this book. This fearless spirit of daring to challenge all authority and despise all sacredness is the basic characteristic of the artistic image of Sun Wukong." This will continue until the entire text is broadcast.
[0076] According to some embodiments, the text to be broadcast further includes: template text for replacing the identified special text, wherein the template text is used to identify the special text.
[0077] In some examples, when special text is detected, it can be filtered out, and the text after removing the special text can be used as the text to be played. Alternatively, the special text can be replaced with a preset template text to form a new text to be played.
[0078] In some examples, the template text can correspond to the type of the special text in order to represent the content identified by the special text.
[0079] For example, suppose the large model's response text is: "Based on your needs, an image has been generated for you:" <img src=’https: newapp-prompt.cdn.bcebos.com bgimg 2024-10 ba3b55742326bfada7d355f5c26e3835.png’> Furthermore, the text blocks output by the large model in streaming mode are shown below:
[0080]
[0081] When broadcasting a message, the text to be played can be "Based on your needs, an image has been generated for you: Image here." The template text "Image here" indicates that this corresponds to a specific text representing an image. This avoids garbled text while ensuring the user clearly understands the content.
[0082] According to an embodiment of this disclosure, as shown in FIG3, a voice broadcasting device 300 based on streaming text is also provided, comprising: an acquisition unit 310 configured to acquire one or more text blocks that have not yet been broadcast via voice and are streamed, wherein the one or more text blocks are at least a portion of a text segment; an identification unit 320 configured to sequentially perform character recognition on the one or more text blocks to identify preset punctuation marks in the one or more text blocks; a first determination unit 330 configured to, in response to identifying the preset punctuation marks and determining that the one or more text blocks contain special text before the preset punctuation marks, acquire first text other than the special text before the preset punctuation marks in the one or more text blocks, and use the first text as the text to be broadcast, wherein the special text is preset text that is not suitable for voice broadcasting; and a broadcasting unit 340 configured to input the text to be broadcast into a voice broadcaster for voice broadcasting.
[0083] Here, the operation of each of the above units 310 to 340 of the streaming text-based voice broadcasting device 300 is similar to the operation of steps 210 to 240 described above, and will not be repeated here.
[0084] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0085] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.
[0086] Referring to Figure 4, a structural block diagram of an electronic device 400 that can serve as a server or client of this disclosure is now described, which is an example of hardware devices that can be applied to various aspects of this disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the disclosure described and / or claimed herein.
[0087] As shown in Figure 4, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. The RAM 403 can also store various programs and data required for the operation of the electronic device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0088] Multiple components in electronic device 400 are connected to I / O interface 405, including: input unit 406, output unit 407, storage unit 408, and communication unit 409. Input unit 406 can be any type of device capable of inputting information to electronic device 400. Input unit 406 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 407 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 408 may include, but is not limited to, a hard disk and an optical disk. Communication unit 409 allows electronic device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0089] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by the computing unit 401, one or more steps of method 200 described above may be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to perform method 200 by any other suitable means (e.g., by means of firmware).
[0090] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0091] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0092] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0093] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0094] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0095] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0096] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0097] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A voice broadcasting method based on streaming text, comprising: The system acquires one or more text blocks that have not yet been voice-broadcast from the streaming output in real time, wherein the one or more text blocks are at least a portion of a text segment; sequentially performs character recognition on the one or more text blocks to identify preset punctuation marks in the one or more text blocks, including: in response to determining that the first sentence of the text segment to be voice-broadcast is to be voice-broadcast, identifying preset punctuation marks representing complete sentences in the one or more text blocks; and in response to determining that the first sentence of the text segment to be voice-broadcast is not to be voice-broadcast, identifying preset punctuation marks representing half sentences in the one or more text blocks; in response to identifying the preset punctuation marks and determining that the one or more text blocks contain special text before the preset punctuation marks, acquiring first text in the one or more text blocks excluding the special text before the preset punctuation marks, and using the first text as the text to be voice-broadcast, wherein the special text is preset text that is not suitable for voice-broadcast; and inputting the text to be voice-broadcast into a voice broadcaster for voice-broadcasting.
2. The method of claim 1, further comprising: When the first sentence of the text to be spoken is determined, in response to the fact that the preset punctuation mark is not recognized in one or more text blocks and the special text is recognized in one or more text blocks, the second text before the special text in the one or more text blocks is obtained, so as to use the second text as the text to be spoken.
3. The method as described in claim 1 or 2, wherein, The text to be broadcast also includes: template text for replacing the identified special text, wherein the template text is used to identify the special text.
4. The method of claim 1, wherein, The preset punctuation marks for representing complete sentences include at least one of the following: period, exclamation mark, and question mark.
5. The method of claim 1, wherein, The default punctuation marks for representing half sentences include at least one of the following: period, exclamation mark, question mark, semicolon, and comma.
6. The method of claim 1, wherein, The special text includes at least one of the following: markup language text, formula, and wherein the special text in the one or more text blocks is identified by recognizing a corresponding identifier.
7. A voice broadcasting device based on streaming text, comprising: An acquisition unit is configured to acquire in real time one or more text blocks that have not yet been voice-broadcast from a streaming output, wherein the one or more text blocks are at least a portion of a text segment; a recognition unit is configured to sequentially perform character recognition on the one or more text blocks to recognize preset punctuation marks in the one or more text blocks, including: a unit for recognizing preset punctuation marks representing complete sentences in the one or more text blocks in response to determining that the first sentence in the text segment to be voice-broadcast is to be voice-broadcast; and a unit for recognizing preset punctuation marks representing half sentences in the one or more text blocks in response to determining that the first sentence in the text segment to be voice-broadcast is not the first sentence; a first determination unit is configured to, in response to recognizing the preset punctuation marks and determining that special text is included before the preset punctuation marks in the one or more text blocks, acquire first text other than the special text before the preset punctuation marks in the one or more text blocks, and use the first text as the text to be broadcast, wherein the special text is preset text that is not suitable for voice broadcast; and a broadcast unit is configured to input the text to be broadcast into a voice broadcaster for voice broadcast.
8. The apparatus of claim 7, further comprising: The second determining unit is configured to, when determining that the first sentence in the text to be voiced is not recognized in the one or more text blocks and the one or more text blocks contain the special text, obtain the second text in the one or more text blocks before the special text, so as to use the second text as the text to be voiced.
9. The apparatus of claim 7 or 8, wherein, The text to be broadcast also includes: template text for replacing the identified special text, wherein the template text is used to identify the special text.
10. The apparatus of claim 7, wherein, The preset punctuation marks for representing complete sentences include at least one of the following: period, exclamation mark, and question mark.
11. The apparatus of claim 7, wherein, The default punctuation marks for representing half sentences include at least one of the following: period, exclamation mark, question mark, semicolon, and comma.
12. The apparatus of claim 7, wherein, The special text includes at least one of the following: markup language text, formula, and wherein the special text in the one or more text blocks is identified by recognizing a corresponding identifier.
13. An electronic device, comprising: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.
15. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method of any one of claims 1-6.
Citation Information
Patent Citations
Text processing, model training and speech synthesis method, device and system and medium
CN114822491A
Text-to-speech conversion method and device
CN117542343A