Method and apparatus for providing voice content, device, and storage medium

By combining virtual characters and multiple tones when generating broadcast information, the problem of low efficiency of traditional voice broadcast is solved, and a higher quality voice content and information acquisition experience is achieved.

WO2025156756A1PCT designated stage Publication Date: 2025-07-31BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Application Number
PCT/CN2024/129046
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Traditional voice broadcasting technology simply reads text aloud, which affects the efficiency of users to obtain information, especially when visual information is difficult to obtain.

Method used

When generating broadcast information, combining the content of the target page and reference content, a virtual character is used to generate multiple voice clips, and the dialogue interaction effect is achieved through multiple tones alternately to provide higher quality voice content.

Benefits of technology

Through virtual role broadcasting technology, the richness of voice content and the efficiency of users to obtain information are improved, and the information acquisition experience is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024129046_31072025_PF_FP_ABST
    Figure CN2024129046_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure relate to a method and apparatus for providing voice content, a device, and a storage medium. The method provided herein comprises: in response to a voice request associated with a target page and on the basis of the target page, generating broadcast information, wherein the broadcast information comprises a plurality of text fragments, and each text fragment is associated with a corresponding virtual character (210); on the basis of the broadcast information, generating a plurality of voice fragments corresponding to the plurality of text fragments, wherein the timbre of each voice fragment is determined on the basis of a virtual character corresponding to the voice fragment (220); and on the basis of the plurality of voice fragments, providing voice content associated with the target page (230). In this way, the embodiments of the present disclosure can provide higher-quality voice broadcast content.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, equipment and storage medium for providing voice content Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for providing voice content. Background Art

[0002] With the development of computer technology, people can use the internet to obtain all kinds of information they need. For example, people can browse the web to read news. Traditionally, the main way terminal devices provide information to people is through visual information, and people need to browse web pages to obtain such content. In some scenarios, people have difficulty obtaining visual information.

[0003] Summary of the Invention

[0004] In a first aspect of the present disclosure, a method for providing voice content is provided. The method comprises: in response to a voice request associated with a target page, generating announcement information based on the target page, the announcement information comprising multiple text segments, each associated with a corresponding virtual character; generating multiple voice segments corresponding to the multiple text segments based on the announcement information, wherein the timbre of each voice segment is determined based on the virtual character corresponding to the voice segment; and providing voice content associated with the target page based on the multiple voice segments.

[0005] In a second aspect of the present disclosure, a device for providing voice content is provided. The device includes: an information generation module configured to generate, in response to a voice request associated with a target page, announcement information based on the target page, the announcement information including multiple text segments, each associated with a corresponding virtual character; a voice generation module configured to generate, based on the announcement information, multiple voice segments corresponding to the multiple text segments, wherein the timbre of each voice segment is determined based on the virtual character corresponding to the voice segment; and a voice providing module configured to provide voice content associated with the target page based on the multiple voice segments.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.

[0008] In a fifth aspect of the present disclosure, a computer program product is provided, which includes computer-executable instructions, which, when executed by a processor, implement the method according to the first aspect of the present disclosure.

[0009] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0011] FIG1 shows a schematic diagram of an example environment in which embodiments according to the present disclosure may be implemented;

[0012] FIG2 illustrates a flow chart of an example process for providing voice content according to some embodiments of the present disclosure;

[0013] FIG3 shows a schematic diagram of an example interface according to some embodiments of the present disclosure;

[0014] FIG4 is a schematic diagram illustrating an example process of providing voice content according to some embodiments of the present disclosure;

[0015] FIG5 is a schematic diagram showing an example process of providing voice content according to other embodiments of the present disclosure;

[0016] FIG6 shows a schematic structural block diagram of an example apparatus for providing voice content according to some embodiments of the present disclosure; and

[0017] FIG7 shows a block diagram of an electronic device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION

[0018] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0019] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.

[0020] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.

[0021] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects shall comply with the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms them. Accordingly, when implementing the various embodiments of the present disclosure, the types, scope of use, and usage scenarios of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method may vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.

[0022] If this specification and the solutions in the examples involve the processing of personal information, such processing will be done only with a legitimate basis (such as with the consent of the subject of personal information or as necessary for the performance of a contract) and only within the prescribed or agreed scope. A user's refusal to process personal information other than that required for basic functions will not affect the user's use of basic functions.

[0023] As discussed above, the primary way terminal devices provide information to people is through visual means, and people need to browse pages to access this content. In some scenarios, accessing visual information can be challenging. For example, when browsing a long news article, users may not have enough time to read the entire article. Alternatively, in some scenarios (e.g., while driving), it may be difficult for users to view the terminal device screen.

[0024] Some traditional solutions offer voice broadcasting capabilities. For example, some web pages support converting webpage content into voice broadcasts using TTS (Text to Speech). However, these broadcasts typically consist of simply reading the corresponding text aloud, which hinders the user's ability to access information.

[0025] The embodiments of the present disclosure propose a solution for providing voice content. According to this solution, in response to a voice request associated with a target page, announcement information is generated based on the target page, the announcement information including multiple text segments, each associated with a corresponding virtual character; based on the announcement information, multiple voice segments are generated corresponding to the multiple text segments, wherein the timbre of each voice segment is determined based on the virtual character corresponding to the voice segment; and based on the multiple voice segments, voice content associated with the target page is provided.

[0026] In this way, the embodiments of the present disclosure can generate a broadcast text associated with a virtual character based on the page content to achieve the effect of dialogue interaction, thereby providing higher quality voice broadcast content.

[0027] Various example implementations of this solution are described in detail below in conjunction with the accompanying drawings.

[0028] Sample Environment

[0029] FIG1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG1 , the example environment 100 may include a terminal device 110 .

[0030] In this example environment 100, a terminal device 110 may run an application 120 that supports interface interaction. The application 120 may be any suitable type of application, including a browser. A user 140 may interact with the application 120 via the terminal device 110 and / or its attached devices.

[0031] In the environment 100 of FIG. 1 , if the application 120 is in an active state, the terminal device 110 may present an interface 150 for supporting interface interaction through the application 120 .

[0032] In some embodiments, the terminal device 110 communicates with the electronic device 130 to enable the provision of services for the application 120. The terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.).

[0033] The electronic device 130 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. The electronic device 130 can include, for example, a computing system / server such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like. The electronic device 130 provides background services for the application 120 that supports content presentation in the terminal device 110.

[0034] A communication connection may be established between the electronic device 130 and the terminal device 110. The communication connection may be established in a wired or wireless manner. The communication connection may include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and the embodiments of the present disclosure are not limited in this respect. In the embodiments of the present disclosure, the electronic device 130 and the terminal device 110 may implement signaling interaction through the communication connection between the two.

[0035] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.

[0036] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0037] Example Process

[0038] 2 shows a flow chart of an example process 200 for providing voice content according to some embodiments of the present disclosure. Process 200 may be implemented at electronic device 130. Process 200 is described below with reference to FIG1.

[0039] As shown in Figure 2, in response to a voice request associated with a target page, the electronic device 130 generates announcement information based on the target page in block 210. In some embodiments, the announcement information includes multiple text segments, each of which is associated with a corresponding virtual character.

[0040] FIG3 shows an example interface 300 according to some embodiments of the present disclosure. As shown in FIG3 , the terminal device 110 may present a target page 305 to the user 140. As an example, the target page 305 may be a web page. Alternatively, the target page 305 may be another suitable page presented by the terminal device 110.

[0041] In some embodiments, the terminal device 110 can obtain a voice request from the user 140 for the target page 305. Such a voice request can also be referred to as a voice broadcast request or a podcast request. As shown in the figure, the terminal device 110 can provide a broadcast control 310 associated with the target page 305. The terminal device 110 can receive a triggering of the broadcast control 310 by the user 140 to obtain the voice request.

[0042] As another example, terminal device 110 may further provide a conversational interface 315 associated with target page 305. This conversational interface 315 may be associated with a virtual object. Such a virtual object may be an appropriate virtual entity implemented based on a generative model, such as an agent or a bot. Such a virtual object may, for example, engage in conversational interaction with a user based on the generative model.

[0043] As shown in the figure, the terminal device 110 can also obtain an input message 325 of the user 140 via the input control 330 in the conversation interface 315. The input message 325 can be determined as a voice request for displaying the target page 305, for example.

[0044] In some embodiments, after receiving the voice request, the terminal device 110 may trigger the electronic device 130 to generate announcement information based on the target page 305. In some embodiments, the electronic device 130 may use an appropriate generative model (e.g., a language model) to generate announcement information associated with the target page.

[0045] Specifically, the terminal device 110 may, for example, obtain context information associated with the target page and provide the obtained context information to the server 130. Furthermore, the server 130 may, for example, process the context information of the target page using a language model to generate the announcement information.

[0046] In some embodiments, such context information may include page content in the target page 305 , such as text content, image content, etc. It should be understood that such page content is not limited to the portion of the target page 305 currently displayed in the terminal device 110 .

[0047] In some embodiments, the terminal device 110 may also determine whether the page content of the target page 305 satisfies a preset condition. For example, such a preset condition may include the number of characters in the page content reaching a preset number. If the number of characters in the target page 305 is less than the preset number, the terminal device 110 may determine that the preset condition is not satisfied.

[0048] Furthermore, the terminal device 110 and / or the server 130 may obtain at least one reference content associated with the target page 305. In some embodiments, the server 130 may, for example, search for other web pages related to the content of the target page 305 as reference content. For example, the server 130 may determine the name of the news event corresponding to the target page 305 and search for other news reports about the news event as reference content.

[0049] Alternatively, the reference content may also include other pages linked or referenced in the target page 305. The server terminal device 110 and / or the server 130 may obtain the corresponding reference content by accessing such pages.

[0050] Furthermore, the terminal device 110 and / or the server 130 may generate the context information based on the page content and the at least one reference content.

[0051] In this way, the embodiments of the present disclosure can provide users with voice broadcast content with richer content, rather than being limited to the page content of the target page.

[0052] In some embodiments, the server 130 may process the acquired context information using a target model to generate corresponding announcement information. In some embodiments, such announcement information may include a data stream output by a language model based on a structured format.

[0053] As an example, the server 130 may indicate to the language model the structured format of the desired output, such as JSON (JavaScript Object Notation). In some embodiments, the server 130 may instruct the language model to stream-output multiple data segments organized in JSON format based on the context information.

[0054] In some embodiments, each data segment may indicate a text segment and a corresponding virtual character. Specifically, the language model may be instructed to generate a broadcast text suitable for a specific virtual character (also referred to as a virtual identity) based on context information.

[0055] In some embodiments, server 130 may provide the language model with a preset set of virtual characters, enabling the language model to determine the virtual character corresponding to each text segment from the set of virtual characters. For example, the set of virtual characters may include a "host" character and a "guest" character. Accordingly, the language model may generate corresponding broadcast text for the "host" character and the "guest" character based on the contextual information.

[0056] In other embodiments, such virtual characters or identities may also be generated by a language model based on contextual information. For example, when the target page content is different, the language model may adaptively generate different characters based on the characteristics of the page content for broadcasting.

[0057] In some embodiments, the language model may stream multiple data segments organized in a structured format. Taking JSON as an example, each data segment may be represented as:

[0058] In each data segment, the field "role" may indicate the virtual role corresponding to the data segment; and the field "content" may indicate the broadcast text corresponding to the virtual role.

[0059] Thus, the server 130 can obtain the data stream output by the language model based on the structured format as the broadcast information.

[0060] 2 , in block 220 , the electronic device 130 generates a plurality of voice segments corresponding to the plurality of text segments based on the broadcast information. Specifically, the timbre of each voice segment is determined based on the virtual character corresponding to the voice segment.

[0061] In some embodiments, the server 130 can parse out the target text segment and the target role corresponding to the target text segment based on the broadcast information. Continuing with the data stream in which the broadcast information is a structured format component as an example, the terminal device 110 and / or the server 130 can parse the target data segment from the data stream based on the structured format. For example, the terminal device 110 and / or the server 130 can parse out the target text segment (i.e., the value of the "content" field) and the corresponding target virtual role (i.e., the value of the "role" field) in the target data segment based on the JSON format defined above.

[0062] In some embodiments, the terminal device 110 and / or the server 130 can stream-parse the data stream output by the language model without requiring the language model to output complete broadcast information, thereby supporting the stream generation of voice segments. In some embodiments, the terminal device 110 and / or the server 130 can implement stream parsing of JSON data based on the definition of the JSON format, for example.

[0063] For example, the terminal device 110 and / or the server 130 can determine the parsing state of starting a new data segment by detecting the symbol ({). Further, the terminal device 110 and / or the server 130 can determine the parsing state of entering a key by detecting the symbol ("). Further, the terminal device 110 and / or the server 130 can determine the parsing state of completing a key by detecting the symbol ("), and can determine the name of the field accordingly.

[0064] After completing the parsing of the key, the terminal device 110 and / or the server 130 can determine the parsing state of the value (value) by detecting the next symbol ("). Further, the terminal device 110 and / or the server 130 can determine the completion of the parsing state of the value (value) by detecting the symbol ("), and can determine the specific value of the field accordingly.

[0065] Additionally, the terminal device 110 and / or the server 130 may determine the end of the parsing state of the data segment by detecting a symbol (}).

[0066] In this way, the embodiments of the present disclosure can implement streaming parsing of JSON data without waiting for the language model to complete the output of the complete result, thereby improving and reducing the user's waiting time.

[0067] In some embodiments, the server 130 may further obtain a target voice segment generated by the voice processing unit based on the target text segment and the target timbre via a target channel with the voice processing unit.

[0068] In some embodiments, the server 130 may establish, for example, a plurality of channels corresponding to a plurality of preset virtual characters with the voice processing unit in response to the voice request, where the plurality of preset virtual characters includes the target virtual character.

[0069] FIG4 illustrates a schematic diagram of an example process 400 for providing voice content according to some embodiments of the present disclosure. As shown, at 410, user 140 may initiate a voice request (e.g., a podcast mode broadcast request) to application 120 (e.g., a browser). Furthermore, at 412, application 120 may establish a WS (web socket) connection with server 130, for example.

[0070] At 414, server 130 may query the cache. At 416, server 130 sends a request to voice processing unit 405 for a JWT token (i.e., a JSON Web Token). At 418, server 130 may obtain the corresponding JWT token from voice processing unit 405. In some embodiments, voice processing unit 405 may be, for example, a service device that deploys a voice model, which may be used to generate voice content with a specified timbre.

[0071] Furthermore, at 420, the server 130 may cache the received JWT token and establish multiple channels with the voice processing unit 405. Taking the example of virtual roles including a "host" role and a "guest" role, at 422, the server 130 may send a request to the voice processing unit 405 to establish a "host" WS connection. At 424, the server 130 may establish a channel corresponding to the "host" role, i.e., a "host" WS connection, with the voice processing unit 405.

[0072] Furthermore, at 426 , the server 130 may send a request to establish a “Guest” WS connection to the voice processing unit 405 . At 428 , the server 130 may establish a channel corresponding to the “Guest” role, ie, a “Guest” WS connection, with the voice processing unit 405 .

[0073] Thus, the server 130 may pre-establish multiple channels with the voice processing unit 405 , and the multiple channels may correspond to multiple different virtual characters.

[0074] Continuing with FIG4 , at 430 , server 130 may send a message to application 120 indicating successful establishment of the associated channel. Accordingly, application 120 may obtain the podcast transcript (i.e., the first information stream) generated by the language model. Furthermore, application 120 may parse the first information stream to send corresponding text segments, i.e., the lines corresponding to different characters, to server 130. In some embodiments, server 130 may also parse the lines of each character directly from the information stream output by the language model, for example.

[0075] After obtaining the text segments corresponding to the avatars, the server 130 can generate audio segments corresponding to the text segments using the target channels corresponding to the avatars. For example, at 434, the server 130 can obtain the host's speech and, at 436, use the "host" WS connection to trigger the voice processing unit 405 to generate an audio segment corresponding to the speech.

[0076] Furthermore, such a generation process can be triggered in a streaming manner. For example, at 438, the server 130 can obtain the guest's speech and, at 440, can use the "guest" WS connection to trigger the voice processing unit 405 to generate an audio clip corresponding to the speech portion. At 442, the server 130 can obtain the host's speech and, at 444, can use the "host" WS connection to trigger the voice processing unit 405 to generate an audio clip corresponding to the speech portion.

[0077] In some embodiments, the voice processing unit 405 can determine the timbre corresponding to different virtual characters and can generate an audio segment corresponding to the text segment based on the corresponding timbre. In some embodiments, the voice processing unit 405 can generate such an audio segment using any appropriate audio generation model.

[0078] Further, such audio clips can also be streamed. As shown, at 446, the application 120 can obtain the streamed audio clips and play the audio clips to the user 140 as the voice content associated with the target page 305.

[0079] In some embodiments, to improve the efficiency of subsequent data processing, the server 130 may also parse the data stream generated by the language model (also referred to as the first data stream) to generate another data stream (also referred to as the second data stream) corresponding to the target format.

[0080] In some embodiments, the target format can utilize one or more predefined identifiers to separate different types of data parts. For example, the target format can be USV (Unicode Separated Values) format, which can use predefined control characters (e.g., invisible characters such as \u001F, \u001E, and \u001D) as delimiters, greatly reducing the probability of conflict with the main text.

[0081] Furthermore, the server 130 may send the second data stream to the terminal device 110. Furthermore, the terminal device 110 may parse the second data stream based on the target format, and may present a broadcast text associated with the voice content based on the parsing result of the second data stream.

[0082] 3 as an example, the terminal device 110 can parse the second data stream in USV format to output the announcement text 320. As an example, the terminal device 110 can output the announcement text 320 in the conversation interface 315 and dynamically update the content of the announcement text 320 based on the parsing result.

[0083] Furthermore, the terminal device 110 may also determine the target text corresponding to the currently playing voice content and may display the portion of the broadcast text corresponding to the target text in a target style. Taking FIG3 as an example, the terminal device 110 may highlight the portion 335 corresponding to the currently broadcast voice content in the conversation interface 315 to facilitate the user's perception of the current broadcast progress.

[0084] FIG5 further illustrates an example process 500 for providing voice content and broadcast text using application 120. As shown in FIG5, at 520, user 120 may send a voice request (i.e., podcast mode broadcast) to content script 505. Further, at 522, content script 505 may send a request to obtain playback resources to background script 510. At 524, background script 510 may return the corresponding audio stream and broadcast text to content script 505.

[0085] Furthermore, at 526, the content script 505 may send the text being read to the background script 510, and the background script 510 may record the text currently being read at 528. At 530, the user 120 may also activate the sidebar component 515 (eg, the conversation interface 315 shown in FIG. 3 ).

[0086] Furthermore, at 532 and 534, the sidebar component 515 can pull chat messages with the virtual object and the currently read text from the background script 510, respectively. Additionally, at 536, the content script 505 can continuously send the currently read text to the background script 510 during the voice blog broadcast. Furthermore, at 538, the background script 510 can continuously push the latest read text to the sidebar component 515.

[0087] At 540, the sidebar component 515 may match the portion of the announcement text that corresponds to the currently read text and update the interface accordingly. For example, the sidebar component 515 may determine the location of the currently read text within the announcement text based on text similarity matching and may highlight the currently read text accordingly. Further, at 542, the sidebar component 515 may display the interface content to the user 120.

[0088] In this way, the embodiments of the present disclosure can achieve the effect of a multi-timbre alternating dialogue based on the multi-timbre channel connection of the voice processing unit, thereby improving the quality of the provided voice content. For example, the broadcast voice content can simulate a conversation scene such as a host interviewing a guest, increasing the richness of the voice content and thus improving the efficiency of users in obtaining information.

[0089] It should be understood that the server 130 mentioned above for generating the first data stream using the language model and the server 130 for establishing a channel with the voice processing unit 405 can be implemented by different physical devices.

[0090] In this way, the embodiments of the present disclosure can generate a broadcast text associated with a virtual character based on the page content to achieve the effect of dialogue interaction, thereby providing higher quality voice broadcast content.

[0091] Example devices and equipment

[0092] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above-described methods or processes. FIG6 shows a schematic block diagram of an example apparatus 600 for providing voice content according to certain embodiments of the present disclosure. Apparatus 600 may be implemented as or included in electronic device 130. Each module / component in apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.

[0093] As shown in Figure 6, the device 600 includes an information generation module 610, which is configured to generate broadcast information based on the target page in response to a voice request associated with the target page, and the broadcast information includes multiple text segments, each text segment is associated with a corresponding virtual character; a voice generation module 620, which is configured to generate multiple voice segments corresponding to the multiple text segments based on the broadcast information, wherein the timbre of each voice segment is determined based on the virtual character corresponding to the voice segment; and a voice providing module 630, which is configured to provide voice content associated with the target page based on the multiple voice segments.

[0094] In some embodiments, the information generation module 610 is further configured to provide the language model with context information associated with the target page to generate announcement information.

[0095] In some embodiments, the context information is generated based on the following process: in response to the page content of the target page not meeting a preset condition, obtaining at least one reference content associated with the target page; and generating context information based on the page content and the at least one reference content.

[0096] In some embodiments, the virtual character corresponding to each text segment is determined by a language model from a preset set of virtual characters.

[0097] In some embodiments, the information generation module 610 is further configured to obtain a data stream output by the language model based on a structured format, where the data stream includes multiple data segments organized in a structured format, each data segment indicating a text fragment and a corresponding virtual character.

[0098] In some embodiments, the speech generation module 620 is further configured to parse a target data segment from the data stream based on a structured format, the target data segment indicating a target text segment and a target virtual character corresponding to the target text segment; and obtain a target speech segment generated by the speech processing unit based on the target text segment and the target timbre via a target channel between the speech processing unit, the target channel corresponding to the target virtual character, and the target timbre determined based on the target virtual character.

[0099] In some embodiments, the apparatus 600 further includes a request processing module configured to establish, in response to a voice request, a plurality of channels corresponding to a plurality of preset virtual characters with the voice processing unit, wherein the plurality of preset virtual characters includes a target virtual character.

[0100] In some embodiments, the data stream is a first data stream, and the device 600 also includes an information processing module, which is configured to generate a second data stream corresponding to a target format by parsing the first data stream, and the target format uses a preset identifier to separate different types of data parts; and send the second data stream to the terminal device to trigger the terminal device to: parse the second data stream based on the target format; and present a broadcast text associated with the voice content based on the parsing result of the second data stream.

[0101] In some embodiments, the terminal device is further configured to: determine a target text corresponding to the currently played voice content; and display the portion of the broadcast text corresponding to the target text in a target style.

[0102] In some embodiments, the terminal device is configured to present the broadcast text in the conversation interface with the virtual object.

[0103] The modules included in the device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the modules in the device 600 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0104] As shown in FIG7 , electronic device 700 is a general-purpose electronic device. Components of electronic device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processing unit 710 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 700.

[0105] The electronic device 700 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 700.

[0106] The electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0107] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 700 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 700 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0108] Input device 750 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 760 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 700 may also communicate with one or more external devices (not shown) via communication unit 740 as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with electronic device 700, or with any device that allows electronic device 700 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0109] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0110] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0111] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0112] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0113] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0114] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for providing voice content, comprising: In response to a voice request associated with a target page, generating broadcast information based on the target page, the broadcast information including a plurality of text segments, each text segment being associated with a corresponding virtual character; Generating a plurality of voice segments corresponding to the plurality of text segments based on the broadcast information, wherein the timbre of each voice segment is determined based on the virtual character corresponding to the voice segment; and Providing voice content associated with the target page based on the plurality of voice segments.

2. The method according to claim 1, wherein generating the broadcast information based on the target page includes: Providing context information associated with the target page to a language model to generate the broadcast information.

3. The method according to claim 2, wherein the context information is generated based on the following process: In response to the page content of the target page not meeting a preset condition, acquiring at least one piece of reference content associated with the target page; and Generating the context information based on the page content and the at least one piece of reference content.

4. The method according to claim 2, wherein the virtual character corresponding to each text segment is determined by the language model from a preset set of virtual characters.

5. The method according to claim 1, wherein generating the broadcast information based on the target page includes: Acquiring a data stream output by a language model based on a structured format, the data stream including a plurality of data segments organized according to the structured format, each data segment indicating a text segment and a corresponding virtual character.

6. The method according to claim 5, wherein generating a plurality of voice segments corresponding to the plurality of text segments based on the broadcast information includes: Parsing a target data segment from the data stream based on the structured format, the target data segment indicating a target text segment and a target virtual character corresponding to the target text segment; And Acquiring, via a target channel between the voice processing unit, a target voice segment generated by the voice processing unit based on the target text segment and the target timbre, the target channel corresponding to the target virtual character, and the target timbre being determined based on the target virtual character.

7. The method according to claim 6, further comprising: In response to the voice request, establishing a plurality of channels corresponding to a preset plurality of virtual characters with the voice processing unit, the preset plurality of virtual characters including the target virtual character.

8. The method according to claim 5, wherein the data stream is a first data stream, and the method further includes: Generating a second data stream corresponding to a target format by parsing the first data stream, the target format using a preset identifier to separate different types of data parts; And Sending the second data stream to a terminal device to trigger the terminal device to: parse The second data stream based on the target format; And presenting a broadcast text associated with the voice content based on the parsing result of the second data stream.

9. The method according to claim 8, wherein the terminal device is further configured to: Determine the target text corresponding to the currently played voice content; and Display the part of the broadcast text corresponding to the target text in the target style.

10. The method according to claim 8, wherein the terminal device is configured to present the broadcast text in a session interface with a virtual object.

11. An apparatus for providing voice content, comprising: An information generation module, configured to generate broadcast information based on the target page in response to a voice request associated with the target page, the broadcast information including a plurality of text segments, each text segment being associated with a corresponding virtual character; A voice generation module, configured to generate a plurality of voice segments corresponding to the plurality of text segments based on the broadcast information, wherein the timbre of each voice segment is determined based on the virtual character corresponding to the voice segment; And A voice providing module, configured to provide voice content associated with the target page based on the plurality of voice segments.

12. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.

14. A computer program product, comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Text-to-speech conversion method, device, electronic equipment and storage medium

    CN112765971A

  • Speech synthesis method and device of text, electronic equipment and storage medium

    CN112908292A

  • Article voice playing method, apparatus and device, and computer readable storage medium

    CN113010138A

  • Automatic audio content generation

    CN113628609A

  • Generating audio rendering from textual content based on character models

    US20190043474A1

Cited By

  • Picture information broadcasting method and device, electronic equipment and storage medium

    CN121418608A