Speech-based text processing method, apparatus, device, and storage medium

By determining the current audio progress using local system time and reference audio progress, the method synchronizes text and audio playback in live streaming, improving the viewing experience.

JP2026522074APending Publication Date: 2026-07-06BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-02
Publication Date
2026-07-06

AI Technical Summary

Technical Problem

In live streaming scenarios, especially during singing performances, the synchronization between lyrics and voice is often off due to network restrictions, leading to a poor viewing experience.

Method used

A method and device that determine the current audio progress based on local system time and the latest reference audio progress to synchronize the display of text content with the playback of the audio stream, using a viewer terminal to render target text content accordingly.

Benefits of technology

Ensures synchronization between the display of text and audio playback, enhancing the audiovisual experience by aligning the progress of text and voice in real-time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026522074000001_ABST
    Figure 2026522074000001_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide speech-based text processing methods, apparatus, devices, and storage media. These methods include obtaining the latest reference audio progress and an audio stream corresponding to a co-streaming terminal, including a guest terminal and / or a streamer terminal; determining the current audio progress of the audio stream based on local system time and the latest reference audio progress; and playing the audio stream and rendering target text content corresponding to the audio stream according to the current audio progress. The speech-based text processing methods provided in the embodiments of the present disclosure improve the alignment effect between text display progress and audio playback progress by determining the current audio progress of the audio stream based on local system time and the latest reference audio progress and rendering target text content corresponding to the audio stream according to the current audio progress.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] , ,

[0001] (Cross - reference to related applications) This application claims priority to a Chinese invention patent application titled "Text Processing Method, Apparatus, Device, and Storage Medium Based on Voice" (application number 202311134444.X) filed on September 4, 2023. The entire content of the said application is incorporated herein by reference.

[0002] (Field of the Invention) Embodiments of the present disclosure relate to the field of live streaming technology, and in particular, to a text processing method, apparatus, device, and storage medium based on voice.

Background Art

[0003] In the scenario of live singing, the singer's client terminal sends the lyrics or the progress of the voice to the streamer, and the streamer transfers the lyrics or the progress of the voice to the viewers. Since the lyrics are restricted during transmission by the network or the streaming system, it is impossible to synchronize the lyrics and the voice, and the viewers feel that the lyrics and the voice are out of sync, which affects the viewing experience. In related technologies, a timer is set on the viewer terminal to accumulate the progress of the voice, but still the synchronization of the lyrics and the voice is off, and the deviation between the lyrics and the voice becomes larger over time.

Summary of the Invention

Problems to be Solved by the Invention

[0004] Embodiments of the present disclosure provide a text processing method, apparatus, device, and storage medium based on voice that improve the alignment effect between the display progress of the text and the playback progress of the voice.

Means for Solving the Problems

[0005] In a first embodiment, one embodiment of the present invention provides an audio-based text processing method performed by a viewer terminal. The method includes obtaining the latest reference audio progress and an audio stream corresponding to a co-streaming terminal, the co-streaming terminal including a guest terminal and / or a streamer terminal; determining the current audio progress of the audio stream based on the local system time and the latest reference audio progress; and playing the audio stream and rendering target text content corresponding to the audio stream according to the current audio progress.

[0006] In a second embodiment, an embodiment of the present invention provides an audio-based text processing device to be placed in a viewer terminal. The device includes a latest reference audio progress acquisition module that acquires the latest reference audio progress and an audio stream corresponding to a co-streaming terminal, the co-streaming terminal comprising: the latest reference audio progress acquisition module which includes a guest terminal and / or a streamer terminal; a current audio progress determination module which determines the current audio progress of the audio stream based on the local system time and the latest reference audio progress; and a target text content rendering module which plays the audio stream and renders target text content corresponding to the audio stream according to the current audio progress.

[0007] In a third embodiment, embodiments of the present disclosure further provide an electronic device comprising one or more processors and a storage device for storing one or more programs, wherein when one or more programs are executed by one or more processors, one or more processors implement a speech-based text processing method as described in embodiments of the present disclosure.

[0008] In a fourth embodiment, embodiments of the present disclosure further provide a storage medium containing computer-executable instructions used to perform a speech-based text processing method described in embodiments of the present disclosure, when executed by a computer processor.

[0009] In a fifth embodiment, embodiments of the present disclosure further provide a computer program product that includes computer executable instructions, which are stored tangibly on a computer storage medium and, when executed by a device, enable the device to perform a speech-based text processing method described in embodiments of the present disclosure.

[0010] Embodiments of the present disclosure disclose a speech-based text processing method, apparatus, device, and storage medium for acquiring the latest reference audio progress and audio stream corresponding to a connected microphone terminal. The connected microphone terminal includes a guest terminal and / or a streamer terminal, which determines the current audio progress of an audio stream based on local system time and the latest reference audio progress, plays the audio stream, and renders target text content corresponding to the audio stream according to the current audio progress. [Brief explanation of the drawing]

[0011] The above and other features, advantages, and aspects of various embodiments of this disclosure will become more apparent upon further detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. Please note that the drawings are schematic and the actual objects and elements are not necessarily drawn to scale. [Figure 1] This is a flowchart of a speech-based text processing method according to one embodiment of the present invention. [Figure 2] This is an illustrative diagram of a live streaming scenario provided by one embodiment of the present disclosure. [Figure 3] This is an example diagram of determining the current audio progress, provided by embodiments of the present disclosure. [Figure 4] This is a schematic diagram of a speech-based text processing device according to one embodiment of the present invention. [Figure 5] This is a schematic diagram of an electronic device according to one embodiment of the present invention. [Modes for carrying out the invention]

[0012] The embodiments of this disclosure will be described in more detail below with reference to the accompanying drawings. While specific embodiments of this disclosure are shown in the accompanying drawings, this disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Rather, these embodiments are provided to enable a more thorough and complete understanding of this disclosure. It should be understood that the drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0013] It should be understood that the various steps described in the embodiments of the method of this disclosure may be performed in different orders and / or in parallel. Furthermore, embodiments of the method may include additional steps and / or omit illustrated steps. The scope of this disclosure is not limited thereto.

[0014] In this specification, the term "including" and its variations have an unlimited meaning; that is, "including, but not limited to." The term "based on" means "based on at least part of." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "several embodiments" means "at least several embodiments." Other terms are defined in the following description.

[0015] It should be noted that the concepts of “first” and “second” as used in this disclosure are used solely to distinguish between different devices, modules, or units, and are not intended to restrict the order or interdependence of the functions performed by these devices, modules, or units.

[0016] Note that the modifiers "one" and "a plurality" mentioned in the present disclosure are illustrative rather than restrictive. Unless clearly indicated in the context, those skilled in the art should note that they should be understood as "one or more".

[0017] In the embodiments of the present disclosure, the names of messages or information exchanged between multiple devices are used only for illustrative purposes and are not used to limit the scope of these messages or information.

[0018] Before using the technical solutions disclosed in various embodiments of the present disclosure, it can be understood that it is necessary to notify the user of the types of personal information included in the present disclosure, the scope of use, usage scenarios, etc., and obtain the user's permission in an appropriate manner in accordance with relevant laws and regulations.

[0019] For example, in response to an active request from the user, a prompt message is sent to the user, and the user is clearly notified that the requested operation requires the acquisition and use of the user's personal information. Thereby, the user can autonomously choose whether to provide personal information to an electronic device, application, server, storage medium, or other software or hardware that executes the operation of the disclosed technical solution based on the prompt message.

[0020] As an optional but not limited implementation, in response to receiving an active request from the user, prompt information is sent to the user in the form of a pop-up window, and the prompt information may be presented in text form. Furthermore, the pop-up window may also include a selection control that allows the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0021] It is understood that the above notification and user authentication process are merely illustrative and do not limit the implementation of the present disclosure. Other methods in accordance with relevant laws and regulations may also be applicable to the implementation of the present disclosure.

[0022] It can be understood that the data related to this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations, and related terms.

[0023] FIG. 1 is a flowchart of a voice-based text processing method according to an embodiment. The embodiments of the present invention are applicable to situations where text is rendered based on voice. This method can be executed by a voice-based text processing device, and the device can be implemented in the form of software and / or hardware. It can also be executed by an electronic device such as a mobile terminal, a PC, or a server.

[0024] As shown in FIG. 1, this method includes the following.

[0025] S110, obtain the latest reference voice progress and the voice stream corresponding to the co-streaming terminal.

[0026] In some embodiments, the co-streaming terminal includes a guest terminal and / or a streamer terminal. The audio stream can be obtained by the singer's (streamer and / or guest's) client terminal by collecting the singer's voice and synthesizing it with background music. For example, Figure 2 is an illustrative diagram of a live streaming scenario in one embodiment of the present invention. As shown in Figure 2, the guest terminal collects the guest's singing voice and sends the collected audio stream to the streamer terminal. The streamer terminal sends the audio stream to the server terminal, which forwards the audio stream to the viewer terminals. If the singer is the streamer, the streamer terminal sends the audio stream directly to the server terminal, which then forwards the audio stream to the viewer terminals. The singer's client terminal sends audio progress at set intervals along with the transmission of the audio stream, and the streamer terminal transmits the audio progress to the viewer terminals via the server terminal. Each time a viewer terminal receives audio progress, it uses the most recently received audio progress as the reference audio progress and determines the viewer terminal's audio progress at each time point based on the most recently received reference audio progress.

[0027] Audio progress can be understood as the progress of audio playback and can be expressed as a timestamp or percentage. For example, if the audio length is 2 minutes and the current playback has reached 1 minute and 20 seconds, the audio progress is 1 minute and 20 seconds. The latest reference audio progress is the audio progress that the listener terminal last received, and the audio progress at each point in time on the listener terminal is determined based on the latest received audio progress; in other words, the latest audio progress is used as the latest reference audio progress.

[0028] As an option, a method for obtaining the latest reference audio progress and audio stream corresponding to the co-streaming terminal can be to obtain the latest media supplement information and audio stream corresponding to the co-streaming terminal, extract the latest reference audio progress from the latest media supplement information, and record the local system time corresponding to the time when the latest media supplement information was received.

[0029] In some embodiments, the latest media supplement information can be understood as Supplemental Enhancement Information (SEI) last received by the listener terminal. Local system time can be measured by the listener terminal's local clock. In this embodiment, audio progress is transmitted to the listener terminal in the form of SEI information, and the listener terminal extracts the audio progress from the SEI information and stores it as the latest reference audio progress. Upon receiving the latest media supplement information, the listener terminal obtains and records the local system time at that time.

[0030] S120: Determine the current audio progress of the audio stream based on the local system time and the most recent reference audio progress.

[0031] In some embodiments, the local system time may be measured by the local clock of the viewer terminal.

[0032] Optionally, a method for determining the current audio progress of an audio stream based on local system time and the latest reference audio progress may include determining the time interval between the current time and the time the latest reference audio progress was received, based on local system time, and determining the current audio progress of the audio stream based on the time interval and the latest reference audio progress.

[0033] In some embodiments, the time interval between the current time and the time the latest reference audio progress was received can be understood as the time interval between the local system time corresponding to the current time and the local system time corresponding to the time the latest reference audio progress was received.

[0034] Specifically, a method for determining the time interval between the current time and the time the latest reference audio progress was received, based on local system time, can involve obtaining a first local system time corresponding to the current time and a second local system time corresponding to the time the latest reference audio progress was received, and then determining the time interval according to the first and second local system times.

[0035] In some embodiments, a method for determining a time interval according to a first local system time and a second local system time can be used to determine the time interval between the current time and the time the latest reference audio progress was received by subtracting the second local system time from the first local system time. For example, assuming the first local system time is T1 and the second local system time is T2, the time interval between the current time and the time the latest reference audio progress was received is ΔT = T1 - T2.

[0036] Specifically, a method for determining the current audio progress of an audio stream according to a time interval and the latest reference audio progress may be to determine the current audio progress of an audio stream based on a time interval and the latest reference audio progress.

[0037] Specifically, the current audio progress of the audio stream is obtained by accumulating the time interval and the latest reference audio progress. In this embodiment, if the current audio progress is t2, the latest reference audio progress is t1, and the time interval is ΔT, then the current audio progress of the audio stream is calculated as follows: t2 = t1 + ΔT. For example, Figure 3 shows an example diagram for determining the current audio progress. As shown in Figure 2, T2 is the local system time corresponding to the time when the latest reference audio progress was received, T1 is the local system time corresponding to the current time, and t1 is the latest reference audio progress. Therefore, the audio progress corresponding to the current time is t2 = t1 + (T1 - T2).

[0038] S130: Play the audio stream and, according to the current audio progress, render the target text content corresponding to the audio stream.

[0039] In this embodiment, the viewer terminal receives the audio stream transmitted from the streamer terminal in real time, decodes the received audio stream, and plays it back. In order to ensure consistency between the displayed text content and the played audio, it is necessary to display the corresponding target text content in accordance with the progress of the audio.

[0040] Specifically, a method for rendering text content corresponding to an audio stream according to the current audio progress can be achieved by obtaining a text file corresponding to the audio stream, extracting the target text content from the text file based on the current audio progress, and then rendering the target text content.

[0041] In some embodiments, the text content in the text file includes audio timestamps. The text content consists of multiple characters, each character containing an audio timestamp. The audio timestamp can be expressed in the format "character:start timestamp, duration". If the audio stream is audio corresponding to a song, the text file is a lyrics text file corresponding to the song. The text file corresponding to the audio stream may be transmitted from the streamer terminal to the viewer terminal, or it may be obtained by the viewer terminal from a pre-configured data source according to the representation of the audio stream.

[0042] Specifically, a method for extracting target text content from a text file based on the current audio progress can use text content within the text file whose audio timestamp matches the current audio progress as the target text content.

[0043] In this embodiment, the text content in a text file is divided into multiple text segments, each segment containing multiple characters. A method for determining the text content that matches the current audio progress of an audio timestamp is to retrieve the characters in the text file that match the timestamp corresponding to the current audio progress of the audio timestamp, and use the text segment in which the matching characters are located, along with the preceding and / or succeeding text segments adjacent to that text segment, as the target text content. After the target text content is determined, it is rendered and displayed on the interface so that the display progress of the text content and the playback progress of the audio stream are synchronized.

[0044] As an option, obtaining the audio stream corresponding to the shared streaming device can be understood as obtaining the corresponding audio stream when a user on the shared streaming device is singing the target song. Therefore, playing the audio stream and rendering the target text content corresponding to the audio stream according to the current audio progress can be understood as playing the audio stream and simultaneously rendering and displaying the target text content corresponding to the audio stream in the interface in the form of subtitles, according to the current audio progress, so that the display of subtitles and the playback of the audio stream are synchronized.

[0045] The solution of this embodiment can be applied to a karaoke scenario using co-streaming. The co-streaming users consist of a streamer and at least one guest. The streamer and at least one guest sing a target song via co-streaming. The co-streaming terminal collects the audio from the co-streaming users, obtains an audio stream, and sends the audio stream to the viewer terminal for playback. While playing the audio stream, the viewer terminal simultaneously displays subtitles corresponding to the audio stream (i.e., lyrics of the target song). This solution ensures synchronization between audio playback and subtitle display.

[0046] A technical solution according to an embodiment of the present invention acquires an audio stream corresponding to the latest reference audio progress and a co-streaming terminal, the co-streaming terminal including guest terminals and / or streamer terminals, i.e., guest terminals and / or streamer terminals participating in the co-streaming, determines the current audio progress of the audio stream based on the local system time and the latest reference audio progress, plays the audio stream, and renders target text content corresponding to the audio stream according to the current audio progress. The speech-based text processing method according to an embodiment of the present invention improves the audiovisual effect of the audio by determining the current audio progress of the audio stream based on the local system time and the latest reference audio progress, and rendering target text content corresponding to the audio stream according to the current audio progress.

[0047] Figure 4 is a schematic diagram showing the structure of a speech-based text processing device according to one embodiment of the present invention. As shown in Figure 4, the device comprises the following:

[0048] The latest reference audio progress acquisition module 210 is configured to acquire the latest reference audio progress and the audio stream corresponding to the co-streaming terminal. The co-streaming terminal includes guest terminals and / or streamer terminals.

[0049] Currently, the audio progress determination module 220 is configured to determine the current audio progress of the audio stream based on the local system time and the most recent reference audio progress.

[0050] The target text content rendering module 230 is configured to play the audio stream and render the target text content corresponding to the audio stream according to the current audio progress.

[0051] Optionally, the current audio progress determination module 220 is further configured to determine the time interval between the current time and the time the latest reference audio progress was received, based on the local system time, and to determine the current audio progress of the audio stream according to the time interval and the latest reference audio progress.

[0052] Optionally, the current voice progress determination module 220 is further configured to obtain a first local system time corresponding to the current time and a second local system time corresponding to the time the latest reference voice progress was received, and to determine a time interval according to the first local system time and the second local system time.

[0053] Optionally, the current audio progress determination module 220 is further configured to determine the current audio progress of the audio stream based on a time interval and the most recent reference audio progress.

[0054] Optionally, the target text content rendering module 230 is further configured to acquire a text file corresponding to the audio stream, include audio timestamps in the text content within the text file, extract target text content from the text file based on the current audio progress, and render the target text content.

[0055] Optionally, the target text content rendering module 230 is further configured to use the text content in a text file whose audio timestamp matches the current audio progress as the target text content.

[0056] Optionally, the latest reference audio progress acquisition module 210 is further configured to receive the latest media supplement information and the audio stream corresponding to the co-streaming terminal, extract the latest reference audio progress from the latest media supplement information, and record the local system time corresponding to the time when the latest media supplement information was received.

[0057] Optionally, the latest reference audio progress acquisition module 210 can be further configured to acquire the corresponding audio stream when a user on a co-streaming terminal is singing the target song.

[0058] Optionally, the target text content rendering module 230 is further configured to render and display the target text content corresponding to the audio stream in the interface as subtitles, according to the current audio progress, while simultaneously playing the audio stream, so that the display of subtitles and the playback of the audio stream are synchronized.

[0059] The speech-based text processing device provided by embodiments of the present invention can perform the speech-based text processing method provided by embodiments of the present invention and has corresponding functional modules and beneficial effects for performing the method.

[0060] It is important to note that the various units and modules included in the above-described device are merely divided according to functional logic, and are not limited to the above division as long as the corresponding functions can be realized. Furthermore, the specific names of the functional units are for convenience to distinguish them from one another and are not used to limit the scope of protection of the embodiments of the present invention.

[0061] Figure 5 is a schematic diagram showing the structure of an electronic device provided by one embodiment of the present invention. Referring to Figure 5, a schematic diagram is shown showing the structure of an electronic device (e.g., terminal device or server in Figure 5) 500 suitable for carrying out one embodiment of the present invention. The terminal devices in embodiments of the present invention include, but are not limited to, mobile devices such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), and PMPs (portable multimedia players), as well as in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital televisions and desktop computers. The electronic device shown in Figure 5 is merely an example and does not limit the functions and scope of use of embodiments of the present invention.

[0062] As shown in Figure 5, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501 that can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 502 or programs loaded from storage device 508 into random access memory (RAM) 503. RAM 503 also stores various programs and data necessary for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An input / output (I / O) interface 505 is also connected to bus 504.

[0063] Typically, the I / O interface 505 is connected to input devices 506, such as a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, and gyroscope; output devices 507, such as a liquid crystal display (LCD), speaker, and vibrator; storage devices 508, such as magnetic tape and hard disk; and communication devices 509. The communication devices 509 enable the electronic device 500 to communicate with other devices wirelessly or via a wired connection to exchange data. Figure 5 shows the electronic device 500 with various devices, but it should be understood that not all of the illustrated devices are required to be implemented or present. Alternatively, more or fewer devices may be implemented or present.

[0064] In particular, according to one embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, one embodiment of the present invention includes a computer program product which includes a computer program recorded on a non-temporary computer-readable medium, and which includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 509, installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, the above-described functions as defined in the method of the embodiment of the present invention are performed.

[0065] In the embodiments of this disclosure, the names of messages or information exchanged between multiple devices are used for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0066] The electronic device according to the embodiment of the present invention and the speech-based text processing method of the above embodiment belong to the same inventive concept. For technical details not fully explained in this embodiment, please refer to the above embodiment. This embodiment has the same beneficial effects as the above embodiment.

[0067] One embodiment of the present disclosure provides a computer storage medium storing a computer program. When this program is executed by a processor, the speech-based text processing method provided in the above embodiment is realized.

[0068] In this disclosure, the computer-readable media described above may be computer-readable signal media, computer-readable storage media, or any combination thereof. Computer-readable storage media include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media include, but are not limited to, electrical connections having one or more conductors, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, computer-readable storage media may be any tangible medium that contains or stores programs that can be used by or in conjunction with instruction execution systems, devices, or equipment. In this disclosure, computer-readable signal media may include data signals that propagate in the baseband or as part of a carrier wave and carry computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium, on which programs used by or in connection with instruction execution systems, devices, or equipment can be transmitted, propagated, or transferred. Program code embodied in a computer-readable medium may be transmitted using any suitable medium, including but not limited to wired, optical, RF (radio frequency), or any suitable combination thereof.

[0069] In some embodiments, client terminals and servers can communicate using any currently known or future-developed network protocol, such as HTTP (HyperTextTransferProtocol), and can be interconnected in any form or medium of digital data communication (e.g., communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), peer-to-peer networks (e.g., ad-hoc peer-to-peer networks), and any other currently known or future-developed networks.

[0070] Computer-readable media may be included in electronic devices or may exist independently without being incorporated into electronic devices.

[0071] The computer-readable medium described above holds one or more programs. When one or more of these programs are executed by an electronic device, the electronic device retrieves the latest reference audio progress and audio stream corresponding to the co-streaming terminal (which includes guest terminals and / or streamer terminals), determines the current audio progress of the audio stream based on the local system time and the latest reference audio progress, plays the audio stream, and renders the target text content corresponding to the audio stream based on the current audio progress.

[0072] Computer program code for performing the operations of this disclosure may be written in one or more programming languages, or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, and traditional procedural programming languages ​​such as "C" or similar programming languages. The program code may run entirely on the user's computer, partially on the user's computer, run as a standalone software package, run partially on the user's computer and partially on a remote computer, or run entirely on a remote computer or server. Where a remote computer is involved, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it may be connected to an external computer (for example, via the Internet through an Internet service provider).

[0073] The flowcharts and block diagrams in the accompanying drawings illustrate possible implementation architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for performing a specified logical function. It should also be noted that in some alternative implementations, the functions shown in the boxes may be executed in a different order than shown in the accompanying drawings. For example, two boxes shown consecutively may actually be executed substantially in parallel, or, depending on the functions they contain, in reverse order. It should also be noted that each box in the block diagrams and / or flowcharts, and any combination of boxes in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or operation, or by a combination of dedicated hardware and computer instructions.

[0074] The units relating to the embodiments described herein may be implemented in software or hardware. In some cases, the names of the units are not limiting to the units themselves. For example, the first acquisition unit may also be described as "a unit for acquiring at least two Internet Protocol addresses."

[0075] At least some of the functions described herein may be performed by one or more hardware logic components. Examples of usable hardware logic components include, but are not limited to, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), and composite programmable logic devices (CPLDs).

[0076] In the context of this disclosure, a machine-readable medium can be a tangible medium capable of storing or storing programs used by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium includes, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatus, or any suitable combination thereof. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0077] The above description represents preferred embodiments of the disclosure and is merely illustrative of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. Other technical solutions formed by any combination of the aforementioned technical features or their equivalents are also included without departure from the scope of this disclosure. For example, this includes technical solutions formed by replacing the aforementioned features with (but not limited to) similarly functional technical features disclosed herein.

[0078] Furthermore, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in a specific illustrated or sequential order. In some situations, multitasking and parallel processing may be advantageous. Similarly, although the above description includes some specific implementation details, these should not be construed as limiting the scope of this disclosure. Some features described in the context of other embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented individually or in any suitable sub-combination mode in multiple embodiments.

[0079] While the subject matter is described using terminology specific to structural features and / or methodological logical actions, it should be understood that the subject matter as defined in the attached claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of how the claims may be implemented.

Claims

1. A speech-based text processing method performed by a viewer terminal, To obtain the latest reference audio progress and the audio stream corresponding to the co-streaming terminal, wherein the co-streaming terminal includes the guest terminal and / or the streamer terminal, Based on the local system time and the most recent reference audio progress, the current audio progress of the audio stream is determined, This includes playing the audio stream and rendering target text content corresponding to the audio stream according to the current audio progress, A method of processing text based on speech.

2. Determining the current audio progress of the audio stream based on the local system time and the latest reference audio progress is: Based on the local system time, determine the time interval between the current time and the time when the latest reference audio progress was received, The process includes determining the current audio progress of the audio stream according to the aforementioned time interval and the most recent reference audio progress, The method according to claim 1.

3. Based on the local system time, determining the time interval between the current time and the time when the latest reference audio progress was received means: To obtain a first local system time corresponding to the current time and a second local system time corresponding to the time when the latest reference audio progress was received. This includes determining a time interval according to the first local system time and the second local system time, The method according to claim 2.

4. Determining the current audio progress of the audio stream according to the aforementioned time interval and the latest reference audio progress means that This includes determining the current audio progress of the audio stream based on the aforementioned time interval and the most recent reference audio progress, The method according to claim 2.

5. Rendering text content corresponding to the audio stream according to the current audio progress is: The process involves obtaining a text file corresponding to the aforementioned audio stream, wherein the text content within the text file includes an audio timestamp. Based on the current audio progress, the target text content is extracted from the text file. This includes rendering the target text content, The method according to claim 1.

6. Extracting the target text content from the text file based on the current audio progress is: This includes using the text content in the text file, where the audio timestamp matches the current audio progress, as the target text content. The method according to claim 5.

7. Obtaining the latest reference audio progress and audio streams corresponding to the shared streaming terminal is, To obtain the latest media supplementary information and audio streams compatible with shared streaming devices, This includes extracting the latest reference audio progress from the latest media supplement information and recording the local system time corresponding to the time when the latest media supplement information was received, The method according to claim 1.

8. Obtaining an audio stream compatible with a shared streaming device is To obtain the corresponding audio stream when a user of a shared streaming device is singing the target song. Therefore, playing the audio stream and rendering the target text content corresponding to the audio stream according to the current audio progress is: The display of subtitles and the playback of the audio stream are synchronized, and the audio stream is played while simultaneously rendering and displaying target text content corresponding to the audio stream in the form of subtitles on the interface according to the current audio progress, including: The method according to claim 1.

9. A speech-based text processing device located within a viewer terminal, A latest reference audio progress acquisition module that acquires the latest reference audio progress and an audio stream corresponding to a co-streaming terminal, wherein the co-streaming terminal includes a guest terminal and / or a streamer terminal, A current audio progress determination module that determines the current audio progress of the audio stream based on the local system time and the most recent reference audio progress, The system includes a target text content rendering module that plays the audio stream and renders target text content corresponding to the audio stream according to the current audio progress. A speech-based text processing device.

10. One or more processors, An electronic device comprising a memory device for storing one or more programs, When the one or more programs are executed by the one or more processors, they cause the one or more processors to execute the speech-based text processing method described in any one of claims 1 to 8. electronic equipment.

11. When executed on a computer processor, it stores computer-executable instructions for performing the speech-based text processing method described in any one of claims 1 to 8. storage medium.

12. A computer program product that is stored in a computer storage medium, and which, when executed by a device, includes computer-executable instructions that cause the device to implement the method described in any one of claims 1 to 8.