Method and apparatus for processing speech data, and device and medium

By providing subtitle text in the audio live broadcast page, the problem of lack of visual content in audio live broadcast is solved, achieving a richer visual experience and better information understanding.

WO2025214266A1PCT designated stage Publication Date: 2025-10-16BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/087302
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-07
Filing Date
2025-04-03
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

The lack of visual content in audio live streaming applications leads to single functionality, and users may not be able to understand the live content due to environmental noise and network transmission issues.

Method used

Subtitle text is provided on the audio live broadcast page. The subtitle text is extracted from the voice data through automatic speech recognition technology and presented on the page, supporting users to distinguish the speeches of different target objects and allowing users to adjust the display settings of subtitles.

Benefits of technology

It enriches the visual effects of audio live broadcasts, helping users better understand the live content, especially in situations of environmental noise and poor network transmission, and improving the convenience of information acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025087302_16102025_PF_FP_ABST
    Figure CN2025087302_16102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a method and apparatus for processing speech data, and a device and a medium. In one method, in response to receiving speech data of a target object (810), subtitle text associated with the speech data is acquired (820); and the subtitle text is provided in an audio live-streaming page (830). By utilizing the exemplary implementation in the present disclosure, more visual elements can be presented in an audio live-streaming application, so that more abundant visual effects are provided, and a user is supported in understanding more information of audio live streaming.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device and medium for processing voice data

[0001] The present application claims priority to the Chinese patent application No. 202410411718.3, filed on April 7, 2024, entitled “Method, apparatus, device and medium for processing voice data”, the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] Exemplary implementations of the present disclosure generally relate to voice data processing, and in particular, to a method, apparatus, device and computer readable storage medium for processing voice data in an audio live streaming application. BACKGROUND

[0003] With the development of computer technology, currently, an audio live streaming application can be provided at a client device. A user can install a live streaming application on the client device and access a live streaming room. In the audio live streaming application, voice data of the live streaming can be provided, at this time, the user can only hear the voice of the host and / or the guest, however, there is a lack of visual content in the audio live streaming application. SUMMARY

[0004] In a first aspect of the present disclosure, a method for processing voice data is provided. In the method, in response to receiving voice data of a target object, caption text associated with the voice data is obtained. The caption text is provided in an audio live streaming page.

[0005] In a second aspect of the present disclosure, an apparatus for processing voice data is provided. The apparatus comprises: an obtaining module configured to, in response to receiving voice data of a target object, obtain caption text associated with the voice data; and a providing module configured to provide the caption text in an audio live streaming page.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device comprises: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.

[0007] In a fourth aspect of the present disclosure, a computer readable storage medium is provided, having stored thereon a computer program, which, when executed by a processor, causes the processor to implement the method according to the first aspect of the present disclosure.

[0008] It is to be understood that the particulars shown herein are by way of example and for purposes of illustrative clarity, not meant to limit the scope of the present disclosure. Other features, objects, and advantages of the present disclosure will be apparent from the following detailed description and drawings. BRIEF DESCRIPTION OF DRAWINGS

[0009] The above and other features, aspects, and advantages of various implementations of the present disclosure will become more apparent from the following detailed description and accompanying drawings. In the drawings:

[0010] FIG. 1 illustrates a block diagram of an application environment according to one example implementation of the present disclosure;

[0011] FIG. 2 illustrates a block diagram for processing speech data according to some implementations of the present disclosure;

[0012] FIG. 3 illustrates a block diagram of a process for adjusting a text box according to some implementations of the present disclosure;

[0013] FIG. 4 illustrates a block diagram for switching a display state of caption text according to some implementations of the present disclosure;

[0014] FIG. 5 illustrates a block diagram of an architecture of various modules according to some implementations of the present disclosure;

[0015] FIG. 6A illustrates a block diagram of a system architecture for processing speech data according to some implementations of the present disclosure;

[0016] FIG. 6B illustrates a block diagram of a structure of a drag component according to some implementations of the present disclosure;

[0017] FIG. 6C illustrates a block diagram of a structure of a scroll component according to some implementations of the present disclosure;

[0018] FIG. 6D illustrates a block diagram of a structure of a switch component according to some implementations of the present disclosure;

[0019] FIGS. 7A to 7C respectively illustrate block diagrams of interactions of various modules according to some implementations of the present disclosure;

[0020] FIG. 8 illustrates a flowchart of a method for processing speech data according to some implementations of the present disclosure;

[0021] FIG. 9 illustrates a block diagram of an apparatus for processing speech data according to some implementations of the present disclosure; and

[0022] FIG. 10 illustrates a block diagram of a device capable of implementing various implementations of the present disclosure. DETAILED DESCRIPTION

[0023] Implementations of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While several implementations of the present disclosure are described, it should be understood that the present disclosure can be embodied in many other forms and should not be construed as limited to the implementations set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art. It should be understood that the drawings and implementations described are for illustrative purposes only and are not intended to limit the scope of the present disclosure.

[0024] In the description of implementations of the present disclosure, the term "includes" and its derivatives mean "including but not limited to". The term "based on" means "based at least in part on". The term "one implementation" or "the implementation" means "at least one implementation". The term "some implementations" means "at least some implementations". Other explicit or implicit definitions can also be included below. As used herein, the term "model" can represent the relationship between various data. For example, the above-mentioned relationship can be obtained based on various technical solutions known at present and / or to be developed in the future.

[0025] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.

[0026] It can be understood that before using the technical solutions disclosed by the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the scene of use, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.

[0027] For example, in response to receiving the active request of the user, the user is sent a prompt message to explicitly prompt the user that the operation requested to be executed will require the acquisition and use of the personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that executes the operation of the technical solutions of the present disclosure according to the prompt message.

[0028] As an optional but non-limiting implementation, in response to receiving the active request of the user, the way of sending the prompt message to the user, for example, can be the way of pop-up window, and the prompt message can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0029] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementations of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementations of the present disclosure.

[0030] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of the execution of a subsequent action performed in response to the event or condition is not necessarily strongly correlated with the time at which the event occurs or the condition is established. For example, in some cases, the subsequent action can be performed immediately upon the occurrence of the event or the establishment of the condition; in other cases, the subsequent action can be performed after a period of time has elapsed since the occurrence of the event or the establishment of the condition.

[0031] Example environment

[0032] It is currently possible to provide an audio live streaming application at a client device. A user can install a live streaming application on a client device and access a live room. Referring to FIG. 1, a block diagram 100 of an application environment according to one example implementation of the present disclosure is shown. As shown in FIG. 1, a server device 110 can be a server end running an audio live streaming application, and users (e.g., users 130, 132, …, and 134) can log in to the audio live streaming application using respective client devices (e.g., client devices 120, 122, …, and 124) and access the server device 110 via a network 112. The users can have respective types (e.g., types 150, 152, …, and 154), and perform respective operations in a live room 140 based on different types.

[0033] For example, the type 150 can be a host of the audio live room 140, which can manage the live room 140 as a host, e.g., provide voice data to other users, invite other users, etc. The type 152 can be a guest of the live room 140, which can provide voice data to other users; and the type 154 can be a spectator who can only listen to voice data of the host and / or the guest.

[0034] In the audio live streaming application, live voice data can be provided, at which time a user can only hear the voice of the host and / or the guest, however there is a lack of visual content in the audio live streaming application. This results in a single function of the audio live streaming application, and due to problems such as environmental noise and / or network transmission, a user can not be able to understand the content of the audio live streaming. At this time, it is desirable to provide more rich visual content in the audio live streaming application, and facilitate a user to obtain more information.

[0035] Summary of processing voice data

[0036] To at least partially address the deficiencies in the prior art, according to one example implementation of the present disclosure, a method for processing voice data is proposed. A summary according to one example implementation of the present disclosure is described with reference to FIG. 2, which shows a block diagram 200 for processing voice data according to some implementations of the present disclosure. As shown in FIG. 2, in an audio live page 210, identifiers of target objects can be presented, e.g., identifiers 130, 132, etc. The target object corresponding to identifier 130 can be, for example, a host of a live room, and the target object corresponding to identifier 132 can be, for example, a guest of the live room. The host and the guest can have a conversation, and other users (e.g., viewers) in the audio live room can listen to the conversation.

[0037] In the audio live application, upon receiving voice data from a target object in the audio live application, subtitle text associated with the voice data can be obtained. According to one example implementation of the present disclosure, the text can be recognized from the voice data based on a variety of ways, and then the subtitle text can be obtained. Further, the subtitle text can be provided in an audio live page in the audio live application. As shown in FIG. 2, subtitle text 220 can be presented in audio live page 210. Alternatively and / or additionally, subtitle text 222 can be presented.

[0038] With the example implementations of the present disclosure, more visual elements can be presented in the audio live page, thereby providing more rich visual effects. Further, more information can be provided to the user, so as to support the user to understand more information of the audio live when the quality of the audio data is suboptimal due to environmental noise and / or network transmission, etc.

[0039] Detailed process of processing audio data

[0040] Having described a summary according to one example implementation of the present disclosure, in the following, more information about the audio live will be provided. As shown in FIG. 2, to present the subtitle text, a text box 230 can be provided in audio live page 210. Here, text box 230 can include identifiers of the target objects and the subtitle text. For example, text box 230 includes identifier “Host” about the host and subtitle text 220 recognized from the audio data of the host, and identifier “Guest” about the guest and subtitle text 222 recognized from the audio data of the guest. In this way, the subtitle text recognized from the audio data of different target objects can be explicitly presented, thereby facilitating the viewers to understand the viewpoints expressed by the respective target objects.

[0041] According to one example implementation of the present disclosure, the subtitle text can be provided to any of the host, the guest, and the audience. Since the audience cannot speak, the subtitle text provided at this time is the subtitle text recognized from the speech data of the host and / or the guest. In this way, it can be ensured that various types of users of the live room application can see the subtitle text of the played audio data, thereby improving the richness of the visual data of the live room application and supporting users to obtain more information.

[0042] It should be understood that although FIG. 2 only schematically shows two target objects, the host and the guest, alternatively and / or additionally, the audio live room can include more target objects, for example, the host can invite more guests, or the audio live room can include more hosts, or the audio live room can include only a single host and not include any guest, etc. At this time, the identifier of each target object can be respectively presented in a single text box, and the subtitle text recognized from the audio data of different target objects can be presented after the respective identifiers.

[0043] According to one example implementation of the present disclosure, assuming that the audio live room includes one host and two guests, at this time, the identifier of each target object can include, for example, “host”, “guest 1”, “guest 2”, etc. Alternatively and / or additionally, the identifier of each target object can be represented by the username, nickname, avatar, and / or in other ways. For example, the subtitle text of each target object can be distinguished in different colors, different fonts, different background colors, etc. In this way, it can be convenient for users to distinguish different target objects, and thus more accurately understand the expressions from each target object.

[0044] According to one example implementation of the present disclosure, in the process of presenting the subtitle text, the subtitle text can be presented in the text box in the time sequence of obtaining the subtitle text. In other words, the subtitle text can be presented in the order of the speech of each target object. As time goes by, the text box 230 can include more lines of subtitle text, and a scroll bar can be displayed at the right side (or other position) of the text box 230. The user can use the scroll bar to browse the subtitle text from each target object, thereby supporting the user to review the previously received information as needed.

[0045] According to one example implementation of the present disclosure, when the user adjusts the scroll bar to browse the previously received subtitle text, the audio data from each target object can be continuously received in real time, and the subtitle text recognized from the newly received audio data can be continuously added in the text box 230 in the time sequence. In this way, it can be supported that the user listens to the currently latest audio data while browsing the previous subtitle text.

[0046] Alternatively and / or additionally, when the user adjusts the scroll bar to browse the previous subtitle text, the audio data associated with the previous subtitle text can be played. Meanwhile, the subtitle text identified from the newly received audio data can be continuously added in the text box 230 in chronological order. In this way, the user can playback the previously received audio data, while the subtitle file of the latest audio data can be stored in real time.

[0047] According to one example implementation of the present disclosure, the user can adjust the settings of the text box as desired. Upon detecting an adjustment action for adjusting the text box, the text box is updated based on the adjustment action. Specifically, the adjustment action is for adjusting at least any of the size of the text box, the position of the text box, the color, font, size, and background color of the subtitle text, etc.

[0048] Specifically, the user can move the position of the text box by a "dragging" action, and can adjust the size of the text box by a "zooming" action. Further, the user can modify the color, font, size, and background color of the subtitle text, etc. For example, the user can set the font color to be consistent with the color tone of the avatar of the target object. Assuming that the avatar tone of the host is warm, and the avatar tone of the guest is cold, the color of the relevant subtitle text of the host can be set to warm, and the color of the relevant subtitle text of the guest can be set to cold. In this way, the user can visually distinguish the speech of each target object.

[0049] According to one example implementation of the present disclosure, the subtitle text can be presented at a position associated with the identifier of the target object. More details are described with reference to FIG. 3, which shows a block diagram 300 of a process for adjusting a text box according to some implementations of the present disclosure. As shown in FIG. 3, a text box 310 can be presented at a position near the identifier 130 of the host, and the text box 310 can include the subtitle text identified from the audio data of the host. Further, a text box 320 can be presented at a position near the identifier 132 of the guest, and the text box 320 can include the subtitle text identified from the audio data of the guest.

[0050] With example implementations of the present disclosure, the subtitle data from different target objects can be displayed in different text boxes, respectively. In this way, the viewpoints from different target objects can be distinguished in a more clear manner, thereby reducing the risk of the user confusing the data from different target objects.

[0051] According to one example implementation of the disclosure, in a case where the target objects include a first target object and a second target object, the first subtitle text of the first target object and the second subtitle text of the second target object can be presented in the text box in a time order in which the first subtitle text and the second subtitle text are acquired.

[0052] In particular, as shown in FIG. 3, in each text box, the corresponding subtitle text can be presented in a time order. For example, the subtitle text recognized from the audio data of the host can be presented in order in the text box 310, and the subtitle text recognized from the audio data of the host can be presented in order in the text box 320. As time elapses, each text box can include more lines of subtitle text, and a scroll bar can be displayed in each text box. The user can utilize the scroll bar to browse the subtitle text from the respective target objects, thereby supporting the user to review the previously received information in a more clear manner according to his or her own needs.

[0053] According to one example implementation of the disclosure, during the live broadcast, the user can switch the display mode of the subtitle text. For example, the user as a spectator can start the subtitle display function, or the user can disable the subtitle display function. More details are described with reference to FIG. 4, which shows a block diagram 400 for switching the display state of the subtitle text according to some implementations of the disclosure.

[0054] As shown in FIG. 4, the audio live broadcast page 210 can further include switching controls 410 and 420. For example, the user can click the control 410 to turn on the subtitle text of the host, and the user can click the control 420 to turn off the subtitle text of the guest. At this time, the text box 430 only includes the subtitle text of the host. According to one example implementation of the disclosure, the user can adjust the switching controls at any time as needed. For example, the user can click the control 420 to turn on the subtitle text of the guest. At this time, the text box 430 will include the text of both the host and the guest. According to one example implementation of the disclosure, the text box 430 can only include the subtitle text of the guest after the turning on operation. Alternatively and / or additionally, the text box 430 can only include the entire subtitle text of the guest, i.e., including the subtitle text before and after the turning on operation.

[0055] According to one example implementation of the present disclosure, the target object can be allowed to select whether to start the caption function. For example, a control for enabling / disabling the caption function can be provided in the audio live application of the host and the guest, and in a case where it is determined that the target object starts the caption function, at least one speech segment can be determined based on the speech data. For example, the respective speech segments can be divided based on a predetermined time length, or for example, the respective speech segments can be divided according to pauses in the speech data, and further, caption text can be extracted from the determined at least one speech segment. In this way, it can be supported that the user decides whether to enable the caption function according to his / her own needs.

[0056] More details of presenting the caption text are described with reference to FIG. 5, which illustrates a block diagram 500 of an architecture of respective modules according to some implementations of the present disclosure. As shown in FIG. 5, a real-time communication (RTC) network 514 can support presenting the caption text at a plurality of clients (e.g., a host client 520, a guest client 522, a guest client 524, and a viewer client 526). Specifically, at the host client 520, the host can start the caption function. At this time, the RTC network 514 can receive speech data 530 from the host client 520, and the RTC network can send the received speech data 531 to an automatic speech recognition (ASR) service 512. The ASR service 512 can perform an automatic speech recognition process to extract caption text from the speech data. The ASR service 512 can transmit the recognized caption text 532 to the RTC network 514.

[0057] Alternatively and / or additionally, a glossary service 510 can perform a text processing operation. Specifically, the RTC network 514 can transmit the caption text 533 to the glossary service 510, and the glossary service 510 can process the caption text, e.g., correct recognition errors, remove sensitive information, etc., and in turn transmit the processed caption text 534 to the RTC network 514. The RTC network 514 can provide the caption text to be presented to the host client 520 through a callback 534. At this time, the host client 520 can present the caption text in a local audio live application.

[0058] Further, the host client 520 can provide the speech data and the caption text to a content distribution network (CDN) 516. For example, the caption text can be added in a supplemental enhancement information (SEI) portion of an audio stream, so that the viewer client 526 can extract caption data 540 from the SEI portion and in turn present the caption text in an audio live application at the viewer client 526. Alternatively and / or additionally, the speech data and the caption text can be provided by the RTC network 514 to the content distribution network (CDN) 516, and in turn the caption text can be added in the SEI portion of the audio stream.

[0059] According to one example implementation of the present disclosure, the subtitle text can be extracted from the supplemental enhancement information of the audio stream associated with the audio data. In this way, the subtitle text will be sent to the spectator client along with the audio data.

[0060] The process of obtaining the audio data of the host and presenting the subtitle data at the host client and the spectator client respectively has been described. The process for the audio data from the guest is also similar, for example, the RTC network 514 can receive the audio data from the guest clients 522 and 524, invoke the ASR service 512 and the word list service 510 to generate the processed subtitle data. The RTC network 514 can provide the subtitle data of the speech uttered by the guest through the callbacks 536 and 537 respectively, so as to be presented at the guest clients 522 and 524 respectively. Further, the corresponding subtitle text can be added to the SEI part of the corresponding audio stream by the guest clients 522 and 524 or the RTC network 514, so as to be presented at the spectator client 526. It should be understood that although FIG. 5 shows the case that the ASR service 512 and the word list service 510 are provided at different locations, alternatively and / or additionally, the above services can be provided inside the RTC network 514.

[0061] According to one example implementation of the present disclosure, in order to obtain the subtitle text, the audio data from the host and the guest can be received by a first service of the live room application (for example, the real-time communication service provided by the RTC network 514), and the subtitle text can be extracted from the audio data by a second service of the live room application (for example, the speech recognition service provided by the ASR service 512). Specifically, as shown in FIG. 5, the real-time communication service provided by the RTC network 514 can receive the audio data 530 from the host. Similarly, assuming that the guest at the guest client 522 or 524 speaks, the audio data from the guest can also be extracted. Then, the subtitle text can be extracted from the audio data by the ASR service. In this way, the services of the live room application can cooperate with each other, and thus obtain the subtitle text.

[0062] According to one example implementation of the present disclosure, the subtitle text is the subtitle text processed by the word list service of the live room application. As shown in FIG. 5, the extracted subtitle text can be processed by the word list service 510. For example, the recognized errors can be corrected, the sensitive information can be removed, and so on, thereby obtaining more accurate subtitle text.

[0063] According to one example implementation of the present disclosure, the subtitle text can be provided to the client of the host. Specifically, the subtitle text can be received at the client of the host (for example, via callback), and then rendered to the audio live page by the client. In this way, the rendering process can be completed by the client device of the host, thereby realizing the presentation of the subtitle text at the client of the host.

[0064] According to one example implementation of the disclosure, the subtitle text can be provided at the clients of the audience and the guests via a "host-side merge" mode, or a "server-side merge" mode. Specifically, in the "host-side merge" mode, the client of the host can generate a rendered page that includes the audio live page and the subtitle text rendered to the audio live page. Further, a content distribution service (e.g., provided by the CDN 516) of the live room application can receive the rendered page, and in turn provide the rendered page to the clients of the audience and the guests for rendering the rendered page at the clients. In this way, the rendered page that has been generated at the client of the host can be reused.

[0065] According to one example implementation of the disclosure, in the "server-side merge" mode, the subtitle text can be received by a first service (e.g., provided by the RTC network 514), and the first service can render the subtitle text to the audio live page to generate a rendered page. Further, the first service can generate a video stream with the rendered page, and send the video stream to a third service. Here, each video frame in the video stream includes the audio live page and the subtitle text rendered to the audio live page. The video stream can be received by the third service (e.g., a content distribution service provided by the CDN 516), and sent to the clients of the audience and the guests for rendering the rendered page at the clients. In this way, the workload at the client of the host can be reduced.

[0066] In the context of the disclosure, the audio live content is rendered at the client device in the form of a video stream. On one hand, the audio portion in the video stream can output the sound signal of the audio live; on the other hand, the picture portion in the video stream can provide the corresponding subtitle text of the sound signal, so as to provide more visual content in the audio live application and facilitate the user to obtain more information.

[0067] According to one example implementation of the disclosure, the speech processing procedure described above can be implemented based on a plurality of classes. Table 1 schematically shows the description of a plurality of core classes involved in processing the speech and rendering the subtitle text in the audio live application.

[0068] Table 1 Description of Core Classes

[0069] The calling method of each class is described with reference to FIG. 6A, which shows a block diagram 600A of a system architecture for processing voice data according to some implementations of the present disclosure. As shown in FIG. 6A, the live captioning business component 644 can call the live captioning business service 640 to render captions. Specifically, data provided by the link manager 642 can be received, e.g., data from the live link component 630, such that it is determined that the callback of the captioning function is enabled or disabled, the captioning callback can be received, and so on. The live captioning business component 644 can receive the live captured stream from the audio live room anchor component 646, such that caption text is rendered. Further, the live captioning business component 644 can call the drag component 610 and the scroll component 620 to support the user’s drag and scroll operations. In this way, rendering / hiding captions and updating captions can be supported in real time.

[0070] FIG. 6B shows a block diagram 600B of the structure of the drag component according to some implementations of the present disclosure. As shown in FIG. 6B, the drag component 610 can include multiple classes: a live captioning tool describer 611, a live captioning tool 612, a live captioning service 613, a live captioning drag parent view 615, a live captioning component 614, and a drag container 616. Specifically, the text box described above can be rendered in the form of a floating window, and the drag function of the text box can be implemented via the above-mentioned classes. For example, the user can place the text box at a desired location through a drag operation, at which time the caption text within the text box will move along with the drag operation.

[0071] FIG. 6C shows a block diagram 600C of the structure of the scroll component according to some implementations of the present disclosure. As shown in FIG. 6C, the scroll component 620 can include multiple classes: a live captioning manager 621, a recycle view 622, a live captioning scroll helper 623, and a live captioning adapter 624. The user can scroll the caption text in the text box, and the above-mentioned classes can be used to support the user’s scroll operation. At this time, the text within the text box will scroll up or down along with the scroll operation.

[0072] FIG. 6D shows a block diagram 600D of the structure of the switch component according to some implementations of the present disclosure. As shown in FIG. 6D, the live link component 630 can include multiple classes: a remote live room event handler 631, a client implementation 632, and a live room implementation 633. The above-mentioned multiple classes can be used to manage the callback function related to enabling / disabling the caption function, such that caption text is rendered or hidden in the audio live room application.

[0073] The various steps for processing video data and presenting caption text have been described above, respectively, and the overall flow of presenting caption text is described below with reference to FIGS. 7A-7C. FIG. 7A illustrates a block diagram 700A of the interaction of various modules, according to some implementations of the present disclosure. As shown in FIG. 7A, the host can operate the host client 722 to open 730 the audio live room application and load the co-streaming caption service component 644. The co-streaming caption service component 644 can call an application programming interface (API) to perform a check function. The server 710 can return an API response, at which point the co-streaming service component 744 can determine 733 whether the host joins the RTC based on the response. If the determination is “yes,” the co-streaming service component 644 can send 734 a start signal (e.g., startSubtitle) to the RTC 712. In turn, the RTC 712 can return 735 an acknowledgement to the co-streaming caption service component 644, and the co-streaming caption service component 644 can present 736 the caption at the host client 722.

[0074] Further, in the case that the caption function is enabled, the co-streaming caption service component 644 can notify 737 the server 710 to provide the caption to the viewer client 720. The server 710 can send 738 the caption text to the viewer client 720 to present the caption at the viewer client 720. The RTC 712 can constantly detect whether new caption text is received, and in the case that new caption text is received, the RTC 712 can update 739 the caption at the host client 722. Further, the host client 722 can write 740 the caption into the SEI part of the speech stream, so that the CDN network 714 transmits 741 the updated caption to the viewer client 720. In this way, the updated caption can be presented at the viewer client 720.

[0075] FIG. 7B illustrates a block diagram 700B of the interaction of various modules, according to some implementations of the present disclosure. As shown in FIG. 7B, the host can operate the host client 722 to switch 750 the caption function (e.g., enable / disable the caption function). The co-streaming caption service component 644 can update 751 the state setting, and the server 710 can return 752 a response to the co-streaming service component 644. In the case that the caption function is enabled, the co-streaming service component 644 can send 753 a start signal to the RTC 712. In turn, the RTC 712 can return 754 an acknowledgement to the co-streaming caption service component 644, and the co-streaming caption service component 644 can present 755 the caption at the host client 722.

[0076] The live-mic captioning service component 644 can notify 756 the server 710 to provide the caption to the spectator client 720. The server 710 can send 757 the caption text to the spectator client 720 for rendering the caption at the spectator client 720. The RTC 712 can constantly detect whether new caption text is received, in the case that new caption text is received, the RTC 712 can update 758 the caption at the anchor client 722. Further, the anchor client 722 can write 759 the caption into the SEI part of the speech stream, so that the CDN network 714 transmits 760 the updated caption to the spectator client 720. In this way, the updated caption can be rendered at the spectator client 720.

[0077] FIG. 7C illustrates a block diagram 700C of the interaction of various modules, according to some implementations of the present disclosure. As shown in FIG. 7C, the anchor can switch 770 the settings of the anchor client 722, for example, switch the radio station (or turn on KTV, cross room, team battle, guest PK, you point me sing), etc., at this time the caption function is disabled. The server 710 can provide an acknowledgement to the anchor client 722, and the live-mic captioning service component 644 can send 772 a message to the RTC 712 to stop the caption. The RTC 712 can return 773 an acknowledgement to the live-mic captioning service component 644, in turn, the live-mic captioning service component 644 can close 774 the caption function, and send 775 a message to the server 710 to close the caption function. In turn, the server 710 can close the caption function at the spectator client 720.

[0078] Although FIGS. 7A-7C illustrate the process of rendering the caption text extracted from the anchor’s speech data at various clients, the caption text extracted from the guest’s speech data can be processed based on similar manners. In this way, the anchor and / or the guest can enable and / or disable the caption function in the audio live room application according to their own needs. With the exemplary implementations of the present disclosure, more visual elements can be presented in the audio live page, thereby providing more rich visual effects. Further, more information is provided to the user, so as to support the user to understand more information of the audio live when the quality of the audio data is suboptimal due to environmental noise and / or network transmission, etc.

[0079] Example process

[0080] FIG. 8 illustrates a flowchart of a method 800 for processing speech data, according to some implementations of the present disclosure. At block 810, it is determined whether speech data of a target object is received. In the case that the speech data is received, the method 800 proceeds to block 820. At block 820, caption text associated with the speech data is obtained. At block 830, the caption text is provided in an audio live page.

[0081] According to one example implementation of the present disclosure, presenting the subtitle text includes presenting the subtitle text at a location associated with the identifier of the target object.

[0082] According to one example implementation of the present disclosure, presenting the subtitle text includes providing a text box in the audio live page, the text box including the identifier of the target object and the subtitle text.

[0083] According to one example implementation of the present disclosure, presenting the subtitle text includes presenting the subtitle text in the text box in a time sequence in which the subtitle text is obtained.

[0084] According to one example implementation of the present disclosure, the method further includes, in response to detecting an adjustment action for adjusting the text box, updating the text box based on the adjustment action, the adjustment action being for adjusting at least any of a size of the text box, a position of the text box, a color, a font, a size, and a background color of the subtitle text.

[0085] According to one example implementation of the present disclosure, the subtitle text is obtained based on: in response to determining that the target object starts a subtitle function, determining at least one speech segment based on the speech data; and extracting the subtitle text from the at least one speech segment.

[0086] According to one example implementation of the present disclosure, obtaining the subtitle text includes extracting the subtitle text from a supplemental enhancement information of an audio stream associated with the audio data.

[0087] According to one example implementation of the present disclosure, the target object includes a first target object and a second target object, and presenting the subtitle text includes presenting, in a text box, a first subtitle text of first speech data of the first target object and a second subtitle text of second speech data of the second target object in a time sequence in which the first subtitle text and the second subtitle text are obtained.

[0088] According to one example implementation of the present disclosure, further including presenting, in the audio live page, a switching control for switching a display mode of the subtitle text; and in response to receiving an interactive operation on the switching control, presenting the subtitle text based on the interactive operation.

[0089] According to one example implementation of the present disclosure, the target object includes at least any of a host and a guest of a live room application; and providing the subtitle text includes providing the subtitle text to at least any of the host, the guest, and a spectator.

[0090] According to one example implementation of the present disclosure, the speech data is speech data from the target object in an audio live application, and providing the subtitle text in the audio live page includes providing the subtitle text in the audio live page in the audio live application.

[0091] According to one example implementation of the disclosure, obtaining the subtitle text includes: receiving, by a first service of the live room application, audio data from the host and the guest; and extracting, by a speech recognition service of the live room application, the subtitle text from the audio data.

[0092] According to one example implementation of the disclosure, providing the subtitle text includes: receiving, at a client of the host, the subtitle text; and rendering, by the client, the subtitle text to the audio live page.

[0093] According to one example implementation of the disclosure, providing the subtitle text includes: providing the subtitle text includes: receiving, by a third service of the live room application, a video stream, each video frame in the video stream including the audio live page and the subtitle text rendered to the audio live page; and presenting, at a client of the viewer and the guest, the video stream.

[0094] According to one example implementation of the disclosure, providing the subtitle text includes: receiving, by the first service, the subtitle text; rendering, by the first service, the subtitle text to the audio live page to generate a rendered page; and presenting, at a client of the viewer and the guest, the rendered page.

[0095] According to one example implementation of the disclosure, the subtitle text is the subtitle text processed by a glossary service of the live room application.

[0096] Example apparatuses and devices

[0097] FIG. 9 illustrates a block diagram of an apparatus 900 for processing speech data, according to some implementations of the disclosure. The apparatus 900 includes an obtaining module 910 configured to, in response to receiving speech data of a target object, obtain subtitle text associated with the speech data; and a providing module 920 configured to provide the subtitle text in an audio live page.

[0098] According to one example implementation of the disclosure, the providing module includes a presenting module configured to present the subtitle text at a location associated with an identifier of the target object.

[0099] According to one example implementation of the disclosure, the presenting module includes a text box providing module configured to provide a text box in the audio live page, the text box including the identifier of the target object and the subtitle text.

[0100] According to one example implementation of the disclosure, the presenting module includes a sequential presenting module configured to present the subtitle text in the text box in a time sequence in which the subtitle text is obtained.

[0101] According to one example implementation of the present disclosure, the apparatus further includes an updating module configured to, in response to detecting the adjustment action for adjusting the text box, update the text box based on the adjustment action, the adjustment action being for adjusting at least any of: a size of the text box, a position of the text box, a color, a font, a size, and a background color of the subtitle text.

[0102] According to one example implementation of the present disclosure, the subtitle text is obtained based on: a segment obtaining module configured to, in response to determining that the target object starts the subtitle function, determine at least one speech segment based on the speech data; and an extracting module configured to extract the subtitle text from the at least one speech segment.

[0103] According to one example implementation of the present disclosure, the obtaining module includes a subtitle extracting module configured to extract the subtitle text from a supplemental enhancement information of an audio stream associated with the audio data.

[0104] According to one example implementation of the present disclosure, the target object includes a first target object and a second target object, and the presenting module includes a text presenting module configured to present, in a time sequence of a first subtitle text of first speech data of the first target object and a second subtitle text of second speech data of the second target object, the first subtitle text and the second subtitle text in the text box.

[0105] According to one example implementation of the present disclosure, the apparatus further includes a switch control presenting module configured to present, in the audio live page, a switch control for switching a display mode of the subtitle text; and a switching module configured to, in response to receiving an interactive operation on the switch control, present the subtitle text based on the interactive operation.

[0106] According to one example implementation of the present disclosure, the target object includes at least any of: a host and a guest of a live room application; and the providing module is further configured to provide the subtitle text to at least any of: the host, the guest, and a spectator.

[0107] According to one example implementation of the present disclosure, the speech data is speech data from the target object in an audio live application, and providing the subtitle text in the audio live page includes providing the subtitle text in an audio live page of the audio live application.

[0108] According to one example implementation of the present disclosure, the obtaining module is further configured to receive, by a first service of a live room application, audio data from the host and the guest; and extract, by a speech recognition service of the live room application, the subtitle text from the audio data.

[0109] According to one example implementation of the present disclosure, the providing module is further configured for: receiving the subtitle text at the client of the host; and rendering, by the client, the subtitle text to the audio live page.

[0110] According to one example implementation of the present disclosure, the providing module is further configured for: receiving, by the third service of the live room application, the video stream, each video frame in the video stream including the audio live page and the subtitle text rendered to the audio live page; and presenting the video stream at the clients of the audience and the guest.

[0111] According to one example implementation of the present disclosure, the providing module is further configured for: receiving, by the first service, the subtitle text; rendering, by the first service, the subtitle text to the audio live page to generate a rendered page; and presenting the rendered page at the clients of the audience and the guest.

[0112] According to one example implementation of the present disclosure, the subtitle text is the subtitle text processed by a glossary service of the live room application.

[0113] FIG. 10 illustrates a block diagram of a device 1000 capable of implementing the implementations of the present disclosure. It should be understood that the computing device 1000 illustrated in FIG. 10 is merely an example and should not be construed to limit the functionality and scope of the implementations described herein. The computing device 1000 illustrated in FIG. 10 can be used to implement the methods described above.

[0114] As shown in FIG. 10, the computing device 1000 is in the form of a general- purpose computing device. The components of the computing device 1000 can include, but are not limited to, one or more processors or processing units 1010, a memory 1020, a storage device 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. The processing unit 1010 can be a real or virtual processor and is capable of executing various processing in accordance with programs stored in the memory 1020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 1000.

[0115] The computing device 1000 typically includes a plurality of computer storage media. Such media can be volatile, nonvolatile, removable, and / or non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. The memory 1020 can be volatile (such as random access memory (RAM)), non-volatile (such as read-only memory (ROM), EEPROM, flash memory, etc.), or some combination of the two. The storage device 1030 can be a removable storage medium and / or a non-removable storage medium implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data.

[0116] The computing device 1000 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 10, a disk drive and / or a CD-ROM drive can be provided for reading from or writing to a removable, non-removable, volatile, or non-volatile media. In these cases, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 1020 can include a computer program product 1025 having one or more program modules configured to carry out the various methods or actions of the implementations of the present disclosure.

[0117] The communication unit 1040 enables communications with other computing devices over a communication media. Additionally, the functionality of the components of the computing device 1000 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating over a communication connection. Thus, the computing device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.

[0118] The input device 1050 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 1060 can be one or more output devices, such as a display, a speaker, a printer, etc. The computing device 1000 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through the communication unit 1040, as needed, a device that enables a user to interact with the computing device 1000, or any device (e.g., a network card, a modem, etc.) that enables the computing device 1000 to communicate with one or more other computing devices. Such communication can be enabled by an input / output (I / O) interface (not shown).

[0119] According to example implementations of the present disclosure, a computer-readable storage medium is provided having computer-executable instructions stored thereon, where the computer-executable instructions are executed by a processor to implement the method described above. According to example implementations of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, where the computer-executable instructions are executed by a processor to implement the method described above. According to example implementations of the present disclosure, a computer program product is provided having a computer program stored thereon, which when executed by a processor implements the method described above.

[0120] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0121] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0122] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0123] The computer program product of the present disclosure can have a signal including said computer program. This signal can be electronic, electromagnetic, optical, or any other suitable type of signal. Such a signal can be provided through a communication connection, such as electrical wiring, optical fiber, wireless interface, etc. Examples of computer program products include computer program implemented on a personal computer, server, or other networked device. A non-transitory computer readable medium, such as a floppy disk, CD-ROM, DVD-ROM, Blu-ray Disc, hard disk, or memory stick, can also be used to implement the present disclosure. The computer program product of the present disclosure can also be provided as a service to download and use the computer program over a network, such as the Internet.

[0124] Having described several implementations of the present disclosure, it will be clear to those skilled in the art that many modifications, additions, and substitutions are possible without departing from the scope and spirit of the described implementations. Many modifications and variations of the present disclosure are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims, the present disclosure can be practiced otherwise than as specifically described. While the present disclosure has been described with reference to the implementation figures, it will be understood by those skilled in the art that various changes can be made and equivalents can be substituted for elements thereof without departing from the scope of the present disclosure. In addition, many modifications can be made to adapt a particular situation or material to the teachings of the disclosure without departing from its scope. Therefore, it is contemplated to cover any and all adaptations and modifications falling within the scope of the appended claims. It should also be noted that, in this document, the term "can" is used to mean various possible constructions, for example, a structure, device, or apparatus can be configured to perform one or more functions.

Claims

1. A method for processing speech data, comprising: In response to receiving the voice data of the target object, obtaining subtitle text associated with the voice data; as well as The subtitle text is provided in the audio live broadcast page.

2. The method of claim 1 , wherein providing the subtitle text comprises: The subtitle text is presented at a location associated with the identification of the target object.

3. The method of claim 2, wherein presenting the subtitle text comprises: A text box is provided in the audio live broadcast page, wherein the text box includes an identifier of the target object and the subtitle text.

4. The method of claim 3, wherein presenting the subtitle text comprises: The subtitle text is presented in the text box according to the time sequence of acquiring the subtitle text.

5. The method according to claim 3, further comprising: In response to detecting an adjustment action for adjusting the text box, the text box is updated based on the adjustment action, wherein the adjustment action is used to adjust at least any one of the following: the size of the text box, the position of the text box, the color, font, size, and background color of the subtitle text.

6. The method according to claim 1, wherein the subtitle text is obtained based on: in response to determining that the target object starts a subtitle function, determining at least one speech segment based on the speech data; and The subtitle text is extracted from the at least one speech segment.

7. The method according to claim 1, wherein obtaining the subtitle text comprises: The subtitle text is extracted from supplemental enhancement information of an audio stream associated with the audio data.

8. The method of claim 3, wherein the target objects include a first target object and a second target object, and presenting the subtitle text comprises: The first subtitle text and the second subtitle text are presented in the text box in a time sequence of acquiring the first subtitle text of the first voice data of the first target object and the second subtitle text of the second voice data of the second target object.

9. The method according to claim 1, further comprising: Presenting a switch control for switching the display mode of the subtitle text in the live audio broadcast page; as well as In response to receiving an interaction operation on the switch control, the subtitle text is presented based on the interaction operation.

10. The method according to claim 1, wherein the target object includes at least any one of the following: the host and the guest of the live broadcast room application; and providing the subtitle text includes providing the subtitle text to at least any one of the following: the host, the guest, and the audience.

11. The method according to claim 10, wherein the voice data is voice data of the target object in an audio live broadcast application, and providing the subtitle text in the audio live broadcast page comprises: The subtitle text is provided in the audio live broadcast page in the audio live broadcast application.

12. The method according to claim 11, wherein obtaining the subtitle text comprises: The first service of the live broadcast room application receives the audio data from the host and the guest; as well as The second service applied by the live broadcast room extracts the subtitle text from the audio data.

13. The method of claim 12, wherein providing the subtitle text comprises: Receiving the subtitle text at the client of the anchor; as well as The client renders the subtitle text to the audio live broadcast page.

14. The method of claim 13, wherein providing the subtitle text comprises: A third service of the live broadcast room application receives a video stream, wherein each video frame in the video stream includes the audio live broadcast page and the subtitle text rendered to the audio live broadcast page; as well as The video stream is presented at the clients of the audience and the guests.

15. The method of claim 12, wherein providing the subtitle text comprises: Receiving the subtitle text by the first service; The first service renders the subtitle text to the live audio broadcast page to generate a rendering page; as well as The rendering page is presented at the clients of the audience and the guest.

16. The method according to claim 12, wherein the subtitle text is subtitle text processed by a vocabulary service applied by the live broadcast room.

17. A device for processing voice data, comprising: an acquisition module configured to acquire, in response to receiving voice data of a target object, subtitle text associated with the voice data; as well as A providing module is configured to provide the subtitle text in the audio live broadcast page.

18. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 16 when executed by the at least one processing unit.

19. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Voice microphone connection interaction method and system of live broadcast room, medium and computer equipment

    CN114007095A

  • Voice subtitle generation method, system and device, storage medium and electronic device

    CN114242058A

  • Data processing method and server

    CN115002502A

  • Live broadcast data processing method and system

    CN115643424A

  • Subtitle processing method and device of live broadcast stream, storage medium and computer equipment

    CN117376593A