Video processing method and apparatus, and electronic device

WO2024131576A9PCT designated stage expired Publication Date: 2026-09-03BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/137590
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-20
Filing Date
2023-12-08
Publication Date
2026-09-03

Smart Images

  • Figure CN2023137590_03092026_PF_FP_ABST
    Figure CN2023137590_03092026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a video processing method and apparatus, and an electronic device. The method comprises: displaying a video interface (S201), the video interface being used for playing back a live video; acquiring first interaction information input by a first user on the video interface (S202), and sending the first interaction information to a server (S203); receiving a video stream sent by the server (S204), and playing back the live video on the video interface according to the video stream (S205), the live video comprising a livestreamer video, a first voice corresponding to the first interaction information, and a second voice corresponding to second interaction information of at least one second user. The complexity of livestreaming interaction is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Video processing methods, devices and electronic equipment

[0001] This application claims priority to Chinese Patent Application No. 202211644421.9, filed on December 20, 2022, entitled “Video Processing Method, Apparatus and Electronic Device”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to the field of video processing technology, and in particular to a video processing method, apparatus, and electronic device. Background Technology

[0003] When a streamer is broadcasting live, viewers can interact with the streamer, thereby enhancing the live broadcast's impact.

[0004] Currently, broadcasters can interact based on text messages sent by users. For example, while watching a broadcaster's live video, users can send text messages in real time, which are displayed on the live video screen. The broadcaster can then see the text and interact with the user. However, in some scenarios (such as outdoor broadcasts or concert broadcasts), broadcasters cannot hold the broadcasting equipment or watch the live video, and viewers can only hear the broadcaster's voice. They cannot obtain interactive information from multiple users within the live video. Broadcasters then need to obtain interactive information from multiple users through other means (such as user reminders during the live broadcast), leading to a higher complexity in live interaction.

[0005] Summary of the Invention

[0006] This disclosure provides a video processing method, apparatus, and electronic device to solve the technical problem of high complexity in live interactive broadcasting in the prior art.

[0007] In a first aspect, this disclosure provides a video processing method, the method comprising:

[0008] The video interface is used to play live videos.

[0009] Obtain the first interactive information input by the first user on the video interface, and send the first interactive information to the server;

[0010] The system receives a video stream sent by the server and plays a live video on the video interface according to the video stream. The live video includes a broadcaster's video, a first voice message corresponding to the first interactive information, and a second voice message corresponding to the second interactive information of at least one second user.

[0011] Secondly, this disclosure provides another video processing method, which includes:

[0012] Receive multiple target interactive messages sent by multiple electronic devices, the target interactive messages including text information and / or voice information;

[0013] Receive broadcast videos sent by the broadcaster's device;

[0014] Based on the broadcaster's video and the multiple target interaction information, a video stream is determined, the video stream including the encoded stream of the broadcaster's video and the encoded stream of the speech associated with the target interaction information;

[0015] The video stream is sent to the broadcasting device and the plurality of electronic devices.

[0016] Thirdly, this disclosure provides a video processing apparatus, which includes a display module, an acquisition module, a transmission module, a receiving module, and a playback module, wherein:

[0017] The display module is used to display a video interface, which is used to play live video.

[0018] The acquisition module is used to acquire first interactive information input by the first user on the video interface;

[0019] The sending module is used to send the first interactive information to the server;

[0020] The receiving module is used to receive the video stream sent by the server;

[0021] The playback module is used to play a live video on the video interface according to the video stream. The live video includes the anchor's video, the first voice corresponding to the first interactive information, and the second voice corresponding to the second interactive information of at least one second user.

[0022] Fourthly, this disclosure provides another video processing apparatus, which includes a receiving module, a determining module, and a transmitting module, wherein:

[0023] The receiving module is used to receive multiple target interactive information sent by multiple electronic devices, the target interactive information including text information and / or voice information;

[0024] The receiving module is also used to receive broadcast videos sent by the broadcasting device;

[0025] The determining module is used to determine a video stream based on the anchor video and the multiple target interaction information, wherein the video stream includes an encoded stream of the anchor video and an encoded stream of the speech associated with the target interaction information;

[0026] The sending module is used to send the video stream to the broadcasting device and the plurality of electronic devices.

[0027] Fifthly, this disclosure provides an electronic device, including: a processor and a memory;

[0028] The memory stores computer-executed instructions;

[0029] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the video processing methods described in the first aspect above and various possible aspects of the first aspect.

[0030] Sixthly, this disclosure provides a server, including: a processor and a memory;

[0031] The memory stores computer-executed instructions;

[0032] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the video processing methods described in the second aspect above and various possible aspects of the second aspect.

[0033] In a seventh aspect, this disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the video processing methods described in the first aspect and various possible embodiments thereof, or implement the video processing methods described in the second aspect and various possible embodiments thereof.

[0034] Eighthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the video processing methods described in the first aspect and various possible aspects thereof, or implements the video processing methods described in the second aspect and various possible aspects thereof.

[0035] This disclosure provides a video processing method, apparatus, and electronic device. The electronic device can acquire a video interface for playing live video, acquire first interactive information input by a first user on the video interface, send the first interactive information to a server, receive a video stream sent by the server, and play the live video on the video interface according to the video stream. The live video includes a broadcaster's video, first audio corresponding to the first interactive information, and second audio corresponding to second interactive information from at least one second user. In this method, when a broadcaster is conducting a live stream, since the live video can include audio corresponding to the interactive information of the watching users, the broadcaster can hear the users' interactive information without watching the live video. Users can also obtain audio interactive information between multiple users and the broadcaster through the live video, thereby reducing the complexity of live interaction. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of this disclosure;

[0038] Figure 2 is a schematic flowchart of a video processing method provided in an embodiment of this disclosure;

[0039] Figure 3 is a schematic diagram of a video display interface provided in an embodiment of this disclosure;

[0040] Figure 4A is a schematic diagram of a process for obtaining first interactive information provided in an embodiment of this disclosure;

[0041] Figure 4B is a schematic diagram of another process for obtaining first interactive information provided in an embodiment of this disclosure;

[0042] Figure 5 is a schematic diagram of another process for obtaining first interactive information provided in an embodiment of this disclosure;

[0043] Figure 6 is a schematic diagram of a scenario for sending first interactive information according to an embodiment of this disclosure;

[0044] Figure 7 is a schematic diagram of another video processing method provided in an embodiment of this disclosure;

[0045] Figure 8 is a schematic diagram of a process for determining target speech provided in an embodiment of this disclosure;

[0046] Figure 9 is a schematic diagram of the structure of a video processing device provided in an embodiment of this disclosure;

[0047] Figure 10 is a schematic diagram of another video processing apparatus provided in an embodiment of this disclosure; and

[0048] Figure 11 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0049] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0050] For ease of understanding, the concepts involved in the embodiments of this disclosure will be explained below.

[0051] Electronic device: A device with wireless transceiver capabilities. Electronic devices can be deployed on land, including indoors or outdoors, handheld, wearable, or vehicle-mounted; they can also be deployed on water (such as on ships). These electronic devices can be mobile phones, tablets, computers with wireless transceiver capabilities, virtual reality (VR) electronic devices, augmented reality (AR) electronic devices, wireless terminals in industrial control, vehicle-mounted electronic devices, wireless terminals in self-driving vehicles, wireless electronic devices in remote medical care, wireless electronic devices in smart grids, wireless electronic devices in transportation safety, wireless electronic devices in smart cities, wireless electronic devices in smart homes, wearable electronic devices, etc. The electronic devices involved in the embodiments of this disclosure can also be referred to as terminals, user equipment (UE), access electronic devices, vehicle-mounted terminals, industrial control terminals, UE units, UE stations, mobile stations, mobile stations, remote stations, remote electronic devices, mobile devices, UE electronic devices, wireless communication devices, UE agents, or UE devices, etc. Electronic devices can be fixed or mobile.

[0052] In related technologies, viewers of live streams can interact with the streamer, thereby improving the streamer's viewing experience. Currently, streamers can interact based on text messages sent by users. For example, while watching a streamer's live video, a user can send the text "Hello" in real time, which can be displayed on the live video screen. After seeing the text "Hello" in the live video, the streamer can reply to the text via voice, thus interacting with the viewers. However, in some scenarios, the streamer cannot hold the live streaming device or watch the live video, and users can only hear the streamer's voice. They cannot obtain interactive information from multiple users within the live video. For example, in a concert live stream, viewers cannot experience the atmosphere of the live concert, and the streamer needs other viewers in the audience to provide interactive information within the live video to interact with the viewers, leading to a higher complexity in live stream interaction.

[0053] To address the aforementioned technical problems, this disclosure provides a video processing method. An electronic device displays a video interface for playing live video, acquires first interactive information input by a first user on the video interface, sends the first interactive information to a server, receives a video stream sent by the server, and acquires a broadcaster's video stream and an audio stream from the video stream. The audio stream is determined based on first voice corresponding to the first interactive information and second voice corresponding to the second interactive information. Based on the broadcaster's video stream and the audio stream, the live video is obtained. Thus, since the live video includes the broadcaster's video, the first voice corresponding to the first interactive information, and the second voice corresponding to the second interactive information of at least one second user, the broadcaster can hear the user's interactive information without watching the live video, and users can also obtain multiple users' voice interaction information with the broadcaster through the live video, thereby reducing the complexity of live interaction.

[0054] The application scenarios of the embodiments of this disclosure will now be described with reference to Figure 1.

[0055] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of this disclosure. Referring to Figure 1, it includes electronic device 1, ..., electronic device N, a broadcasting device, and a server. Electronic devices 1, ..., and N can send interactive information to the server, and the broadcasting device can send broadcast videos to the server. The electronic devices generate live video based on the voice and broadcast video corresponding to the interactive information, and send the live video to electronic devices 1, ..., N, and the broadcasting device.

[0056] Please refer to Figure 1. In the network structure described above, viewers can directly send text and cheers to the broadcaster. Viewers can hear the broadcaster's singing and background noise, which may include cheers from multiple viewers. When multiple viewers sing through multiple electronic devices, the broadcaster can hear the viewers singing along, and viewers can also hear the broadcaster singing and singing along. Furthermore, viewers and broadcaster can directly communicate via voice, thus enabling direct interaction and reducing the complexity of live streaming interaction.

[0057] It should be noted that Figure 1 is merely an exemplary illustration of the application scenarios of the embodiments of this disclosure, and is not intended to limit the application scenarios of the embodiments of this disclosure.

[0058] The technical solutions of this disclosure and how they solve the aforementioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this disclosure will now be described with reference to the accompanying drawings.

[0059] Figure 2 is a flowchart illustrating a video processing method provided in an embodiment of this disclosure. Referring to Figure 2, the method may include:

[0060] S201, Display video interface.

[0061] The execution subject of this disclosure can be an electronic device or a video processing device installed in an electronic device. Optionally, the video processing device can be implemented by software, or by a combination of software and hardware.

[0062] Optionally, the video interface is used to play live video. For example, the video interface can display a live video including the streamer's video. Optionally, the electronic device can display the video interface in response to a user's triggering action on the live streaming application. For example, when a user clicks on the live streaming application, the electronic device can display a live streaming page, which may include links to multiple streamers' live streams. When the user clicks on any of these live stream links, the electronic device can display the video interface, which can play the live video including that streamer's video.

[0063] The process of displaying the video interface will be explained below with reference to Figure 3.

[0064] Figure 3 is a schematic diagram of a video display interface provided in an embodiment of this disclosure. Referring to Figure 3, the device includes an electronic device. The display page of the electronic device includes controls for a live streaming application. When a user clicks on a control for the live streaming application, the electronic device can display a page corresponding to the live streaming application, which includes live video controls for broadcaster 1, broadcaster 2, and broadcaster 3. When a user clicks on the live video control for broadcaster 2, the electronic device can display a video interface and play the live video corresponding to broadcaster 2 in the video interface.

[0065] S202. Obtain the first interactive information entered by the first user on the video interface.

[0066] Optionally, the first interactive information may include text information and voice information. For example, if a user enters the text "Hello" in the video interface, the electronic device can identify that text as the first interactive information; if a user enters the voice message "Go for it!" in the video interface, the electronic device can identify that voice message as the first interactive information.

[0067] Optionally, the electronic device can acquire the first interactive information input by the first user on the video interface based on the following two feasible implementation methods:

[0068] One feasible implementation method:

[0069] In response to a trigger operation on a text editing control in a video interface, a text window for inputting text is displayed. For example, when an electronic device plays a live video on a video interface, the video interface may include a text editing control, and when a user clicks the text editing control, the video interface may display a text window for inputting text. For example, the text editing control may be a text input bar, and when the user clicks the text input bar, the user can input relevant text into the text input bar via the electronic device.

[0070] Acquire the text content edited in the text window, and determine the text content as first interaction information. For example, when a user clicks a text editing control in a video interface, the video interface may pop up a text window. If the user inputs the text "Go for it" in the text window and clicks a confirmation control, the electronic device may determine the text "Go for it" as the first interaction information. For example, after the user inputs the text "Go for it" in the text input bar and clicks a send control, the electronic device may determine the text "Go for it" as the first interaction information.

[0071] In the following, the process of acquiring first interaction information in this implementation will be described with reference to FIG. 4A to FIG. 4B.

[0072] FIG. 4A is a schematic diagram of a process for acquiring first interaction information provided by an embodiment of the present disclosure. Referring to FIG. 4A, the process involves an electronic device. A display page of the electronic device is a video interface, and the video played in the video interface is a live video. The video interface includes a text editing control. When a user clicks the text editing control, the video interface pops up a text window. The user inputs the text "Go for it" into the text input bar of the text window and clicks the send control, and the electronic device determines that the acquired first interaction information is the text "Go for it".

[0073] FIG. 4B is a schematic diagram of another process for acquiring first interaction information provided by an embodiment of the present disclosure. Referring to FIG. 4B, the process involves an electronic device. A display page of the electronic device is a video interface, and the video played in the video interface is a live video. When the user clicks any position on the video interface, a text input bar may pop up at the bottom of the video interface. The user inputs the text "Go for it" into the text input bar and clicks the send control, and the electronic device determines that the acquired first interaction information is the text "Go for it".

[0074] Another feasible implementation mode:

[0075] In response to a trigger operation on a voice acquisition control in a video interface, the voice of a first user is acquired, and the voice of the first user is determined as first interaction information. For example, when an electronic device plays a live video on a video interface, the video interface may include a voice acquisition control. When the user clicks the voice acquisition control, the electronic device can collect the voice of the first user in real time and determine the voice as the first interaction information. For example, if the voice of the first user collected by the electronic device is the voice "come on", the electronic device can determine the voice "come on" as the first interaction information.

[0076] In the following, with reference to FIG. 5, the process of acquiring the first interaction information in this feasible implementation manner will be described.

[0077] FIG. 5 is a schematic diagram of another process of acquiring first interaction information provided by an embodiment of the present disclosure. Referring to FIG. 5, an electronic device is included. Wherein, the display page of the electronic device is a video interface, and the video played in the video interface is a live video. A voice acquisition control is included in the video interface. When the user long-presses the voice acquisition control, the electronic device can collect surrounding voice. If the voice input by the user to the electronic device is "come on", when the user stops touching the voice acquisition control, the electronic device determines that the acquired first interaction information is the voice "come on".

[0078] S203: Send first interaction information to a server.

[0079] Optionally, after the electronic device acquires the first interaction information input by the first user, the electronic device can send the first interaction information to the server. For example, the first interaction information acquired by the electronic device can be text information and voice information, and the electronic device can send the acquired text information and voice information to the server. It should be noted that when the first interaction information is text information, the electronic device can send the text information to the server, or convert the text information into voice information and then send the voice information to the server, which is not limited in the embodiments of the present disclosure.

[0080] In the following, with reference to FIG. 6, the process of sending interaction information to the server will be described.

[0081] FIG. 6 is a schematic diagram of a scenario for sending first interaction information provided by an embodiment of the present disclosure. Referring to FIG. 6, the diagram includes: a first viewer, a second viewer, a third viewer, and a server. When the three viewers watch the live video, the first viewer can send text A to the server via the electronic device used by the first viewer (not shown in FIG. 6), the second viewer can send text B to the server via the electronic device used by the second viewer, and the third viewer can send voice to the server via the electronic device used by the third viewer. In this way, the server can receive the interaction information sent by the three viewers, and then generate a live video stream based on the interaction information sent by the three viewers and the acquired anchor video.

[0082] It should be noted that, in the embodiment shown in FIG. 6, after the texts of the first viewer and the second viewer are converted into speech by the electronic device, the speech corresponding to the texts may be sent to the server; alternatively, the texts may be sent to the server, and the server converts the texts into speech. This is not limited in the embodiments of the present disclosure.

[0083] S204: Receive the video stream sent by the server.

[0084] Optionally, after the electronic device sends the first interaction information to the server, the electronic device may receive the video stream sent by the server.

[0085] S205: Play the live video on the video interface according to the video stream.

[0086] Optionally, the live video includes a anchor video, a first speech corresponding to the first interaction information, and a second speech corresponding to second interaction information of at least one second user. Optionally, the second user may be another user other than the first user. For example, in actual application, the server may receive interaction information sent by multiple users, and therefore the live video may include speeches corresponding to interaction information of multiple users.

[0087] Optionally, the second interaction information may be interaction information input by the second user on the video interface. For example, the second interaction information may include text information and speech information. The method for the second user to input the second interaction information on the video interface is the same as the method for the first user to input the first interaction information on the video interface, and details are not described herein again in the embodiments of the present disclosure.

[0088] Optionally, the first speech may be a speech corresponding to the first interaction information. For example, when the first interaction information is text information, the first speech may be a speech corresponding to the text information; when the first interaction information is speech information, the electronic device may determine the speech information as the first speech. For example, if the first interaction information is the text "Go for it", the first speech may be the speech "Go for it"; if the first interaction information is the speech "Go for it", the first speech may be the speech "Go for it".

[0089] Optionally, the second speech may be a speech corresponding to the second interaction information. For example, when the second interaction information is text information, the second speech may be a speech corresponding to the text information; when the second interaction information is speech information, the electronic device may determine the speech information as the first speech. For example, if the second interaction information is the text "Hello", the second speech may be the speech "Hello"; if the second interaction information is the speech "Hello", the second speech may be the speech "Hello".

[0090] Optionally, the video stream may include a broadcaster video stream and an audio stream. Electronic devices can play live video on the video interface using the following feasible implementation methods: Obtaining the broadcaster video stream and audio stream from the video stream. The broadcaster video stream can be obtained based on the broadcaster video. For example, the server can encode the broadcaster video to obtain the corresponding broadcaster video stream. The audio stream can be determined based on a first speech and a second speech. For example, when the server obtains the first speech and the second speech, it can generate synthesized speech based on the first speech and the second speech, and encode the synthesized speech to obtain the audio stream.

[0091] The live video is obtained by analyzing the broadcaster's video and audio streams. For example, after receiving a video stream from a server, an electronic device can decode the broadcaster's video and audio streams to obtain the live video, which may include the broadcaster's video, first audio, and second audio.

[0092] Optionally, the live video may also include images of a first user and a second user. For example, when the first user inputs first interactive information into the video interface, the electronic device can acquire the first user's image. When the electronic device sends the first interactive information to the server, it can also send the first user's image to the server. The server can then add the first user's image to the live video, thus including the first user's image in the live video. Similarly, when a second user inputs second interactive information into the video interface, the electronic device can acquire the second user's image. When the electronic device sends the second interactive information to the server, it can also send the second user's image to the server. The server can then add the second user's image to the live video, thus including the second user's image in the live video. For example, when the live video is a concert live stream, the audience area of ​​the live video received by the electronic device can include images of both the first and second users, thereby improving the live video's quality.

[0093] Embodiments of the present disclosure provide a video processing method. An electronic device displays a video interface for playing a live video, displays a text window for inputting text in response to a trigger operation on a text editing control in the video interface, acquires text content edited in the text window, and determines the text content as first interaction information; or, in response to a trigger operation on a voice acquisition control in the video interface, acquires a voice of a first user, and determines the voice of the first user as first interaction information. The electronic device can send the first interaction information to a server, receive a video stream sent by the server, and play the live video on the video interface according to the video stream. In this way, since the live video includes a host's video, a first voice corresponding to the first interaction information, and a second voice corresponding to second interaction information of at least one second user, the host can hear the users' interaction information without watching the live video, and users can also obtain voice interaction information between multiple users and the host through the live video, thereby reducing the complexity of live interaction.

[0094] Based on the embodiment shown in FIG. 2, another video processing method will be described below with reference to FIG. 7.

[0095] FIG. 7 is a schematic diagram of another video processing method provided by an embodiment of the present disclosure. Referring to FIG. 7, the flow of the method includes:

[0096] S701: Receive a plurality of target interaction information sent by an electronic device.

[0097] The execution subject of the embodiments of the present disclosure may be a server, or may be a video processing apparatus disposed in a server. Optionally, the video processing apparatus may be implemented by software, or may be implemented by a combination of software and hardware, which is not limited in the embodiments of the present disclosure.

[0098] Optionally, the target interaction information may include text information and / or voice information. For example, the target interaction information may be text information input by a user watching a live broadcast, may be voice information input by a user watching a live broadcast, or may be both text information and voice information input by a user watching a live broadcast. For example, when a user watches a live video, if the user inputs the text "Come on" on the video interface, the target interaction information received by the server is the text "Come on"; if the user inputs the voice "Come on" on the video interface, the target interaction information received by the server is the voice "Come on".

[0099] Optionally, the electronic device may be an electronic device of a first user or a second user. For example, the electronic device may be an electronic device of a user watching a live video. When the user watches the live video, the user can input text information and voice information into the electronic device, and the server can receive the text information and voice information sent by the electronic device.

[0100] S702, receiving anchor video sent by the anchor device.

[0101] Optionally, the anchor video may be a video including the anchor, and the anchor device may be a device that captures the anchor video. For example, in a live streaming scenario, the anchor video in the live streaming video is provided by the anchor device. After the anchor device captures the video including the anchor, it may send the live streaming video including the anchor video to the server, and the server may send the live streaming video to electronic devices of users watching the live stream, so that users can watch the live streaming video.

[0102] S703, determining a video stream based on the anchor video and a plurality of target interaction information.

[0103] Optionally, the video stream includes an encoded stream of the anchor video and an encoded stream of speech associated with the target interaction information. For example, if the target interaction information is text information, the server may convert the text information into speech information. For example, if the target interaction information is the text "Go for it", the server may convert the text "Go for it" into the speech "Go for it".

[0104] It should be noted that, when the server receives text information sent by an electronic device, and the server converts the text information into speech information, if the server can determine the user's timbre through the speech sent by the electronic device, then when the server converts the text information into speech information, the timbre associated with the speech information may be the user's timbre.

[0105] Optionally, the server may generate the video stream according to the following feasible implementation: generating an audio stream based on the plurality of target interaction information, and generating the video stream based on the anchor video and the audio stream. For example, the server encodes the anchor video to obtain an anchor video stream, merges the anchor video stream and the audio stream, and then obtains the video stream. Optionally, the audio stream is obtained based on speech corresponding to the plurality of target interaction information. For example, the audio stream may be an encoded stream associated with speech corresponding to the plurality of target interaction information.

[0106] Optionally, the electronic device may generate the audio stream according to the following feasible implementation: acquiring a plurality of third speeches associated with the plurality of target interaction information, determining a target speech from the third speeches, and generating the audio stream based on the target speech. Optionally, the third speech may be a speech corresponding to the target interaction information. For example, if the target interaction information is the text "Go for it", the third speech may be the speech "Go for it"; if the target interaction information is the speech "Hello", the third speech may be the speech "Hello".

[0107] Optionally, the number of target speech items is greater than or equal to a first threshold. For example, in practical applications, each speech item includes corresponding semantic information. Therefore, the server can classify the third speech item into multiple types based on the semantic information, and then determine the target speech item based on the number of each type of speech item.

[0108] Optionally, the server determines the target speech from the third-party speech. Specifically, this involves: clustering the semantics of multiple third-party speech samples to obtain at least one class of third-party speech samples; obtaining the number of speech samples in each class; and identifying the class of third-party speech samples whose number of speech samples is greater than or equal to a first threshold as the target speech. For example, if the server obtains 100 third-party speech samples and the first threshold is 20, and if the server clusters the semantics of these 100 speech samples to obtain 30 speech samples corresponding to semantic A, 60 speech samples corresponding to semantic B, and 10 speech samples corresponding to semantic C, then the electronic device can identify the 30 speech samples corresponding to semantic A and the 60 speech samples corresponding to semantic B as the target speech.

[0109] The process of determining the target speech will be explained below with reference to Figure 8.

[0110] Figure 8 is a schematic diagram illustrating a process for determining target speech according to an embodiment of this disclosure. Referring to Figure 8, it includes 100 third-party speech samples acquired by the server. The server (not shown in Figure 8) can perform semantic clustering on the 100 third-party speech samples based on the semantics of each sample, resulting in two categories of third-party speech. One category of third-party speech has the semantic meaning of "cheer up," and the other category has the semantic meaning of "pleasant to listen to." The number of third-party speech samples with the semantic meaning of "cheer up" is 30, and the number of third-party speech samples with the semantic meaning of "pleasant to listen to" is 70. If the first threshold is 50, the server determines that the target speech has the semantic meaning of "pleasant to listen to," that is, it identifies the 70 third-party speech samples with the semantic meaning of "pleasant to listen to" as the target speech.

[0111] Optionally, the server generates an audio stream based on the target speech, specifically by determining the speech volume corresponding to each type of target speech based on the number of speech items in each type. Optionally, the server can obtain a first preset relationship and determine the speech volume corresponding to the target speech based on the first preset relationship and the number of speech items in the target speech. The first preset relationship may include at least one number of speech items and the speech volume corresponding to each number of speech items. For example, the first preset relationship may be as shown in Table 1:

[0112] Table 1

[0113] It should be noted that Table 1 is only an example to illustrate the first preset relationship, and is not a limitation on the first preset relationship.

[0114] For example, if the server determines that the number of target voices is 1, then the server determines the corresponding voice volume as volume a; if the server determines that the number of target voices is 2, then the server determines the corresponding voice volume as volume b; if the server determines that the number of target voices is 3, then the server determines the corresponding voice volume as volume c.

[0115] An audio stream is generated based on the target speech and its corresponding volume. For example, the server can generate synthesized speech based on the target speech, where the semantics of the synthesized speech are the same as those of the target speech. The synthesized speech can have the effect of multiple users speaking the target speech. The server sets the volume of the synthesized speech to the volume of the target speech and encodes the synthesized speech to obtain the audio stream.

[0116] Optionally, the server can also receive user images sent by electronic devices, and can add at least one user's image to the live stream video. For example, in a concert live stream scenario, the server can add at least one user's image to the audience area of ​​the acquired live stream video, thereby improving the quality of the live stream video.

[0117] S704, Send video streams to broadcast devices and electronic devices.

[0118] Optionally, after the server receives the video stream, it can send the video stream to the broadcaster's device and the electronic devices of multiple users watching the live stream. In this way, the broadcaster can obtain interactive information from users through the audio of the live video and respond to the interactive information. Users watching the live stream can also experience the effect of the live broadcast through the live video (such as the effect of multiple audience members singing together, the effect of asking questions to the broadcaster on the spot, etc.), thereby improving the interactive effect of the live broadcast.

[0119] This disclosure provides a video processing method that receives multiple target interaction messages sent by multiple electronic devices, receives a broadcast video sent by a broadcast device, determines a video stream based on the broadcast video and the multiple target interaction messages, and sends the video stream to the broadcast device and the multiple electronic devices. In this way, the broadcaster can hear the user's interaction messages without watching the live video, and the more interaction messages there are, the better the chorus effect, improving the live broadcast effect. Users can also obtain multiple voice interaction messages between users and the broadcaster through the live video, thereby reducing the complexity of live broadcast interaction.

[0120] Figure 9 is a schematic diagram of a video processing device provided in an embodiment of this disclosure. Referring to Figure 9, the video processing device 10 includes a display module 11, an acquisition module 12, a sending module 13, a receiving module 14, and a playback module 15, wherein:

[0121] The display module 11 is used to display a video interface, which is used to play live video.

[0122] The acquisition module 12 is used to acquire first interactive information input by the first user on the video interface;

[0123] The sending module 13 is used to send the first interactive information to the server;

[0124] The receiving module 14 is used to receive the video stream sent by the server;

[0125] The playback module 15 is used to play a live video on the video interface according to the video stream. The live video includes the anchor's video, the first voice corresponding to the first interactive information, and the second voice corresponding to the second interactive information of at least one second user.

[0126] According to one or more embodiments of this disclosure, the acquisition module 12 is specifically used for:

[0127] In response to a trigger operation on the text editing control in the video interface, a text window for inputting text is displayed;

[0128] The text content edited within the text window is obtained and identified as the first interactive information.

[0129] According to one or more embodiments of this disclosure, the acquisition module 12 is specifically used to: acquire the voice of the first user in response to a trigger operation on the voice acquisition control in the video interface;

[0130] The first user's voice is identified as the first interactive information.

[0131] According to one or more embodiments of this disclosure, the live video also includes images of a first user and a second user.

[0132] The video processing apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.

[0133] Figure 10 is a schematic diagram of another video processing apparatus provided in an embodiment of this disclosure. Referring to Figure 10, the video processing apparatus 20 includes a receiving module 21, a determining module 22, and a sending module 23, wherein:

[0134] The receiving module 21 is used to receive multiple target interactive information sent by multiple electronic devices, the target interactive information including text information and / or voice information;

[0135] The receiving module 21 is also used to receive broadcast videos sent by the broadcasting device;

[0136] The determining module 22 is used to determine a video stream based on the anchor video and the multiple target interaction information, wherein the video stream includes an encoded stream of the anchor video and an encoded stream of the speech associated with the target interaction information;

[0137] The sending module 23 is used to send the video stream to the broadcasting device and the plurality of electronic devices.

[0138] According to one or more embodiments of this disclosure, the determining module 22 is specifically used for:

[0139] An audio stream is generated based on the multiple target interaction information, wherein the audio stream is obtained based on the speech corresponding to the multiple target interaction information.

[0140] The video stream is generated based on the broadcaster's video and the audio stream.

[0141] According to one or more embodiments of this disclosure, the determining module 22 is specifically used for:

[0142] Acquire multiple third-party voices associated with multiple target interaction information;

[0143] The target speech is determined from the third speech, and the number of the target speech is greater than or equal to a first threshold.

[0144] The audio stream is generated based on the target speech.

[0145] According to one or more embodiments of this disclosure, the determining module 22 is specifically used for:

[0146] The semantics of the multiple third speech words are clustered to obtain at least one class of third speech words;

[0147] The number of third speech items in each category is obtained, and the category of third speech items whose number of speech items is greater than or equal to a first threshold is determined as the target speech item.

[0148] According to one or more embodiments of this disclosure, the determining module 22 is specifically used for:

[0149] The volume of each type of target speech is determined based on the number of speech items in each category.

[0150] The audio stream is generated based on the target speech and the corresponding speech volume.

[0151] According to one or more embodiments of this disclosure, at least one user's image is added to the broadcast video.

[0152] The video processing apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.

[0153] Figure 11 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Referring to Figure 11, it shows a schematic diagram of the structure of an electronic device 1100 suitable for implementing an embodiment of this disclosure. The electronic device 1100 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. The electronic device shown in Figure 11 is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this disclosure.

[0154] As shown in Figure 11, the electronic device 1100 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 1101, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1102 or a program loaded from storage device 1108 into random access memory (RAM) 1103. The RAM 1103 also stores various programs and data required for the operation of the electronic device 1100. The processing unit 1101, ROM 1102, and RAM 1103 are interconnected via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0155] Typically, the following devices can be connected to I / O interface 1105: input devices 1106 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 1107 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1108 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1109. Communication device 1109 allows electronic device 1100 to communicate wirelessly or wiredly with other devices to exchange data. Although FIG11 shows electronic device 1100 with various devices, it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0156] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1109, or installed from storage device 1108, or installed from ROM 1102. When the computer program is executed by processing device 1101, it performs the functions defined in the methods of embodiments of this disclosure.

[0157] This disclosure also includes a server, which may include a processor and a memory, the memory storing computer execution instructions, and the processor executing the computer execution instructions stored in the memory, causing the processor to perform the video processing method described in any of the above embodiments.

[0158] This disclosure also includes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the video processing method as described in any of the above embodiments.

[0159] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0160] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0161] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.

[0162] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0163] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0164] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0165] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0166] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0167] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0168] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0169] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0170] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0171] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0172] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0173] It is understood that the data involved in this technical solution (including but not limited to the data itself, its acquisition, or its use) shall comply with the requirements of relevant laws, regulations, and provisions. Data may include information, parameters, and messages, such as flow control instructions.

[0174] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0175] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0176] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

[0177] In a first aspect, this disclosure provides one or more embodiments of a video processing method, the method comprising:

[0178] The video interface is used to play live videos.

[0179] Obtain the first interactive information input by the first user on the video interface, and send the first interactive information to the server;

[0180] The system receives a video stream sent by the server and plays a live video on the video interface according to the video stream. The live video includes a broadcaster's video, a first voice message corresponding to the first interactive information, and a second voice message corresponding to the second interactive information of at least one second user.

[0181] According to one or more embodiments of this disclosure, obtaining first interactive information input by a first user on a video interface includes:

[0182] In response to a trigger operation on the text editing control in the video interface, a text window for inputting text is displayed;

[0183] The text content edited within the text window is obtained and identified as the first interactive information.

[0184] According to one or more embodiments of this disclosure, obtaining first interactive information input by a first user on a video interface includes:

[0185] In response to a trigger operation on the voice acquisition control in the video interface, the voice of the first user is acquired;

[0186] The first user's voice is identified as the first interactive information.

[0187] According to one or more embodiments of this disclosure, the live video also includes images of a first user and a second user.

[0188] Secondly, this disclosure provides one or more embodiments of another video processing method, the method comprising:

[0189] Receive multiple target interactive messages sent by multiple electronic devices, the target interactive messages including text information and / or voice information;

[0190] Receive broadcast videos sent by the broadcaster's device;

[0191] Based on the broadcaster's video and the multiple target interaction information, a video stream is determined, the video stream including the encoded stream of the broadcaster's video and the encoded stream of the speech associated with the target interaction information;

[0192] The video stream is sent to the broadcasting device and the plurality of electronic devices.

[0193] According to one or more embodiments of this disclosure, determining a video stream based on the broadcaster's video and the plurality of target interaction information includes:

[0194] An audio stream is generated based on the multiple target interaction information, wherein the audio stream is obtained based on the speech corresponding to the multiple target interaction information.

[0195] The video stream is generated based on the broadcaster's video and the audio stream.

[0196] According to one or more embodiments of this disclosure, generating an audio stream based on the plurality of target interaction information includes:

[0197] Acquire multiple third-party voices associated with multiple target interaction information;

[0198] The target speech is determined from the third speech, and the number of the target speech is greater than or equal to a first threshold.

[0199] The audio stream is generated based on the target speech.

[0200] According to one or more embodiments of this disclosure, determining the target speech in the third speech includes:

[0201] The semantics of the multiple third speech words are clustered to obtain at least one class of third speech words;

[0202] The number of third speech items in each category is obtained, and the category of third speech items whose number of speech items is greater than or equal to a first threshold is determined as the target speech item.

[0203] According to one or more embodiments of this disclosure, generating the audio stream based on the target speech includes:

[0204] The volume of each type of target speech is determined based on the number of speech items in each category.

[0205] The audio stream is generated based on the target speech and the corresponding speech volume.

[0206] According to one or more embodiments of this disclosure, the method further includes:

[0207] Add an image of at least one user to the broadcaster's video.

[0208] Thirdly, this disclosure provides one or more embodiments of a video processing apparatus, which includes a display module, an acquisition module, a transmission module, a receiving module, and a playback module, wherein:

[0209] The display module is used to display a video interface, which is used to play live video.

[0210] The acquisition module is used to acquire first interactive information input by the first user on the video interface;

[0211] The sending module is used to send the first interactive information to the server;

[0212] The receiving module is used to receive the video stream sent by the server;

[0213] The playback module is used to play a live video on the video interface according to the video stream. The live video includes the anchor's video, the first voice corresponding to the first interactive information, and the second voice corresponding to the second interactive information of at least one second user.

[0214] According to one or more embodiments of this disclosure, the acquisition module is specifically used for:

[0215] In response to a trigger operation on the text editing control in the video interface, a text window for inputting text is displayed;

[0216] The text content edited within the text window is obtained and identified as the first interactive information.

[0217] According to one or more embodiments of this disclosure, the acquisition module is specifically used for:

[0218] In response to a trigger operation on the voice acquisition control in the video interface, the voice of the first user is acquired;

[0219] The first user's voice is identified as the first interactive information.

[0220] According to one or more embodiments of this disclosure, the live video also includes images of a first user and a second user.

[0221] Fourthly, according to one or more embodiments of this disclosure, another video processing apparatus is provided, comprising a receiving module, a determining module, and a transmitting module, wherein:

[0222] The receiving module is used to receive multiple target interactive information sent by multiple electronic devices, the target interactive information including text information and / or voice information;

[0223] The receiving module is also used to receive broadcast videos sent by the broadcasting device;

[0224] The determining module is used to determine a video stream based on the anchor video and the multiple target interaction information, wherein the video stream includes an encoded stream of the anchor video and an encoded stream of the speech associated with the target interaction information;

[0225] The sending module is used to send the video stream to the broadcasting device and the plurality of electronic devices.

[0226] According to one or more embodiments of this disclosure, the determining module is specifically used for:

[0227] An audio stream is generated based on the multiple target interaction information, wherein the audio stream is obtained based on the speech corresponding to the multiple target interaction information.

[0228] The video stream is generated based on the broadcaster's video and the audio stream.

[0229] According to one or more embodiments of this disclosure, the determining module is specifically used for:

[0230] Acquire multiple third-party voices associated with multiple target interaction information;

[0231] The target speech is determined from the third speech, and the number of the target speech is greater than or equal to a first threshold.

[0232] The audio stream is generated based on the target speech.

[0233] According to one or more embodiments of this disclosure, the determining module is specifically used for:

[0234] The semantics of the multiple third speech words are clustered to obtain at least one class of third speech words;

[0235] The number of third speech items in each category is obtained, and the category of third speech items whose number of speech items is greater than or equal to a first threshold is determined as the target speech item.

[0236] According to one or more embodiments of this disclosure, the determining module is specifically used for:

[0237] The volume of each type of target speech is determined based on the number of speech items in each category.

[0238] The audio stream is generated based on the target speech and the corresponding speech volume.

[0239] According to one or more embodiments of this disclosure, at least one user's image is added to the broadcast video.

[0240] Fifthly, this disclosure provides an electronic device, including: a processor and a memory;

[0241] The memory stores computer-executed instructions;

[0242] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the video processing methods described in the first aspect above and various possible aspects of the first aspect.

[0243] Sixthly, this disclosure provides a server, including: a processor and a memory;

[0244] The memory stores computer-executed instructions;

[0245] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the video processing methods described in the second aspect above and various possible aspects of the second aspect.

[0246] In a seventh aspect, this disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the video processing methods described in the first aspect and various possible embodiments thereof, or implement the video processing methods described in the second aspect and various possible embodiments thereof.

[0247] Eighthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the video processing methods described in the first aspect and various possible aspects thereof, or implements the video processing methods described in the second aspect and various possible aspects thereof.

[0248] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0249] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0250] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A video processing method, comprising: The video interface is used to play live videos. Obtain the first interactive information input by the first user on the video interface, and send the first interactive information to the server; The system receives a video stream sent by the server and plays a live video on the video interface according to the video stream. The live video includes a broadcaster's video, a first voice message corresponding to the first interactive information, and a second voice message corresponding to the second interactive information of at least one second user.

2. The method according to claim 1, wherein obtaining the first interactive information input by the first user on the video interface includes: In response to a trigger operation on the text editing control in the video interface, a text window for inputting text is displayed; The text content edited within the text window is obtained and identified as the first interactive information.

3. The method according to claim 1, wherein obtaining the first interactive information input by the first user on the video interface includes: In response to a trigger operation on the voice acquisition control in the video interface, the voice of the first user is acquired; The first user's voice is identified as the first interactive information.

4. The method according to any one of claims 1-3, wherein the live video further includes images of the first user and the second user.

5. A video processing method, comprising: Receive multiple target interactive messages sent by an electronic device, the target interactive messages including text messages and / or voice messages; Receive broadcast videos sent by the broadcaster's device; Based on the broadcaster's video and the multiple target interaction information, a video stream is determined, the video stream including the encoded stream of the broadcaster's video and the encoded stream of the speech associated with the target interaction information; The video stream is sent to the broadcasting device and the plurality of electronic devices.

6. The method according to claim 5, wherein determining the video stream based on the anchor video and the plurality of target interaction information includes: An audio stream is generated based on the multiple target interaction information, wherein the audio stream is obtained based on the speech corresponding to the multiple target interaction information. The video stream is generated based on the broadcaster's video and the audio stream.

7. The method of claim 6, wherein generating an audio stream based on the plurality of target interaction information includes: Acquire multiple third-party voices associated with multiple target interaction information; The target speech is determined from the third speech, and the number of the target speech is greater than or equal to a first threshold. The audio stream is generated based on the target speech.

8. The method of claim 7, wherein determining the target speech in the third speech comprises: The semantics of the multiple third speech words are clustered to obtain at least one class of third speech words; The number of third speech items in each category is obtained, and the category of third speech items whose number of speech items is greater than or equal to a first threshold is determined as the target speech item.

9. The method according to claim 7 or 8, wherein generating the audio stream based on the target speech comprises: The volume of each type of target speech is determined based on the number of speech items in each category. The audio stream is generated based on the target speech and the corresponding speech volume.

10. The method according to any one of claims 5-9, wherein the method further comprises: Add an image of at least one user to the broadcaster's video.

11. A video processing apparatus, comprising a display module, an acquisition module, a transmission module, a receiving module, and a playback module, wherein: The display module is used to display a video interface, which is used to play live video. The acquisition module is used to acquire first interactive information input by the first user on the video interface; The sending module is used to send the first interactive information to the server; The receiving module is used to receive the video stream sent by the server; The playback module is used to play a live video on the video interface according to the video stream. The live video includes the anchor's video, the first voice corresponding to the first interactive information, and the second voice corresponding to the second interactive information of at least one second user.

12. A video processing apparatus, comprising a receiving module, a determining module, and a transmitting module, wherein: The receiving module is used to receive multiple target interactive information sent by an electronic device, the target interactive information including text information and / or voice information; The receiving module is also used to receive broadcast videos sent by the broadcasting device; The determining module is used to determine a video stream based on the anchor video and the multiple target interaction information, wherein the video stream includes an encoded stream of the anchor video and an encoded stream of the speech associated with the target interaction information; The sending module is used to send the video stream to the broadcasting device and the plurality of electronic devices.

13. An electronic device, comprising: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the video processing method as described in any one of claims 1 to 4.

14. A server, comprising: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the video processing method as described in any one of claims 5 to 10.

15. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the video processing method as described in any one of claims 1 to 4, or implement the video processing method as described in any one of claims 5 to 10.