Live-streaming processing method, and device and storage medium

By synthesizing the current video frame of the first live streaming client and the latest video frames of other live streaming clients as target video frames in a multi-person live streaming scenario, and displaying them with a single rendering view, the problem of live streaming smoothness and performance in multi-person live streaming is solved, and the live streaming smoothness is improved.

WO2025002428A9PCT designated stage expired Publication Date: 2026-01-22BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/102664
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-06-28
Filing Date
2024-06-28
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

In multi-person live streaming scenarios, the smoothness of the live stream and the performance of the live streaming client are severely affected, and overheating may occur.

Method used

The system receives the live stream from the second live streaming client through the first live streaming client and caches the latest video frame. When the current video frame is obtained, the latest cached video frame is read. Based on a single rendered view, the current video frame and the latest video frame are combined into a target video frame for display according to the preset layout information.

Benefits of technology

It reduces the resource consumption of the live streaming client, reduces performance pressure, and improves the smoothness of live streaming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024102664_22012026_PF_FP_ABST
    Figure CN2024102664_22012026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a live-streaming processing method, and a device and a storage medium. The live-streaming processing method comprises: a first live-streaming video chat client receiving a live stream of a second live-streaming video chat client, and caching the latest video frame of the live stream of the second live-streaming video chat client; when the current video frame of the first live-streaming video chat client is acquired, reading the cached latest video frame; on the basis of a single rendering view, synthesizing the current video frame and the latest video frame into a target video frame according to preset layout information; and displaying the target video frame. In a multi-person video chat live-streaming scenario, the current video frame of a first live-streaming video chat client and the latest video frame of a live stream of another, a second live-streaming video chat client are synthesized into a target video frame, and rendering can be realized only by using a single rendering view, thereby reducing the occupancy of resources, reducing the performance pressure on a live-streaming video chat client, and improving the live-streaming fluency.
Need to check novelty before this filing date? Find Prior Art

Description

Live processing method, device and storage medium

[0001] The present application claims priority to the Chinese patent application No. 202310782876.5, filed on June 28, 2023, entitled "Live processing method, device and storage medium", the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The embodiments of the present disclosure relate to the technical field of computer and network communication, in particular to a live processing method, device and storage medium. BACKGROUND

[0003] In the process of network live, multi-connection is a way that anchors often use to interact with other anchors for live, based on which the anchor of the current live room can invite other anchors to live together.

[0004] However, with the increase of the number of anchors in the multi-connection live scene, the live fluency and the performance of the live co-connection client are seriously affected, and may be seriously overheated.

[0005] SUMMARY

[0006] The embodiments of the present disclosure provide a live processing method, device and storage medium to reduce the performance pressure on the live co-connection client and improve the live fluency in the multi-connection live scene.

[0007] In a first aspect, the embodiments of the present disclosure provide a live processing method applied to a first live co-connection client, the method comprising: receiving a live stream of a second live co-connection client, and caching a latest video frame of the live stream of the second live co-connection client; reading the cached latest video frame when a current video frame of the first live co-connection client is acquired; combining the current video frame and the latest video frame into a target video frame based on a single rendering view according to preset layout information; and displaying the target video frame.

[0008] In a second aspect, the embodiments of the present disclosure provide a live processing device applied to a first live co-connection client, the device comprising: a receiving unit configured to receive a live stream of a second live co-connection client; a caching unit configured to cache a latest video frame of the live stream of the second live co-connection client; read the cached latest video frame when a current video frame of the first live co-connection client is acquired; a rendering unit configured to combine the current video frame and the latest video frame into a target video frame based on a single rendering view according to preset layout information; and a display unit configured to display the target video frame.

[0009] In a third aspect, an electronic device is provided, and the electronic device includes at least one processor and a memory. The memory stores computer-executable instructions. The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the live processing method according to the first aspect and various possible designs of the first aspect.

[0010] In a fourth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the live processing method according to the first aspect and various possible designs of the first aspect is implemented.

[0011] In a fifth aspect, a computer program product is provided, and the computer program product includes computer-executable instructions. When a processor executes the computer-executable instructions, the live processing method according to the first aspect and various possible designs of the first aspect is implemented. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor.

[0013] FIG. 1 is a scene example diagram of a live processing method according to an embodiment of the present disclosure;

[0014] FIG. 2 is a flow diagram of a live processing method according to an embodiment of the present disclosure;

[0015] FIG. 3 is a flow diagram of a live processing method according to another embodiment of the present disclosure;

[0016] FIG. 4a is an interface diagram of a layout according to an embodiment of the present disclosure;

[0017] FIG. 4b is an interface diagram of a layout according to another embodiment of the present disclosure;

[0018] FIG. 5 is a diagram of a preset list view control aligned with a layer corresponding to a latest video frame of each live stream according to an embodiment of the present disclosure;

[0019] FIG. 6 is a flow diagram of a live processing method according to another embodiment of the present disclosure;

[0020] FIG. 7 is a diagram of a target template according to an embodiment of the present disclosure;

[0021] FIG. 8 is a flowchart of a live processing method according to another embodiment of the present disclosure;

[0022] FIG. 9 is a schematic diagram of determining a target template according to an embodiment of the present disclosure;

[0023] FIG. 10 is a structural block diagram of a live processing device according to an embodiment of the present disclosure;

[0024] FIG. 11 is a schematic diagram of a hardware structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] In order to make the objects, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present disclosure.

[0026] In the process of network live streaming, multi-connection is a way that anchors often use to interact with other anchors for live streaming. Based on this way, the anchor in the current live streaming room can invite other anchors to live stream together.

[0027] In the multi-connection live streaming scenario, for any first live streaming client, different rendering views, different GL rendering threads, different frame buffer objects, and different surface views are used for rendering the current video frame of the first live streaming client and the video frame of each other second live streaming client, respectively. The rendering view, the GL rendering thread, the frame buffer object, and the surface view cannot be reused, thus occupying a large amount of resources of the live streaming client. However, as the number of anchors in the multi-connection live streaming scenario increases, the live streaming fluency is seriously affected, the performance of the live streaming client is seriously affected, and the live streaming client may be seriously heated.

[0028] To solve the above technical problems, the present disclosure provides a live processing method. For any first live intercom client in a multi-person live intercom room, the live processing method comprises: receiving, by a live intercom client, a live stream of a second live intercom client, and caching a latest video frame of each live stream; when a current video frame of the first live intercom client is acquired, reading the cached latest video frame; based on a single rendering view, synthesizing the current video frame and the latest video frame into a target video frame according to preset layout information; and displaying the target video frame. In a multi-person live intercom scene, the current video frame of the first live intercom client and the latest video frame of the live stream of the second live intercom client are synthesized into a target video frame, and only a single rendering view is needed to realize rendering, which reduces resource occupation, reduces the performance pressure on the live intercom client, and improves the smoothness of the live stream.

[0029] The live processing method provided by the present disclosure is applied in the scene as shown in FIG. 1, and the execution subject is a first live intercom client. The first live intercom client can receive a live stream of one or more second live intercom clients, and cache a latest video frame of each live stream to a corresponding layer Layer. The first live intercom client can collect a current video frame as an initial layer OriginLayer. Then, when the first live intercom client collects the current video frame, the cached latest video frame of each live stream is read, the current video frame and the latest video frame of each live stream are synthesized into a target video frame, and the target video frame is displayed.

[0030] It should be noted that the user information and data involved in the present application are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0031] The live processing method of the present disclosure will be described in detail below in combination with specific embodiments.

[0032] Referring to FIG. 2, FIG. 2 is a flowchart of a live processing method provided by an embodiment of the present disclosure. The method of the present embodiment can be applied in any live intercom client (denoted as a first live intercom client) in a multi-person live intercom room. The live processing method comprises:

[0033] S201, receiving a live stream of a second live intercom client, and caching a latest video frame of the live stream of the second live intercom client.

[0034] In the embodiment, in the multi-person live streaming connection scene, the first live streaming connection client can invite one or more second live streaming connection clients to join the live streaming, so that the first live streaming connection client and the second live streaming connection client can simultaneously view the picture of the first live streaming connection client and the picture of the one or more second live streaming connection clients in the interface of the first live streaming connection client and the second live streaming connection client. In the multi-person live streaming connection scene, the first live streaming connection client can receive the live streaming of the one or more second live streaming connection clients. Specifically, the one or more second live streaming connection clients can upload the live streaming to a server, and the server can send the live streaming of the one or more second live streaming connection clients to the first live streaming connection client.

[0035] Further, the live streaming of each second live streaming connection client is analyzed and processed to obtain a video frame in the live streaming, and the latest video frame is cached. Specifically, the latest video frame of each live streaming can be drawn and cached in a corresponding layer.

[0036] S202, when the current video frame of the first live streaming connection client is obtained, the cached latest video frame is read.

[0037] In the embodiment, the first live streaming connection client can collect the picture on the side of the first live streaming connection client in real time to obtain a current video frame, and drive the synthesis of the current video frame and the latest video frame of the live streaming of each second live streaming connection client based on the current video frame. Since the current video frame collected by the first live streaming connection client does not need to be transmitted through the network, the output of the current video frame collected by the first live streaming connection client is stable, and thus driving by the current video frame collected by the first live streaming connection client can ensure the stability of the live streaming picture. When the current video frame collected by the first live streaming connection client is obtained, the cached latest video frame of each live streaming can be read. Specifically, the latest video frame of each live streaming can be read from each layer for the synthesis of the target video frame.

[0038] S203, based on a single rendering view, the current video frame and the latest video frame of each live streaming are synthesized into a target video frame according to preset layout information.

[0039] In the embodiment, after obtaining the current video frame collected by the first live co-host client and the latest video frame of the live stream of each second live co-host client, the current video frame collected by the first live co-host client and the latest video frame of each live stream are synthesized and drawn on the same texture to obtain a target video frame, and a common rendering view (RenderView) is used to render the current video frame and the latest video frame of each live stream in one view, without using different rendering views for the current video frame of the first live co-host client and the video frame of the live stream of each second live co-host client, thereby reducing the performance pressure on the first live co-host client. In the synthesis process, the preset layout information can be used to determine how the current video frame and the latest video frame of each live stream are laid out.

[0040] In addition, when the current video frame and the latest video frame of each live stream are synthesized into a target video frame based on a single rendering view and according to preset layout information, only one GL rendering thread is needed, without using one GL rendering thread for the current video frame of the first live co-host client and the video frame of the live stream of each second live co-host client, thereby further reducing the performance pressure on the first live co-host client.

[0041] Optionally, in the embodiment, the target video frame can be rendered into a single frame buffer object (FBO), thereby avoiding the use of different FBOs for buffering the current video frame of the first live co-host client and the video frame of the live stream of each second live co-host client, and further reducing the performance pressure on the first live co-host client.

[0042] S204, displaying the target video frame.

[0043] In the embodiment, after the target video frame is synthesized and rendered, the target video frame is displayed in the display interface of the first live co-host client.

[0044] Since the current video frame of the first live co-host client and the latest video frame of the live stream of each second live co-host client are synthesized into a target video frame, in the embodiment, only one surface view (SurfaceView) is needed to display the target video frame, without using different SurfaceViews for displaying the current video frame of the first live co-host client and the video frame of the live stream of each second live co-host client, thereby further reducing the performance pressure on the first live co-host client.

[0045] Optionally, if the target video frame is rendered into a single FBO, the rendered target video frame buffered in the FBO is displayed by using a single SurfaceView.

[0046] The live processing method provided by the embodiment receives a live stream of a second live microphone-connection client by a first live microphone-connection client, and caches a latest video frame of the live stream of the second live microphone-connection client; when a current video frame of the first live microphone-connection client is acquired, the cached latest video frame is read; based on a single rendering view, the current video frame and the latest video frame are combined into a target video frame according to preset layout information; and the target video frame is displayed. In a multi-person online live scenario, the current video frame of the first live microphone-connection client and the latest video frame of the live stream of the other second live microphone-connection client are combined into a target video frame, only a single rendering view is needed to realize rendering, resource occupation is reduced, the performance pressure on the live microphone-connection client is reduced, and the live stream smoothness is improved.

[0047] On the basis of the above-mentioned embodiments, as shown in FIG. 3, the step of combining the current video frame and the latest video frame into a target video frame according to preset layout information in S203 comprises:

[0048] S301, determining sizes and positions of a first layer corresponding to the current video frame and a second layer corresponding to the latest video frame in the target video frame according to preset layout information.

[0049] S302, drawing the current video frame into the first layer, drawing the latest video frame into the second layer respectively, and combining the first layer and the second layer into the target video frame.

[0050] In the embodiment, according to the preset layout information, how the current video frame of the first live co-host client and the latest video frame of each second live co-host client live stream are laid out in the target video frame can be determined during the synthesis. For example, as shown in FIG. 4a, the first layer of the current video frame of the first live co-host client is located at the left side in the target video frame, and displays the picture of the current video frame of the first live co-host client, and the second layer of the latest video frame of each second live co-host client live stream is arranged in a column vertically and located at the right side in the target video frame, and displays the picture of each second live co-host client; for another example, as shown in FIG. 4b, the first layer of the current video frame of the first live co-host client is displayed full screen in the target video frame, and displays the picture of the current video frame of the first live co-host client, and the second layer of the latest video frame of each second live co-host client live stream is in the form of a floating window and located above the current video frame of the first live co-host client, and the picture of each second live co-host client is suspended above the picture of the current video frame; of course, other layout manners can also be used, and in each layout manner, the size and position of the first layer corresponding to the current video frame and the second layer corresponding to the latest video frame of each live stream in the target video frame can be determined according to the preset layout information, and then the current video frame is drawn into the first layer corresponding to the current video frame, and the latest video frame of each live stream is drawn into the second layer corresponding to the latest video frame of each live stream, so as to synthesize the target video frame from the layers.

[0051] Optionally, before the step S203 of synthesizing the current video frame and the latest video frame into a target video frame based on a single rendering view according to preset layout information, the method further comprises: obtaining the size and position of each view control in a preset list view control, taking the size and position of any view control as the size and position of the second layer corresponding to the latest video frame of any second live co-host client live stream in the target video frame, and obtaining preset layout information so that the second layer corresponding to the latest video frame of the live stream is aligned with the view control in the preset list view control.

[0052] In the embodiment, since the target video frame finally adopts a single surface view, and each live stream in the target video frame needs to be interacted with the UI, a preset list view control (RecyclerView) can be set in the display interface of the multi-person live connection scene, the preset list view control includes a plurality of view controls (ViewHolders), for example, the preset list view control RecyclerView shown in FIG. 5 is a nine-square list view control, and includes nine view controls ViewHolder. The second layer corresponding to the latest video frame of each second live intercommunication client live stream is respectively aligned with each view control ViewHolder, and interaction with the corresponding live stream can be realized through each view control ViewHolder. In order to realize that the second layer corresponding to the latest video frame of each second live intercommunication client live stream is respectively aligned with each view control, the size and position of each view control in the preset list view control are obtained, and the size and position of each view control are respectively taken as the size and position of the layer corresponding to the latest video frame of each live stream in the target video frame, thereby serving as the preset layout information. Then, during composition, the current video frame and the latest video frame of each second live intercommunication client live stream are composed into a target video frame according to the preset layout information, so that the second layer corresponding to the latest video frame of each live stream in the target video frame is respectively aligned with each view control in the preset list view control.

[0053] On the basis of any of the above embodiments, during composition, the current video frame of the first live intercommunication client and the latest video frame of the second live intercommunication client live stream can have a certain overlapping area, and therefore a display level value needs to be set, the highest level is placed on the top layer, the lowest level is placed on the bottom layer, and if there is an overlapping area, the video frame with a higher level covers the video frame with a lower level. Optionally, as shown in FIG. 6, S203 described above that the current video frame and the latest video frame are composed into a target video frame according to the preset layout information includes:

[0054] S401, determining a target stencil in a fragment shader rendering pipeline for stencil testing, wherein a depth value at a region corresponding to the current video frame in the target stencil is a level value corresponding to the current video frame, and a depth value at a region corresponding to the latest video frame is a level value corresponding to the latest video frame;

[0055] S402, performing stencil testing according to the level value of the current video frame, the level value of the latest video frame, and the target stencil, and composing the current video frame and the latest video frame based on the stencil testing result.

[0056] In the embodiment, a stencil test link is included in the rendering pipeline of the OpenGL fragment shader. The stencil test is based on a stencil to determine whether a fragment is retained or discarded. If a fragment does not pass the stencil test, the fragment is discarded, i.e., not drawn. If a fragment passes the stencil test, the fragment is retained. In the embodiment, a target stencil for the stencil test in the fragment shader rendering pipeline can be preconfigured. In the target stencil, the depth value at the region corresponding to the current video frame and the latest video frame of each second live-mic client live stream is the level value zOrder of the region. As shown in FIG. 7, the higher the level value zOrder, the larger the number, i.e., the higher the depth value. For example, the latest video frame of the second live-mic client live stream is suspended above the current video frame of the first live-mic client, i.e., the level value zOrder of the latest video frame of the second live-mic client live stream is higher than the level value zOrder of the current video frame of the first live-mic client. The level value zOrder of the latest video frame of the second live-mic client live stream can be set to 3, and the level value zOrder of the current video frame of the first live-mic client can be set to 2. In the target stencil, the depth value at each position in the region (in the box) corresponding to the latest video frame of the live stream is 3, and the depth value at each position in the region not covered by the current video frame of the first live-mic client is 2.

[0057] In the synthesis of the current video frame of the first live-mic client and the latest video frame of the second live-mic client live stream, the stencil test is performed according to the level value of the current video frame, the level value of the latest video frame, and the target stencil. The current video frame is rendered and drawn at the level value of the current video frame and the position of the latest video frame that passes the stencil test. The current video frame is not rendered and drawn at the position that does not pass the stencil test. Finally, the target video frame after synthesis is obtained.

[0058] In some embodiments, the specific process of the stencil test can include:

[0059] The level value at any first position of the current video frame is determined as the actual depth value at the first position. The actual depth value at the first position of the current video frame is compared with the depth value at the corresponding position in the target stencil. If the actual depth value at the first position is higher than or equal to the depth value at the corresponding position in the target stencil, the stencil test at the first position passes, and the first position of the current video frame can be rendered. If the actual depth value at the first position is lower than the depth value at the corresponding position in the target stencil, the stencil test at the first position does not pass, and the first position of the current video frame is not rendered.

[0060] determining a level value at any second position of a latest video frame of any second live streaming of a second live-mic client as an actual depth value at the second position, comparing the actual depth value at the second position of the latest video frame with a depth value at a corresponding position in the target template, if the actual depth value at the second position is higher than or equal to the depth value at the corresponding position in the target template, the template test at the second position passes, and the second position of the latest video frame can be rendered; if the actual depth value at the second position is lower than the depth value at the corresponding position in the target template, the template test at the second position fails, and the second position of the latest video frame is not rendered.

[0061] by rendering all first positions passing the template test in the current video frame, and rendering all second positions passing the template test in the latest video frame of each live streaming, the hierarchical relationship in the synthesis is realized, and after the positions passing the template test in the current video frame and the latest video frame of each second live-mic client live streaming are rendered, the target video frame after the synthesis is obtained.

[0062] On the basis of the above embodiment, if the level value of the latest video frame of the any live streaming is higher than the level value of the current video frame, the target template for the template test in the fragment shader rendering pipeline determined in S401 can include:

[0063] S501, setting the depth value at all positions of the target template as the level value of the current video frame;

[0064] S502, updating the depth value at the corresponding region of the latest video frame in the target template as the level value of the latest video frame according to the level value of the latest video frame and the preset layout information.

[0065] In the embodiment, the preset level value zOrder of the current video frame is acquired, it is assumed that the level value zOrder of the current video frame is 2, the depth value of all positions of the target template is set as the level value of the current video frame, that is, the depth value of all positions of the target template is 2, as shown in the upper side of FIG. 9; the preset level value zOrder of the latest video frame of any second live streaming of the live streaming client is acquired, it is assumed that the level value zOrder of the latest video frame of any second live streaming of the live streaming client is 3, and the position and size of the latest video frame of the live streaming in the target video frame can be obtained from the preset layout information, thus, the corresponding area of the latest video frame of any second live streaming of the live streaming client in the target template can be determined, that is, the square area in the figure, further, the depth value of the corresponding area of the latest video frame of any second live streaming of the live streaming client in the target template is updated as the level value of the latest video frame of the second live streaming of the live streaming client, that is, the depth value of each position in the square area of the target template is updated as 3, as shown in the lower side of FIG. 9, and finally the target template is obtained.

[0066] In some embodiments, the target template is configured with a buffer (cache), and S501-S502 are implemented in the buffer of the target template, that is, the depth value of all positions of the target template is set as the level value of the current video frame in the buffer of the target template; the depth value of the corresponding area of the latest video frame of any second live streaming of the live streaming client in the target template is updated as the level value of the latest video frame of the second live streaming of the live streaming client in the buffer of the target template.

[0067] The live processing method, device and storage medium provided by the embodiments of the present disclosure, through the first live streaming client, the live stream of the second live streaming client is received, and the latest video frame of the live stream of the second live streaming client is cached; when the current video frame of the first live streaming client is acquired, the cached latest video frame is read; based on a single rendering view, the current video frame and the latest video frame are combined into a target video frame according to the preset layout information; and the target video frame is displayed. In the multi-person live streaming scene, the current video frame of the first live streaming client and the latest video frame of the live stream of the second live streaming client are combined into a target video frame, only a single rendering view is needed to realize rendering, the resource occupation is reduced, the performance pressure of the live streaming client is reduced, and the live stream smoothness is improved.

[0068] Corresponding to the live processing method of the above embodiment, FIG. 10 is a structural block diagram of a live processing apparatus provided by an embodiment of the present disclosure. For ease of illustration, only parts related to the embodiments of the present disclosure are shown. Referring to FIG. 10, the live processing apparatus 600 includes a receiving unit 601, a caching unit 602, a rendering unit 603, and a display unit 604.

[0069] In some embodiments, the receiving unit 601 is configured to receive a live stream of a second live co-host client; the caching unit 602 is configured to cache a latest video frame of the live stream of the second live co-host client; the cached latest video frame is read when a current video frame of the first live co-host client is acquired; the rendering unit 603 is configured to synthesize the current video frame and the latest video frame into a target video frame according to preset layout information based on a single rendering view; and the display unit 604 is configured to display the target video frame.

[0070] In one or more embodiments of the present disclosure, when synthesizing the current video frame and the latest video frame into a target video frame according to preset layout information, the rendering unit 603 is configured to: determine sizes and positions of a first layer corresponding to the current video frame and a second layer corresponding to the latest video frame in the target video frame according to the preset layout information; draw the current video frame into the first layer, draw the latest video frame into the second layer respectively, and synthesize the first layer and the second layer into the target video frame.

[0071] In one or more embodiments of the present disclosure, when synthesizing the current video frame and the latest video frame into a target video frame according to preset layout information, the rendering unit 603 is configured to: determine a target stencil in a fragment shader rendering pipeline for stencil testing, wherein a depth value at a region corresponding to the current video frame in the target stencil is a level value corresponding to the current video frame, and a depth value at a region corresponding to the latest video frame in the target stencil is a level value corresponding to the latest video frame; perform stencil testing according to the level value of the current video frame, the level value of the latest video frame, and the target stencil, and synthesize the current video frame and the latest video frame based on a stencil testing result.

[0072] In one or more embodiments of the present disclosure, the rendering unit 603, when performing template testing according to the level value of the current video frame, the level value of the latest video frame of each live stream, and the target template, is configured to: determine the level value at any first position of the current video frame as the actual depth value at the first position, compare the actual depth value at the first position of the current video frame with the depth value at the corresponding position in the target template, and if the actual depth value at the first position is higher than or equal to the depth value at the corresponding position in the target template, the template testing at the first position passes; determine the level value at any second position of the latest video frame as the actual depth value at the second position, compare the actual depth value at the second position of the latest video frame with the depth value at the corresponding position in the target template, and if the actual depth value at the second position is higher than or equal to the depth value at the corresponding position in the target template, the template testing at the second position passes.

[0073] In one or more embodiments of the present disclosure, the rendering unit 603, when synthesizing the current video frame and the latest video frame based on the template testing result, is configured to: if the template testing at any first position of the current video frame passes, render the first position of the current video frame; if the template testing at any second position of the latest video frame passes, render the second position of the latest video frame; and after the positions in the current video frame and the latest video frame that pass the template testing are rendered, obtain the synthesized target video frame.

[0074] In one or more embodiments of the present disclosure, the level value of the latest video frame is higher than the level value of the current video frame.

[0075] In one or more embodiments of the present disclosure, the rendering unit 603, when determining the target template for template testing in a fragment shader rendering pipeline, is configured to: set the depth value at all positions of the target template as the level value of the current video frame; and according to the level value of the latest video frame and the preset layout information, update the depth value at the corresponding region of the latest video frame in the target template as the level value of the latest video frame.

[0076] In one or more embodiments of the present disclosure, before the rendering unit 603 composites the current video frame and the latest video frame of each live stream into a target video frame according to preset layout information based on a single rendering view, the rendering unit 603 is further configured to: obtain the size and position of each view control in a preset list view control, and take the size and position of any view control as the size and position of a second layer corresponding to the latest video frame in the target video frame, to obtain the preset layout information, so that the second layer corresponding to the latest video frame is aligned with the any view control in the preset list view control.

[0077] In one or more embodiments of the present disclosure, when the rendering unit 603 composites the current video frame and the latest video frame into a target video frame according to preset layout information based on a single rendering view, the rendering unit 603 is configured to: composite the current video frame and the latest video frame into a target video frame according to preset layout information based on a single rendering view and a single rendering thread.

[0078] In one or more embodiments of the present disclosure, when the display unit 604 displays the target video frame, the display unit 604 is configured to: display the target video frame using a single surface view.

[0079] In one or more embodiments of the present disclosure, when the rendering unit 603 composites the current video frame and the latest video frame into a target video frame according to preset layout information based on a single rendering view, the rendering unit 603 is further configured to: render the target video frame into a single frame buffer object; and when the display unit 604 displays the target video frame using a single surface view, the display unit 604 is configured to: display the rendered target video frame cached in the frame buffer object using a single surface view.

[0080] The live processing apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments, and has similar implementation principles and technical effects, which will not be described here again in this embodiment.

[0081] Referring to FIG. 11, a structural diagram of an electronic device 700 suitable for implementing embodiments of the disclosure is illustrated, which can be a terminal device or a server. The terminal device can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a Personal Digital Assistant (PDA), a Portable Android Device (PAD), a Portable Media Player (PMP), a car terminal (e.g., a car navigation terminal), and the like, and a stationary terminal such as a digital TV, a desktop computer, and the like. The electronic device illustrated in FIG. 11 is merely an example and should not impose any limitation on the functions and use range of embodiments of the disclosure.

[0082] As illustrated in FIG. 11, the electronic device 700 can include a processing device (e.g., a central processor, a graphic processor, or the like) 701 that can perform various appropriate actions and processes according to a program stored in a Read Only Memory (ROM) 702 or a program loaded into a Random Access Memory (RAM) 703 from a storage device 708. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An Input / Output (I / O) interface 705 is also connected to the bus 704.

[0083] Generally, the following devices can be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 707 including, for example, a Liquid Crystal Display (LCD), a speaker, a vibrator, and the like; a storage device 708 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 709. The communication device 709 can allow the electronic device 700 to communicate with other devices wirelessly or wired to exchange data. Although FIG. 11 illustrates the electronic device 700 having various devices, it should be understood that all of the illustrated devices are not required to be implemented or possessed. More or less devices can be alternatively implemented or possessed.

[0084] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0085] Note that the computer readable medium described above in the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium may, for example, be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program used or used in conjunction with an instruction execution system, device, or apparatus. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take on many forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The computer readable signal medium can also be any computer readable medium that can send, propagate, or transfer a program for use by or in connection with an instruction execution system, device, or apparatus. The program code contained on the computer readable medium can be transmitted using any suitable medium, including but not limited to wire, cable, fiber optic, RF (radio frequency), or any suitable combination of the foregoing.

[0086] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and not be assembled into the electronic device.

[0087] The computer readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the embodiments described above.

[0088] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0089] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may

[0090] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself. For example, the first obtaining unit can also be described as a unit for obtaining at least two Internet protocol addresses.

[0091] The functions described above in the detailed description of embodiments of the present disclosure can be implemented in at least in part by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0092] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage media can include, without limitation, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0093] In a first aspect, according to one or more embodiments of the present disclosure, a live processing method is provided, applied to a first live co-host client, the method comprising: receiving a live stream of a second live co-host client, and caching a latest video frame of the live stream of the second live co-host client; reading the cached latest video frame when a current video frame of the first live co-host client is acquired; synthesizing the current video frame and the latest video frame into a target video frame based on a preset layout information; and displaying the target video frame.

[0094] According to one or more embodiments of the present disclosure, the synthesizing the current video frame and the latest video frame into a target video frame based on a preset layout information comprises: determining sizes and positions of a first layer corresponding to the current video frame and a second layer corresponding to the latest video frame in the target video frame according to the preset layout information; drawing the current video frame into the first layer, drawing the latest video frame into the second layer respectively, and synthesizing the first layer and the second layer into the target video frame.

[0095] According to one or more embodiments of the present disclosure, the synthesizing the current video frame and the latest video frame into a target video frame based on a preset layout information comprises: determining a target stencil for stencil testing in a fragment shader rendering pipeline, wherein a depth value at a region corresponding to the current video frame in the target stencil is a level value corresponding to the current video frame, and a depth value at a region corresponding to the latest video frame in the target stencil is a level value corresponding to the latest video frame; performing stencil testing according to the level value of the current video frame, the level value of the latest video frame, and the target stencil, and synthesizing the current video frame and the latest video frame based on a stencil testing result.

[0096] According to one or more embodiments of the present disclosure, the template test based on the level value of the current video frame, the level value of the latest video frame and the target template comprises: determining the level value at any first position of the current video frame as an actual depth value at the first position, comparing the actual depth value at the first position of the current video frame with a depth value at a corresponding position in the target template, and if the actual depth value at the first position is higher than or equal to the depth value at the corresponding position in the target template, the template test at the first position is passed; determining the level value at any second position of the latest video frame as an actual depth value at the second position, comparing the actual depth value at the second position of the latest video frame with a depth value at a corresponding position in the target template, and if the actual depth value at the second position is higher than or equal to the depth value at the corresponding position in the target template, the template test at the second position is passed.

[0097] According to one or more embodiments of the present disclosure, the synthesizing of the current video frame and the latest video frame based on the template test result comprises: rendering the first position of the current video frame if the template test at the first position of the current video frame is passed; rendering the second position of the latest video frame if the template test at the second position of the latest video frame is passed; and obtaining the synthesized target video frame after the positions of the current video frame and the latest video frame that pass the template test are rendered.

[0098] According to one or more embodiments of the present disclosure, the level value of the latest video frame is higher than the level value of the current video frame.

[0099] According to one or more embodiments of the present disclosure, the determining of the target template for template test in the fragment shader rendering pipeline comprises: setting a depth value at all positions of the target template as the level value of the current video frame; and updating the depth value at a corresponding region of the latest video frame in the target template as the level value of the latest video frame according to the level value of the latest video frame and the preset layout information.

[0100] According to one or more embodiments of the present disclosure, before the synthesizing of the current video frame and the latest video frame into a target video frame based on a single rendering view according to preset layout information, the method further comprises: obtaining the size and position of each view control in a preset list view control, taking the size and position of any view control as the size and position of a second layer corresponding to the latest video frame in the target video frame, and obtaining the preset layout information so that the second layer corresponding to the latest video frame is aligned with the any view control in the preset list view control.

[0101] According to one or more embodiments of the present disclosure, the synthesizing the current video frame and the latest video frame into a target video frame based on the single rendering view according to preset layout information comprises: synthesizing the current video frame and the latest video frame into a target video frame based on the single rendering view and a single rendering thread according to preset layout information.

[0102] According to one or more embodiments of the present disclosure, the displaying the target video frame comprises: displaying the target video frame by using a single surface view.

[0103] According to one or more embodiments of the present disclosure, the synthesizing the current video frame and the latest video frame into a target video frame based on the single rendering view according to preset layout information further comprises: rendering the target video frame into a single frame buffer object; and the displaying the target video frame by using a single surface view comprises: displaying the rendered target video frame cached in the frame buffer object by using a single surface view.

[0104] In a second aspect, according to one or more embodiments of the present disclosure, a live processing device is provided, which is applied to a first live co-voice client, and the device comprises: a receiving unit configured to receive a live stream of a second live co-voice client; a caching unit configured to cache a latest video frame of the live stream of the second live co-voice client; when a current video frame of the first live co-voice client is acquired, the cached latest video frame is read; a rendering unit configured to synthesize the current video frame and the latest video frame into a target video frame based on a single rendering view according to preset layout information; and a display unit configured to display the target video frame.

[0105] According to one or more embodiments of the present disclosure, when synthesizing the current video frame and the latest video frame into a target video frame according to preset layout information, the rendering unit is configured to: determine sizes and positions of a first layer corresponding to the current video frame and a second layer corresponding to the latest video frame in the target video frame according to preset layout information; draw the current video frame into the first layer, draw the latest video frame into the second layer respectively, and synthesize the first layer and the second layer into the target video frame.

[0106] According to one or more embodiments of the present disclosure, the rendering unit, when compositing the current video frame and the latest video frame into a target video frame according to preset layout information, is configured to: determine a target stencil in a fragment shader rendering pipeline for stencil testing, wherein a depth value at a region corresponding to the current video frame in the target stencil is a level value corresponding to the current video frame, and a depth value at a region corresponding to the latest video frame in the target stencil is a level value corresponding to the latest video frame; and perform stencil testing according to the level value of the current video frame, the level value of the latest video frame, and the target stencil, and composite the current video frame and the latest video frame based on a stencil testing result.

[0107] According to one or more embodiments of the present disclosure, when the rendering unit performs stencil testing according to the level value of the current video frame, the level value of the latest video frame of each live stream, and the target stencil, the rendering unit is configured to: determine a level value at any first position of the current video frame as an actual depth value at the first position, compare the actual depth value at the first position of the current video frame with a depth value at a corresponding position in the target stencil, and if the actual depth value at the first position is higher than or equal to the depth value at the corresponding position in the target stencil, stencil testing at the first position is passed; determine a level value at any second position of the latest video frame as an actual depth value at the second position, compare the actual depth value at the second position of the latest video frame with a depth value at a corresponding position in the target stencil, and if the actual depth value at the second position is higher than or equal to the depth value at the corresponding position in the target stencil, stencil testing at the second position is passed.

[0108] According to one or more embodiments of the present disclosure, when the rendering unit composites the current video frame and the latest video frame based on a stencil testing result, the rendering unit is configured to: if stencil testing at any first position of the current video frame is passed, render the first position of the current video frame; if stencil testing at any second position of the latest video frame is passed, render the second position of the latest video frame; and after rendering is completed at positions in the current video frame and the latest video frame that pass stencil testing, obtain the target video frame after compositing.

[0109] According to one or more embodiments of the present disclosure, the level value of the latest video frame is higher than the level value of the current video frame.

[0110] According to one or more embodiments of the present disclosure, the rendering unit, when determining a target stencil for stencil test in a fragment shader rendering pipeline, is configured to: set a depth value at all positions of the target stencil to a level value of the current video frame; and update a depth value at a region corresponding to the latest video frame in the target stencil to the level value of the latest video frame according to the level value of the latest video frame and the preset layout information.

[0111] According to one or more embodiments of the present disclosure, before the rendering unit composites the current video frame and the latest video frame of each live stream into a target video frame based on a single rendering view and according to preset layout information, the rendering unit is further configured to: obtain a size and a position of each view control in a preset list view control; and set the size and the position of any view control as a size and a position of a second layer corresponding to the latest video frame in the target video frame, to obtain the preset layout information, so that the second layer corresponding to the latest video frame is aligned with the any view control in the preset list view control.

[0112] According to one or more embodiments of the present disclosure, when the rendering unit composites the current video frame and the latest video frame into a target video frame based on a single rendering view and according to preset layout information, the rendering unit is configured to: composite the current video frame and the latest video frame into the target video frame based on a single rendering view and a single rendering thread and according to the preset layout information.

[0113] According to one or more embodiments of the present disclosure, when the display unit displays the target video frame, the display unit is configured to: display the target video frame using a single surface view.

[0114] According to one or more embodiments of the present disclosure, when the rendering unit composites the current video frame and the latest video frame into a target video frame based on a single rendering view and according to preset layout information, the rendering unit is further configured to: render the target video frame into a single frame buffer object; and when the display unit displays the target video frame using a single surface view, the display unit is configured to: display the rendered target video frame cached in the frame buffer object using the single surface view.

[0115] In a third aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, including: at least one processor and a memory; the memory stores computer-executable instructions; and the at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the live processing method according to the first aspect and various possible designs of the first aspect.

[0116] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer readable storage medium is provided, and the computer readable storage medium has stored therein computer executing instructions which, when executed by a processor, implement the live processing method as described in the first aspect above and various possible designs of the first aspect.

[0117] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, and the computer program product includes computer executing instructions which, when executed by a processor, implement the live processing method as described in the first aspect above and various possible designs of the first aspect.

[0118] The above description merely illustrates the preferred embodiments of the present disclosure and the principles of the applied technologies. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or equivalent features, without departing from the above disclosed concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features disclosed in the present disclosure (but not limited to) having similar functions.

[0119] In addition, although each operation is depicted in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.

[0120] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A live processing method applied to a first live co-host client, the method comprising: receiving a live stream of a second live co-host client and buffering a latest video frame of the live stream of the second live co-host client; reading the buffered latest video frame when a current video frame of the first live co-host client is obtained; compositing the current video frame and the latest video frame into a target video frame according to preset layout information based on a single rendering view; displaying the target video frame.

2. The method of claim 1, wherein, The compositing the current video frame and the latest video frame into a target video frame according to preset layout information comprises: determining sizes and positions of a first layer corresponding to the current video frame and a second layer corresponding to the latest video frame in the target video frame according to preset layout information; drawing the current video frame into the first layer, drawing the latest video frame into the second layer respectively, and compositing the first layer and the second layer into the target video frame.

3. The method of claim 1, wherein, The compositing the current video frame and the latest video frame into a target video frame according to preset layout information comprises: determining a target stencil for stencil test in a fragment shader rendering pipeline, wherein a depth value at a region corresponding to the current video frame in the target stencil is a level value of the current video frame, and a depth value at a region corresponding to the latest video frame in the target stencil is a level value of the latest video frame; performing stencil test according to the level value of the current video frame, the level value of the latest video frame, and the target stencil, and compositing the current video frame and the latest video frame based on a stencil test result.

4. The method of claim 3, wherein, The performing stencil test according to the level value of the current video frame, the level value of the latest video frame, and the target stencil comprises: determining a level value at any first position of the current video frame as an actual depth value at the first position of the current video frame, comparing the actual depth value at the first position with a depth value at a corresponding position in the target stencil, and if the actual depth value at the first position is higher than or equal to the depth value at the corresponding position in the target stencil, stencil test at the first position is passed; determining a level value at any second position of the latest video frame as an actual depth value at the second position of the latest video frame, comparing the actual depth value at the second position with a depth value at a corresponding position in the target stencil, and if the actual depth value at the second position is higher than or equal to the depth value at the corresponding position in the target stencil, stencil test at the second position is passed.

5. The method of claim 4, wherein, The compositing the current video frame and the latest video frame based on the stencil test result comprises: if stencil test at any first position of the current video frame is passed, rendering the first position of the current video frame; if stencil test at any second position of the latest video frame is passed, rendering the second position of the latest video frame; after rendering at positions of the current video frame and the latest video frame that pass stencil test is completed, obtaining the composited target video frame.

6. The method according to any one of claims 3-5, wherein, The level value of the latest video frame is higher than the level value of the current video frame.

7. The method of claim 6, wherein, The determining includes: Setting the depth value at all positions of the target stencil to the level value of the current video frame; According to the level value of the latest video frame and the preset layout information, updating the depth value of the latest video frame at the corresponding region of the target stencil to the level value of the latest video frame.

8. The method of claim 1, wherein, Before the synthesizing, the method further includes: Obtaining the size and position of each view control in the preset list view control, taking the size and position of any view control as the size and position of the second layer corresponding to the latest video frame in the target video frame, obtaining the preset layout information, so that the second layer corresponding to the latest video frame is aligned with the any view control in the preset list view control.

9. The method according to any one of claims 1-5, wherein, The synthesizing includes: According to the preset layout information, synthesizing the current video frame and the latest video frame into a target video frame based on a single rendering view. The displaying includes:

10. The method of claim 9, wherein, Displaying the target video frame by using a single surface view. The synthesizing includes:

11. The method of claim 10, wherein, Rendering the target video frame into a single frame buffer object. The displaying includes: Displaying the rendered target video frame cached in the frame buffer object by using a single surface view. 12.A live processing apparatus applied to a first live co-host client, the apparatus comprising: a receiving unit configured to receive a live stream of a second live co-host client; a caching unit configured to cache a latest video frame of the live stream of the second live co-host client; when a current video frame of the first live co-host client is obtained, reading the cached latest video frame; a rendering unit configured to synthesize the current video frame and the latest video frame into a target video frame based on a single rendering view according to preset layout information; a display unit configured to display the target video frame. at least one processor and a memory; 13. An electronic device comprising: the memory stores computer-executed instructions; the at least one processor executes the computer-executed instructions stored in the memory, so that the at least one processor performs the method according to any one of claims 1-11. 14.A computer-readable storage medium, the computer-readable storage medium storing computer-executed instructions, when a processor executes the computer-executed instructions, the method according to any one of claims 1-11 is implemented. 15.A computer program product, the computer program product comprising computer-executed instructions, when a processor executes the computer-executed instructions, the method according to any one of claims 1-11 is implemented. ​