Video stream processing method, apparatus, and electronic device
By adaptively creating the main window layout for the video session and obtaining the target video stream on the requesting end, the problem of a single client-side screen and high server-side load in existing technologies is solved, enabling flexible configuration of the video session system and improving user experience.
Patent Information
- Application Number
- CN202211414772.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-11
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2042-11-11
AI Technical Summary
Existing multi-user video conferencing systems suffer from limited client-side visuals, high server load, high costs, and a limited number of supported clients.
The requesting end adaptively creates the layout information of the main video session window based on user needs and preferences, and provides the corresponding target video stream through the server and the provider, thereby realizing flexible configuration of the main video session window.
It improves user experience, reduces data processing volume, lowers computing load and cost, and supports more target video streams.
Smart Images

Figure CN115801992B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of information processing technology, and in particular to a video stream processing method, apparatus and electronic device. Background Technology
[0002] A multi-user video conferencing system includes various participating devices such as servers and clients. The server may include multipoint control units (MCUs). Clients acquire video streams and send them to the MCU on the server. The MCU mixes the received video streams to generate a single video stream, which is then sent to each client for display.
[0003] In related technologies, video conferencing systems involving multiple participants suffer from problems such as a single client-side view, high server load, and high cost, and can only support a limited number of clients. Summary of the Invention
[0004] This disclosure provides a video stream processing method, apparatus, and electronic device to enable the requesting end to adaptively subscribe to the main video session window according to its needs and preferences, thereby achieving flexible configuration of the main video session window.
[0005] In a first aspect, embodiments of this disclosure provide a video processing method applied to a requesting end, the method comprising:
[0006] Get the layout information of the main video session window, which includes multiple video sub-windows;
[0007] Obtain the main window video stream based on the layout information;
[0008] The main video stream is displayed in the main video session window. The main video stream is generated based on the target video stream corresponding to at least one video sub-window. The target video stream is determined based on the layout information of the main video session window.
[0009] Secondly, embodiments of this disclosure provide a video stream processing method applied to a server, the method comprising:
[0010] Get the layout information of the requesting client's main video session window, which includes multiple video sub-windows;
[0011] The target video stream is provided to the requesting end so that the requesting end can display the main window video stream in the main video session window; the main window video stream is generated based on the target video stream corresponding to at least one video sub-window, and the target video stream is determined based on the layout information of the main video session window.
[0012] Thirdly, embodiments of this disclosure provide a video stream processing method applied to a providing end. The method includes: providing a target video stream corresponding to a video sub-window to a requesting end according to an instruction from a server, so that the requesting end can display the main window video stream in the main window of the video session; the main window video stream is generated based on the target video stream corresponding to at least one video sub-window, and the target video stream is determined based on the layout information of the main window of the video session.
[0013] Fourthly, embodiments of this disclosure provide a video stream processing apparatus applied to a requesting end, the apparatus comprising:
[0014] The first layout information acquisition module is used to acquire the layout information of the main video session window, which includes multiple video sub-windows.
[0015] The main window video stream acquisition module is used to acquire the main window video stream based on the layout information;
[0016] The display module is used to display the main window video stream in the main video session window. The main window video stream is generated based on the target video stream corresponding to at least one video sub-window. The target video stream is determined based on the layout information of the main video session window.
[0017] Fifthly, embodiments of this disclosure provide a video stream processing apparatus applied to a server, the apparatus comprising:
[0018] The second layout information acquisition module is used to acquire the layout information of the video session main window of the requesting end. The video session main window includes multiple video sub-windows.
[0019] The first video stream providing module is used to provide the target video stream to the requesting end so that the requesting end can display the main window video stream in the main window of the video session; the main window video stream is generated based on the target video stream corresponding to at least one video sub-window, and the target video stream is determined based on the layout information of the main window of the video session.
[0020] In a sixth aspect, embodiments of this disclosure provide a video stream processing apparatus applied to a providing end. The apparatus includes: a second video stream providing module, configured to provide a target video stream corresponding to a video sub-window to a requesting end according to an instruction from a server, so that the requesting end can display the main window video stream in the main video session window; the main window video stream is generated based on the target video stream corresponding to at least one video sub-window, and the target video stream is determined based on the layout information of the main video session window.
[0021] In a seventh aspect, embodiments of this disclosure provide an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor, when executing the computer program, implements the method provided in any embodiment of this disclosure.
[0022] Eighthly, embodiments of this disclosure provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method provided in any embodiment of this disclosure.
[0023] The video stream processing method disclosed herein allows the requesting end to adaptively create the layout information of the main video session window according to its needs and preferences, achieving flexible configuration of the main video session window and improving user experience. Furthermore, different providers do not uniformly provide the same large-volume video stream to the requesting end; instead, they provide the requesting end with the target video stream corresponding to the video sub-window. This reduces the amount of data processed, lowers computational load and cost, and can support more target video streams.
[0024] The above overview is for illustrative purposes only and is not intended to be limiting in any way. Further aspects, embodiments, and features of this disclosure will become readily apparent from the accompanying drawings and the following detailed description, in addition to the illustrative aspects, embodiments, and features described above. Attached Figure Description
[0025] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments according to this disclosure and should not be construed as limiting the scope of this disclosure.
[0026] Figure 1 This is a scene diagram of a video conference;
[0027] Figure 2 This is a schematic diagram illustrating an application scenario of a video stream processing method according to an embodiment of this disclosure;
[0028] Figure 3 The main window of the video session as viewed by three different participants in one embodiment of this disclosure is shown.
[0029] Figure 4 This is a flowchart illustrating a video stream processing method according to an embodiment of the present disclosure;
[0030] Figure 5 This is a schematic diagram of the display screen on the requesting end after user operation in one embodiment of the present disclosure;
[0031] Figure 6 This is a flowchart illustrating a video stream processing method according to another embodiment of the present disclosure;
[0032] Figure 7A The video session main window layout information set for the requesting client M;
[0033] Figure 7BThis is a diagram illustrating the effect of displaying the main window video stream to the requesting client M.
[0034] Figure 8 This is a schematic diagram showing the main window video stream displayed for requesting clients K, A, and E respectively.
[0035] Figure 9 This is a schematic diagram of the main video session window provided in one embodiment of the present disclosure;
[0036] Figure 10 This is a structural block diagram of a video stream processing apparatus according to an embodiment of the present disclosure;
[0037] Figure 11 This is a structural block diagram of a video stream processing apparatus according to another embodiment of the present disclosure;
[0038] Figure 12 This is a block diagram of an electronic device used to implement embodiments of the present disclosure. Detailed Implementation
[0039] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of this disclosure. Therefore, the drawings and description are to be considered exemplary in nature and not restrictive.
[0040] To facilitate understanding of the technical solutions of the embodiments of this disclosure, the related technologies of the embodiments of this disclosure are described below. The following related technologies are optional solutions and can be combined with the technical solutions of the embodiments of this disclosure in any way, and they all fall within the protection scope of the embodiments of this disclosure.
[0041] This article may involve the following concepts: video mixing server and multiple people in the same frame. A video mixing server is a video transcoding server that supports multi-channel video decoding and encoding, arranging videos in a grid according to requirements. Multiple people in the same frame refers to arranging videos of multiple people together on a single screen.
[0042] Figure 1 This is a scene diagram of a video conference. (Example) Figure 1As shown, A, B, C, D, E, F, and J are participants in the same video conference. During a video conference, a participant can request to view the videos of other participants. For example, participant J can request to view the videos of participants A, B, C, D, E, and F. Then, participants A, B, C, D, E, and F will provide their video streams to participant J. Participant J's terminal will display the received video stream, thus allowing participant J to view the images of participants A, B, C, D, E, and F. Here, the party requesting the video stream is called the requesting party, such as requesting party J; and the party providing the video stream is called the providing party, such as providing parties A, B, C, D, E, and F.
[0043] In related technologies, in video sessions involving multiple participants, such as video conferencing, the server, such as a video mixing server, receives high-resolution (e.g., 720p) video streams from multiple providers and then sends the multiple high-resolution video streams to the requesting end for display. The requesting end displays the screen of the requested provider according to a predetermined layout.
[0044] In this approach, regardless of a participant's request, the server provides that participant with high-resolution video streams from all other participants, allowing that participant to view the feeds of all other participants. For example, as... Figure 1 As shown, client J may not want to view the video feed from provider A. However, after receiving client J's connection request, the server will still provide client J with the video feed from provider A. For example, client J may only want to view a smaller portion of provider A's video feed, but after receiving client J's connection request, the server will still send a high-resolution video stream of A to client J, causing client J to see the larger portion of provider A's video feed. Therefore, this method wastes server computing power and bandwidth.
[0045] On the other hand, each requesting client displays the same screen size. For example, the server sends video streams from clients A, B, C, D, E, and F to requesting client J. In the main window displayed on requesting client J, the screen sizes of A, B, C, D, E, and F are identical. Similarly, the server sends video streams from clients A, J, C, D, E, and F to requesting client B. In the main window displayed on requesting client B, the screen sizes of A, J, C, D, and E are identical. Furthermore, the screen sizes displayed on requesting clients J and B are also the same. This approach results in a monotonous display on each client, making it impossible to flexibly configure screen size and layout according to the client's needs, thus degrading the user experience.
[0046] To address the above situation, this disclosure provides a corresponding solution. On the requesting end, the user can create the layout information of the main video session window according to their preferences. The main video session window includes multiple video sub-windows. After obtaining the layout information of the main video session window, the requesting end sends the layout information to the server. Upon receiving the layout information from the requesting end, the server issues an instruction to the provider, which then provides the requesting end with the target video stream corresponding to the video sub-windows according to the server's instruction. Thus, the requesting end obtains the main window video stream based on the layout information. The requesting end displays the main window video stream in the main video session window, allowing the user on the requesting end to view the provider's subscribed content. The main window video stream is generated based on the target video stream corresponding to at least one video sub-window, and the target video stream is determined based on the layout information of the main video session window.
[0047] Figure 2 This diagram illustrates an application scenario of a video stream processing method according to an embodiment of this disclosure. The video stream processing method of this embodiment can be applied to video conferencing. The requesting end and the providing end are participants in the same video conference.
[0048] For example, A, B, C, D, E, F, and J are participants in the same video conference. For instance, the user at requesting end J wants to view the second-sized images from providers A and B, and the first-sized images from providers C, D, E, and F, where the second size is larger than the first size. After obtaining the layout information of the main window of the video session created by the user, requesting end J sends the layout information to the server. Upon receiving the layout information, the server instructs providers A, B, C, D, E, and F to provide the requesting end J with the target video streams corresponding to the video sub-windows. The target video streams provided by providers A and B correspond to the second-sized images, and the target video streams provided by providers C, D, E, and F correspond to the first-sized images. After obtaining the main window video stream, requesting end J displays the main window video stream in the main video session window. In the main window displayed by requesting end J, the images of A and B are both second-sized images, while the images of C, D, E, and F are both first-sized images.
[0049] For example, a user on requesting end B wants to view a second-sized image from provider A, and also wants to view first-sized images from providers A, C, D, E, and F, with the second size being larger than the first size. After obtaining the layout information of the main window of the video session created by the user, requesting end B sends the layout information to the server. Upon receiving the layout information, the server issues instructions to providers A, C, D, E, and F, which then provide the requesting end B with the target video streams corresponding to the video sub-windows. The target video stream provided by provider A corresponds to the second-sized image, while the target video streams provided by providers C, D, E, and F correspond to the first-sized image. After obtaining the main window video stream, requesting end B displays the main window video stream in the main video session window. In the main window displayed by requesting end J, the image of A is the second-sized image, while the images of C, D, E, and F are all the first-sized images.
[0050] pass Figure 2 As can be seen, the layout information of requesting client J and requesting client B is different, and therefore, the main video session window displayed by requesting client J and requesting client B is also different.
[0051] Through the video stream processing method in this embodiment, the requesting user can adaptively create the layout information of the main video session window according to their needs and preferences. After obtaining the layout information of the main video session window, the requesting end can obtain the main window video stream based on the layout information and display the main window video stream. In this way, the main video session window displayed by the requesting end meets the user's needs and preferences.
[0052] The video stream processing method disclosed herein allows the requesting end to adaptively create the layout information of the main video session window according to its needs and preferences, achieving flexible configuration of the main video session window and improving user experience. Furthermore, different providers do not uniformly provide the same large-volume video stream to the requesting end; instead, they provide the requesting end with the target video stream corresponding to the video sub-window. This reduces the amount of data processed, lowers computational load and cost, and can support more target video streams.
[0053] The video stream processing method of this disclosure can be applied to video conferencing scenarios. In existing video conferencing, the main window layout viewed by each participant is the same, and the size of each sub-window in the main window is also the same. Using the technical solution of this disclosure, each participant can create the layout of the main video session window and the participants they wish to see according to their preferences or needs. Therefore, the main window view viewed by different participants can be different, the layout of the main window can be different, and / or the participant views displayed in the sub-windows of the main window can be different. Figure 3The following is an illustration of the main video session window viewed by three different participants in one embodiment of this disclosure. For example, the main video session window layout of requesting end K and requesting end A is the same, but the participants' screens and screen sizes displayed in each video sub-window are different; the participants displayed in the main video session windows of requesting end A and requesting end E are the same, but the participants' screens and screen sizes are different.
[0054] The video stream processing method of this disclosure can be applied to live streaming scenarios. In existing live streaming scenarios, viewers can usually only see the live stream feed and not the feeds of other viewers participating in the live stream. Using the technical solution of this disclosure, viewers can subscribe to other viewers' video streams according to their preferences. Therefore, the video sub-window displayed on the viewer's end can show not only the live stream feed but also the feeds of other viewers. Furthermore, the viewer can adjust the size and position of the live stream feed and the feeds of other viewers by setting layout information.
[0055] The technical solution disclosed herein can also be applied to other video conversation scenarios that require multiple participants, which will not be listed here.
[0056] The implementation schemes provided in the embodiments of this disclosure will be described in detail below.
[0057] Figure 4 This is a flowchart illustrating a video stream processing method according to an embodiment of the present disclosure. The method is applied to the requesting end. The method includes steps S41 to S43.
[0058] In step S41, the layout information of the main video session window is obtained. The main video session window includes multiple video sub-windows.
[0059] On the requesting side, users can create the layout information for the main video session window according to their needs and preferences. After the user creates the layout information, the requesting side can obtain the layout information of the main video session window based on the user's submission instructions and send the layout information to the server. The main video session window can include multiple video sub-windows.
[0060] In step S42, the main window video stream is obtained based on the layout information.
[0061] After obtaining the layout information of the requesting client's main video session window, the server provides the target video stream to the requesting client. For example, the target video stream is determined based on the layout information of the main video session window. The main window video stream is generated based on the target video stream corresponding to at least one video sub-window. Thus, the requesting client obtains the main window video stream corresponding to the layout information.
[0062] In step S43, the main window video stream is displayed in the main video session window. After the requesting end obtains the main window video stream, it displays the main window video stream in the main video session window, so that the corresponding image is displayed in the video sub-window.
[0063] In the video stream processing method of this disclosure, the requesting end does not passively receive the main window video stream uniformly returned by the server. The user on the requesting end can create layout information for the main video session window according to their needs and preferences. After obtaining the layout information, the requesting end sends it to the server, and the server provides the requesting end with the main window video stream corresponding to the layout information. After obtaining the main window video stream, the requesting end displays the main window video stream in the main video session window. Thus, the user on the requesting end can view the video feed from the provider they have subscribed to in the main window, satisfying user needs and preferences and improving the user experience.
[0064] This approach allows the requesting end to adaptively create the layout information of the main video session window based on its needs and preferences, enabling flexible configuration of the main video session window and improving the user experience. Furthermore, different providers do not uniformly provide the same large-volume video stream to the requesting end; instead, they provide the provider with the target video stream corresponding to the video sub-window. This reduces the amount of data processed, lowers the computational load and cost, and allows for support of more target video streams.
[0065] In one embodiment, the layout information may include the identification information of the video stream provider corresponding to the video sub-window and the size of the video sub-window.
[0066] In one embodiment, obtaining the layout information of the main video session window includes: providing editing controls for video sub-windows on the initial page of the main video session window and displaying identifiers corresponding to the participants; and determining the size of the video sub-window for editing and the identifier information of the corresponding subscription provider based on user operations.
[0067] For example, the requesting end can display the initial page of the main video session window. This initial page can provide editing controls for the video sub-windows and display identification information corresponding to the participants. Users can edit the size of the video sub-windows by manipulating the editing controls. For instance, users can draw the main video session window using the editing controls, creating the size of each video sub-window during the drawing process.
[0068] In one implementation, providing editing controls for video sub-windows on the initial page of the main video session window includes: pre-arranging the display modes of the video sub-windows on the initial page of the main video session window, including picture-in-picture mode or grid mode. For example, the initial page may provide options for the display modes of the video sub-windows, which may include picture-in-picture mode and grid mode. Users can select a display mode according to their preferences or needs. For example, if a user selects a grid mode, the requesting client will display the layout style of the selected display mode after the user's selection. Users can manipulate the layout style of the display mode displayed on the requesting client to adjust the size of the video sub-windows, thereby creating the size of each video sub-window.
[0069] Users can subscribe to providers for video sub-windows by specifying identifier information for each sub-window; each identifier corresponds to one provider. For example, a user can drag and drop a participant's identifier into a video sub-window, thus subscribing that sub-window to a provider. Users can also fill in the participant's identifier into the video sub-window, which also subscribes to a provider for that sub-window.
[0070] For example, each participant has corresponding identification information. By specifying identification information for each video sub-window, the user can assign the corresponding participant's identification information to each video sub-window, thus subscribing to the corresponding provider for that video sub-window. Therefore, the server can instruct the corresponding provider to provide the target video stream based on the identification information in the layout information. In this embodiment, letters such as "A" or "B" are used to represent the provider's identification information; it can be understood that the provider's identification information can be a string.
[0071] In one embodiment, the video stream processing method may further include: receiving user operations, determining the arrangement order of multiple video sub-windows, the layout information further including the arrangement order of the multiple video sub-windows; and providing the layout information to the server.
[0072] Figure 5 This is a schematic diagram of the display screen of the requesting end after user operation in one embodiment of this disclosure. For example, after determining the size of each video sub-window, the user can specify identification information for each video sub-window, with each identification information corresponding to a provider end. For example, the identification information of the participant can be dragged into the video sub-window. Thus, the screen of the participant corresponding to the identification information will be displayed in the video sub-window, thereby determining the arrangement order of multiple video sub-windows.
[0073] For example, users can specify identifiers for each video sub-window and then edit the size of each sub-window. This also allows them to determine the arrangement order of multiple video sub-windows.
[0074] This approach allows users on the requesting end to flexibly choose their preferred display mode for the main video session window and to edit the size and arrangement of each video sub-window, better meeting user needs and preferences and improving user experience.
[0075] For example, the main video session window includes at least two video sub-windows of different sizes. Thus, users can choose to view a larger or smaller view of a particular participant, depending on their preference.
[0076] In this embodiment, the layout information can be presented in the form of a picture or a table. The requesting end can send the picture or table corresponding to the layout information to the server, and the server can parse the layout information of the requesting end from the picture or table. It is understood that in other embodiments, the requesting end can also display the layout information of the main video session window in the form of text, etc. Thus, the requesting end can send the corresponding text to the server, and the server can parse the layout information of the requesting end from the text corresponding to the layout information.
[0077] In one embodiment, the target video stream corresponding to the video sub-window corresponds to video rating information, which is determined based on the size of the video sub-window.
[0078] For example, video rating information can be determined based on the area proportion of the video sub-window within the main video session window. For instance, the main video session window may fill the entire display screen of the requesting client, and the screen resolution is fixed. Based on the area proportion of the video sub-window within the main video session window, the resolution of the video sub-window can be determined, thereby determining the video rating information corresponding to that resolution. For example, the resolution of the video sub-window can be determined based on the ratio of its length and width to the length and width of the main video session window, thus determining the video rating information. In this disclosure, a higher video rating corresponds to a larger resolution and a larger video stream data volume.
[0079] After determining the video rating information of the video sub-window, the server can send the video rating information to the provider, and the provider can then provide the target video stream corresponding to the video rating information.
[0080] In one embodiment, the provider can provide the target video stream to the server.
[0081] For example, obtaining the main window video stream based on layout information may include: sending the layout information to the server; and obtaining the main window video stream generated from at least one target video stream from the server. In other words, the provider provides the target video stream corresponding to the video level information to the server, the server mixes at least one target video stream into a single main window video stream, and provides the main window video stream to the requesting end, which then directly obtains the main window video stream from the server.
[0082] For example, obtaining the main window video stream based on layout information may include: sending the layout information to the server; obtaining the target video streams of each video sub-window from the server; and mixing at least one target video stream into the main window video stream. In this embodiment, the requesting end obtains at least one target video stream from the server and mixes at least one target video stream into the main window video stream.
[0083] In one embodiment, the provider can provide the target video stream to the requesting end. Obtaining the main window video stream based on layout information may include: sending the layout information to the server; obtaining the target video stream from the providers corresponding to each video sub-window; and mixing at least one target video stream with the main window video stream. In this embodiment, after receiving the video level information, the provider, under the instruction of the server, provides the target video stream corresponding to the video level information to the corresponding requesting end. After obtaining the target video stream from at least one provider, the requesting end mixes at least one target video stream into the main window video stream.
[0084] Figure 6 This is a flowchart illustrating a video stream processing method according to another embodiment of the present disclosure. The method is applied to a server, such as a video mixing server. The method includes steps S61 to S62.
[0085] In step S61, the layout information of the video session main window of the requesting end is obtained. The video session main window includes multiple video sub-windows.
[0086] In step S62, a target video stream is provided to the requesting end so that the requesting end can display the main window video stream in the main video session window; the main window video stream is generated based on the target video stream corresponding to at least one video sub-window, and the target video stream is determined based on the layout information of the main video session window.
[0087] In one embodiment, providing a target video stream to a requesting end includes: sending video level information to at least one providing end based on layout information, and instructing the providing end to provide the target video stream based on the video level information. The target video stream is either forwarded to the requesting end by the server or sent to the requesting end by the providing end.
[0088] For example, the layout information includes the identifier of the video stream provider corresponding to the video sub-window, and the size of the video sub-window. The target video stream corresponding to the video sub-window corresponds to video rating information, which is determined based on the size of the video sub-window provided by the provider.
[0089] For example, after obtaining the layout information, the server can determine the video rating information based on the size of the video sub-window. The server then sends the video rating information corresponding to the video sub-window to the provider corresponding to the video sub-window. Upon receiving the video rating information, the provider provides the corresponding target video stream.
[0090] For example, the server can instruct the provider to offer a target video stream, and the server will then forward the target video stream to the requesting client. For instance, instructing the provider to offer a target video stream based on video rating information could include: obtaining the target video stream corresponding to the video rating information returned by the provider according to the instruction, and forwarding the target video stream to the requesting client. In this approach, after obtaining the target video stream, the requesting client mixes at least one target video stream into the main window video stream. The mixing process is completed by the requesting client, and the server forwards the target video stream, which can further reduce the server's computing power and load.
[0091] For example, instructing the provider to provide a target video stream based on video rating information may include: receiving at least one target video stream returned by the provider based on video rating information, and performing a mixing operation on at least one target video stream; and providing the main window video stream generated by the mixing operation to the requesting client. In this approach, the server performs the mixing operation, which can reduce the load and computing power on the requesting client.
[0092] Table 1 shows... Figure 3 The video streams requested by requester K and requester A
[0093]
[0094] Note: ① indicates that the video rating information is level one, and ② indicates that the video rating information is level two.
[0095] For example, performing a mixing operation on at least one target video stream should be understood as performing a mixing operation on the target video stream returned by the provider in response to a request from the requesting end. For instance, if both requesting end K and requesting end A submit the layout information of the main video session window to the server, then the providers include providers A, B, C, D, E, F, G, H, M, N, J, and K. Table 1 shows the video level information of the target video streams provided by each provider in response to a request from requesting end K, and also shows the video level information of the target video streams provided by each provider in response to a request from requesting end A.
[0096] For requesting client K, providing clients A, B, C, D, E, F, G, H, M, and N provide target video streams corresponding to the video rating information. After receiving each target video stream, the server performs a mixing operation on the multiple target video streams and generates a main window video stream, which is then sent to requesting client K. Alternatively, requesting client K receives each target video stream and performs a mixing operation on the multiple target video streams to generate the main window video stream.
[0097] For requesting client A, providing clients B, C, D, F, G, H, M, N, J, and K provide target video streams corresponding to the video rating information. After receiving each target video stream, the server performs a mixing operation on the multiple target video streams and generates a main window video stream, which is then sent to requesting client A. Alternatively, requesting client A receives each target video stream and performs a mixing operation on the multiple target video streams to generate the main window video stream.
[0098] For example, the server can instruct the provider to offer the target video stream to the requesting client. Instructing the provider to offer the target video stream based on video rating information can include instructing the provider to offer the target video stream to the requesting client based on the video rating information of the target video stream. In this way, the server does not need to receive the target video stream, nor does it need to perform mixing operations; instead, it instructs the provider to offer the target video stream to the requesting client, which can further reduce the server's computing power and load.
[0099] In this embodiment, the video level information can reflect the video stream data volume. For example, a first-level video stream can be a video stream with a first data volume, and a second-level video stream can be a video stream with a second data volume, where the second data volume is greater than the first data volume. In Table 1, the first level is represented by "①" and the second level is represented by "②".
[0100] For example, a first-level video stream can be a 180p resolution video stream, and a second-level video stream can be a 720p resolution video stream.
[0101] For example, the video rating information is matched to the size of the video sub-window on the providing end. For example, Figure 2 In the above scenario, if the size of the video sub-window that client B subscribes to is large, client A needs to provide the target video stream corresponding to the large-sized video sub-window. The video stream level information sent by the server to client A can be the second level. If the size of the video sub-window that client B subscribes to is small, client C needs to provide the target video stream corresponding to the small-sized video sub-window. The video stream level information sent by the server to client C can be the first level.
[0102] For example, when the video rating information received by the provider is at level 1, the provider provides the target video stream corresponding to level 1; when the video rating information received by the provider is at level 2, the provider provides the target video stream corresponding to level 2. Different providers can provide video streams of different ratings.
[0103] For example, there can be one or more requesting clients, and one requesting client can request one or more providers. Each requesting client can send the layout information of the video session main window to the server at any time to request a video stream. Based on the layout information of the requesting client's video session main window, the server sends video level information to at least one provider and instructs the provider to provide the target video stream corresponding to the video level information.
[0104] In one embodiment, different requesting clients can request different levels of video streams from the same provider. For example... Figure 3 In the layout information of requester K, the video sub-window corresponding to B is of the second size; therefore, the video level information of B subscribed by requester K is of the second level. In the layout information of requester A, the video sub-window corresponding to B is of the first size; therefore, the video level information of B subscribed by requester A is of the first level. After obtaining the layout information of requester K and requester A, the server sends the second-level and first-level video level information to provider B. For example, provider B can provide a target video stream of the first level and a target video stream of the second level. Requester K obtains the target video stream of the second level from provider B, and requester A obtains the target video stream of the first level from provider B.
[0105] In one embodiment, provider B provides both the first-level target video stream and the second-level target video stream to the server. The server provides the second-level target video stream of provider B to the requesting client K according to the layout information of the requesting client K. The server provides the first-level target video stream of provider B to the requesting client A according to the layout information of the requesting client A.
[0106] In one embodiment, sending video level information to at least one provider based on layout information may include: determining at least one provider based on the identification information of the video sub-window; determining the video level information based on the size of the video sub-window; and sending the corresponding video level information to the provider corresponding to the video sub-window. In the layout information, the video sub-window has identification information corresponding to the provider. After receiving the layout information, the server determines the provider corresponding to the video sub-window based on the identification information and determines the corresponding video level information based on the size of the video sub-window.
[0107] In one embodiment, the layout information of the main video session window is different for at least two requesting ends. For example... Figure 3 In this scenario, the layout information of the main video session window sent by requester K and requester A to the server is different. The server provides the target video stream to requester K based on its layout information and to requester A based on its layout information. Because the layout information of requester K and requester A is different, the main window video streams displayed by requester K and requester A will look different, such as... Figure 3 As shown.
[0108] In practical applications, when setting the main video session window, the user on the requesting end may choose to include a number of video sub-windows that are greater than the number of providers subscribed to by the requesting end. In one embodiment, providing the target video stream to the requesting end further includes: if the main video session window includes idle video sub-windows, then providing the requesting end with a filler video stream that matches the size of the idle video sub-windows.
[0109] For example, an idle video sub-window can be understood as the requesting end not having subscribed to the provider's video sub-window when setting up the main video session window.
[0110] Mixing at least one target video stream may include: if the main video session window includes an idle video sub-window, selecting a filler video stream from a pool of candidate video streams whose size matches that of the idle video sub-window; and mixing the filler video stream with at least one target video stream.
[0111] Figure 7A The video session main window layout information set for the requesting client M. Figure 7B This is a schematic diagram illustrating the effect of displaying the main window video stream to the requesting client M. For example... Figure 7A As shown, the layout information of the video session main window of the requesting end M includes three video sub-windows. Two of the video sub-windows are subscribed to by the providers A and B respectively, and the third video sub-window is not subscribed to by any provider. Therefore, the third video sub-window is an idle video sub-window.
[0112] For example, requesting client M enters the video session after requesting clients K and A. Requesting client M sends the layout information of the main window of the video session to the server. For an idle video sub-window a, the server can select a filler video stream from the acquired candidate video streams that match the size of the idle video sub-window a. For example, the candidate video streams include the target video stream provided by provider C to requesting client A and the target video stream provided by provider K to requesting client A. The server can select either the second-level target video stream provided by provider C or the second-level target video stream provided by provider K as the filler video stream; for example, it can select the second-level target video stream provided by provider K as the filler video stream. After mixing the filler video stream with at least one target video stream, a main window video stream is generated, which is then displayed by requesting client M. The effect of the main window video stream displayed by requesting client M is as follows: Figure 7B As shown.
[0113] In one embodiment, the layout information also includes the arrangement order of the video sub-windows corresponding to multiple providers; the target video streams corresponding to multiple providers are sorted according to the arrangement order of the video sub-windows indicated by the layout information.
[0114] For example, performing a mixing operation on at least one target video stream includes: sorting multiple target video streams according to the arrangement order of video sub-windows indicated by layout information; and performing a mixing operation on the sorted multiple target video streams.
[0115] For example, for requester K, after receiving the target video streams from providers A, B, C, D, E, F, G, H, M, and N, the server or requester K sorts the multiple target video streams according to the arrangement order of the video sub-windows. Then, it performs a mixing operation on the sorted target video streams to generate the main window video stream. Requester K displays the main window video stream; the screen displayed by requester K is as follows: Figure 3 As shown.
[0116] In one implementation, the server stores a subscription status information table, which records the video rating information of the target video stream requested by the requesting client. For example, the server may store the subscription status information table in the form of a subscription status information table. After obtaining the layout information of the requesting client's main video session window, the server records the content of the layout information in the subscription status information table according to the layout information. Alternatively, the server may create the subscription status information table based on the layout information after obtaining the layout information of the requesting client's main video session window. Table 2 shows a subscription status information table in one embodiment of this disclosure.
[0117] Table 2 is a subscription status information table in one embodiment of this disclosure.
[0118]
[0119] Note: ① indicates the video rating is Level 1, ② indicates the video rating is Level 2, and × indicates the requesting end did not request the corresponding video stream from the provider.
[0120] As shown in Table 2, the subscription status information table records the requesting end, the providers corresponding to each video sub-window in the requesting end's video session main window, and the video level information of the target video stream requested by the requesting end. For example, requesting end B requests video streams from providers A, C, D, F, and F. The video level information requested by requesting end B from providers A, C, D, F, and F are ②, ①, ①, ①, and ①, respectively. Providers A, C, D, F, and F provide the target video streams corresponding to the video level information to requesting end B, or providers A, C, D, F, and F provide the target video streams corresponding to the video level information to the server, and the server forwards the target video streams to requesting end B. Requesting end B did not request video streams from B, G, H, M, N, J, and K.
[0121] As can be seen from the subscription status information table in Table 2, requesting ends K, A, and E all request the video stream from providing end M. Requesting end K requests the first-level target video stream from providing end M, requesting end A requests the first-level target video stream from providing end M, and requesting end E requests the second-level target video stream from providing end M.
[0122] It should be noted that the subscription status information table is not limited to the form shown in Table 2. It can also be other forms of information table, as long as it records layout-related information.
[0123] When the provider directly offers the target video stream to the requesting client, if the provider can offer a video stream that meets the highest video stream quality standard, the provider can process the video stream and offer the corresponding target video stream to the requesting client based on the video quality standard. For example, in Table 2, provider M can offer the highest video stream quality. Provider M can process the acquired video stream and offer a first-level target video stream to both requesting clients K and A, and a second-level target video stream to requesting client E. In this way, the server instructs the provider to offer the corresponding target video stream to the requesting client based on the video quality standard. The server does not need to process the video stream, further reducing the server's load and computing power.
[0124] In practice, if the provider performs video stream processing and delivers different target video streams of different levels to different requesting clients based on video rating information, this will increase the provider's load and network bandwidth requirements. It's understandable that providers are typically terminal devices with limited load and network bandwidth capabilities; if they perform video stream processing, it will usually cause buffering or stuttering on the provider's device.
[0125] In one implementation, sending video level information to at least one provider based on layout information includes: querying a subscription status information table, determining the highest-level video level information from the video level information corresponding to multiple requesting clients requesting the same target video stream, and sending the determined highest-level video level information to the provider.
[0126] For example, in Table 2, requesting clients K, A, and E all request the target video stream from provider M. Querying the subscription status information table, requesting client K requests the target video stream from provider M with a first-level video rating, requesting client A requests the target video stream from provider M with a first-level video rating, and requesting client E requests the target video stream from provider M with a second-level video rating. From the video rating information corresponding to requesting clients K, A, and E requesting the target video stream from provider M, the highest-level video rating is determined to be the second-level. Therefore, the server sends the determined highest-level video rating to provider M, which is the second-level video rating. After receiving the second-level video rating information, provider M provides the target video stream corresponding to the second-level video rating to the server.
[0127] In this way, the provider does not need to provide multiple different target video streams, but only needs to provide one level of target video stream, thereby reducing the load on the provider and network bandwidth.
[0128] In one implementation, providing the target video stream to the requesting end further includes: if the highest determined video level information is greater than the video level information corresponding to the layout information of the requesting end; then before providing the target video stream to the requesting end, processing the obtained video stream corresponding to the highest determined video level information to obtain the target video stream corresponding to the layout information of the requesting end.
[0129] For example, in Table 2, provider M provides the target video stream corresponding to the second level to the server. For requesting client E, requesting client E requests the target video stream of the second level from provider M, the server can directly provide the target video stream of the second level from provider M to requesting client E.
[0130] For requester K, requesting a first-level target video stream from provider M, if the second level is higher than the first, then before providing the target video stream from provider M to requester K, the server processes the acquired second-level target video stream from provider M to obtain the first-level target video stream from provider M. For example, the server can segment the second-level target video stream from provider M to obtain the first-level target video stream from provider M. Then, the server provides the first-level target video stream from provider M to requester K.
[0131] For requester A, the server will provide the first-level target video stream from provider M to requester A.
[0132] In practice, after the provider receives the highest video level information determined by the server, the provider may not be able to provide the server with the target video stream corresponding to the highest video level information.
[0133] In one embodiment, upon receiving the highest-level video rating information, the provider can compare its maximum video streaming capacity with the highest-level video rating information. If the provider's maximum video streaming capacity is greater than or equal to the highest-level video rating information, the provider provides the server with the target video stream corresponding to the highest-level video rating information. If the provider's maximum video streaming capacity is less than the highest-level video rating information, the provider provides the server with the target video stream corresponding to its maximum capacity.
[0134] For example, in Table 2, the server sends the second-level video stream to the provider M. If the maximum video stream capacity of provider M is the second level or greater, then provider M provides the server with the target video stream corresponding to the second level. If provider M's device capacity is limited, and its maximum video stream capacity is the first level, then provider M provides the server with the target video stream corresponding to the first level. Correspondingly, the server provides the requesting client E with the target video stream corresponding to the first level.
[0135] In one implementation, the video stream processing method on the server side may further include: updating the subscription status information table when a newly submitted layout information is obtained or a request from a client to exit a multi-person video session is received.
[0136] For example, the newly submitted layout information may be the new layout information resubmitted to the server by the requesting end that has already requested the video stream after changing the main window of the video session, or it may be the layout information of the main window of the video session sent to the server by a new requesting end.
[0137] When the server receives newly submitted layout information, it updates the subscription status information table. For example, client C submits the layout information of the main video session window to the server, requesting the first-level target video stream from providers A and C, and the second-level target video stream from providers B and D. The server updates the subscription status information table based on the layout information from client C. Table 3 shows the updated subscription status information table.
[0138] Table 3 shows the subscription status information after the requesting client C submits layout information in one embodiment of this disclosure.
[0139]
[0140] Note: ① indicates the video rating is Level 1, ② indicates the video rating is Level 2, and × indicates the requesting end did not request the corresponding video stream from the provider.
[0141] In one embodiment, on the server side, if the level of the first provider in the newly submitted layout information request is lower than the maximum level of the first provider already existing in the subscription status information table, the server processes the target video stream of the maximum level of the existing first provider to obtain the target video stream corresponding to the newly submitted layout information and provides it to the corresponding requesting client. That is, if the level of the target video stream of the first provider in the newly submitted layout information request is lower than the maximum level of the target video stream of the first provider already existing in the subscription status information table, the server processes the target video stream of the maximum level of the existing first provider to obtain the target video stream of the first provider corresponding to the video level information in the newly submitted layout information and provides it to the requesting client.
[0142] For example, after obtaining the layout information of the requesting client C, the server determines that the highest level of the target video stream of provider A already existing in the subscription status information table is the second level, and the first level is lower than the second level, for the target video stream of provider A requested by client C at the first level. In this case, the server does not need to send the first level video stream information to provider A. Instead, the server processes the existing second-level target video stream of provider A to obtain the first-level target video stream of provider A, and then provides the first-level target video stream of provider A to the requesting client C.
[0143] In one embodiment, on the server side, if the level of the target video stream of the first provider in the newly submitted layout information request is equal to the level of the target video stream of the first provider that already exists in the subscription status information table, the server provides the target video stream of the existing first provider to the requesting party.
[0144] For example, if the server determines that the second-level target video stream of provider B already exists in the subscription status information table, the server does not need to send the second-level video level information to provider B again. Instead, the server can directly provide the existing second-level target video stream of provider B to the requesting client C.
[0145] For example, if requester C requests a first-level target video stream from provider C, the server can directly provide the existing first-level target video stream from provider C to requester C. Alternatively, when requester C requests its own video stream, the corresponding video sub-window of requester C's video session main window can directly display its own image without interacting with the server.
[0146] In one embodiment, on the server side, if the level of the first provider in the newly submitted layout information request is greater than the maximum level of the first provider already existing in the subscription status information table, the server sends the video level information of the first provider in the newly submitted layout information to the first provider, so that the first provider can provide the corresponding target video stream. That is, if the level of the target video stream in the newly submitted layout information request is greater than the maximum level of the target video stream of the first provider already existing in the subscription status information table, the server sends the video level information of the first provider in the newly submitted layout information to the first provider, so that the first provider can provide the corresponding target video stream.
[0147] For example, for a second-level target video stream requested by client C from provider D, the server determines that the highest level of the target video stream already existing in provider D's subscription status information table is the first level, and the second level is higher than the first level. In this case, the server needs to send the second-level video level information to provider D so that provider D can provide the second-level target video stream to the server. After receiving the second-level target video stream from provider D, the server provides the second-level target video stream from provider D to client C.
[0148] When the server receives a request from a client to exit a multi-user video session, the server updates the subscription status information table. For example, if client A in Table 2 exits the multi-user video session, the server updates the subscription status information table, as shown in Table 4.
[0149] Table 4 shows the subscription status information after requester A exits in one embodiment of this disclosure.
[0150]
[0151] Note: ① indicates the video rating is Level 1, ② indicates the video rating is Level 2, and × indicates the requesting end did not request the corresponding video stream from the provider.
[0152] After requesting client A exits the multi-user video session, the target video stream required by provider C changes from level two to level one, and the target video stream required by provider K also changes from level two to level one. The target video streams required by providers B, D, F, G, H, M, N, and J are not affected by requesting client A's exit.
[0153] In one embodiment, when a request is received from the second requesting party to exit the multi-person video session, the video level information is resent to the corresponding provider according to the updated subscription status information table, so that the provider provides the target video stream corresponding to the newly received video level information.
[0154] For example, in Table 4, after requester A exits the multi-user video session, the maximum video stream level of provider C changes from level two to level one, and the maximum video stream level of provider K also changes from level two to level one. The server resends the video stream level information to the corresponding providers based on the updated subscription status information table. Specifically, the server resends level one video stream to provider C and provider K. Providers C and K then provide the target video streams corresponding to the level one video stream.
[0155] In practical applications, a provider may leave a video session for some reason, and the server will no longer be able to receive the video stream from that provider. For example, when a provider leaves a multi-person video session, the corresponding video sub-window of the provider in the requesting end may be displayed as a black screen or a white screen.
[0156] In one implementation, when multiple requesters request the same target video stream, the server can provide the highest-ranking target video stream to each requester.
[0157] For example, in Table 2, requesting clients K, A, and E all request the target video stream from provider M. Requesting client K requests the first-level video rating information for the target video stream from provider M, requesting client A requests the first-level video rating information for the target video stream from provider M, and requesting client E requests the second-level video rating information for the target video stream from provider M. From the video rating information corresponding to requesting clients K, A, and E requesting the target video stream from provider M, the highest-level video rating information is determined to be the second-level. Therefore, the server sends the determined highest-level video rating information to provider M as the second-level. After receiving the second-level video rating information, provider M provides the target video stream corresponding to the second-level to the server. The server then provides the target video stream corresponding to the second-level of provider M to requesting clients K, A, and E.
[0158] In this approach, the server only needs to receive a target video stream of a certain level from the provider and then provide the received target video stream to each requesting client. The server does not need to process the received target video stream according to the video level information of the requesting client, which further reduces the server's load and computing power.
[0159] After receiving the target video stream, if the target video stream's resolution is higher than the requested resolution, the requesting client can compress the target video stream before displaying it. Understandably, since the target video stream received by the requesting client has a higher resolution, compression and display will not reduce the resolution of the corresponding video sub-window, thus ensuring image clarity.
[0160] For example, Figure 8 This is a schematic diagram showing the main window video stream displayed on requesting clients K, A, and E respectively, as follows: Figure 8 As shown, requesting clients K, A, and E simultaneously request video streams from provider M. The server provides the determined second-level target video stream to requesting clients K, A, and E. Since the target video stream requested by requesting clients K and A from provider M has a lower level than the second level, requesting clients K and A compress the received second-level target video stream from provider M and display it in the corresponding video sub-window, as shown. Figure 8 As shown.
[0161] In one implementation, the subscription status information table may record the subscription quantity of each provider, the number of target video streams of each level for each provider, and the target video level information of the target video streams provided by the providers to the server. Here, the highest-level video level information determined from the multiple video level information of the target video stream requesting provider C is the target video level information. Table 5 is a subscription status information table in another embodiment of this disclosure, and Table 5 corresponds to Table 2.
[0162] Table 5 Subscription Status Information Table in Another Embodiment of This Disclosure
[0163] Provider Target video rating information Second level First level A Second level 3 0 B Second level 2 2 C Second level 1 4 D First level 0 5 E First level 0 3 F Second level 1 4 G First level 0 3 H First level 0 3 M Second level 1 2 N First level 0 3 J First level 0 2 K Second level 1 1
[0164] As shown in Table 5, the subscription status information table records each provider, the number of video streams requested for each provider at different levels, and the level of the target video stream currently being provided by the provider to the server. For example, for provider C, the number of clients requesting the first-level target video stream is 4, and the number of clients requesting the second-level target video stream is 1. Table 5 determines that the target video stream level for provider C is second-level. Therefore, provider C provides the second-level target video stream to the server, and the server provides the second-level target video stream to clients K, A, E, B, and J respectively.
[0165] Whether it is obtaining newly submitted layout information or receiving a request from a client to exit a multi-user video session, the server will update the subscription status information table. The server can consult the subscription status information table to re-obtain the target video level information of each provider and send the target video level information to the provider so that the provider can provide the server with the target video stream corresponding to the target video level information.
[0166] For example, when client C submits new layout information to the server and updates the subscription status information table, the subscription status information table is shown in Table 6. Table 6 is the subscription status information table after client C joins the service.
[0167] Table 6 shows the subscription status information after the requesting client C submits layout information in one embodiment of this disclosure.
[0168] Provider Target video rating information Second level First level A Second level 3 1 B Second level 3 2 C Second level 1 5 D Second level 1 5 E First level 0 3 F Second level 1 4 G First level 0 3 H First level 0 3 M Second level 1 2 N First level 0 3 J First level 0 2 K Second level 1 1
[0169] As can be seen from Table 6, the number of subscriptions for the first level of provider A was updated from 0 to 1, the number of subscriptions for the second level of provider B was updated from 2 to 3, the number of subscriptions for the first level of provider C was updated from 4 to 5, the number of subscriptions for the second level of provider D was updated from 2 to 1, and the target video level of provider C was updated from the first level to the second level.
[0170] In one embodiment, the server may store the link addresses of each provider, and these link addresses may contain the target video stream of that provider. When the requesting client sends layout information to the server, the server can provide the requesting client with the link address of the requested provider based on the layout information. The requesting client then obtains the target video stream through the link address.
[0171] The technical solutions of this disclosure can be applied to video sessions with more participants, such as more than 100 participants in the same video session. In related technologies, for video conferences with more than 100 participants, since the server provides high-resolution (e.g., 720p) video streams to each requesting end, and the requesting end simultaneously receives high-resolution video streams from more than 100 people, the simultaneous display of 100 people in the same frame consumes a large amount of network bandwidth and CPU usage.
[0172] In this disclosed technical solution, the requesting end can request video streams from more than one hundred providers. The main video session window of the requesting end can include more than one hundred small-sized video sub-windows, and the video level information corresponding to the small-sized video sub-windows is the first level. Therefore, the data volume of the target video stream provided by each provider is the data volume corresponding to the first level (e.g., the data volume for a resolution of 180p). After the server mixes more than one hundred target video streams, it sends the main window video stream to the requesting end. After the requesting end displays the main window video stream, the user on the requesting end can view the video session screen with one hundred people in the same frame, with 100 video streams centrally arranged on one screen. Therefore, the technical solution of this disclosed embodiment can realize one hundred people in the same frame and can support one hundred people mixing with extremely low load, with CPU usage approximately 20% of that in related technologies for one hundred people in the same frame; it also reduces network bandwidth requirements, with network bandwidth usage approximately 10% of that in related technologies for one hundred people in the same frame scenario.
[0173] In this article, the video sub-window is sized in two ways: large and small. The large size is, for example... Figure 3 The requesting end A provides the size of the video sub-window of end C, the smaller size being, for example... Figure 3 The requesting end A provides the size of the video sub-window from end B. A larger size corresponds to the second-level video stream quality information, with a data volume of 720p resolution; a smaller size corresponds to the first-level video stream quality information, with a data volume of 180p resolution. It is understood that in other embodiments, the size of the video sub-window is not limited to... Figure 3 While the video sub-window is limited to two sizes, its size can be set to many more. Correspondingly, the video rating information is not limited to the first and second ratings, but can include more ratings. The video rating information can be matched with the size of the video sub-window as needed.
[0174] The larger the size of the video sub-window, the larger the amount of data corresponding to the matched video rating information. For example, a first size threshold and a second size threshold can be set for the size of the video sub-window, with the second size threshold being greater than the first size threshold. When the size of the video sub-window is less than or equal to the first size threshold, the matched video rating information can be of the first level; when the size of the video sub-window is greater than the first size threshold but less than or equal to the second size threshold, the matched video rating information can be of the second level; and when the size of the video sub-window is greater than the second size threshold, the matched video rating information can be of the third level. This limits the video rating of the target video stream, reduces the number of rating types in the target video stream, and reduces the computing power required by the server.
[0175] Figure 9 This is a schematic diagram of the main video session window provided in one embodiment of the present disclosure. In one embodiment, the main video session window includes a grid mode, such as... Figure 3 As shown, the video sub-windows are arranged in a grid. In another embodiment, the main video session window may include a picture-in-picture mode, such as... Figure 9 As shown, the screen of provider A is displayed in a large window, while the screens of providers B, C, and D are displayed in smaller windows, and the screens of providers B, C, and D are located within the screen of provider A.
[0176] In one implementation, the video stream processing method is applied to the providing end. The video stream processing method may include: providing a target video stream corresponding to a video sub-window to the requesting end according to the instruction of the server, so that the requesting end can display the main window video stream in the main video session window; the main window video stream is generated based on the target video stream corresponding to at least one video sub-window, and the target video stream is determined based on the layout information of the main video session window.
[0177] In one embodiment, providing the target video stream according to the server's instructions includes: providing the requesting client with the target video stream corresponding to the received video rating information; the video rating information is determined based on the size of the video sub-window.
[0178] For example, after receiving the video rating information sent by the server, the provider provides the server with the target video stream corresponding to the video rating information, and the server provides the target video stream to the requesting client. Alternatively, after receiving the video rating information sent by the server, the provider pushes the target video stream corresponding to the video rating information in the layout information to the requesting client.
[0179] It should be noted that the video stream may include at least one of audio data and video data. The audio acquisition device at the providing end can acquire the voice of the user at the providing end and generate audio data. The camera device at the providing end can acquire the image of the user at the providing end and generate video data.
[0180] For example, when the provider cannot capture the user's voice, it can provide video data corresponding to the video rating information. The requesting end cannot hear the user's voice but can view the user's image. Similarly, when the provider cannot capture the user's image, it can provide audio data. The requesting end cannot view the user's image but can hear the user's voice.
[0181] For example, the requesting end and / or the providing end can be terminal devices such as mobile phones, computers, and conference room equipment.
[0182] Figure 10 This is a structural block diagram of a video stream processing apparatus according to one embodiment of this disclosure. Corresponding to the application scenarios and methods provided in the embodiments of this disclosure, this disclosure also provides a video stream processing apparatus applied to a requesting end, such as... Figure 10 As shown, the video stream processing device may include a first layout information acquisition module 101, a main window video stream acquisition module 102, and a display module 103.
[0183] The first layout information acquisition module 101 is used to acquire the layout information of the main video session window, which includes multiple video sub-windows.
[0184] The main window video stream acquisition module 102 is used to acquire the main window video stream based on the layout information.
[0185] The display module 103 is used to display the main window video stream in the main video session window. The main window video stream is generated based on the target video stream corresponding to at least one video sub-window. The target video stream is determined based on the layout information of the main video session window.
[0186] In one embodiment, the first layout information acquisition module 101 is used to: provide editing controls for video sub-windows on the initial page of the main video session window and display the identifiers corresponding to the participants; and determine the size of the video sub-window for editing and the identifier information of the corresponding subscription provider based on user operations.
[0187] In one embodiment, the video stream processing apparatus further includes: an order determination module, configured to receive user operations and determine the arrangement order of multiple video sub-windows, wherein the layout information further includes the arrangement order of the multiple video sub-windows; and a layout information providing module, configured to provide the layout information to the server.
[0188] Figure 11 This is a structural block diagram of a video stream processing apparatus according to another embodiment of this disclosure. Corresponding to the application scenarios and methods provided in the embodiments of this disclosure, this disclosure also provides a video stream processing apparatus applied to a server, such as... Figure 11As shown, the video stream processing device may include a second layout information acquisition module 111 and a first video stream providing module 112.
[0189] The second layout information acquisition module 111 is used to acquire the layout information of the video session main window of the requesting end. The video session main window includes multiple video sub-windows.
[0190] The first video stream providing module 112 is used to provide a target video stream to the requesting end so that the requesting end can display the main window video stream in the main window of the video session; the main window video stream is generated based on the target video stream corresponding to at least one video sub-window, and the target video stream is determined based on the layout information of the main window of the video session.
[0191] In one embodiment, the first video stream providing module 112 is configured to: send video level information to at least one providing end according to layout information, and instruct the providing end to provide a target video stream according to the video level information, wherein the target video stream is forwarded by the server to the requesting end or sent by the providing end to the requesting end.
[0192] For example, the first video stream providing module 112 is further configured to: receive at least one target video stream returned by a providing end according to video level information, and perform a mixing operation on at least one target video stream; and provide the main window video stream generated by the mixing to the requesting end.
[0193] For example, the server stores a subscription status information table, which records the video rating information of the target video stream requested by the requesting end. The first video stream providing module 112 is further configured to: query the subscription status information table, determine the highest-rated video rating information from the video rating information corresponding to multiple requesting ends requesting the same target video stream; and send the determined highest-rated video rating information to the providing end, so that the providing end provides the target video stream with the determined highest rating.
[0194] For example, the first video stream providing module 112 is further configured to: if the determined highest video level information is greater than the video level information corresponding to the layout information of the requesting end; then before providing the target video stream to the requesting end, process the obtained video stream corresponding to the highest video level information to obtain the target video stream corresponding to the layout information of the requesting end.
[0195] In one embodiment, the apparatus further includes an information table update module, used to update the subscription status information table when a newly submitted layout information is obtained or a request from a client to exit a multi-person video session is received.
[0196] In one embodiment, the first video stream providing module 112 is further configured to: if the main video session window includes an idle video sub-window, provide the requesting end with a filler video stream that matches the size of the idle video sub-window.
[0197] This disclosure also provides a video stream processing apparatus applied to a providing end. The apparatus includes: a second video stream providing module, configured to provide a target video stream corresponding to a video sub-window to a requesting end according to an instruction from a server, so that the requesting end can display the main window video stream in the main video session window; the main window video stream is generated based on the target video stream corresponding to at least one video sub-window, and the target video stream is determined based on the layout information of the main video session window.
[0198] For example, the second video stream providing module is used to provide the requesting end with a target video stream corresponding to the received video rating information. The video rating information is determined based on the size of the video sub-window.
[0199] Figure 12 This is a block diagram of an electronic device used to implement embodiments of the present disclosure. For example... Figure 12 As shown, the electronic device includes a memory 1210 and a processor 1220. The memory 1210 stores a computer program that can run on the processor 1220. When the processor 1220 executes the computer program, it implements the method described in the above embodiments. The number of memories 1210 and processors 1220 can be one or more.
[0200] The electronic device also includes:
[0201] The communication interface 1230 is used to communicate with external devices and exchange and transmit data.
[0202] If the memory 1210, processor 1220, and communication interface 1230 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 12 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0203] Optionally, in a specific implementation, if the memory 1210, processor 1220 and communication interface 1230 are integrated on a single chip, the memory 1210, processor 1220 and communication interface 1230 can communicate with each other through an internal interface.
[0204] This disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods provided in this disclosure.
[0205] This disclosure also provides a chip, which includes a processor for calling and executing instructions stored in a memory, causing a communication device on which the chip is installed to perform the methods provided in this disclosure.
[0206] This disclosure also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method provided in the application embodiment.
[0207] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.
[0208] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0209] In the above embodiments, implementation can be achieved, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to this disclosure is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.
[0210] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0211] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.
[0212] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process. Furthermore, the scope of the preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functionality involved.
[0213] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).
[0214] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware, the program being stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiments.
[0215] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a disk, or an optical disk, etc.
[0216] The above are merely specific embodiments of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in this disclosure, and these should all be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A video stream processing method, characterized in that, Applied to the requesting end, the method includes: The layout information of the main video session window is obtained and sent to the server so that the server can issue an instruction to the provider; the main video session window includes multiple video sub-windows; the layout information is created by the user of the requesting end. The main window video stream provided by the provider according to the instructions is obtained based on the layout information. The main window video stream is displayed in the main video session window, so that the user on the requesting end can see the subscribed provider's screen. The main window video stream is generated based on the target video stream corresponding to at least one of the video sub-windows. The target video stream is determined based on the layout information of the main video session window. The layout information includes the identifier information of the video stream provider corresponding to the video sub-window, as well as the size and arrangement order of the video sub-window; the target video stream corresponding to the video sub-window corresponds to the video level information, which is determined according to the size of the video sub-window.
2. The method according to claim 1, characterized in that, The video rating information is determined based on the area ratio of the video sub-window in the main video session window.
3. The method according to claim 1, characterized in that, The process of obtaining the layout information of the main video session window includes: The initial page of the main video session window provides editing controls for the video sub-windows and displays identifiers corresponding to the participants; Based on user actions, the size of the video sub-window to be edited and the identification information of the corresponding subscription provider are determined.
4. A video stream processing method, characterized in that, Applied to the server side, the method includes: Obtain the layout information of the requesting end's main video session window, which includes multiple video sub-windows; the layout information is created by the user of the requesting end. An instruction is sent to the provider to provide the target video stream to the requesting end, so that the requesting end can display the main window video stream in the main window of the video session, so that the user of the requesting end can view the subscribed provider's screen; the main window video stream is generated based on the target video stream corresponding to at least one of the video sub-windows, and the target video stream is determined based on the layout information of the main window of the video session; The layout information includes the identification information of the video stream provider corresponding to the video sub-window, as well as the size and arrangement order of the video sub-window, the target video stream corresponding to the video sub-window and the video level information, and the video level information is determined according to the size of the video sub-window of the provider.
5. The method according to claim 4, characterized in that, The step of providing the target video stream to the requesting end includes: Based on the layout information, video level information is sent to at least one provider, and the provider is instructed to provide the target video stream based on the video level information. The target video stream is forwarded by the server to the requesting end or sent by the provider to the requesting end.
6. The method according to claim 5, characterized in that, The server stores a subscription status information table, which records the video rating information of the target video stream requested by the requesting end.
7. The method according to claim 6, characterized in that, Sending video level information to at least one provider based on the layout information includes: Query the subscription status information table to determine the highest video level information from the video level information corresponding to multiple requesting ends that request the same target video stream; Send the highest determined video level information to the provider so that the provider can provide the target video stream of the highest determined level.
8. The method according to claim 7, characterized in that, The step of providing the target video stream to the requesting end also includes: If the highest determined video level information is greater than the video level information corresponding to the layout information of the requesting end; Before providing the target video stream to the requesting end, the video stream corresponding to the highest video level information is processed to obtain the target video stream corresponding to the layout information of the requesting end.
9. The method according to claim 6, characterized in that, The method further includes at least one of the following steps: If the level of the first provider in the newly submitted layout information request is lower than the maximum level of the first provider already existing in the subscription status information table, the target video stream of the maximum level of the existing first provider is processed to obtain the target video stream corresponding to the newly submitted layout information and provided to the corresponding requesting end. If the level of the first provider in the newly submitted layout information request is greater than the maximum level of the first provider already existing in the subscription status information table, the video level information of the first provider in the newly submitted layout information is sent to the first provider so that the first provider can provide the corresponding target video stream. When a request is received from the second requesting party to exit the multi-person video session, the video level information is resent to the corresponding provider based on the updated subscription status information table, so that the provider provides the target video stream corresponding to the newly received video level information.
10. A video stream processing method, characterized in that, Applied to the provider, the method includes: providing a target video stream corresponding to a video sub-window to the requesting end according to an instruction from the server, so that the requesting end can display the main window video stream in the main video session window, enabling the user of the requesting end to view the subscribed provider's screen; the main window video stream is generated based on the target video stream corresponding to at least one of the video sub-windows, and the target video stream is determined based on the layout information of the main video session window; the layout information is created by the user of the requesting end; the instruction is issued by the server. The layout information includes the identification information of the video stream provider corresponding to the video sub-window, as well as the size and arrangement order of the video sub-window; the target video stream corresponding to the video sub-window corresponds to the video level information, which is determined according to the size of the video sub-window of the provider.
11. A video stream processing apparatus, characterized in that, Applied to the requesting end, the device includes: The first layout information acquisition module is used to acquire the layout information of the main video session window and send the layout information to the server so that the server can issue an instruction to the provider; the main video session window includes multiple video sub-windows; the layout information is created by the user of the requesting end; The main window video stream acquisition module is used to acquire the main window video stream provided by the provider according to the instruction based on the layout information. The display module is used to display the main window video stream in the main window of the video session, so that the user of the requesting end can view the subscribed provider's screen. The main window video stream is generated based on the target video stream corresponding to at least one of the video sub-windows. The target video stream is determined based on the layout information of the main window of the video session. The layout information includes the identification information of the video stream provider corresponding to the video sub-window, as well as the size and arrangement order of the video sub-window; The target video stream corresponding to the video sub-window corresponds to the video rating information, which is determined based on the size of the video sub-window.
12. A video stream processing apparatus, characterized in that, Applied to the server side, the device includes: The second layout information acquisition module is used to acquire the layout information of the main video session window of the requesting end, wherein the main video session window includes multiple video sub-windows; the layout information is created by the user of the requesting end. The first video stream providing module is used to send an instruction to the providing end to provide a target video stream to the requesting end, so that the requesting end can display the main window video stream in the main window of the video session, so that the user of the requesting end can watch the subscribed providing end screen; the main window video stream is generated according to the target video stream corresponding to at least one of the video sub-windows, and the target video stream is determined according to the layout information of the main window of the video session; The layout information includes the identification information of the video stream provider corresponding to the video sub-window, as well as the size and arrangement order of the video sub-window, the target video stream corresponding to the video sub-window and the video level information, and the video level information is determined according to the size of the video sub-window of the provider.
13. A video stream processing apparatus, characterized in that, Applied to the providing end, the device includes: a second video stream providing module, configured to provide a target video stream corresponding to a video sub-window to a requesting end according to an instruction from the server, so that the requesting end can display the main window video stream in the main video session window, enabling the user of the requesting end to view the subscribed providing end screen; the main window video stream is generated based on the target video stream corresponding to at least one of the video sub-windows, and the target video stream is determined based on the layout information of the main video session window; the layout information is created by the user of the requesting end; the instruction is issued by the server; The layout information includes the identification information of the video stream provider corresponding to the video sub-window, as well as the size and arrangement order of the video sub-window; the target video stream corresponding to the video sub-window corresponds to the video level information, which is determined according to the size of the video sub-window of the provider.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor, when executing the computer program, implements the method of any one of claims 1-10.
Citation Information
Patent Citations
Method and system for adjusting video streaming image resolution ratio and code stream
CN101848382A
At least one display window layout adjustment and touch automatic calibration method and system
CN109144304A