Video conferencing method, apparatus, system and device for cloud-edge-terminal collaborative computing
Through the video conferencing method of cloud-edge collaborative computing, combined with the preset allocation model, the problems of difficulty in configuration, poor terminal user experience and high operating costs in the existing technology are solved, and efficient video stream processing and distribution are achieved, improving user experience and reducing operating costs.
Patent Information
- Application Number
- CN202211268790.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-17
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-10-17
AI Technical Summary
In the existing video conferencing technology, each architecture has shortcomings in system configuration, end user experience and operation costs, especially the defects of the SFU architecture relying on high-performance terminals, resulting in poor user experience and high operating costs.
The video conferencing method of cloud edge collaborative computing is adopted. By obtaining the operating parameters and video stream demand information of each node, combining the preset allocation model, the video stream processing allocation task is determined, and tasks are sent to the target node to achieve effective processing and allocation of video streams.
It reduces the difficulty of system configuration, improves the end user experience, reduces the investment cost of operators, overcomes the shortcomings of SFU architecture relying on high-performance terminals, and improves the overall quality of video conferencing.
Smart Images

Figure CN115665364B_ABST
Abstract
Description
Technical Field
[0001] This document belongs to the field of emerging information technologies, and specifically relates to a video conferencing method, apparatus, system, and device for cloud-edge-end collaborative computing. Background Art
[0002] With the development of information technologies, video conferencing has become a tool for people's daily communication. WebRTC enables web-based video conferencing, and real-time communication capabilities can be achieved through a browser. Therefore, the use of WebRTC technology in video conferencing is becoming increasingly widespread.
[0003] Currently, there are mainly three architectures for video conferencing based on WebRTC technology. The first is the Mesh architecture, where each participant is connected to each other in a P2P manner, and data exchange basically does not pass through a central server. Due to the absence of a media server, it is difficult to perform additional processing on videos in the Mesh network structure, and operations such as video recording, video encoding, and video mixing are not supported. The second is the MCU (Multipoint Control Unit) architecture. The MCU is a traditional centralized network structure, and participants are only connected to the central MCU media server. The MCU media server combines all participants' video streams to generate a video stream containing all participants' images. Participants only need to pull the mixed video, and the MCU server is responsible for all complex operations such as video encoding, transcoding, decoding, and mixing. The server-side pressure is relatively high, and a high configuration is required. At the same time, due to the fixed mixed video, the interface layout is not very flexible. The last is the SFU (Selective Forwarding Unit) architecture. In the SFU network structure, although there is a central node media server, it only forwards and does not perform media processing tasks with high resource consumption such as mixing and transcoding. Therefore, the server pressure is much smaller.
[0004] Each of these three architectures has its own advantages and disadvantages. The Mesh architecture has the lowest requirements for the backend server, but the highest requirements for the terminal configuration and terminal network. At the same time, many functions are difficult to implement, and the functional scalability is low. Therefore, this architecture is rarely used currently. The MCU architecture has the highest requirements for the backend, the lowest requirements for the terminal configuration, and the lowest requirements for the terminal network. However, since current video conferencing operations are mostly provided in the form of SaaS, this requires operators to purchase backend servers or rent MCU services, which is too heavy an asset burden for operators. The SFU architecture has a media server that only forwards, with relatively low requirements for the backend server, relatively high requirements for the terminal configuration, and moderate requirements for the terminal network. Summary of the Invention
[0005] In view of the above problems in the prior art, the purpose of this paper is to provide a video conferencing method, device, system and equipment for cloud-edge-end collaborative computing, which can reduce the configuration difficulty of the system and improve the terminal user experience.
[0006] To solve the above technical problems, the specific technical solutions of this paper are as follows:
[0007] Adopting the above technical solution, a video conferencing method for cloud-edge-end collaborative computing in this paper is applied to a video conferencing system. The method includes:
[0008] Obtain the operating parameters of each node in the video conferencing system and the video stream demand information in the video conference. The nodes at least include cloud nodes, edge nodes and terminal nodes, and the video stream demand information includes the video stream information that all terminal nodes need to display;
[0009] According to the video stream demand information, the operating parameters and a preset allocation model, determine the video stream processing allocation task. The video stream processing allocation task includes at least one target node corresponding to the video stream demand information and the allocation task corresponding to each target node;
[0010] According to the video stream processing allocation task, send the allocation task to the target node so that the target node completes the video stream demand information in the video conference.
[0011] Further, the operating parameters at least include static parameters and dynamic parameters;
[0012] The obtaining of the operating parameters of each node in the video conferencing system includes:
[0013] Obtain the static parameters obtained by testing when each node is installed or in a silent idle state. The static parameters at least include node decoding ability, decoding consumption time and CPU memory consumption;
[0014] During the meeting, obtain the dynamic parameters of each node in real time. The dynamic parameters at least include the network bandwidth, CPU computing power and Mem memory of the node.
[0015] Further, the preset allocation model includes:
[0016] P = Min(∑ (i,j,k) (CT i,k X i,j,k ) + ∑( i,j,k )(NT i,j X i,j,k ))
[0017] Wherein,
[0018] ∑ j,k (Ci,k X i,j,k ) < C i ;
[0019] ∑ j,k (N i,k X i,j,k ) < N i ;
[0020] ∑ j,k (CPU i,k X i,j,k ) < CPU i ;
[0021] ∑ j,k (Mem i,k X i,j,k ) < Mem i ;
[0022] ∑ i X i,j,k (i) = 1;
[0023] ∑ i,k X i,j,k (k) = 1;
[0024] X i,j,k ∈ {0, 1};
[0025] Among them, P is the video processing allocation task, and X i,j,k represents the k streams viewed by the j-th user processed by the i-th node; CT i,k represents the time required for the i-th node to process k streams; NT i,j represents the network time required to transmit to the j-th user after the i-th node finishes processing the stream; C i,k represents the capacity consumed by the i-th node to process k streams, and C i represents the total stream processing capacity of the i-th node; N i,k represents the network bandwidth required for the i-th node to process k streams, and N i represents the bandwidth limit that the i-th node can use in total; CPU i,k represents the CPU computing power consumed by the i-th node to process k streams, and CPU i represents the maximum computing power that the i-th node can use; Mem i,k represents the Mem memory situation consumed by the i-th node to process k streams, and Mem i represents the maximum memory situation that the i-th node can use.
[0026] Furthermore, the preset allocation model includes the video stream processing capabilities of each node, and the video stream processing capabilities include the time required for the node to process the video stream, network bandwidth, CPU computing power, and the consumed Mem memory;
[0027] The video stream processing capability is obtained through the following steps:
[0028] According to a preset encoding format, under the conditions of a preset resolution, bit rate, frame rate, and number of bitstreams, obtain the video stream processing capabilities of each node, or,
[0029] During the installation or idle time of the video conferencing system, test the preset resolution, bit rate, and frame rate of the encodings supported by the system to obtain the video stream processing capabilities of each node.
[0030] Further, determining the video stream processing allocation task according to the video stream requirement information, the operating parameters, and a preset allocation model includes:
[0031] Obtain the historical allocation list of each node, where the historical allocation list includes the video stream processing allocation tasks that have been processed by the node;
[0032] Match the historical allocation list with the video stream requirement information to obtain a matching result;
[0033] If the matching result is at least partially successful, the node corresponding to the at least successfully matched video stream requirement information is used as the target node;
[0034] Determine the video stream processing allocation task corresponding to the video stream requirement information that fails to match according to the video stream requirement information that fails to match, the operating parameters, and a preset allocation model to obtain the target node corresponding to the video stream requirement information that fails to match.
[0035] Further, the video conferencing system includes an SFU module, and the SFU node is used to receive the video stream data of all terminal nodes;
[0036] Issuing the allocated processing task to the target node according to the video stream processing allocation task, so that the target node completes the video stream requirement information in the video conference, includes:
[0037] After the target node receives the allocated task, subscribe to the video stream data from the SFU module, so that the SFU module issues the corresponding video stream data to the target node;
[0038] The target node processes the received video stream and sends the processed data to the SFU module, so that the SFU module forwards the processed data to the terminal nodes that need to display.
[0039] Further, the method further includes:
[0040] After the meeting ends, obtain all the assigned tasks of each node and the operating parameters of the node during the execution of each assigned task;
[0041] When the running duration of the running parameters meeting the preset conditions exceeds the preset value, load the assigned task into the historical assignment list of the node.
[0042] On the other hand, this article also provides a video conferencing device for cloud-edge-terminal collaborative computing, which is applied to a video conferencing system. The device includes:
[0043] A data acquisition module, configured to acquire the operating parameters of each node in the video conferencing system and the video stream demand information in the video conference. The nodes at least include a cloud node, an edge node, and a terminal node. The video stream demand information includes the video stream information that all terminal nodes need to display;
[0044] A determination module, configured to determine a video stream processing assignment task according to the video stream demand information, the operating parameters, and a preset assignment model. The video stream processing assignment task includes at least one target node corresponding to processing the video stream demand information and the assignment task corresponding to each target node;
[0045] An assignment module, configured to send an assignment task to the target node according to the video stream processing assignment task, so that the target node completes the video stream demand information in the video conference.
[0046] On the other hand, this article also provides a video conferencing system for cloud-edge-terminal collaborative computing. The system includes:
[0047] Nodes, where the nodes at least include a cloud node, an edge node, and a terminal node;
[0048] A task assignment module, which acquires the operating parameters of each node in the video conferencing system and the video stream demand information in the video conference. The video stream demand information includes the video stream information that all terminal nodes need to display;
[0049] Determine a video stream processing assignment task according to the video stream demand information, the operating parameters, and a preset assignment model. The video stream processing assignment task includes at least one target node corresponding to processing the video stream demand information and the assignment task corresponding to each target node;
[0050] Send an assignment task to the target node according to the video stream processing assignment task, so that the target node completes the video stream demand information in the video conference
[0051] An SFU module, configured to send video stream data to a target node according to the subscription information of the target node, and send the processed video stream data received from the target node to the terminal node.
[0052] Finally, the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described above when executing the computer program.
[0053] This article provides a video conferencing method, device, system and equipment for cloud-edge collaborative computing, which fully utilizes the video stream processing capabilities of various devices and delegates video processing to other devices when the user terminal does not require too high configuration. This can solve the defect of the SFU architecture in the prior art that it needs to rely on high-performance terminals, improve the experience of terminal users, and adopt the minimum processing time and transmission time as the goal. Therefore, under the constraints, it tends to allocate tasks to terminals or edge devices, overcoming the defect of the pure MCU architecture that is heavily dependent on central media services, saving the operator's investment costs.
[0054] In order to make the above and other purposes, features and advantages of this article more obvious and easy to understand, the following specifically cites preferred embodiments and describes them in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of this article or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of this article. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0056] Figure 1 A schematic diagram of an implementation framework of a video conferencing method for cloud-edge-device collaborative computing provided by this description embodiment is shown;
[0057] Figure 2 A detailed implementation framework diagram of a video conferencing method for cloud-edge-device collaborative computing provided by this description embodiment is shown;
[0058] Figure 3 A schematic diagram of the steps of a video conferencing method for cloud-edge-device collaborative computing provided by this description embodiment is shown;
[0059] Figure 4 A schematic diagram of a process flow of a video conferencing method for cloud-edge-device collaborative computing provided by an embodiment of the present description is shown;
[0060] Figure 5 A schematic diagram of the structure of a video conferencing device for cloud-edge-device collaborative computing provided by an embodiment of the present description is shown;
[0061] Figure 6Shows a schematic structural diagram of the computer device provided herein.
[0062] Explanation of the reference numerals in the drawings:
[0063] 10. Terminal node;
[0064] 20. Edge node;
[0065] 30. Cloud node;
[0066] 100. Data acquisition module;
[0067] 200. Determination module;
[0068] 300. Allocation module;
[0069] 602. Computer device;
[0070] 604. Processor;
[0071] 606. Memory;
[0072] 608. Driving mechanism;
[0073] 610. Input / output module;
[0074] 612. Input device;
[0075] 614. Output device;
[0076] 616. Presentation device;
[0077] 618. Graphical user interface;
[0078] 620. Network interface;
[0079] 622. Communication link;
[0080] 624. Communication bus. Detailed implementation manners
[0081] Next, the technical solutions in the embodiments herein will be clearly and completely described in conjunction with the accompanying drawings in the embodiments herein. Obviously, the described embodiments are only a part of the embodiments herein, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments herein without creative efforts shall fall within the scope of protection herein.
[0082] It should be noted that the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings of this article are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this article described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.
[0083] In the prior art, there are three types of network architectures for video conferencing, namely Mesh architecture, MCU architecture and SFU architecture. Each of these three architectures has its own advantages and disadvantages. The Mesh architecture has the least requirements for the backend server, but the highest requirements for the terminal configuration and the terminal network. At the same time, many functions are difficult to implement and the function scalability is low. Therefore, this architecture is rarely used currently. The MCU architecture has the highest requirements for the backend, the lowest requirements for the terminal configuration and the terminal network. However, since current video conferencing operations are mostly provided in the form of SaaS, this requires the operator to purchase a backend server or rent an MCU service, which is too heavy an asset burden for the operator. The SFU architecture has a media server that only forwards, with relatively low requirements for the backend server, relatively high requirements for the terminal configuration, and moderate requirements for the terminal network.
[0084] To solve the above problems, a video conferencing method with low configuration requirements and fast function implementation is provided, that is, a video conferencing method for cloud-edge-end collaborative computing, as Figure 1 shown, is the network environment for the implementation of the method. The network environment applies a video conferencing system, including various types of devices and an SFU module, such as terminal node 10, edge node 20 and cloud node 30. Among them, the edge node covers some terminal nodes according to its deployment characteristics and can communicate with the terminal nodes it covers. Both the terminal node and the edge node can communicate with the cloud node. The terminal node can select to send a corresponding viewing task to the network according to the user's video demand, such as displaying corresponding video stream information. The video stream information can include video and audio. The terminal node, the edge node and the cloud node each have an information collection module, a task scheduling interface and a video processing module. Among them, the information collection module can collect the CPU, memory, network, etc. of the node. The task scheduling interface is used for task and data interaction with other external nodes. The video processing module is used to process the video stream. Processing the video stream includes, but is not limited to, decoding, transcoding, merging, etc.
[0085] Furthermore, a task allocation module is provided in the cloud node. The task allocation model determines the relative video stream processing tasks according to the user's required video stream and the processing capabilities of each device node, as well as the preset allocation model, and allocates the video stream processing tasks to each device through the task scheduling interface. After receiving the assigned tasks, each node subscribes to the video stream from the SFU module, and also feeds back the video stream after its own processing results to the SFU module. The SFU model then sends the processed video stream to the user so that the user can display the corresponding video stream. This article reduces the configuration requirements of a single terminal and improves the user experience through the collaborative work of the cloud, edge and end.
[0086] For example, Figure 2 The figure shows an implementation framework diagram of a video conferencing system for cloud-edge-end collaborative computing in an embodiment of this specification.
[0087] Furthermore, the embodiments of this article provide a video conferencing method with cloud-edge-end collaborative computing, which can utilize the video processing capabilities of various devices to improve user experience. Figure 3 This is a step diagram of a video conferencing method for cloud-edge-device collaborative computing provided in the embodiments of this article. This specification provides method operation steps as described in the embodiments or flowcharts, but may include more or fewer operation steps based on conventional or non-creative labor. The order of steps listed in the embodiments is only one way of executing the steps among many orders, and does not represent the only order of execution. When the actual system or device product is executed, it can be executed in the order of the methods shown in the embodiments or drawings or in parallel. Specifically, Figure 3 As shown, the method may include:
[0088] S201: Obtaining operating parameters of each node in the video conferencing system and video stream demand information in the video conferencing, wherein the nodes include at least cloud nodes, edge nodes, and terminal nodes, and the video stream demand information includes video stream information that all terminal nodes need to display;
[0089] S202: Determine a video stream processing allocation task according to the video stream demand information and the operating parameters, and a preset allocation model, wherein the video stream processing allocation task includes processing at least one target node corresponding to the video stream demand information and an allocation task corresponding to each target node;
[0090] S203: According to the video stream processing allocation task, the allocation task is issued to the target node, so that the target node completes the video stream demand information in the video conference.
[0091] It can be understood that in the video conferencing process, according to the real-time operating status of different nodes and combined with the corresponding preset allocation model, the optimal video stream processing allocation task is determined, and this task is then sent to the corresponding processing nodes, thus realizing the collaborative work of cloud-edge-end nodes. Compared with the traditional method of placing video processing tasks on a single node, the method provided in the embodiments of this specification can fully release the processing capabilities of all device nodes, thereby reducing the configuration requirements of the nodes, lowering costs, and at the same time ensuring the quality of video conferencing and improving the user experience.
[0092] In the embodiments of this specification, by integrating the SFU architecture and the MCU architecture, the processing capabilities of various devices for video can be fully utilized. When the capabilities of user terminals are limited or the configurations are relatively low, video processing (including but not limited to multi-channel video transcoding) can be placed on edge nodes or cloud nodes, overcoming the defect that pure SFU depends on high-performance terminals and enhancing the user experience of terminal users. At the same time, the investment of video conferencing operators is reduced. When cloud nodes are limited or the number of participants is huge, more tasks can be placed on the edge or terminal users. Also, if the processing capabilities of terminal devices or edge devices close to terminal users are powerful.
[0093] In some embodiments of this specification, the operating parameters at least include static parameters and dynamic parameters;
[0094] The obtaining of the operating parameters of each node in the video conferencing system includes:
[0095] Obtaining the static parameters tested when each node is installed or in a silent idle state, where the static parameters at least include the node decoding ability, decoding consumption time, and CPU memory consumption;
[0096] During the meeting, the dynamic parameters of each node are obtained in real time, where the dynamic parameters at least include the network bandwidth, CPU computing power, and Mem memory of the node.
[0097] It can be understood that the static parameters can be the static video processing capabilities of the nodes. Further, they can be the efficiency of processing a single video stream (such as a preset encoding format) and the hardware consumption of the nodes, such as the decoding speed of the video stream, the computing power consumption, CPU memory consumption, and bandwidth consumption involved in the decoding process, which are mainly related to their configuration performance. The dynamic parameters can be the changes in hardware data during the video stream processing in the video process of the nodes. Through the static parameters and dynamic parameters, the maximum processing capabilities and processing efficiencies of the nodes can be determined in real time.
[0098] During a video conference, since different nodes may be assigned different tasks, that is, the number of video streams involved is different. Based on determining the static video processing capabilities of different nodes, data such as the efficiency and consumption of different nodes in processing the assigned video streams can be determined. Therefore, the capabilities of different nodes in processing different assigned tasks can be roughly determined through their static and dynamic parameters. Furthermore, a task assignment for the nodes can be selected to determine a better assignment scheme, so as to determine the scheme with the shortest processing time and the highest efficiency, improving the user experience.
[0099] Furthermore, since it involves users entering and leaving the conference, as well as video speeches of different users, etc., the video requirements of the terminal nodes may also change in real time. Correspondingly, the assigned tasks determined according to the preset assignment model may also change. Therefore, the same node (any one of the terminal node, edge node, and cloud node) can receive different assigned tasks at different times. Therefore, as an option:
[0100] During the conference, the dynamic parameters of each node can be obtained in real time, or;
[0101] During the conference, the dynamic parameters of each node can be obtained at a preset frequency.
[0102] It can be understood that obtaining the dynamic parameters of the nodes in real time can obtain accurate and reliable data. Correspondingly, it will also increase certain hardware and software costs. When the number of users is relatively large, high-performance hardware support is also required on the cloud node side; moreover, during a video conference, the changes in the user terminals and video requirements do not change very frequently. Collecting dynamic data in real time will result in certain invalid data, wasting network resources. By setting a preset frequency for collection, the collection frequency can be reduced, improving the utilization rate of the data. As an option, the preset frequency can be set according to the actual situation, and the specific setting method is not limited in the embodiments of this specification.
[0103] It should be noted that the dynamic parameters of the nodes correspond to the video stream demand information of the video conference, that is, the video stream demand information also changes continuously. The bitstream decoding and display that the user terminal needs to view are dynamically changing. Correspondingly, the dynamic parameters of different nodes will change synchronously during the task processing.
[0104] In some embodiments of this specification, the preset assignment model includes:
[0105] P = Min(∑ (i,j,k) (CT i,k X i,j,k ) + ∑ (i,j,k) (NT i,j X i,j,k )), (1)
[0106] Among them, the constraints are as follows:
[0107] ∑ j,k (C i,k X i,j,k ) < C i ;
[0108] ∑ j,k (N i,k X i,j,k ) < N i ;
[0109] ∑ j,k (CPU i,k X i,j,k ) < CPU i ;
[0110] ∑ j,k (Mem i,k X i,j,k ) < Mem i ;
[0111] ∑ i X i,j,k (i) = 1;
[0112] ∑ i,k X i,j,k (k) = 1;
[0113] X i,j,k ∈ {0, 1};
[0114] Among them, P is the video processing allocation task, and X i,j,k represents the k streams viewed by the j-th user processed by the i-th node; CT i,k represents the time required for the i-th node to process k streams; NT i,j represents the network time required for the i-th node to transmit the processed streams to the j-th user; C i,k represents the capacity consumed by the i-th node to process k streams, and C i represents the total capacity of the i-th node to process streams; N i,k represents the network bandwidth required for the i-th node to process k streams, and N i represents the bandwidth limit that the i-th node can use in total; CPU i,k represents the CPU computing power consumed by the i-th node to process k streams, and CPU i represents the maximum computing power that the i-th node can use; Mem i,k represents the Mem memory situation consumed by the i-th node to process k streams, and Mem i represents the maximum memory situation that the i-th node can use.
[0115] It can be understood that in the embodiments of this specification, the preset allocation model is modeled and analyzed with the goal of minimizing the processing time, so as to allocate the video processing tasks of the terminal nodes, cloud nodes, and edge nodes, and then achieve the collaborative work of the cloud-edge-terminal. Through the above formula (1) combined with its corresponding constraints, the video stream demand information can be disassembled and divided. For example, the video demand of each user terminal (i.e., the terminal node) is used as a sub-task. In this way, multiple sub-tasks are used as a task set, and all nodes are used as a node set. The candidate task set can be obtained in a permutation and combination manner. By counting the processing time of each candidate task in the candidate task set (including the time required to process the video stream and the time required to transmit the processed video stream), the candidate task with the minimum processing time is used as the final video stream processing allocation task.
[0116] It should be noted that when the node for processing the video stream is the same as the node that needs to view the video stream, that is, when i and j are the same, NT ij is 0. When the i device is an edge node device, if the j user does not belong to the service range of this edge node, NT ij can be set to a very large number. Correspondingly, this edge node is obviously not the optimal choice for processing the video stream of the i device.
[0117] It should be noted again that in the constraints, the required video stream of each terminal node can be processed as a percentage, and the CPU i,k , Mem i,k can be processed as a percentage. Through the normalization process and the 0-1 allocation, the unity of different candidate allocation tasks during comparison can be ensured, and the accuracy and reliability of determining the allocation task are improved. Optionally, in the same candidate allocation task, the tasks corresponding to different nodes can be normalized according to the number of video streams, so as to ensure that the total amount of tasks received by all allocated nodes in the same candidate allocation task is 1.
[0118] In some embodiments of this specification, the preset allocation model includes the video stream processing capabilities of each node. The video stream processing capabilities include the time required for the node to process the video stream, network bandwidth, CPU computing power, and the consumed Mem memory;
[0119] The video stream processing capabilities are obtained through the following steps:
[0120] According to the preset coding format, under the conditions of preset resolution, bit rate, frame rate, and number of bitstreams, the video stream processing capabilities of each node are obtained, or,
[0121] During the installation or idle time of the video conferencing system, the encoding supported by the system is tested for preset resolution, bit rate, and frame rate to obtain the video stream processing capabilities of each node.
[0122] It can be understood that the above video stream processing capabilities are all inherent performances of the nodes, that is, static parameters. By quantifying the static parameters, it is convenient for the subsequent visualization of different allocation tasks and the processing capabilities of different nodes under the same allocation task, improving the accuracy and reliability of the determination of the allocation tasks.
[0123] It should be noted that the process of determining the above video stream processing capabilities can be designed according to the actual situation. Since there will be certain losses in the hardware of the nodes during the operation of the system, the calculated video stream processing capabilities are rough or approximate data, not completely accurate data.
[0124] In some embodiments of this specification, determining the video stream processing allocation task according to the video stream demand information, the operating parameters, and a preset allocation model includes:
[0125] Obtain the historical allocation list of each node, where the historical allocation list includes the video stream processing allocation tasks that have been processed by the node;
[0126] Match the historical allocation list with the video stream demand information to obtain a matching result;
[0127] If the matching result is at least partially successfully matched, the node corresponding to the at least successfully matched video stream demand information is used as the target node;
[0128] Determine the video stream processing allocation task corresponding to the video stream demand information that fails to match, using the video stream demand information that fails to match, the operating parameters, and a preset allocation model, so as to obtain the target node corresponding to the video stream demand information that fails to match.
[0129] It can be understood that by setting the historical allocation list of each node, it is possible to determine the video stream processing tasks that have been processed by different nodes and are relatively efficient and stable; when the video demands of each node in the video stream demand information correspond to the historical allocation tasks of at least one node, the node corresponding to the historical allocation task can be directly used as the target node for the corresponding video demand, and then the above step S202 of the allocation task is performed on the remaining video demand information, thereby reducing the number of video demand information, reducing the calculation amount in the process of determining the allocated personnel, and improving the allocation efficiency.
[0130] Exemplarily, in a video conference, there are 10 users (i.e., terminal nodes), a total of 15 devices (i.e., the sum of terminal nodes, edge nodes, and cloud nodes). By obtaining the video requirements of 10 users, video stream requirement information is obtained. The 15 devices correspond to 15 historical allocation lists. By matching, it is determined that the video requirements of 2 users are in the corresponding historical allocation lists of 2 devices. Then, these 2 devices can be used as the target nodes for the video requirements of the 2 users. Subsequently, calculations for the video stream processing allocation tasks are performed on the remaining 13 devices and the video requirements of 8 users, reducing the computational amount, thereby enabling faster determination of all allocation tasks and improving the allocation efficiency.
[0131] In some embodiments of this specification, the video conferencing system includes an SFU module, and the SFU node is used to receive the video stream data of all terminal nodes;
[0132] According to the video stream processing allocation task, sending the allocated processing task to the target node so that the target node completes the video stream requirement information in the video conference, including:
[0133] After the target node receives the allocated task, subscribing to the video stream data from the SFU module so that the SFU module sends the corresponding video stream data to the target node;
[0134] The target node processes the received video stream and sends the processed data to the SFU module so that the SFU module forwards the processed data to the terminal nodes that need to display it.
[0135] It can be understood that by setting the SFU module in the embodiments of this specification, the characteristics of the SFU module can be utilized to achieve fast reception and forwarding of video streams, without encoding and decoding, consuming very little CPU resources. The direct forwarding method also greatly reduces latency and improves real-time performance. Additionally, by integrating edge computing technology, cloud-edge-terminal collaborative work is achieved.
[0136] In some embodiments of this specification, the method further includes:
[0137] After the meeting ends, obtaining all the allocation tasks of each node and the operating parameters of the node during the execution of each allocation task;
[0138] When the duration for which the operating parameters meet the preset conditions exceeds the preset value, loading the allocation task into the historical allocation list of the node.
[0139] It can be understood that although there are many solutions for the 0-1 programming, in the worst case, its time complexity is O(2^n), so the computational amount is very large. Therefore, after each meeting, the tasks assigned to each node are recorded and evaluated. As an alternative emperor, if the task runs stably for more than a certain empirical threshold (such as 300 seconds) under a certain coding format and meets certain CPU and memory usage conditions, this task can be put into the historical allocation list and directly used in the next allocation. This can gradually optimize the efficiency of task allocation. During the long-term use of video conferencing, the cloud-edge-end collaborative working mode can be determined faster and more efficiently, improving the user experience.
[0140] In video conferencing, the encoding, transcoding, and decoding of video streams are basic but computationally intensive. The three current architectures based on WebRTC each have their own advantages and disadvantages. The present invention combines the characteristics of SFU and MCU, integrates edge computing technology, and establishes a cloud-edge-end collaborative computing mode. Its advantages and effects are as follows:
[0141] Improve the user experience. The present invention can make full use of the video processing capabilities of various devices. When the capabilities of the user terminal are limited or the configuration is low, video processing (including but not limited to multi-channel video encoding, transcoding, and decoding) can be placed on the edge node or the central node, overcoming the defect that pure SFU needs to rely on high-performance terminals and improving the experience of end users.
[0142] Reduce the investment of video conferencing operators. In the case where the central node is limited or the number of participants is huge, more tasks can be placed on the edge or end users. At the same time, if the processing capabilities of the terminal device or the edge device close to the end user are strong, the scheduling model of the present invention aims at minimizing the processing time and transmission time. Therefore, under the condition of meeting the constraints, it tends to allocate tasks to the terminal or edge device, overcoming the defect that the pure MCU architecture seriously depends on the central media service and saving the investment cost of the operator.
[0143] In some embodiments of this specification, a video conferencing method for cloud-edge-end collaborative computing is also provided. As Figure 4 shown, it is the processing logic diagram of the method:
[0144] Test static parameters such as the terminal decoding ability, decoding consumption time, CPU and memory consumption during the installation or silent idle state of the video conferencing;
[0145] During the meeting, collect dynamic information such as the video stream requirements to be displayed by the terminal and the usage conditions of the CPU, memory, and network of each device;
[0146] According to the collected data, check the common allocation list of the device. If there is one, directly assign tasks to the device;
[0147] In the model calculation, directly set the nodes that have been assigned tasks to 1, then calculate according to the model type, and distribute the calculation results to each node;
[0148] Determine whether the meeting has ended;
[0149] If the meeting has not ended, return to the step of collecting dynamic information;
[0150] If the meeting has ended, put the task descriptions that meet the conditions into the common allocation list according to the node operation status.
[0151] Based on the above-provided method, an embodiment of this specification further provides a video conferencing device for cloud-edge-terminal collaborative computing, which is applied to a video conferencing system, as Figure 5 shown, the device includes:
[0152] A data acquisition module 100, configured to acquire the operation parameters of each node in the video conferencing system and the video stream demand information in the video conference. The nodes at least include cloud nodes, edge nodes, and terminal nodes, and the video stream demand information includes the video stream information that all terminal nodes need to display;
[0153] A determination module 200, configured to determine a video stream processing allocation task according to the video stream demand information, the operation parameters, and a preset allocation model. The video stream processing allocation task includes at least one target node corresponding to processing the video stream demand information and the allocation task corresponding to each target node;
[0154] An allocation module 300, configured to send an allocation task to the target node according to the video stream processing allocation task, so that the target node completes the video stream demand information in the video conference.
[0155] Furthermore, an embodiment of this specification further provides a video conferencing system for cloud-edge-terminal collaborative computing, and the system includes:
[0156] Nodes, the nodes at least include cloud nodes, edge nodes, and terminal nodes;
[0157] A task allocation module, which acquires the operation parameters of each node in the video conferencing system and the video stream demand information in the video conference. The video stream demand information includes the video stream information that all terminal nodes need to display;
[0158] According to the video stream demand information, the operation parameters, and a preset allocation model, determine a video stream processing allocation task. The video stream processing allocation task includes at least one target node corresponding to processing the video stream demand information and the allocation task corresponding to each target node;
[0159] Allocate tasks according to the video stream processing, and send the allocated tasks to the target node, so that the target node can complete the video stream requirement information in the video conference.
[0160] The SFU module is used to send video stream data to the target node according to the subscription information of the target node, and send the processed video stream data received from the target node to the terminal node.
[0161] As Figure 6 As shown, a computer device provided by an embodiment of this article. The device in this article can be the computer device in this embodiment, and execute the method in this article above. The computer device 602 may include one or more processors 604, such as one or more central processing units (CPUs), and each processing unit may implement one or more hardware threads. The computer device 602 may also include any memory 606, which is used to store any kind of information such as code, settings, data, etc. Non-limiting, for example, the memory 606 may include any one or more combinations of the following: any type of RAM, any type of ROM, flash memory devices, hard disks, optical discs, etc. More generally, any memory can use any technology to store information. Further, any memory can provide volatile or non-volatile retention of information. Further, any memory may represent a fixed or removable component of the computer device 602. In one case, when the processor 604 executes the associated instructions stored in any memory or combination of memories, the computer device 602 may perform any operation of the associated instructions. The computer device 602 also includes one or more drive mechanisms 608 for interacting with any memory, such as a hard disk drive mechanism, an optical disc drive mechanism, etc.
[0162] The computer device 602 may also include an input / output module 610 (I / O), which is used to receive various inputs (via the input device 612) and provide various outputs (via the output device 614)). A specific output mechanism may include a presentation device 616 and an associated graphical user interface (GUI) 618. In other embodiments, the input / output module 610 (I / O), the input device 612, and the output device 614 may not be included, and it is only used as a computer device in the network. The computer device 602 may also include one or more network interfaces 620, which are used to exchange data with other devices via one or more communication links 622. One or more communication buses 624 couple the components described above together.
[0163] The communication link 622 can be implemented in any way, for example, through a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 622 can include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc. governed by any protocol or combination of protocols.
[0164] Corresponding to Figure 1 In the method of, embodiments herein also provide a computer-readable storage medium having a computer program stored thereon, and when the computer program is run by a processor, it executes the steps of the above method.
[0165] Embodiments herein also provide a computer-readable instruction, wherein when the processor executes the instruction, the program therein causes the processor to execute the method as Figure 1 shown.
[0166] It should be understood that in various embodiments herein, the magnitudes of the sequence numbers of the above processes do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments herein.
[0167] It should also be understood that in the embodiments herein, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this text generally represents an "or" relationship between the associated objects before and after.
[0168] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this text.
[0169] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0170] In several embodiments provided in this document, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings, direct couplings, or communication connections between each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be electrical, mechanical, or other forms of connection.
[0171] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments in this document.
[0172] In addition, each functional unit in the various embodiments of this document can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0173] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution in this document, or the part that contributes to the prior art, or all or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this document. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs and other media that can store program codes.
[0174] Specific embodiments are used in this document to elaborate on the principles and implementation methods of this document. The descriptions of the above embodiments are only used to help understand the methods and their core ideas in this document; at the same time, for those of ordinary skill in the art, according to the ideas in this document, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to this document.
Claims
1. A video conferencing method for cloud-edge-terminal collaborative computing, which is applied to a video conferencing system, and is characterized in that, The method includes: Obtaining the operation parameters of each node in the video conferencing system and the video stream demand information in the video conference. The nodes at least include a cloud node, an edge node, and a terminal node. The video stream demand information includes the video stream information that all terminal nodes need to display; Determining a video stream processing allocation task according to the video stream demand information, the operation parameters, and a preset allocation model. The video stream processing allocation task includes at least one target node corresponding to the video stream demand information and the allocation task corresponding to each target node; Issuing an allocation task to the target node according to the video stream processing allocation task, so that the target node completes the video stream demand information in the video conference; The determining of the video stream processing allocation task according to the video stream demand information, the operation parameters, and the preset allocation model includes: Obtaining the historical allocation list of each node. The historical allocation list includes the video stream processing allocation tasks that have been processed by the node and whose operation duration with operation parameters meeting the preset conditions exceeds the preset value; Matching the historical allocation list with the video stream demand information to obtain a matching result; If the matching result is at least partially successfully matched, the node corresponding to the at least successfully matched video stream demand information is used as the target node; Determining the video stream processing allocation task corresponding to the video stream demand information that is not successfully matched according to the video stream demand information that is not successfully matched, the operation parameters, and the preset allocation model, so as to obtain the target node corresponding to the video stream demand information that is not successfully matched. The method further includes: After the meeting ends, obtaining all the allocation tasks of each node and the operation parameters of the node operation during the execution of each allocation task; When the operation duration of the allocation task whose operation parameters meet the preset conditions exceeds the preset value, loading the allocation task into the historical allocation list of the node. The preset condition is that the CPU and memory meet the preset usage rate.
2. The method according to claim 1, wherein The operation parameters at least include static parameters and dynamic parameters; The obtaining of the operation parameters of each node in the video conferencing system includes: Obtaining the static parameters tested when each node is installed or in a silent idle state. The static parameters at least include the node decoding ability, decoding consumption time, and CPU memory consumption; During the meeting process, obtaining the dynamic parameters of each node in real time. The dynamic parameters at least include the network bandwidth, CPU computing power, and Mem memory of the node.
3. The method according to claim 1, wherein The preset allocation model includes: , Wherein, ; where P is the video processing assignment task, represents the i th node processing j the stream viewed by the user k ; represents the time required for the i th node to process k streams; represents the network time required for the i th node to transfer the processed stream to j the user; represents the capacity consumed by node i processing k streams, represents the total capacity of node i to process streams; represents the network bandwidth required for node i to process k streams, represents the total available bandwidth limit of node i ; represents the CPU computing power consumed by node i processing k streams, represents the maximum available computing power of node i ; represents the Mem memory situation consumed by node i processing k streams, represents the maximum available memory situation of node i .
4. The method according to claim 3, characterized in that The preset allocation model includes the video stream processing capabilities of each node. The video stream processing capabilities include the time required for the node to process the video stream, network bandwidth, CPU computing power, and the consumed Mem memory; The video stream processing capabilities are obtained through the following steps: Obtaining the video stream processing capabilities of each node according to a preset encoding format under preset resolution, bit rate, frame rate, and number of bitstreams, or, Testing the preset resolution, bit rate, and frame rate for the encodings supported by the system when the video conferencing system is installed or idle, so as to obtain the video stream processing capabilities of each node.
5. The method according to claim 1, characterized in that, The video conferencing system includes an SFU module, and the SFU node is used to receive the video stream data of all terminal nodes; According to the video stream to process and allocate tasks, and send the allocated processing tasks to the target nodes, so that the target nodes complete the video stream requirement information in the video conference, including: After the target node receives the allocated task, it subscribes to the video stream data from the SFU module, so that the SFU module sends the corresponding video stream data to the target node; The target node processes the received video stream and sends the processed data to the SFU module, so that the SFU module forwards the processed data to the terminal nodes that need to display.
6. A video conferencing device for cloud-edge-end collaborative computing, which is applied to a video conferencing system, and is characterized in that The device includes: A data acquisition module, which is used to acquire the operation parameters of each node in the video conferencing system and the video stream requirement information in the video conference. The nodes at least include cloud nodes, edge nodes and terminal nodes, and the video stream requirement information includes the video stream information that all terminal nodes need to display; A determination module, which is used to determine the video stream processing allocation task according to the video stream requirement information, the operation parameters, and a preset allocation model. The video stream processing allocation task includes at least one target node corresponding to the video stream requirement information and the allocation task corresponding to each target node. Specifically: Obtain the historical allocation list of each node. The historical allocation list includes the video stream processing allocation tasks that have been processed by the node and the operation duration of which the operation parameters meet the preset conditions exceeds the preset value; Match the historical allocation list with the video stream requirement information to obtain a matching result; If the matching result is at least partially successfully matched, the nodes corresponding to the at least successfully matched video stream requirement information are used as target nodes; Determine the video stream processing allocation task corresponding to the video stream requirement information that fails to match according to the video stream requirement information that fails to match, the operation parameters, and a preset allocation model, so as to obtain the target nodes corresponding to the video stream requirement information that fails to match; An allocation module, which is used to send the allocation task to the target node according to the video stream processing allocation task, so that the target node completes the video stream requirement information in the video conference, and; After the meeting ends, obtain all the allocation tasks of each node and the operation parameters of the node during the execution of each allocation task; When the operation duration of the allocation task whose operation parameters meet the preset conditions exceeds the preset value, load the allocation task into the historical allocation list of the node. The preset condition is that the CPU and memory meet the preset usage rate conditions.
7. A video conferencing system for cloud-edge-terminal collaborative computing, characterized in that, The system includes: Nodes, the nodes at least include cloud nodes, edge nodes and terminal nodes; A task allocation module, which acquires the operation parameters of each node in the video conferencing system and the video stream requirement information in the video conference. The video stream requirement information includes the video stream information that all terminal nodes need to display; Determine a video stream processing allocation task according to the video stream demand information, the operating parameters, and a preset allocation model. The video stream processing allocation task includes at least one target node corresponding to the video stream demand information and an allocation task corresponding to each target node. Specifically: Obtain the historical allocation list of each node. The historical allocation list includes video stream processing allocation tasks that have been processed by the node and whose operating parameters meet the preset conditions and the running duration exceeds the preset value. Match the historical allocation list with the video stream demand information to obtain a matching result. If the matching result is at least partially successfully matched, the node corresponding to the at least successfully matched video stream demand information is used as the target node. Determine the video stream processing allocation task corresponding to the video stream demand information that fails to match by using the video stream demand information that fails to match, the operating parameters, and a preset allocation model, so as to obtain the target node corresponding to the video stream demand information that fails to match. According to the video stream processing allocation task, send an allocation task to the target node, so that the target node completes the video stream demand information in the video conference, and; After the conference ends, obtain all the allocation tasks of each node and the operating parameters of the node during the execution of each allocation task. When the running duration of the allocation task whose operating parameters meet the preset conditions exceeds the preset value, load the allocation task into the historical allocation list of the node. The preset conditions are that the CPU and memory meet the preset usage rate conditions. The SFU module is used to send video stream data to the target node according to the subscription information of the target node, and send the video stream data processed by the received target node to the terminal node.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Cloud side-end collaborative resource management method and system
CN113452566A
Optimal energy consumption task unloading method, device and system based on cloud edge collaboration
CN115051999A