Video conference data collaboration method and system
By merging multi-screen code streams on the multi-point control unit and adopting real-time adaptive decoding and group description technology, the terminal burden and compatibility issues caused by multi-screen display in existing video conferencing systems are solved, and efficient multi-screen display and system management are achieved.
Patent Information
- Application Number
- CN202210945663.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-08
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-08-08
AI Technical Summary
In existing video conferencing systems, multi-screen display functions are usually handled by the terminal, which increases the terminal burden, causes poor stability and compatibility, and makes it difficult to achieve unified management and system upgrades.
Multi-image synthesis is realized on the multi-point control unit, and multiple QCIF-sized H.261 encoded streams are merged into one CIF-sized stream. Data packets are transmitted through the channel between the multi-point control unit and the data provider. Real-time adaptive decoding and packet description technology are used to realize the free combination and transmission of audio mixed data.
It reduces the burden on terminals, improves system stability and compatibility, facilitates unified management and system upgrades, and does not add extra burden in terms of bandwidth occupancy and signal processing, thus achieving efficient multi-screen display.
Smart Images

Figure CN115314666B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote video conferencing, and in particular to a video conferencing data collaboration method and system. Background Art
[0002] A video conferencing system is a method of holding meetings over a communications network. It not only transmits the audio and video of participants, but also transmits documents, charts, and other information. Through video conferencing systems, people can hold meetings and discuss issues in real time, eliminating unnecessary travel. Network video conferencing systems are a promising industry, widely applicable in areas such as administrative meetings, e-commerce, and distance education. The H.323 protocol, developed by the ITU-T, the International Telecommunication Union's telecommunications standardization organization, is a multimedia communication standard based on packet-switched networks and a universal standard for video conferencing across all types of packet-switched networks. A multipoint control unit (MPCU) is a crucial component of a video conferencing system, providing centralized control and processing of audio, video, and data packets. It switches, connects, and mixes various media, reducing the channel bandwidth requirements of each user terminal and the burden of processing multimedia information on each terminal, significantly lowering network load. While the MCU handles audio mixing, few articles have examined its use in video mixing. During group discussions and free speeches during a video conference, it may be necessary to view the content from multiple venues. To enable each participant to simultaneously view the content from all other venues, multiple venue images must be displayed simultaneously on the local computer. Implementing multi-screen display and picture-in-picture technology within the multipoint control unit's video MP allows users to display images from any of the sub-venues, providing users with a more realistic, natural, and convenient visual experience.
[0003] Many current video conferencing products implement multi-view display functionality on the terminal, using multiple decoders to process the streams from multiple venues and display the images from multiple venues. This approach is simple to implement, but it increases the burden on the terminal, resulting in poor stability and compatibility. Product upgrades require upgrading all terminals. True video MP multi-view functionality should be implemented on the multipoint control unit. While there are many approaches and methods for multi-view synthesis, this article focuses on combining four QCIF-sized H.261 encoded streams into a single CIF-sized stream. The resulting output is identical to a standard H.261 CIF-format stream. This process of combining the multi-view streams on the multipoint control unit does not impose any additional bandwidth or signal processing burden on the terminal compared to receiving a single video signal. Besides reducing terminal complexity and stability as mentioned above, it also facilitates unified management and system upgrades during video conferencing. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention proposes a video conference data collaboration method, comprising:
[0005] The multipoint control unit sends an invitation message including an identification message to the data providing device;
[0006] Establishing a channel between the multipoint control unit and the data providing device, and receiving a data packet corresponding to the identification message sent by the data providing device through the channel;
[0007] The data providing device feeds back a data packet of a corresponding identification type to the multipoint control unit, where the data packet of the corresponding identification type carries attributes of the video conference participant;
[0008] When a video type data packet needs to be transmitted, the multipoint control unit sends a video type identification message to the data providing device, and the data providing device transmits the video type data packet through the channel according to the video type identification message;
[0009] When a data type data packet needs to be transmitted, the multipoint control unit sends a data type identification message to the data providing device, and the data providing device transmits the data type data packet through the channel according to the data type identification message.
[0010] Furthermore, the video conference data input modes include: single-point single-channel audio input mode and single-point multi-channel audio input mode;
[0011] In the single-point single-channel audio input mode, data from M conference points are mixed in a multi-point video conference. Each conference point has only one channel of audio data, for a total of M channels of audio input data.
[0012] In the single-point multi-channel audio input mode, data from M conference points are mixed in a multi-point video conference, and each conference point has N channels of input audio data, for a total of M channels of audio input data.
[0013] Furthermore, in the single-point single-channel audio input mode, the audio input data of the k-th conference side at time t is a k (t), the range of value is [-2 Q-1 , 2 Q-1 -1], where Q is the number of sample quantization bits, k = 1, 2, ..., M;
[0014] The audio input data of the first k points on the conference side are mixed and the output audio mixed data is b k (t), then b M (t) is the total audio mixing data of the input audio data of all M conference points participating in the mixing.
[0015] Furthermore, corresponding to the single-point single-channel audio input mode, a real-time adaptive decoding method is used to decode the audio mixed data, and the audio mixed data is b k (t) The corresponding weight after decoding is w k (t):
[0016]
[0017] The audio mixed data is b k (t) The proportion of the decoded audio in the mixed audio output P k (t), definition:
[0018]
[0019] Furthermore, in the single-point multi-channel audio input mode, each conference point has N-channel audio input data. Let the n-channel audio input data of the k-th conference point at time t be a n (t), n=1,2,...,N;k=1,2,...,M, respectively mix the n-th audio input data of each conference side, and output M-channel audio mixed data after mixing, then b n (t) is the audio mixed data output after mixing the n-th audio input data on the conference side of all points.
[0020] Furthermore, the output M channels of audio mixed data are grouped, and the grouped data packets are described, wherein the description is a category characteristic of the group corresponding to each channel of audio mixed data; and the description is carried in a control protocol related to the audio mixed data.
[0021] The present invention also proposes a video conference data collaboration system for implementing the above-mentioned video conference data collaboration method, comprising: a multipoint control unit and a data providing device;
[0022] The multipoint control unit sends an invitation message including an identification message to the data providing device;
[0023] After receiving the invitation message including the identification message, the data providing device transmits a data packet of a type corresponding to the identification message on a channel according to the identification message.
[0024] Furthermore, it also includes a grouping unit, which groups the audio mixed data to form a plurality of data packets, and performs group description on the plurality of data packets, and then transmits them through a channel.
[0025] Furthermore, the multipoint control unit includes multiple control modules, each control module simultaneously supports K completely independent conferences, each conference corresponds to an independent audio processing module, and each audio processing module has K inputs I1, I2...Ik and K outputs O1, O2...Ok.
[0026] Compared with the existing technology, this application has the following beneficial technical effects:
[0027] A channel is established between the multipoint control unit and the data providing device, and the multipoint control unit sends an invitation message including an identification message to the data providing device; and receives a data packet corresponding to the identification message sent by the data providing device through the channel; it can better coordinate the transmission of data packets of corresponding types. When a video type data packet needs to be transmitted, the multipoint control unit sends a video type identification message to the data providing device, and the data providing device transmits the video type data packet through the channel according to the video type identification message; when a data type data packet needs to be transmitted, the multipoint control unit sends a data type identification message to the data providing device, and the data providing device transmits the data type data packet through the channel according to the data type identification message, so as to facilitate subsequent high-speed transmission and decoding processing.
[0028] The data providing device feeds back a data packet of the corresponding identification type to the multipoint control unit. The data packet of the corresponding identification type carries attributes of the video conference participants, thereby better distinguishing the conference organizing role, the conference system controlling role and the conference participating role.
[0029] The audio mixed data to be output is grouped, and a description of the groups is provided for the multiple data packets after grouping. The group description is a category characteristic corresponding to the data in the multiple data packets; this includes: carrying the category description in a control protocol related to the audio mixed data; and transmitting the multiple data packets after grouping and the category characteristic description corresponding to the group through a channel. This enables the transmission of multiple-channel audio mixed data using fewer data packets, and realizes the free combination and transmission of the multiple-channel audio mixed data. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0031] Figure 1 is a flow chart of the video conferencing data collaboration method of the present invention;
[0032] Figure 2 It is a schematic diagram of the structure of the multi-point control unit of the present invention. DETAILED DESCRIPTION
[0033] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0034] In the drawings of the specific embodiments of the present invention, in order to better and more clearly describe the working principles of the various components in the system, the connection relationship of the various parts in the device is shown, which only clearly distinguishes the relative position relationship between the various components, and does not constitute a limitation on the signal transmission direction, connection sequence and size, dimension and shape of the components or structures.
[0035] like Figure 1 The figure shows a flow chart of the video conferencing data collaboration method of the present application.
[0036] The invitation message sent by the multipoint control unit to the central control unit includes an identification message, which is used to identify the type of data packet to be transmitted and the purpose of the identification channel. The identification message includes a video type identification and a data type identification.
[0037] According to the identification message, the data packet corresponding to the identification message is transmitted on the channel. After receiving the identification message, the data providing device controls the channel to transmit the video type data packet or the data type data packet.
[0038] When a video data packet needs to be transmitted, the multipoint control unit sends a video type identification message to the central control unit, and the central control unit controls the channel to transmit the video type data packet according to the video type identification message;
[0039] When a data type data packet needs to be transmitted, the multipoint control unit sends a data type identification message to the central control unit, and the central control unit controls the channel to transmit the data type data packet according to the data type identification message.
[0040] In a preferred embodiment, both the data type data packet and the video type data packet contain video conference participant attributes, which are assigned by a data providing device. The video conference participant attributes are used to characterize the role of the corresponding participant in the data packet, specifically including one or more of the following: conference organizer role, conference system controller role, and conference participant role.
[0041] The data providing device collects data through a camera and a microphone, prepares the data into a data packet of a corresponding type, and assigns participant attributes to the data packet. The participant attributes assigned here can be pre-configured.
[0042] A channel is established between the multipoint control unit and the data providing device, and a data packet corresponding to the identification message sent by the data providing device is received through the channel.
[0043] In the multi-point video conference of this embodiment, intermediate processing by the multi-point control unit is required. Taking the communication between two video conference sites as an example, the multi-point control unit is responsible for inviting, receiving and forwarding data packets. At this time, there is the establishment of a channel between the first data providing device and the multi-point control unit and the establishment of a media channel between the multi-point control unit and the second data providing device.
[0044] The conversation between the first data providing device and the second data providing device enables the data providing device to select data packets according to the attributes of the participants, and to control the collaborative presentation of data packet types in the multi-point video conference in a targeted manner, thereby completing the control and transmission of conference data packet types while saving media processing resources and network transmission bandwidth.
[0045] like Figure 2 The figure shows the structure of the multi-point control unit. The multi-point control unit includes multiple control modules. Each control module can simultaneously support K completely independent conferences. Each conference corresponds to an independent audio processing module. Each audio processing module has K inputs I1, I2...Ik and K outputs O1, O2...Ok.
[0046] In a multipoint video conference, each data providing device establishes a unicast-based connection with the multipoint control unit, and sends and receives data packets to and from the multipoint control unit in real time.
[0047] In a multipoint video conference, the multipoint control unit is also responsible for call signaling processing, conference control, video core switching, audio mixing, video and audio adaptation, and screen splitting. The data provider also includes a mixing module, decoding module, signaling module, control module, configuration module, and other functional modules. These modules are primarily responsible for encoding the video and audio streams captured by the local camera group and microphone into corresponding data packets according to identification messages and sending them to the multipoint control unit. Simultaneously, they decode corresponding data packets fed back by other data providers through the multipoint control unit and output them to the displays and speakers of the local conference terminals.
[0048] The mixing module of the data providing device inputs K outputs O1, O2...Ok of the audio processing module, and its output is a data packet processed according to the identification message in the invitation.
[0049] Assume that in a multipoint conference, data from M conference points are mixed, and each conference point has only one channel of audio data, that is, there are M channels of audio input data, that is, a single-point single-channel audio input mode.
[0050] At time t, let the audio input data of point k (k = 1, 2, ..., M) be a k (t), its value range is [-2 Q-1 , 2 Q -1 -1], where Q is the number of sample quantization bits.
[0051] Assume that after audio mixing, there are M channels of output audio mixed data, where the audio mixed data output after mixing the first k (k = 1, 2, ..., M) points of audio input data is b k (t). For example, when k=2, the first two audio input data are mixed, and the output audio mixed data is b2(t); then b M (t) is the total audio mixed data of all M channels of input audio data involved in the mixing.
[0052] For the embodiment of the single-point single-channel audio input mode, a real-time adaptive decoding method is introduced. Specifically, the data after the k-th (k=1, 2, ..., M) point speech decoding is b k The corresponding weight of (t) is w k (t):
[0053]
[0054] Considering the characteristics of the multi-point audio signals involved in the mixing, their own proportions are used as weights to determine their proportions in the output of the audio mix. k (t), definition:
[0055]
[0056] In another embodiment, it is assumed that in a multi-point conference, data from M conference points are mixed, and each conference point has N channels of input audio data, that is, a single-point multi-channel audio input mode.
[0057] At time t, let the n-th audio input data of point k (k = 1, 2, ..., M) be a n (t), n (n = 1, 2, ..., N); respectively mix the n-th audio input data of each point, and after audio mixing, there are M-channel audio mixed data outputs, then b n (t) is the audio mixed data output after mixing the n-th audio input data of all points.
[0058] In the embodiment of the single-point multi-channel audio input mode, the M-channel audio mixed data to be output may be grouped to form multiple data packets, which are then transmitted through the channel.
[0059] The multiple data packets are grouped together, where the grouping descriptions are the class characteristics corresponding to the data in the multiple data packets. This includes: carrying the class descriptions in a control protocol related to the audio mixed data; and transmitting the multiple data packets and the class characteristic descriptions corresponding to the groups over a channel. This enables the transmission of multiple-channel audio mixed data using fewer data packets, enabling the free combination and transmission of multiple-channel audio mixed data.
[0060] The present invention also proposes a video conference data collaboration system for implementing the above-mentioned video conference data collaboration method, comprising: a multipoint control unit and a data providing device;
[0061] The multipoint control unit establishes a channel with the data providing device and sends an invitation message including an identification message to the data providing device;
[0062] After receiving the invitation message including the identification message, the data providing device transmits a data packet of a type corresponding to the identification message on a channel according to the identification message.
[0063] The video conference data collaboration system further comprises a grouping unit, which groups the audio mixed data to form a plurality of data packets, and performs group description on the plurality of data packets before transmitting them through a channel.
[0064] The multipoint control unit includes multiple control modules, each of which supports K completely independent conferences at the same time. Each conference corresponds to an independent audio processing module, and each audio processing module has K inputs I1, I2...Ik and K outputs O1, O2...Ok.
[0065] In the process of mixing the audio data of participants collected at the multipoint video conference site and continuously sending them from the data providing device to the multipoint control unit, the RTP protocol is preferably used as the data encapsulation protocol.
[0066] Due to factors such as last-come-first-served situations and packet loss in network transmission, as well as the uneven statistical characteristics of the signals generated by the participants themselves as audio signal sources, transmission jitter and incorrect sorting and data loss of the encoded bit stream after transmission are caused. In order to effectively solve this problem, jitter buffering technology is preferably used to solve it.
[0067] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0068] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A video conference data collaboration method, characterized in that: include: The multipoint control unit sends an invitation message including an identification message to the data providing device; A channel is established between the multipoint control unit and the data providing device, and a data packet corresponding to an identification message sent by the data providing device is received through the channel; the data providing device feeds back a data packet of a corresponding identification type to the multipoint control unit, the data packet of the corresponding identification type carrying attributes of the video conference participant; when a video type data packet needs to be transmitted, the multipoint control unit sends a video type identification message to the data providing device, and the data providing device transmits the video type data packet through the channel based on the video type identification message; when a data type data packet needs to be transmitted, the multipoint control unit sends a data type identification message to the data providing device, and the data providing device transmits the data type data packet through the channel based on the data type identification message; Video conference data input modes include: single-point single-channel audio input mode and single-point multi-channel audio input mode; In the single-point single-channel audio input mode, in a multi-point video conference, the data from the M-point conference side is mixed. Each point conference side has only one channel of audio data, and there are M channels of audio input data in total. Suppose the audio input data of the k-th point conference side at time t is , the value range is , Q is the number of sampling quantization bits, k=1,2,..., M; the audio input data of the first k points on the conference side are mixed and the output audio mixed data is ; Corresponding to the single-point single-channel audio input mode, the real-time adaptive decoding method is used to decode the audio mixed data. The corresponding weight after decoding is : ; Audio mixed data The proportion of the decoded audio in the mixed audio output ; In the single-point multi-channel audio input mode, in a multi-point video conference, data from M conference points are mixed. Each conference point has N channels of input audio data, for a total of M channels of audio input data. Suppose the n-th channel of audio input data from the k-th conference point at time t is , n=1,2,..., N;k=1,2,..., M, respectively mix the n-th audio input data of each conference side, and output M-channel audio mixed data after mixing. The audio mixed data is output after mixing the n-th audio input data on all conference sides; the output M-channel audio mixed data are grouped, and the grouped data packets are described, where the description is the category characteristic of the group corresponding to each channel of audio mixed data; the description is carried in the control protocol related to the audio mixed data.
2. A video conference data collaboration system, used to implement the video conference data collaboration method according to claim 1, characterized in that: include: Multipoint control unit, channel, grouping unit and data providing device; The multipoint control unit establishes a channel with the data providing device and sends an invitation message including an identification message to the data providing device; After receiving the invitation message including the identification message, the data providing device transmits a data packet of a type corresponding to the identification message on a channel according to the identification message; The grouping unit groups the audio mixed data into multiple data packets, and performs group description on the multiple data packets before transmitting them through the channel.
3. The video conferencing data collaboration system according to claim 2, characterized in that: The multipoint control unit includes multiple control modules, each of which supports K completely independent conferences at the same time. Each conference corresponds to an independent audio processing module, and each audio processing module has K inputs I1, I2...Ik and K outputs O1, O2...Ok.
Citation Information
Patent Citations
Video and audio processing method, multi-point control unit and video conference system
CN101370114A
Multi-picture mixing method and apparatus for video meeting system
CN101478642A
Method, apparatus and system for multipath media stream transmission and reception
CN101489090A
Data transmission method and system for video session
CN103634562A
Multi-channel audio and video integration method and system and computer readable storage medium
CN111711835A