A conference control method, apparatus and computer readable storage medium

By using multi-channel video data acquisition and caching synthesis technology on the video network mobile terminal, the problem of insufficient single-channel video images in video network conferences has been solved, and the synchronous synthesis of multiple video data streams and the enhancement of visual information have been achieved.

CN113194278BActive Publication Date: 2025-12-12VISIONVERA INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110310969.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-23
Publication Date
2025-12-12
Estimated Expiration
2041-03-23

AI Technical Summary

Technical Problem

Existing video network mobile terminals can only capture and transmit one video feed during a meeting, resulting in insufficient visual information received by participating users.

Method used

Multiple video data streams are simultaneously acquired through the built-in camera, external camera, and interconnected electronic devices of the video network mobile terminal. After being cached in the buffer area, the video images are synthesized to generate a single video data stream, which is then sent to the participating terminal.

Benefits of technology

It increases the amount of visual information received by participating users, enhances the video call experience, and ensures the synchronous synthesis of various video data streams through the buffer area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113194278B_ABST
    Figure CN113194278B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a conference control method, device and computer readable storage medium, wherein the method comprises: after the establishment of a visual networking conference is successful, starting each built-in camera and an external camera of a mobile terminal to collect first video data streams; acquiring a second video data stream sent by an interconnected external electronic device; respectively buffering the first video data streams collected by each camera in a corresponding first buffer area and buffering the second video data stream in a second buffer area; extracting video images from each first buffer area and the second buffer area respectively, and performing video image synthesis to combine the first video data streams and the second video data stream into one video data stream; and after encoding the synthesized video data stream, sending the video data stream to a visual networking server through the visual networking, so as to send more visual information to each visual networking terminal participating in the conference and improve the video conference experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of communication technology, in particular to a conference control method and device and a computer readable storage medium. BACKGROUND

[0002] The current video networking mobile terminal mainly includes conference control or joining an existing conference when conducting a video networking conference. The video networking mobile terminal joins the conference and is commonly used in live interviews, reports, learning and training, etc. Currently, the mobile terminal defaults to opening the front camera to collect video pictures after joining the conference, transmits the collected video pictures to the video networking server, and reaches each video networking terminal after being transferred by the video networking server. During the conference, if the user of the video networking mobile terminal wants to collect video pictures of other angles, the user needs to manually switch the camera, and only one camera can be opened to collect video pictures each time.

[0003] It can be seen that the existing video networking conference control scheme can only collect one video picture and send it to each video networking terminal for the users of each video networking terminal, and the amount of visual information received is small. SUMMARY

[0004] In view of the above problems, the embodiments of the present application are proposed to provide a conference control method and a corresponding conference control device to overcome the above problems or at least partially solve the above problems.

[0005] To solve the above problems, the embodiments of the present application disclose a conference control method, wherein the conference control method is applied to a video networking mobile terminal, and the method comprises:

[0006] After the video networking conference is successfully established, each built-in camera and an external camera of the mobile terminal are opened to collect first video data streams;

[0007] A second video data stream sent by an interconnected external electronic device is acquired, wherein the second video data stream is a video data stream collected by the external electronic device through a camera, or a second video data stream obtained by the external electronic device by screen capture;

[0008] Each first video data stream collected by each camera is cached in a corresponding first cache area, and the second video data stream is cached in a second cache area, wherein the number of the first cache areas is the same as the total number of the opened built-in cameras and external cameras;

[0009] Video images are extracted from each first cache area and the second cache area respectively, and video image synthesis is performed to combine the first video data stream and the second video data stream into one video data stream;

[0010] Encode the synthesized video data stream, and send to the video conference server through the video conference, so as to send to each video conference terminal participating in the video conference through the video conference server.

[0011] Optionally, after the step of extracting video images from the first cache area and the second cache area respectively and synthesizing the video images, the method further comprises: displaying the synthesized video images locally.

[0012] Optionally, before the step of starting each built-in camera and external camera of the mobile terminal to collect the first video data stream, the method further comprises:

[0013] After the video conference is successfully established, it is determined whether to start the multi-path video participation mode according to the historical video conference record;

[0014] In the case of determining to start the multi-path video participation mode, the step of starting each built-in camera and external camera of the mobile terminal to collect the first video data stream is performed.

[0015] Optionally, before the step of starting each built-in camera and external camera of the mobile terminal to collect the first video data stream, the method further comprises:

[0016] After the video conference is successfully established, output a selection prompt box, wherein the selection prompt box is used to prompt the user to start the multi-path video participation mode;

[0017] Receive a first input of the selection prompt box by the user, wherein the first input is used to start the multi-path video participation mode;

[0018] In response to the first input, the step of starting each built-in camera and external camera of the mobile terminal to collect the first video data stream is performed.

[0019] Optionally, the step of extracting video images from the first cache area and the second cache area respectively and synthesizing the video images comprises:

[0020] Extracting the first video image with the shortest storage time from each first cache area respectively;

[0021] Extracting the second video image with the shortest storage time from the second cache area;

[0022] Scaling and splicing each first video image and the second video image to generate a synthesized video image with a preset resolution.

[0023] Optionally, before the step of scaling and splicing each of the first video images and the second video images to generate a synthesized video image of a preset resolution, the method further comprises:

[0024] According to the number of each of the first video images and the second video images, a plurality of synthesized video image styles are recommended to be displayed;

[0025] A user's selection operation on a target synthesized video image style is received.

[0026] Optionally, the method further comprises: capturing a screenshot of the local screen at a preset frequency to generate a first video data stream.

[0027] To solve the above problems, the embodiments of the present application disclose a conference control device, wherein the conference control device is applied to a vision Internet mobile terminal, and the device comprises:

[0028] An opening module is configured to open each built-in camera and an external camera of the mobile terminal to collect a first video data stream after a vision Internet conference is successfully established;

[0029] An acquisition module is configured to acquire a second video data stream sent by an interconnected external electronic device, wherein the second video data stream is a video data stream collected by a camera of the external electronic device or a second video data stream obtained by the external electronic device through a screenshot;

[0030] A cache module is configured to cache each first video data stream collected by each camera in a corresponding first cache area and cache the second video data stream in a second cache area, wherein the number of the first cache areas is the same as the total number of the opened built-in cameras and external cameras;

[0031] A synthesis module is configured to extract video images from each first cache area and the second cache area respectively, perform video image synthesis, and combine the first video data stream and the second video data stream into a video data stream;

[0032] A sending module is configured to encode the synthesized video data stream and send the video data stream to a vision Internet server through the vision Internet, so as to send the video data stream to each vision Internet terminal participating in the conference through the vision Internet server.

[0033] Optionally, the device further comprises:

[0034] A display module is configured to display the synthesized video image locally after the synthesis module extracts video images from each first cache area and the second cache area respectively and performs video image synthesis.

[0035] Optionally, the device further comprises:

[0036] a mode determining module, configured to determine whether to start the multi-path video attending mode according to historical video conference records before the starting module starts the built-in camera and the external camera of the mobile terminal to collect the first video data stream respectively after the Web Real-Time Communication conference is successfully established;

[0037] a first calling module, configured to call the starting module to start the built-in camera and the external camera of the mobile terminal to collect the first video data stream respectively in the case of determining to start the multi-path video attending mode.

[0038] Optionally, the apparatus further comprises:

[0039] an output module, configured to output a selection prompt box before the starting module starts the built-in camera and the external camera of the mobile terminal to collect the first video data stream respectively after the Web Real-Time Communication conference is successfully established, wherein the selection prompt box is used to prompt the user to start the multi-path video attending mode;

[0040] an input receiving module, configured to receive a first input of the user on the selection prompt box, wherein the first input is used to start the multi-path video attending mode;

[0041] a second calling module, configured to call the starting module to start the built-in camera and the external camera of the mobile terminal to collect the first video data stream respectively in response to the first input.

[0042] Optionally, the synthesizing module comprises:

[0043] a first sub-module, configured to extract the first video image with the shortest storage time from each of the first cache areas respectively;

[0044] a second sub-module, configured to extract the second video image with the shortest storage time from the second cache area;

[0045] a third sub-module, configured to scale and splice each of the first video images and the second video image to generate a synthesized video image with a preset resolution.

[0046] Optionally, the synthesizing module further comprises:

[0047] a fourth sub-module, configured to recommend display of a plurality of synthesized video image styles according to the number of the first video images and the second video image before the third sub-module scales and splices each of the first video images and the second video image to generate a synthesized video image with a preset resolution;

[0048] a fifth sub-module, configured to receive a selection operation of a target synthesized video image style of the user.

[0049] Optionally, the apparatus further comprises:

[0050] a screenshot module configured to take screenshots of the local screen at a preset frequency to generate a first video data stream.

[0051] To solve the above problems, the embodiments of the present application disclose a conference control apparatus, comprising:

[0052] one or more processors; and

[0053] one or more machine-readable media having stored thereon instructions, which when executed by the one or more processors, cause the apparatus to perform any of the conference control methods described above.

[0054] To solve the above problems, the embodiments of the present application disclose a computer readable storage medium, wherein the computer program stored therein causes the processor to perform any of the conference control methods described above.

[0055] The embodiments of the present application have the following advantages:

[0056] The embodiments of the present application can send more visual information to the participating video networking terminals and improve the video call experience by collecting multiple video data streams through the built-in camera, the external camera and the interconnected external electronic device of the video networking mobile terminal and sending the video data streams to the participating video networking terminals during the video networking conference, compared with the prior art of collecting a single video data stream in the video networking conference control method. In addition, the embodiments of the present application can ensure the synchronous synthesis of the multiple video data streams by buffering the multiple video data streams in the buffer area and extracting the video images from the buffer area for synthesis. BRIEF DESCRIPTION OF DRAWINGS

[0057] Figure 1 is a structural block diagram of a video networking mobile terminal of the present application;

[0058] Figure 2 is a step flow chart of a conference control method embodiment of the present application;

[0059] Figure 3 is a structural block diagram of a conference control apparatus embodiment of the present application;

[0060] Figure 4 is a networking schematic diagram of a video networking of the present application;

[0061] Figure 5 is a hardware structural schematic diagram of a node server of the present application;

[0062] Figure 6 is a hardware structural schematic diagram of an access switch of the present application;

[0063] Figure 7 Figure 1 is a schematic diagram of a hardware structure of an Ethernet gateway according to the present application. DETAILED DESCRIPTION

[0064] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0065] Videoconferencing system: The videoconferencing system is based on high-definition audio and video transmission based on the videoconferencing network, and is constructed by corresponding management software and client. It supports the access of various special terminals and mobile terminals. The main functions include: organizing a conference, video call, live broadcast, watching live broadcast, etc. Its related applications include: Pamir conference control terminal, conference scheduling server, conference management network background, etc. The hardware terminals it supports include: Jupiter / Enming series hardware terminals, palm video mobile terminal, Pamir mobile terminal conference tablet, etc.

[0066] Pamir: The Pamir is a videoconferencing scheduling system client, a client software running on a personal computer platform, used for conference reservation, conference process management and control, starting and stopping a conference, switching a speaker, and split-screen mode. The front-end main operation module of the conference system completely controls the entire process of the conference.

[0067] Pamir mobile terminal: The Pamir mobile terminal is a videoconferencing mobile conference terminal, which is a Pamir running on a mobile platform after partial function simplification. The hardware platform is usually an Android system tablet or mobile phone, etc. It can be connected to the videoconferencing system management end through a wireless network to perform conference control and participate in conference operation, and also has functions such as live broadcast and video phone. The videoconferencing tablet or mobile device described below refers to the Pamir mobile terminal by default.

[0068] The present application is used in the Pamir mobile conference terminal.

[0069] The videoconferencing mobile terminal of the present application embodiment, i.e. the Pamir mobile conference terminal, includes at least two cameras, i.e. a front camera and a rear camera, and the videoconferencing mobile terminal is externally connected with an electronic device and / or a camera. A schematic structural diagram of an exemplary videoconferencing mobile terminal is shown in Figure 1 .

[0070] As shown in Figure 1 , the videoconferencing mobile terminal includes: a front camera, a rear camera, an externally connected external camera, a local screen sharing module, four video capture modules, a multi-picture synthesis module, a video encoding module, a video output module, a local video output module, and a conference service control module.

[0071] The conference control method provided by the embodiment of the application is applied to Figure 1 The Pamir mobile terminal software installed on the video networking mobile terminal is controlled, the software simultaneously opens two paths of front and rear cameras of the mobile terminal, two video capture modules are called to capture two paths of video pictures captured by the front camera and the rear camera respectively, one video capture module is called to capture one path of video picture captured by the external camera, one video capture module is called to capture one path of video picture shared by the local screen, each path of camera video picture is synthesized into one path of video picture by a multi-picture synthesis module, the video data stream is output, the video data stream after synthesis is encoded by a video encoding module, and finally the encoded video data stream is output by an output module. After the video data stream is received by the opposite end device, the multi-path video picture can be displayed simultaneously after the received video data stream is video decoded, so that the multi-view video conference is realized. The innovative video conference control technology can greatly improve the practical effect of the video networking mobile terminal video conference. The conference control process provided by the embodiment of the application is as follows:

[0072] With reference to Figure 2 , a step flowchart of a conference control method embodiment of the application is shown, the method can be applied to a video networking mobile terminal, and specifically can include the following steps:

[0073] Step 201: After the video networking conference is established successfully, each built-in camera and external camera of the mobile terminal is opened to capture the first video data stream.

[0074] The built-in camera of the mobile terminal can include a front camera and a rear camera, the number of front cameras can be one or more, and the number of rear cameras can also be one or more. In actual opening of the built-in camera, only one front camera and one rear camera can be opened, or all front cameras and rear cameras can be opened. Each camera corresponds to one path of video data stream, for example: two built-in cameras are opened, and two paths of first video data stream are corresponded. The external camera can also be one or more, and each external camera corresponds to one path of video data stream. Therefore, the total number of first video data streams is the sum of the number of opened built-in cameras and external cameras.

[0075] Step 202: The second video data stream sent by the interconnected external electronic device is acquired.

[0076] The second video data stream is the video data stream captured by the external electronic device through the camera, or the second video data stream obtained by the external electronic device by screen capture.

[0077] The built-in camera, the external camera and the external electronic device can provide multiple video stream data, compared with only one camera collecting one video stream data, the visual information collection amount can be increased.

[0078] Step 203: respectively cache the first video data stream collected by each camera into the corresponding first cache area, and cache the second video data stream into the second cache area.

[0079] The number of the first cache areas is the same as the total number of the built-in cameras and the external cameras.

[0080] Each first cache area and the second cache area is an independent cache area, and the size of each cache area can be set by the person skilled in the art according to the actual demand, which is not specifically limited in the embodiment of the application.

[0081] Step 204: respectively extract video images from each first cache area and the second cache area, and perform video image synthesis to merge the first video data stream and the second video data stream into one video data stream.

[0082] Since there may be an asynchronous problem in collecting each video data stream, the collected video data stream is cached into the cache area and then extracted to solve the problem that each video data stream cannot be synchronized.

[0083] When synthesizing each frame of video image, the video image can be synthesized into any appropriate picture mode, which is not specifically limited in the embodiment of the application. For example, each frame of video image is scaled and superimposed on the same picture to form a picture mode with different sizes or the same size.

[0084] In the embodiment of the application, the video data stream collected by each camera is cached in the cache area, and the video image is extracted from the cache area for synthesis, so that the synchronization of each video data stream can be ensured.

[0085] Step 205: after encoding the synthesized video data stream, the video data stream is sent to the video live streaming server through the video live streaming, so as to be sent to each video live streaming terminal participating in the conference through the video live streaming server.

[0086] The synthesized video data stream is one video data stream, and the encoding mode can be any existing encoding mode, and the specific encoding mode is not specifically limited in the embodiment of the application.

[0087] After each video live streaming terminal participating in the conference receives the encoded video data stream sent by the video live streaming server, the video data stream is decoded, and the decoded video stream picture is displayed on the display screen. The displayed picture contains each video picture collected by the mobile terminal sending the video data stream.

[0088] In the embodiment of the application, multiple video data streams are collected by the built-in camera, the external camera and the external electronic device connected to the mobile terminal of the video conference in the process of the video conference, and are sent to each video conference terminal. Compared with the prior art of collecting only a single video data stream, more visual information can be sent to each video conference terminal, and the video call experience is improved. In addition, in the embodiment of the application, each video data stream is cached in the cache area, and the video images are extracted from the cache area for synthesis, so that the synchronization of the synthesis of each video data stream can be ensured.

[0089] In an optional embodiment, after the video images are extracted from the first cache area and the second cache area respectively, and the video image synthesis is performed, the synthesized video image can also be displayed locally.

[0090] The optional conference control method facilitates the user of the mobile terminal of the video conference to view the overall effect of the synthesized video image.

[0091] In an optional embodiment, before the respective built-in cameras and the external camera of the mobile terminal are started to collect the first video data stream, the embodiment of the application can further include the following process:

[0092] After the video conference is successfully established, whether to start the multi-path video conference mode is determined according to the historical video conference record. In the case of determining to start the multi-path video conference mode, the step of starting the respective built-in cameras and the external camera of the mobile terminal is performed.

[0093] When determining whether to start the multi-path video conference mode according to the historical video conference record, whether the multi-path video conference mode is started during the last video conference can be determined according to the historical video conference record. If yes, the multi-path video conference mode is determined to be started. Otherwise, the multi-path video conference mode is not started. The probability of starting the multi-path video conference mode during the recent video conference can also be counted according to the historical video conference record. In the case of the probability being greater than a preset probability, the multi-path video conference mode is determined to be started. Otherwise, the multi-path video conference mode is not started. In the case of determining not to start the multi-path video conference mode, it is defaulted that only the front camera is started to collect the video data stream.

[0094] In the optional embodiment, whether to start the multi-path video conference mode can be determined according to the historical use habits of the user, and the video conference experience of the user can be improved.

[0095] In an optional embodiment, before the respective built-in cameras and the external camera of the mobile terminal are started to collect the first video data stream, the embodiment of the application can further include the following process:

[0096] First, output a selection prompt box after the Web real-time collaboration conference connection is successfully established;

[0097] The selection prompt box is used to prompt the user to start the multi-path video participation mode, and the selection prompt box can be provided with a first option and a second option. The first option is used to prompt the user to start the multi-path video participation mode, and the second option is used to prompt the user to use the single-path video participation mode for video call.

[0098] Secondly, a first input of the user to the selection prompt box is received;

[0099] The first input is used to start the multi-path video participation mode, and the first input can be a selection operation of the user to the first option, for example, long pressing the first option, double-clicking the first option, or single-clicking the first option, and the specific operation of the first input is not limited in the embodiment of the application.

[0100] Finally, in response to the first input, a step of starting each built-in camera and an external camera of the mobile terminal to collect first video data streams is performed.

[0101] In the case where the second input of the user to the selection prompt box is received, the front camera of the mobile terminal is started to perform the video conference.

[0102] In the optional embodiment, the selection prompt box is output to prompt the user that the multi-path video participation mode can be started, and the user can select whether to start according to actual needs, thereby improving the video conference experience of the user.

[0103] In an optional embodiment, the video images are extracted from each first cache area and second cache area respectively, and the video image synthesis is performed in the following manner:

[0104] First, the first video image with the shortest storage time is extracted from each first cache area, and the second video image with the shortest storage time is extracted from the second cache area.

[0105] The video image with the shortest storage time is extracted from the cache area, which can ensure the timeliness of the extracted video image.

[0106] Secondly, the first video image and the second video image are scaled and spliced to generate a synthesized video image with a preset resolution.

[0107] The preset resolution is the resolution of a single-frame video image.

[0108] When the multiple video images are scaled and spliced, the scaling ratio and splicing manner of each frame of video image can be set by the user according to actual needs, and it is only necessary to ensure that the resolution of the spliced synthesized video image is the preset resolution.

[0109] By scaling and splicing multiple video images to generate a frame of synthetic video image of a preset resolution, the traffic consumption in the video conference process can be maintained without additional increase.

[0110] In an optional embodiment, before scaling and splicing each first video image and second video image to generate a synthetic video image of a preset resolution, the video conference method of the embodiment of the application can further include the following steps:

[0111] First, according to the number of each first video image and second video image, a plurality of synthetic video image styles are recommended to be displayed;

[0112] The synthetic video image style is used to limit the display mode of the first video image and the second video image after being synthesized, and the synthetic video image style can include but is not limited to: multi-region average up-down distribution display, multi-region average left-right distribution display, each first video image being reduced and then superimposed on a second video image at a preset position, and the like.

[0113] Secondly, a selection operation of a target synthetic video image style by a user is received.

[0114] Each synthetic video image style can correspond to a template, and the user can select a target synthetic video image style by selecting a target template.

[0115] In this optional manner, a plurality of synthetic video image styles are provided for the user to select, so as to meet the personalized needs of the user.

[0116] In an optional embodiment, the conference control method further includes: taking a screenshot of the local screen at a preset frequency to generate a first video data stream.

[0117] In this optional conference control method, the content displayed on the local screen can be shared to each video networking terminal participating in the conference without manual screenshot sharing by the user.

[0118] It should be noted that when the local screen screenshot is added as a first video data stream to the synthetic video image, the final synthetic video image does not need to be displayed on the local terminal to avoid the problem of video circulation.

[0119] In an optional embodiment, after the synthetic video data stream is encoded, the video networking is used to send the synthetic video data stream to the video networking server in the following manner:

[0120] First, the synthetic video image is read at a preset frame rate;

[0121] The preset frame rate can be set by those skilled in the art according to actual needs, and the embodiment of the application does not make a specific limitation on this.

[0122] Secondly, the read synthetic video image is encoded according to a preset encoding mode, and then is sent to the video live streaming server through the video live streaming.

[0123] When the synthetic video image is encoded, H264 can be used for encoding. H.264 is a highly compressed digital video codec standard proposed by the Joint Video Team of ITU-T Video Coding Experts Group and ISO / IEC Moving Picture Experts Group. This standard is commonly referred to as H.264 / AVC (or AVC / H.264 or H.264 / MPEG-4 AVC or MPEG-4 / H.264 AVC) to explicitly indicate its two developers.

[0124] In actual implementation, the encoded synthetic video image can be sent to the video live streaming server through the V2V video live streaming, and is forwarded to the opposite end device of the video live streaming by the video live streaming server.

[0125] Relying on the structural security of the V2V video live streaming technology, data cannot be threatened by viruses, hackers, illegal rebroadcasting, illegal insertion and other hidden dangers. At the same time, all data packets are not broadcasted in the sending and transmission process, and are not opened and read for each packet on each video live streaming server or router. Only an independent channel is established between the required points and point transmission, and only the first data packet is opened and read by the video live streaming server to ensure network security.

[0126] It should be noted that, for the method embodiments, in order to simply describe, they are all described as a series of action combinations, but those skilled in the art should know that the embodiments of the present application are not limited to the order of the described actions, because according to the embodiments of the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present application.

[0127] Referring to Figure 3 , a structural block diagram of a conference control device embodiment of the present application is shown, which can be applied to a video live streaming mobile terminal, and specifically can include the following modules:

[0128] The starting module 301 is configured to start the built-in camera and the external camera of the mobile terminal to collect the first video data stream after the video live streaming conference is successfully established.

[0129] The acquisition module 302 is configured to acquire a second video data stream sent by an interconnected external electronic device, wherein the second video data stream is a video data stream collected by the external electronic device through a camera, or a second video data stream obtained by the external electronic device by screen capture.

[0130] The cache module 303 is configured to cache the first video data streams collected by the cameras in corresponding first cache areas respectively, and cache the second video data stream in a second cache area, wherein the number of the first cache areas is the same as the total number of the opened built-in cameras and external cameras.

[0131] The synthesis module 304 is configured to extract video images from the first cache areas and the second cache area respectively, and perform video image synthesis to combine the first video data streams and the second video data stream into one video data stream.

[0132] The sending module 305 is configured to encode the synthesized video data stream, and send the encoded video data stream to the video live streaming server through the video live streaming, so as to send the video data stream to each video live streaming terminal through the video live streaming server.

[0133] In an optional embodiment, the conference control apparatus of the embodiment of the present application further includes the following modules:

[0134] The display module is configured to display the synthesized video images locally after the synthesis module extracts the video images from the first cache areas and the second cache area respectively and performs video image synthesis.

[0135] In an optional embodiment, the conference control apparatus of the embodiment of the present application further includes the following modules:

[0136] The mode determination module is configured to determine whether to open the multi-path video conference mode according to historical video conference records before the opening module opens the built-in cameras and the external cameras of the mobile terminal to collect the first video data streams.

[0137] The first calling module is configured to call the opening module to open the built-in cameras and the external cameras of the mobile terminal to collect the first video data streams in the case of determining to open the multi-path video conference mode.

[0138] In an optional embodiment, the conference control apparatus of the embodiment of the present application further includes the following modules:

[0139] The output module is configured to output a selection prompt box before the opening module opens the built-in cameras and the external cameras of the mobile terminal to collect the first video data streams, wherein the selection prompt box is used to prompt the user to open the multi-path video conference mode.

[0140] The input receiving module is configured to receive a first input of the selection prompt box by the user, wherein the first input is used to open the multi-path video conference mode.

[0141] The second calling module is configured to call the starting module to start the built-in camera and the external camera to collect the first video data stream in response to the first input.

[0142] In an optional embodiment, the synthesizing module comprises the following sub-modules:

[0143] The first sub-module is configured to extract the first video image with the shortest storage time from each of the first cache areas;

[0144] The second sub-module is configured to extract the second video image with the shortest storage time from the second cache area;

[0145] The third sub-module is configured to scale and splice each of the first video image and the second video image to generate a synthesized video image with a preset resolution.

[0146] In an optional embodiment, the synthesizing module further comprises the following sub-modules:

[0147] The fourth sub-module is configured to recommend a plurality of synthesized video image styles according to the number of the first video images and the second video image before the third sub-module scales and splices each of the first video image and the second video image to generate a synthesized video image with a preset resolution;

[0148] The fifth sub-module is configured to receive a selection operation of a target synthesized video image style by a user.

[0149] Optionally, the apparatus further comprises:

[0150] The screenshot module is configured to take screenshots of the local screen at a preset frequency to generate a first video data stream.

[0151] For the apparatus embodiment, it is basically similar to the method embodiment, so the description is relatively simple, and the related parts are described in the method embodiment.

[0152] The vision internet is an important milestone of network development, is a real-time network, and can realize real-time transmission of high-definition video and push a large number of Internet applications to high-definition video and high-definition face-to-face.

[0153] The vision internet adopts real-time high-definition video switching technology, and can integrate services such as high-definition video conference, video monitoring, intelligent monitoring and analysis, emergency command, digital broadcast television, time-delay television, network teaching, live broadcast, VOD, television mail, individual recording (PVR), internal network (self-operated) channel, intelligent video broadcast, information release and other video, voice, picture, text, communication and data services on a network platform. High-definition video is played through a television or a computer.

[0154] In order for those skilled in the art to better understand the embodiments of the present application, the vision internet is introduced as follows:

[0155] Some technologies applied in the vision internet are described as follows:

[0156] Network technology

[0157] The network technology innovation of the vision internet improves the traditional Ethernet, so as to face the huge video flow on the network. Unlike pure network packet switching or network circuit switching, the vision internet technology adopts packet switching to meet the streaming demand. The vision internet technology has the flexibility, simplicity and low price of packet switching, and has the quality and safety guarantee of circuit switching, realizes the full-network switching type virtual circuit and the seamless connection of data format.

[0158] Switching technology

[0159] The vision internet adopts the two advantages of asynchronous and packet switching of Ethernet, eliminates the defects of Ethernet under the premise of full compatibility, has full-network end-to-end seamless connection, directly passes through the user terminal and directly carries IP data packets. User data does not need any format conversion in the full-network range. The vision internet is a higher form of Ethernet, which is a real-time switching platform, can realize the full-network large-scale real-time transmission of high-definition video which cannot be realized by the current Internet, and promotes many network video applications to high-definition and unification.

[0160] Server technology

[0161] The server technology on the Live-Streaming Network and the Unified Video Platform is different from the traditional server. The streaming media transmission is based on the connection-oriented basis. Its data processing capacity is irrelevant to the traffic and communication time. A single network layer can contain signaling and data transmission. For voice and video services, the complexity of the Live-Streaming Network and the Unified Video Platform streaming media processing is much simpler than data processing. The efficiency is more than 100 times higher than the traditional server.

[0162] Storage Technology

[0163] The super-speed storage technology of the Unified Video Platform adopts the most advanced real-time operating system to adapt to the super large capacity and super large flow of media content. The program information in the server instruction is mapped to the specific hard disk space. The media content does not pass through the server and is directly sent to the user terminal in an instant. The user waiting time is generally less than 0.2 seconds. The optimal sector distribution greatly reduces the mechanical movement of the hard disk head seeking. The resource consumption is only 20% of the same level of IP Internet. But it generates more than 3 times the concurrent flow of the traditional hard disk array. The comprehensive efficiency is improved by more than 10 times.

[0164] Network Security Technology

[0165] The structural design of the Live-Streaming Network completely eliminates the network security problems that plague the Internet from the structure through the separate licensing of each service, complete isolation of device and user data, etc. Generally, there is no need for an antivirus program and a firewall. The attacks of hackers and viruses are eliminated. The Live-Streaming Network provides a worry-free secure network for users from the structure.

[0166] Service Innovation Technology

[0167] The Unified Video Platform integrates services and transmission. Whether it is a single user, a private network user, or a network integration, it is only an automatic connection. The user terminal, set-top box, or PC is directly connected to the Unified Video Platform to obtain various forms of rich and colorful multimedia video services. The Unified Video Platform uses a "recipe" configuration table mode to replace the traditional complex application programming. A very small amount of code can realize complex applications and realize "unlimited" new service innovation.

[0168] The networking of the Live-Streaming Network is as follows:

[0169] The Live-Streaming Network is a centralized control network structure. The network can be a tree network, a star network, a ring network, etc. But on this basis, there needs to be a centralized control node in the network to control the entire network.

[0170] AsFigure 4 As shown, the video network is divided into two parts: an access network and a metropolitan area network.

[0171] The devices in the access network part can be mainly divided into three categories: node servers, access switches, and terminals (including various set-top boxes, coding boards, memories, etc.). The node servers are connected to the access switches, and the access switches can be connected to multiple terminals and can be connected to an Ethernet.

[0172] Among them, the node server is a node in the access network that has a centralized control function and can control the access switch and the terminal. The node server can be directly connected to the access switch or directly connected to the terminal.

[0173] Similarly, the devices in the metropolitan area network part can also be divided into three categories: metropolitan area servers, node switches, and node servers. The metropolitan area servers are connected to the node switches, and the node switches can be connected to multiple node servers.

[0174] Among them, the node server is the node server in the access network part, that is, the node server belongs to both the access network part and the metropolitan area network part.

[0175] The metropolitan area server is a node in the metropolitan area network that has a centralized control function and can control the node switch and the node server. The metropolitan area server can be directly connected to the node switch or directly connected to the node server.

[0176] As can be seen, the entire video network is a hierarchical and centralized control network structure, and the network controlled by the node server and the metropolitan area server can be of various structures such as tree type, star type, and ring type.

[0177] In other words, the access network part can form a unified video platform (the part enclosed by the dashed line), and multiple unified video platforms can form a video network; each unified video platform can be interconnected through the metropolitan area and the wide area video network.

[0178] Classification of video network devices

[0179] 1.1 The devices in the video network of the embodiments of the application can be mainly divided into three categories: servers, switches (including Ethernet gateways), and terminals (including various set-top boxes, coding boards, memories, etc.). The video network as a whole can be divided into a metropolitan area network (or a national network, a global network, etc.) and an access network.

[0180] 1.2 The devices in the access network part can be mainly divided into three categories: node servers, access switches (including Ethernet gateways), and terminals (including various set-top boxes, coding boards, memories, etc.).

[0181] The specific hardware structure of each access network device is as follows:

[0182] Node server:

[0183] As shown in Figure 5 , mainly including network interface module 501, switching engine module 502, CPU module 503, disk array module 504;

[0184] Among them, network interface module 501, CPU module 503, disk array module 504 into the package into switching engine module 502; switching engine module 502 to the incoming packet address table 505 operation, thereby obtaining the package guide information; and according to the package guide information into the corresponding packet buffer 506 queue; if the queue of packet buffer 506 is close to full, then discard; switching engine module 502 polling all packet buffer queue, if meet the following conditions for forwarding: 1) the port sending buffer is not full; 2) the queue packet counter is greater than zero. Disk array module 504 mainly realizes the control of hard disk, including the initialization of hard disk, read and write operation; CPU module 503 is mainly responsible for the protocol processing between the access switch, terminal (not shown in the figure), the configuration of address table 505 (including downstream protocol packet address table, upstream protocol packet address table, data packet address table), and the configuration of disk array module 504.

[0185] Access switch:

[0186] As shown in Figure 6 , mainly including network interface module (downlink network interface module 601, uplink network interface module 602), switching engine module 603 and CPU module 604;

[0187] Among them, the package (uplink data) into the downlink network interface module 601 into the packet detection module 605; packet detection module 605 detects the destination address (DA), source address (SA), data packet type and packet length, if it meets the requirements, if it meets the requirements, allocate the corresponding stream identifier (stream-id), and enter the switching engine module 603, otherwise discard; the package (downlink data) into the uplink network interface module 602 into the switching engine module 603; the data packet into the CPU module 604 into the switching engine module 603; switching engine module 603 to the incoming packet address table 606 operation, thereby obtaining the package guide information; if the package into the switching engine module 603 is the downlink network interface to the uplink network interface, then combined with the stream identifier (stream-id) into the corresponding packet buffer 607 queue; if the queue of packet buffer 607 is close to full, then discard; if the package into the switching engine module 603 is not the downlink network interface to the uplink network interface, then according to the package guide information, the data packet into the corresponding packet buffer 607 queue; if the queue of packet buffer 607 is close to full, then discard.

[0188] The switch engine module 603 polls all the packet buffer queues, which are divided into two cases in the embodiment of the present application:

[0189] If the queue is from the downlink network interface to the uplink network interface, the following conditions are met for forwarding: 1) the sending buffer of the port is not full; 2) the queue packet counter is greater than zero; 3) a token generated by the code rate control module is obtained;

[0190] If the queue is not from the downlink network interface to the uplink network interface, the following conditions are met for forwarding: 1) the sending buffer of the port is not full; 2) the queue packet counter is greater than zero.

[0191] The code rate control module 608 is configured by the CPU module 604 to generate tokens for all the packet buffer queues from the downlink network interface to the uplink network interface in a programmable interval, so as to control the code rate of the uplink forwarding.

[0192] The CPU module 604 is mainly responsible for the protocol processing with the node server, the configuration of the address table 606, and the configuration of the code rate control module 608.

[0193] Ethernet translation gateway :

[0194] As shown in Figure 7 , it mainly includes a network interface module (a downlink network interface module 701, an uplink network interface module 702), a switch engine module 703, a CPU module 704, a packet detection module 705, a code rate control module 708, an address table 706, a packet buffer 707, and a MAC adding module 709 and a MAC deleting module 710.

[0195] Among them, the data packet coming into the downlink network interface module 701 enters the packet detection module 705; the packet detection module 705 detects whether the Ethernet MAC DA, the Ethernet MAC SA, the Ethernet length or frame type, the video on the network destination address DA, the video on the network source address SA, the video on the network packet type and the packet length of the data packet meet the requirements, and if they meet the requirements, a corresponding stream identifier (stream-id) is allocated; then, the MAC DA, the MAC SA and the length or frame type (2 bytes) are subtracted by the MAC deleting module 710, and enter the corresponding receiving buffer, otherwise they are discarded;

[0196] The downlink network interface module 701 detects the sending buffer of the port, and if there is a packet, the Ethernet MAC DA of the corresponding terminal is learned according to the video on the network destination address DA of the packet, the Ethernet MAC DA of the terminal, the Ethernet MAC SA of the Ethernet conversion gateway and the Ethernet length or frame type are added, and then the packet is sent.

[0197] The functions of other modules in the Ethernet gateway are similar to those of the access switch.

[0198] Terminal:

[0199] The network interface module, the service processing module and the CPU module are mainly included; for example, the set top box mainly includes the network interface module, the audio and video codec engine module, the CPU module; the encoding board mainly includes the network interface module, the audio and video codec engine module, the CPU module; the memory mainly includes the network interface module, the CPU module and the disk array module.

[0200] 1.3 The equipment of the metropolitan area network part can be mainly divided into two categories: the node server, the node switch and the metropolitan area server. The node switch mainly includes the network interface module, the switch engine module and the CPU module; the metropolitan area server mainly includes the network interface module, the switch engine module and the CPU module.

[0201] 2, Web real-time communication data packet definition

[0202] 2.1 Access network data packet definition

[0203] The data packet of the access network mainly includes the following parts: destination address (DA), source address (SA), reserved bytes, payload (PDU) and CRC.

[0204] As shown in the following table, the data packet of the access network mainly includes the following parts:

[0205] DA SA Reserved Payload CRC

[0206] Among them:

[0207] The destination address (DA) is composed of 8 bytes (byte), the first byte indicates the type of the data packet (such as various protocol packets, multicast data packets, unicast data packets, etc.), and there are 256 possibilities at most, the second byte to the sixth byte is the metropolitan area network address, and the seventh and eighth bytes are the access network address;

[0208] The source address (SA) is also composed of 8 bytes (byte), and the definition is the same as the destination address (DA);

[0209] The reserved bytes are composed of 2 bytes;

[0210] The payload part has different lengths according to different data report types. If it is various protocol packets, it is 64 bytes, if it is a single multicast data packet, it is 32+1024=1056 bytes, and of course it is not limited to the above two;

[0211] CRC is composed of 4 bytes, and its calculation method follows the standard Ethernet CRC algorithm.

[0212] 2.2 Metropolitan area network packet definition

[0213] The topology of the metropolitan area network is a graph, and there can be two or even more than two connections between two devices, i.e., node switches and node servers, node switches and node switches, and node switches and node servers. However, the metropolitan area network address of the metropolitan area network device is unique, and in order to accurately describe the connection relationship between the metropolitan area network devices, a parameter, label, is introduced in the embodiment of the application to uniquely describe a metropolitan area network device.

[0214] The definition of the label in the specification is similar to the definition of the label of MPLS (Multi-Protocol Label Switching). Assuming that there are two connections between device A and device B, there are two labels for the data packet from device A to device B, and there are also two labels for the data packet from device B to device A. The labels are divided into in-label and out-label. Assuming that the label of the data packet entering device A (in-label) is 0x0000, the label of the data packet leaving device A (out-label) can become 0x0001. The network entry process of the metropolitan area network is a network entry process under centralized control, which means that the address allocation and label allocation of the metropolitan area network are dominated by the metropolitan area server, and the node switch and the node server are only passive executors. This is different from the label allocation of MPLS, which is the result of mutual negotiation between switches and servers.

[0215] As shown in the following table, the metropolitan area network packet mainly includes the following parts:

[0216] DA SA Reserved Tag Payload CRC

[0217] i.e., destination address (DA), source address (SA), reserved byte (Reserved), label, payload (PDU), and CRC. Among them, the format of the label can refer to the following definition: the label is 32 bits, of which the high 16 bits are reserved, and only the low 16 bits are used. Its position is between the reserved byte and the payload of the data packet.

[0218] Based on the above characteristics of the video networking, one of the core ideas of the embodiment of the application is proposed, which follows the protocol of the video networking and transmits the video data stream synthesized by the video networking mobile terminal to each video networking terminal participating in the conference by the video networking server.

[0219] Each embodiment in the specification adopts a progressive manner for description, and each embodiment focuses on the different places from other embodiments. The same and similar parts of each embodiment can be referred to each other.

[0220] Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, apparatus, or computer program product. Accordingly, embodiments of the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, embodiments of the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, and the like) embodying computer readable program code.

[0221] Embodiments of the present application are described herein with reference to the drawings, which are as follows: Figure 1 one or more functions specified by one or more blocks Figure 1 an apparatus with one or more functions specified by one or more blocks

[0222] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the Figure 1 one or more functions specified by one or more blocks Figure 1 an apparatus with one or more functions specified by one or more blocks

[0223] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the Figure 1 one or more functions specified by one or more blocks Figure 1 Figure 1 an apparatus with one or more functions specified by one or more blocks

[0224] While preferred embodiments of the present application have been described, additional variations and modifications can be made to these embodiments by those skilled in the art once they have the benefit of the foregoing description. Therefore, the appended claims are intended to encompass within their scope all such variations and modifications as are included within the scope of the present application.

[0225] Finally, it is to be understood that the phraseology or terminology such as "first" and "second" etc. used herein is merely intended to differentiate one entity or operation from another entity or operation, without necessarily requiring or implying any actual such relationship or order between such entities or operations. Moreover, the terms "comprising", "including", or any other closure, are intended to cover the non-exclusive inclusion such that a process, method, article, or apparatus that comprises a list of elements does not include those elements alone but can include other elements not expressly listed or even include elements inherent in such process, method, article, or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus including the element.

[0226] The conference control method and the conference control device provided by the present application are described in detail above, and the principles and implementation modes of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed, and the above description should not be understood as a limitation of the present application.

Claims

1. A conference control method characterized by, The conference control method is applied to a WebRTC mobile terminal, and the method comprises: After the WebRTC conference is successfully established, each built-in camera and external camera of the mobile terminal is started to collect first video data streams; A second video data stream sent by an interconnected external electronic device is acquired, wherein the second video data stream is a video data stream collected by a camera of the external electronic device, or a second video data stream obtained by the external electronic device by screen capturing; Each first video data stream collected by each camera is cached in a corresponding first cache area, and the second video data stream is cached in a second cache area, wherein the number of the first cache areas is the same as the total number of the started built-in cameras and external cameras; Video images are extracted from each first cache area and the second cache area respectively, and video image synthesis is performed to combine the first video data streams and the second video data stream into one video data stream; After the synthesized video data stream is encoded, it is sent to a WebRTC server through WebRTC, and then sent to each WebRTC terminal participating in the conference through the WebRTC server; The step of extracting video images from each first cache area and the second cache area respectively and performing video image synthesis comprises: A first video image with the shortest storage time is extracted from each first cache area; A second video image with the shortest storage time is extracted from the second cache area; Each first video image and the second video image are scaled and spliced to generate a synthesized video image with a preset resolution; when the first video data stream includes a video data stream generated by screen capturing the local screen at a preset frequency, the synthesized video image is not displayed locally; when the first video data stream does not include a video data stream generated by screen capturing the local screen at a preset frequency, the synthesized video image is displayed locally.

2. The method of claim 1, wherein, Before the step of starting each built-in camera and external camera of the mobile terminal to collect first video data streams, the method further comprises: After the WebRTC conference is successfully established, it is determined whether to start a multi-path video participation mode according to historical video conference records; If it is determined to start the multi-path video participation mode, the step of starting each built-in camera and external camera of the mobile terminal to collect first video data streams is performed.

3. The method of claim 1, wherein, Before the step of starting each built-in camera and external camera of the mobile terminal to collect first video data streams, the method further comprises: After the WebRTC conference is successfully established, a selection prompt box is output, wherein the selection prompt box is used to prompt a user to start a multi-path video participation mode; A first input of the user to the selection prompt box is received, wherein the first input is used to start the multi-path video participation mode; In response to the first input, the step of starting each built-in camera and external camera of the mobile terminal to collect first video data streams is performed.

4. The method of claim 1, wherein, Before the step of scaling and splicing each of the first video image and the second video image to generate a synthetic video image of a preset resolution, the method further comprises: According to the number of each of the first video image and the second video image, a plurality of synthetic video image styles are recommended to be displayed; Receiving a user selection operation on a target synthetic video image style.

5. A conference control apparatus characterized by comprising: The conference control device is applied to a WebRTC mobile terminal, and the device comprises: An opening module configured to open each built-in camera and an external camera of the mobile terminal to collect first video data streams after a WebRTC conference is successfully established; An acquisition module configured to acquire second video data streams sent by interconnected external electronic devices, wherein the second video data streams are video data streams collected by the external electronic devices through cameras or second video data streams obtained by the external electronic devices through screen capture; A cache module configured to cache the first video data streams collected by each camera in corresponding first cache areas and cache the second video data streams in a second cache area, wherein the number of the first cache areas is the same as the total number of the opened built-in cameras and external cameras; A synthesis module configured to extract video images from each of the first cache areas and the second cache area, perform video image synthesis, and combine the first video data streams and the second video data streams into one video data stream; A sending module configured to encode the synthesized video data stream and send it to a WebRTC server through the WebRTC, so as to send it to each WebRTC terminal participating in the conference through the WebRTC server; The synthesis module comprises: A first sub-module configured to extract the first video image with the shortest storage time from each of the first cache areas; A second sub-module configured to extract the second video image with the shortest storage time from the second cache area; A third sub-module configured to scale and splice each of the first video image and the second video image to generate a synthetic video image of a preset resolution; A display module configured to not display the synthesized video image on the local terminal when the first video data stream includes a video data stream generated by capturing the local screen at a preset frequency, and display the synthesized video image on the local terminal when the first video data stream does not include a video data stream generated by capturing the local screen at a preset frequency.

6. The apparatus of claim 5, wherein, The device further comprises: A mode determination module configured to determine whether to open a multi-path video participation mode according to historical video conference records before the opening module opens each built-in camera and external camera of the mobile terminal to collect first video data streams after the WebRTC conference is successfully established; A first calling module configured to call the opening module to open each built-in camera and external camera of the mobile terminal to collect first video data streams when it is determined to open the multi-path video participation mode.

7. The apparatus of claim 5, wherein, The device further comprises: The output module is configured to output a selection prompt box before the start module starts the built-in camera and the external camera of the mobile terminal to collect the first video data stream after the Web conference is successfully established, wherein the selection prompt box is configured to prompt a user to start the multi-path video attending mode. The input receiving module is configured to receive a first input of the selection prompt box by the user, wherein the first input is configured to start the multi-path video attending mode. The second calling module is configured to call the start module to start the built-in camera and the external camera of the mobile terminal to collect the first video data stream in response to the first input.

8. The apparatus of claim 5, wherein, The synthesis module further includes: The fourth sub-module is configured to recommend a plurality of synthesis video image styles according to the number of the first video images and the second video images before the third sub-module scales and splices the first video images and the second video images to generate a synthesis video image with a preset resolution. The fifth sub-module is configured to receive a selection operation of a target synthesis video image style by the user.

9. The apparatus of claim 5, wherein, The apparatus further includes: The screenshot module is configured to take screenshots of the local screen at a preset frequency to generate a first video data stream.

10. A video call device, characterized by The apparatus includes: One or more processors; and One or more machine readable media having instructions stored thereon that, when executed by the one or more processors, cause the apparatus to perform the conference control method of any one of claims 1 to 4.

11. A computer readable storage medium, characterized in that, The computer program stored therein causes the processor to perform the conference control method of any one of claims 1 to 4.

Citation Information

Patent Citations

  • A method and system for processing a composite video image

    CN101828392A

  • Real-time splicing method and apparatus of multiple videos

    CN106454256A

  • Video communication method, terminal and computer readable storage medium

    CN109672843A

  • Video telephone device and communication method

    JP2013026782A