Method, device and equipment for processing video conference based on ultra-low bandwidth and medium
By detecting key video areas in video conferencing and synthesising video streams, the problem of video stream quality degradation due to terminal performance differences and bandwidth fluctuations is solved, and the transmission of high-definition video streams is realized, improving the picture quality of video conferencing.
Patent Information
- Application Number
- CN202311839647.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-25
AI Technical Summary
In video conferencing, due to the performance differences and bandwidth fluctuations of the conference terminals, the video streams uploaded by some terminals are not high-definition, resulting in the problem of lag.
By detecting the video key areas in the video stream, video stream synthesis is performed, video streams are generated to replace the video key areas, and sent to other conference terminals to improve the video stream quality.
It effectively compensates for the decline in video stream quality caused by non-high-definition shooting and bandwidth fluctuations, ensures that the video stream can meet the needs and improves the picture quality of video conferencing.
Smart Images

Figure CN120378568A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and in particular, to a method, apparatus, device, and medium for processing video conferences based on ultra-low bandwidth. Background Art
[0002] With the development of communication technologies, users have higher and higher requirements for the quality and efficiency of communication, and the demands are becoming more and more diverse and differentiated. In the video conferencing scenario, users are not satisfied with merely being able to see real-time video images, and the demand for high-definition, high-quality, and high-stability video conferencing services is becoming stronger and stronger.
[0003] In a video conference, the video stream captured by the conference terminal is uploaded to the server, and then distributed by the server to the display devices at each venue for display. However, due to the different performances of the conference terminals, some conference terminals may upload high-definition video streams, while some may not. If all conference terminals are replaced with high-definition cameras, video conference images may freeze during video stream transmission due to large bandwidth fluctuations or other reasons. Summary of the Invention
[0004] In view of the above problems, a method, apparatus, device, and medium for processing video conferences based on ultra-low bandwidth are provided to overcome or at least partially solve the above problems, including:
[0005] A method for processing a video conference based on ultra-low bandwidth, the method includes:
[0006] When a trigger event for a target conference terminal is detected, determining a video key area in the video stream sent by the target conference terminal;
[0007] For the video key area, performing video stream synthesis to obtain a video stream for replacing the video key area, and sending the video stream to other conference terminals.
[0008] Optionally, the performing video stream synthesis for the video key area to obtain a video stream for replacing the video key area includes:
[0009] Obtaining a data stream at least including audio data collected by the target conference terminal;
[0010] According to the data stream at least including audio data, performing video stream synthesis to obtain a video stream for replacing the video key area, and sending the video stream to other conference terminals.
[0011] Optionally, the performing video stream synthesis according to the data stream at least including audio data to obtain a video stream for replacing the video key area includes:
[0012] Obtain a target image corresponding to the key video region, and perform video stream synthesis based on the data stream containing at least audio data and the target image to obtain a video stream for replacing the key video region.
[0013] Optionally, the data stream further includes video feature attribute information generated based on the video stream collected by the target conference terminal. The performing video stream synthesis based on the data stream containing at least audio data and the target image to obtain a video stream for replacing the key video region includes:
[0014] Perform video stream synthesis based on the data stream containing at least audio data and video feature attribute information, and the target image, to obtain a video stream for replacing the key video region.
[0015] Optionally, when the target conference terminal is in the first shooting mode, the data stream contains audio data; when the target conference terminal is in the second shooting mode, the data stream contains audio data and video feature attribute information.
[0016] Optionally, before sending the video stream to other conference terminals, it further includes:
[0017] Create a simulated user for the target conference terminal, and send a notification message to the other conference terminals; wherein, the notification message is used to control the other conference terminals to modify the subscription to the original user of the target conference terminal to a subscription to the simulated user;
[0018] The sending the video stream to other conference terminals includes:
[0019] Send the video stream to the other conference terminals through the simulated user.
[0020] Optionally, the trigger event for the target conference terminal includes any one of the following:
[0021] Receiving a video synthesis request sent by the target conference terminal, the bandwidth fluctuation amplitude of the channel connected to the target conference terminal being greater than a threshold, and there being a key video region in the video stream sent by the target conference terminal with a clarity less than the threshold.
[0022] An apparatus for processing video conferences based on ultra-low bandwidth, the apparatus includes:
[0023] A key video region determination module, configured to determine a key video region in the video stream sent by the target conference terminal when detecting a trigger event for the target conference terminal;
[0024] A video stream synthesis module is configured to perform video stream synthesis for the video key area, obtain a video stream for replacing the video key area, and send the video stream to other conference terminals.
[0025] An electronic device includes a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the method for processing a video conference based on ultra-low bandwidth as described above is implemented.
[0026] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the method for processing a video conference based on ultra-low bandwidth as described above is implemented.
[0027] The embodiments of the present invention have the following advantages:
[0028] In the embodiments of the present invention, when a trigger event for a target conference terminal is detected, the video key area in the video stream sent by the target conference terminal is determined. Then, for the video key area, video stream synthesis is performed to obtain a video stream for replacing the video key area, and the video stream is sent to other conference terminals. This realizes the detection of the video key area in a video conference and the video synthesis for the video key area, and can thus make up for the situation where the video stream is difficult to meet the requirements due to reasons such as non-high-definition shooting and large bandwidth fluctuations, improving the quality of the video stream. Description of the Drawings
[0029] To more clearly illustrate the technical solutions of the present invention, the drawings required for the description of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0030] Figure 1 is a flowchart of the steps of a method for processing a video conference based on ultra-low bandwidth provided by some embodiments of the present invention;
[0031] Figure 2 is a system architecture diagram of a video conference provided by some embodiments of the present invention;
[0032] Figure 3 is a flowchart of the steps of another method for processing a video conference based on ultra-low bandwidth provided by some embodiments of the present invention;
[0033] Figure 4 is a structural block diagram of a device for processing a video conference based on ultra-low bandwidth provided by some embodiments of the present invention. Detailed Embodiments
[0034] To make the above objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0035] Referring to Figure 1 , a flowchart of steps of a method for processing video conferences based on ultra-low bandwidth provided by some embodiments of the present invention is shown. This method can be applied to a server in a video conference.
[0036] In some examples, the server can be a single server or multiple servers. For example, Figure 2 , the server can include a conference server (such as an all-in-one machine, XMCU, etc.) and a video post-processing server (which can also be an AI server). The conference server and the video post-processing server are connected to each other. This method can be partially executed on the conference server and partially executed on the video post-processing server. For example, the relevant content of video stream synthesis can be executed by the video post-processing server. The relevant content of video stream synthesis can include AI processing (such as determining digital portraits, super-resolution reconstruction, etc.) and streaming media processing (such as audio encoding and decoding, video encoding and decoding). Other content can be executed by the conference server.
[0037] Among them, the video post-processing server (AI server) is a data server that can provide artificial intelligence. It can be used to support local applications and web pages, and can also provide complex AI models and services for cloud and local servers.
[0038] In some examples, the video post-processing server can be provided with an AI conference management module, a user management module, and a proxy service module. Among them, the AI conference management module can be used to manage the establishment of a communication channel with the conference server, calculate and classify the received data stream, and at the same time encode and decode the audio stream and video stream to synthesize a high-definition digital portrait video stream. The user management module can be used to manage user information in a video conference. When a conference terminal needs to create a high-definition digital portrait video stream, user information of a conference terminal in a certain video conference will be created in this user management module. The user name, portrait data, etc. used in this user information can be synchronized from the conference server to the AI server when the conference terminal starts the conference, mainly including the conference ID, user ID when logging in to the conference terminal, image ID, identification of the conference terminal, etc. The proxy module can be used to establish a communication channel with the conference server, and then can implement functions such as file upload and file download. After the proxy module establishes a communication channel with the conference server, it can forward files.
[0039] Specifically, the following steps may be included:
[0040] Step 101, when a trigger event for a target conference terminal is detected, determine the video key area in the video stream sent by the target conference terminal.
[0041] Wherein, the target conference terminal and other conference terminals are participating terminals in the same video conference.
[0042] In a video conference, some conference terminals may use non-high-definition shooting, or due to reasons such as large bandwidth fluctuations during upload, the video stream is difficult to meet the requirements, such as low video clarity. Then, when a trigger event is detected, the video key area in the video stream can be determined for optimization.
[0043] In some examples, when a trigger event is detected, it can be switched from the general processing mode (that is, directly uploading the video stream collected by the target conference terminal to the server, and then the server sending it to other conference terminals) to the AI processing mode, and then the video key area is optimized in the AI processing mode.
[0044] In some examples, the video key area may be a video frame in the video stream with a clarity less than a threshold (such as a resolution less than 1024), that is, a low-definition area. For example, the video key area may include the portrait area of the speaker and the environmental area (i.e., the venue environment) around the portrait area of the speaker.
[0045] In practical applications, the server can monitor the video stream sent by the target conference terminal, such as detecting which video frames have a clarity less than the threshold, and then the video key area can be determined from the video stream. In some examples, the conference server can forward the video stream to the video post-processing server, and then the video post-processing server monitors the video key area in the video stream.
[0046] In some embodiments of the present invention, the trigger event for the target conference terminal includes any one of the following: receiving a video synthesis request sent by the target conference terminal, the bandwidth fluctuation amplitude of the channel connected to the target conference terminal being greater than the threshold, and the video stream sent by the target conference terminal having a video key area with a clarity less than the threshold.
[0047] In some embodiments, an interactive interface can be provided to the user on the target conference terminal. The user can perform manual operations through the interactive interface, and then generate and send a video synthesis request to the server through the target conference terminal. When the server receives the video synthesis request, a trigger event is detected. For example, when the camera connected to the target conference terminal can only capture a video stream with low clarity and cannot capture a high-definition video stream, the user can control the target conference terminal to generate and send a video synthesis request through the interactive interface; another example is that due to large bandwidth fluctuations or other reasons, the video stream uploaded by the target conference terminal has a situation of video conference picture freezing, and the user can control the target conference terminal to generate and send a video synthesis request through the interactive interface.
[0048] In some other embodiments, the server can establish communication channels with each participating terminal. The server can be provided with a bandwidth fluctuation monitoring module, which can detect the bandwidth fluctuation conditions of each communication channel, and a certain bandwidth threshold can be set for each channel. After the video conference is started, the server can detect whether the bandwidth fluctuation amplitude of the channel connected to the target conference terminal is greater than the threshold (such as 100K) through the bandwidth fluctuation monitoring module. When it is detected that the bandwidth fluctuation amplitude of the channel connected to the target conference terminal is greater than the threshold, a trigger event is detected.
[0049] In some other embodiments, the server can perform real-time monitoring on the video stream sent by the target conference terminal, such as detecting which video frames have a clarity less than the threshold. When it is monitored that there is a video key area in the video stream with a clarity less than the threshold, a trigger event is detected. In some examples, the conference server can forward the video stream to the video post-processing server, and then the video post-processing server monitors the video key area in the video stream, and then can notify the conference server for subsequent processing.
[0050] In some examples, a scheduling platform can be set up. The scheduling platform can be used to control the number of cameras and microphones that allow users to turn on and off, and also needs to schedule whether to allow users to turn on the AI function. In some examples, affected by the computing power of the background GPU cluster, it may not be allowed for each user to randomly turn on and off the AI enhancement and synthesis functions.
[0051] In some examples, the server can be initialized first. Then, the scheduling platform, the conference server, and the video post - processing server can establish TCP Socket connections. The participating terminals can perform user logins. After the user login is successful, the conference terminal in charge of hosting can start a conference through the scheduling platform. The scheduling platform can then control the server to open a conference room and feedback conference information to the conference terminal in charge of hosting. Other participating terminals can join the conference room through the conference information, and the scheduling platform can control the participating terminals to turn on the cameras and microphones and can also control the activation of the video synthesis function.
[0052] Step 102: Perform video stream synthesis for the video key area to obtain a video stream for replacing the video key area, and send the video stream to other conference terminals.
[0053] After determining the video key area, video stream synthesis can be performed on the video key area. The video key area in the original video stream is replaced with the synthesized video stream, and the new video stream is sent to other conference terminals, thereby ensuring that the video stream can meet the requirements.
[0054] In some examples, this video stream synthesis process is a super - resolution reconstruction process. Super - resolution reconstruction is to improve the resolution of the original image through hardware or software methods. The process of obtaining a high - resolution image from a series of low - resolution images is super - resolution reconstruction. Specifically, when a video conference encounters lags or when the clarity of the video stream captured by a conference terminal is low, it is switched to a high - definition video stream to improve the quality of the played video and ensure the normal progress of the video conference.
[0055] In some embodiments of the present invention, the step of performing video stream synthesis for the video key area to obtain a video stream for replacing the video key area includes:
[0056] Obtain a data stream collected by the target conference terminal that at least includes audio data; perform video stream synthesis based on the data stream that at least includes audio data to obtain a video stream for replacing the video key area, and send the video stream to other conference terminals.
[0057] In a video conference, the target conference terminal can collect a video stream and upload it to the server. Then, the server sends the video stream to other conference terminals for display through a display device.
[0058] In the case where the video stream sent by the target conference terminal does not meet the requirements, that is, when it is necessary to synthesize the video stream for the key video area, the server can switch from obtaining the video stream collected by the target conference terminal (collected by the microphone and camera) to obtaining the data stream collected by the target conference terminal that contains at least audio data (collected by the microphone).
[0059] If a data stream containing at least audio data is obtained, the server can use the data stream containing at least audio data to synthesize the video stream, obtain the video stream for replacing the key video area, and can send the synthesized video stream to other conference terminals to ensure the high definition of the video stream in the video conference.
[0060] In some examples, when synthesizing the video stream, since the server only uses the data stream collected by the target conference terminal that contains at least audio data and no longer uses the video stream, the target conference terminal can stop sending the video stream to the server and only send the data stream that contains at least audio data.
[0061] In some embodiments of the present invention, it further includes:
[0062] Sending a notification message to the target conference terminal; wherein, the notification message is used to control the target conference terminal to stop sending the video stream and send the data stream that contains at least audio data.
[0063] After detecting the trigger event, the server can send a notification message to the target conference terminal. After receiving the notification message, the target conference terminal can stop sending the video stream (it can also stop collecting the video stream), and then can send the data stream that contains at least audio data (the data stream that contains at least audio data does not contain video data) collected to the server.
[0064] In some examples, such as Figure 2 , the conference server can detect whether a trigger event occurs. When detecting the trigger event, the conference server sends a notification message to the target conference terminal, establishes a communication connection between the conference server and the video post-processing server, sends the data stream that contains at least audio data sent by the target conference terminal to the video post-processing server, and the video post-processing server performs relevant operations on video synthesis according to the data stream that contains at least audio data.
[0065] In the embodiments of the present invention, by synthesizing the video stream according to the audio data uploaded by the conference terminal, the requirements for bandwidth and other aspects of the data uploaded by the conference terminal are reduced. It is possible to upload data with a lower bandwidth and keep the high-definition video conference images displayed on other terminals, avoiding the situation of video conference image jamming due to large fluctuations in bandwidth and other reasons.
[0066] In some examples, the target conference terminal can encapsulate the data stream containing at least audio data obtained by collection in a signaling message, and send it to the conference server through the signaling message. The conference server forwards it to the video post-processing server. The video post-processing server calls the streaming media library for decoding processing, and sends the decoded data stream to the AI module. The AI module performs AI synthesis processing, and then can send the synthesized video stream to other conference terminals, and display the synthesized video stream on other conference terminals.
[0067] In some examples, the target conference terminal needs to be authorized by the scheduling platform before sending the data stream to the server. This platform can be set to automatically trigger the authorization process when certain conditions are met, or it can be defaulted that the scheduling platform has been authorized. For example, when the conference terminal sends a request authorization signaling message to the conference server, the scheduling platform needs to confirm the authorization. When the conference server sends an authorization notification signaling message to the conference terminal, the scheduling platform can default the authorization.
[0068] In some embodiments of the present invention, the synthesizing a video stream according to the data stream containing at least audio data to obtain a video stream for replacing the video key area includes:
[0069] Obtain a target image corresponding to the video key area, and perform video stream synthesis according to the data stream containing at least audio data and the target image to obtain a video stream for replacing the video key area.
[0070] In practical applications, the target conference terminal can upload the target image to the server, and can receive the data stream containing audio data, decode it by the media library and convert it into text information. Then, the target image and the text information can be input into the super-resolution reconstruction model to generate a video stream for replacing the video key area, that is, output a video stream of a human figure speaking.
[0071] In some examples, the video stream for replacing the video key area can be used to present a digital human figure. A digital human is a humanoid image made by computer technology or a product made by computer software. They have the appearance or behavior pattern of a human, but they are not a video of a certain person in the real world. They can run and exist independently. The ontology of the digital human exists in a computing device (such as a computer, a mobile phone, a VR headset, etc.), and is presented through a display device, so that humans can see it with their eyes or can have voice interaction. They have an independent personality setting, a specific identity, an independent name, an independent character image, and an independent knowledge base to answer specific questions.
[0072] Among them, the super-resolution reconstruction model is a semi-supervised model (a model trained by inputting portrait data and semantic information and after manual intervention, which is more accurate than full-supervised models and supervised models), synthesizes the generated video stream of the portrait speaking and the audio stream to generate an audio-visual stream of a high-definition digital portrait, and encodes it at the same time.
[0073] In some examples, the target image may include the image of the venue and the portrait image of the participants corresponding to the target conference terminal. The target image may be a real portrait taken before the meeting starts. Then, the video stream for replacing the key area of the video may include replacing the following content:
[0074] The real portrait captured in real time in the original video stream → replaced with → the real portrait taken before the meeting starts → the video stream for replacing the key area of the video
[0075] The meeting scene captured in real time in the original video stream → replaced with → the meeting scene taken before the meeting starts → the video stream for replacing the key area of the video
[0076] In practical applications, the target conference terminal can determine the target image and synchronize it to the server. For example, the target conference terminal sends the high-definition venue image and the portrait pictures of the participants (the conference terminal will classify and number the received photo set) to the conference server, and the conference server forwards them to the video post-processing server, which is temporarily stored in the media library (along with information such as the conference identifier).
[0077] In some examples, the high-definition scene and the portrait pictures of the participants can be directly captured by the camera of the target conference terminal, or when the participants sign in, the signing device takes the portrait pictures and then sends them to the target conference terminal.
[0078] In some examples, after the target conference terminal obtains the target image, it can classify and number the picture set so that the server can quickly find the corresponding image when synthesizing the video stream.
[0079] For example, the numbering rule is: conference scene ID + secondary number (classification) + tertiary number.
[0080] The one-digit secondary number represents the scene, the two-digit secondary number represents the portrait (one number for one person can be used), and the tertiary number is a refinement of the secondary number. For example, the conference scene ID: 10001-1, 10001 represents the first scene under the corresponding conference number 10001, and the person image ID: 10001-01, 10001-01 represents the first person under the conference number 10001.
[0081] In some related technologies, the server can find high-definition pictures in the media library in non-conference scenarios or replace pictures of other people. The video stream synthesized by this direct replacement method is separated from the real conference scenario and does not seem like a conference held in a real scenario, resulting in a poor viewing experience for users.
[0082] In the embodiments of the present invention, by pre-shooting high-definition venue images and venue portraits, when the switching condition is met, the AI processing mode is triggered. The AI server will directly use the pre-shot high-definition venue images and venue portraits, and according to the real conference scenario, replace them with a video stream corresponding to the current real conference scenario / real portrait, maximizing the restoration of the real conference scenario, making the viewing users feel that the digital portrait has not been replaced, and at the same time improving the quality of the played video.
[0083] In some embodiments of the present invention, the data stream collected by the target conference terminal and containing at least audio data may further include video feature attribute information generated according to the video stream collected by the target conference terminal. The video feature attribute information may be generated by the target conference terminal according to the video stream collected by itself and is sent to the server together with the video data in the data stream.
[0084] In some examples, the video feature attribute information may include real-time facial micro-expression information for motion compensation and reconstruction, important area texture light and shadow information, hand motion capture data information, etc. In some examples, audio feature attribute information may also be generated according to the audio data, so as to better integrate the audio and video, including one-dimensional timing information for multi-modal alignment and video generation, audio compression / restoration inference compensation information, and keyword information for predicting and perceiving key speech time points, etc.
[0085] In some embodiments of the present invention, when the target conference terminal is in the first shooting mode, the data stream collected by the target conference terminal and uploaded to the server may include audio data, that is, only audio data is sent (video capture is turned off and only audio capture is turned on). For example, the portrait of the real speaker is sitting and only the part above the shoulders is exposed, and there are basically no body limb movements, that is, only the part above the shoulders of the user is photographed (half-body shooting mode). In this shooting mode, the synthesized digital portrait also only has the upper body part, so only audio data needs to be sent.
[0086] In some embodiments of the present invention, when the target conference terminal is in the second shooting mode, the data stream collected and uploaded by the target conference terminal to the server may include audio data and video feature attribute information (audio collection and video collection are enabled, but the collected video is only used to generate video feature attribute information and the video stream is not uploaded). For example, if the real person speaking is standing and has many body movements, that is, it is necessary to shoot the body movements of the user (full body shooting mode), the synthesized digital portrait in this shooting mode can be a full body portrait, that is, it is necessary to send audio data and video feature attribute information.
[0087] In some embodiments of the present invention, before the meeting starts, the conference terminal responsible for hosting can manually input: half body or full body shooting, or it can be determined by the conference server by obtaining the video conference screen as half body or full body shooting.
[0088] In some embodiments of the present invention, the synthesizing a video stream based on the data stream at least including audio data and the target image to obtain a video stream for replacing the video key area includes: synthesizing a video stream based on the data stream at least including audio data and video feature attribute information, and the target image, to obtain a video stream for replacing the video key area.
[0089] When the data stream sent by the target conference terminal includes audio data and video feature attribute information, the target image synchronized when the user created the meeting can be found, and the received data stream including audio data can be decoded by the media library and converted into text information, and then the target image, text information, and video feature attribute information are input into the super-resolution reconstruction model to generate a video stream for replacing the video key area, that is, output a video stream of the person speaking.
[0090] In some embodiments of the present invention, before sending the video stream to other conference terminals, it further includes: creating a simulated user for the target conference terminal and sending a notification message to the other conference terminals; wherein, the notification message is used to control the other conference terminals to modify the subscription to the original user of the target conference terminal to a subscription to the simulated user;
[0091] In some embodiments of the present invention, the sending the video stream to other conference terminals includes:
[0092] Sending the video stream to the other conference terminals through the simulated user.
[0093] In the case where no trigger event is detected, the server sends the video stream collected by the target conference terminal to other conference terminals, and other conference terminals subscribe to the data of the user corresponding to the target conference terminal. After detecting the trigger event, since the target conference terminal no longer sends the video stream, a simulated user for the target conference terminal can be created, and a notification message (broadcast) can be sent to other conference terminals. The notification message can include information about the simulated user. After receiving the notification message, other conference terminals can modify the subscription to the original user of the target conference terminal to a subscription to the simulated user.
[0094] After synthesizing the video stream based on the data stream sent by the target conference terminal, the server can publish the video stream in the name of the simulated user, and other conference terminals subscribing to the simulated user can then receive the video stream sent by the simulated user through the server.
[0095] In the embodiments of the present invention, by determining the video key area in the video stream sent by the target conference terminal when detecting a trigger event for the target conference terminal, then performing video stream synthesis on the video key area to obtain a video stream for replacing the video key area, and sending the video stream to other conference terminals, it realizes detecting the video key area in a video conference and performing video synthesis on the video key area, thereby being able to make up for the situation where the video stream is difficult to meet the requirements due to reasons such as non-high-definition shooting and large bandwidth fluctuations, and improving the quality of the video stream.
[0096] Refer to Figure 3 , which shows a flowchart of steps of another method for processing video conferences based on ultra-low bandwidth provided by some embodiments of the present invention, and specifically may include the following steps:
[0097] Step 301, when detecting a trigger event for the target conference terminal, determine the video key area in the video stream sent by the target conference terminal.
[0098] Step 302, obtain the data stream collected by the target conference terminal that at least includes audio data.
[0099] Step 303, based on the data stream that at least includes audio data, perform video stream synthesis to obtain a video stream for replacing the video key area, and send the video stream to other conference terminals.
[0100] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0101] Referring to Figure 4 , a schematic structural diagram of a device for processing video conferences based on ultra-low bandwidth provided by some embodiments of the present invention is shown, which may specifically include the following modules:
[0102] A video key area determination module 401, configured to determine a video key area in a video stream sent by the target conference terminal when a trigger event for the target conference terminal is detected;
[0103] A video stream synthesis module 402, configured to perform video stream synthesis on the video key area to obtain a video stream for replacing the video key area, and send the video stream to other conference terminals.
[0104] In some embodiments of the present invention, the performing video stream synthesis on the video key area to obtain a video stream for replacing the video key area includes:
[0105] Obtaining a data stream collected by the target conference terminal and containing at least audio data;
[0106] Performing video stream synthesis according to the data stream containing at least audio data to obtain a video stream for replacing the video key area, and sending the video stream to other conference terminals.
[0107] In some embodiments of the present invention, the performing video stream synthesis according to the data stream containing at least audio data to obtain a video stream for replacing the video key area includes:
[0108] Obtaining a target image corresponding to the video key area, and performing video stream synthesis according to the data stream containing at least audio data and the target image to obtain a video stream for replacing the video key area.
[0109] In some embodiments of the present invention, the data stream further contains video feature attribute information generated according to the video stream collected by the target conference terminal. The performing video stream synthesis according to the data stream containing at least audio data and the target image to obtain a video stream for replacing the video key area includes:
[0110] Perform video stream synthesis based on a data stream that includes at least audio data and video feature attribute information, and the target image, to obtain a video stream for replacing the video key area.
[0111] In some embodiments of the present invention, when the target conference terminal is in the first shooting mode, the data stream includes audio data; when the target conference terminal is in the second shooting mode, the data stream includes audio data and video feature attribute information.
[0112] In some embodiments of the present invention, it further includes:
[0113] A notification message sending module, configured to create a simulated user for the target conference terminal and send a notification message to the other conference terminals; wherein, the notification message is used to control the other conference terminals to modify the subscription to the original user of the target conference terminal to a subscription to the simulated user.
[0114] In some embodiments of the present invention, the sending the video stream to other conference terminals includes:
[0115] A simulated user sending module, configured to send the video stream to the other conference terminals through the simulated user.
[0116] In some embodiments of the present invention, the trigger event for the target conference terminal includes any one of the following:
[0117] Receiving a video synthesis request sent by the target conference terminal, the bandwidth fluctuation amplitude of the channel connected to the target conference terminal being greater than a threshold, and the video key area of the video stream sent by the target conference terminal having a clarity less than the threshold.
[0118] In an embodiment of the present invention, by determining the video key area in the video stream sent by the target conference terminal when a trigger event for the target conference terminal is detected, then performing video stream synthesis for the video key area to obtain a video stream for replacing the video key area, and sending the video stream to other conference terminals, it realizes detecting the video key area in a video conference and performing video synthesis on the video key area, and thus can make up for the situation where the video stream is difficult to meet the requirements due to reasons such as non-high-definition shooting and large bandwidth fluctuations, and improves the quality of the video stream.
[0119] Some embodiments of the present invention further provide an electronic device, which may include a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, it implements the method for processing a video conference based on ultra-low bandwidth as described above.
[0120] Some embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for processing a video conference based on ultra-low bandwidth as described above is implemented.
[0121] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments.
[0122] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to select to authorize or reject.
[0123] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0124] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0125] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the method, terminal device (system), and computer program product according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0126] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified in one block or multiple blocks.
[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operational steps are performed on the computer or other programmable terminal device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified in one block or multiple blocks.
[0128] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0129] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the above elements.
[0130] The above has introduced in detail the method, apparatus, device and medium for processing video conferences based on ultra-low bandwidth. Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for processing video conferencing based on ultra-low bandwidth, characterized in that, The method includes: When a trigger event for a target conference terminal is detected, determining a video key area in the video stream sent by the target conference terminal; For the video key area, performing video stream synthesis to obtain a video stream for replacing the video key area, and sending the video stream to other conference terminals.
2. The method according to claim 1, wherein The performing video stream synthesis for the video key area to obtain a video stream for replacing the video key area includes: Obtaining a data stream collected by the target conference terminal and containing at least audio data; According to the data stream containing at least audio data, performing video stream synthesis to obtain a video stream for replacing the video key area, and sending the video stream to other conference terminals.
3. The method according to claim 2, wherein The performing video stream synthesis according to the data stream containing at least audio data to obtain a video stream for replacing the video key area includes: Obtaining a target image corresponding to the video key area, and performing video stream synthesis according to the data stream containing at least audio data and the target image to obtain a video stream for replacing the video key area.
4. The method according to claim 3, characterized in that, The data stream further contains video feature attribute information generated according to the video stream collected by the target conference terminal. The performing video stream synthesis according to the data stream containing at least audio data and the target image to obtain a video stream for replacing the video key area includes: Performing video stream synthesis according to the data stream containing at least audio data and video feature attribute information, and the target image, to obtain a video stream for replacing the video key area.
5. The method according to claim 4, wherein When the target conference terminal is in the first shooting mode, the data stream contains audio data; when the target conference terminal is in the second shooting mode, the data stream contains audio data and video feature attribute information.
6. The method according to any one of claims 1 to 5, characterized in that Before sending the video stream to other conference terminals, it further includes: Creating a simulated user for the target conference terminal, and sending a notification message to the other conference terminals; wherein, the notification message is used to control the other conference terminals to modify the subscription to the original user of the target conference terminal to a subscription to the simulated user; The sending the video stream to other conference terminals includes: Sending the video stream to the other conference terminals through the simulated user.
7. The method according to claim 1, wherein The trigger event for the target conference terminal includes any one of the following: Receiving a video synthesis request sent by the target conference terminal, the bandwidth fluctuation amplitude of the channel connected to the target conference terminal being greater than a threshold, and there being a video key area with a clarity less than the threshold in the video stream sent by the target conference terminal.
8. An apparatus for processing video conferencing based on ultra-low bandwidth, characterized in that, The device includes: A video key area determination module, configured to determine a video key area in the video stream sent by the target conference terminal when a trigger event for the target conference terminal is detected; A video stream synthesis module, configured to perform video stream synthesis for the video key area to obtain a video stream for replacing the video key area, and send the video stream to other conference terminals.
9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, it implements the method for processing a video conference based on ultra-low bandwidth as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the method for processing a video conference based on ultra-low bandwidth as described in any one of claims 1 to 7.