A video conference control method, device and computing device cluster

By using a cloud-based conferencing system to replace real-time video data distribution with a replacement video when the video upload status is abnormal, the problem of abnormal video images in video conferences has been solved, and the viewing experience of participants has been improved.

CN122457735APending Publication Date: 2026-07-24HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-01-22
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In video conferencing, sudden abnormalities in the participants' images (such as lag or black screen) can negatively impact the viewing experience.

Method used

When the cloud conferencing system detects an abnormal video upload status, it generates a video replacement instruction and uses the replacement video to replace the real-time video data distribution, ensuring smooth video playback.

Benefits of technology

To prevent video stuttering or black screen when the video upload status is abnormal, and to improve the viewing experience for participants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122457735A_ABST
    Figure CN122457735A_ABST
Patent Text Reader

Abstract

A video conference control method applied to a cloud conference system, the cloud conference system comprising a first client and at least one participant client, the first client and the participant client establishing a data connection through the cloud, the method comprising: in a multi-person video conference scene, the cloud conference system detecting a video replacement instruction and determining a replacement video, wherein the video replacement instruction is used to instruct to use the replacement video to replace video data acquired from the first client in real time and distribute to the participant client; and sending the replacement video to the participant client and no longer sending the video data acquired from the first client in real time to the participant client. Through the method, when, for example, a video uploading state is an abnormal state or a user of the first client does not want to upload his / her current screen, no screen freezing or black screen phenomenon occurs in a related sub-window, and the viewing experience of a participant is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video conferencing technology, and in particular to a video conferencing control method, apparatus, and computing device cluster. Background Technology

[0002] With the development of communication technology, video conferencing applications are becoming increasingly widespread, not only in traditional office meetings but also in telemedicine, education and training, remote maintenance, emergency command, and video customer service. However, during video conferences, situations frequently arise where one or more participants' screens suddenly become abnormal (such as freezing or going black), which can negatively impact their viewing experience. Therefore, how to ensure an uninterrupted viewing experience during video conferences, especially in the event of unforeseen circumstances, is a pressing technical problem that needs to be solved. Summary of the Invention

[0003] This application provides a video conferencing control method, apparatus, computing device cluster, computer storage medium, and computer product that can ensure the viewing experience of participants in the meeting when the video upload status of the client is abnormal.

[0004] In a first aspect, this application provides a video conferencing control method applied to a cloud-based conferencing system. The cloud-based conferencing system includes a first client and at least one participating client. The first client and the participating clients establish a data connection through the cloud. The method includes: in a multi-person video conferencing scenario, the cloud-based conferencing system detects a video replacement instruction and determines a replacement video, wherein the video replacement instruction is used to instruct the replacement video to be used instead of distributing the video data obtained in real time from the first client to the participating clients; and the replacement video is sent to the participating clients, and the video data obtained in real time from the first client is no longer sent to the participating clients.

[0005] In this implementation, when sending the replacement video, the transmission of video data acquired in real time from the first client is paused simultaneously, and the sending target (participating client) of the previous video data is retained when sending the replacement video. This implementation makes the transition from video data to the replacement video smoother, and the screen in the sub-window related to the first client is more fluid, further enhancing the viewing experience for participants.

[0006] In one possible implementation of the first aspect, the video upload status of the first client is obtained; if the video upload status of the first client is detected to be abnormal, a video replacement instruction is generated.

[0007] For example, video upload status includes: the network status when the participant's video is uploaded from their client and the on / off status of the image acquisition device after the participant activates the video capture function. Network status can include metrics such as network latency, jitter, or packet loss rate between the client and server. These metrics can be compared to the thresholds set for each network parameter. When the value of any network parameter exceeds the preset threshold, the uplink network status is considered poor, and the video upload status is considered abnormal. Alternatively, the video upload status can also be determined by the operating status of the image acquisition device.

[0008] In this way, if the video uploaded by the first client experiences stuttering or a black screen, a replacement video related to the first client can be used promptly. Therefore, when the video upload status is abnormal, the video in this sub-window will not experience stuttering or a black screen, improving the viewing experience for participants.

[0009] In one possible implementation of the first aspect, an abnormal video upload status means that the network status of the uplink data channel of the first client is abnormal or the uplink data is interrupted due to the first client turning off the camera.

[0010] For example, video upload status includes: the network status when the participant's video is uploaded from their client and the on / off status of the image acquisition device after the participant activates the video capture function. Network status can include metrics such as network latency, jitter, or packet loss rate between the client and server. These metrics can be compared to the thresholds set for each network parameter. When the value of any network parameter exceeds the preset threshold, the uplink network status is considered poor, and the video upload status is considered abnormal. Alternatively, the video upload status can also be determined by the operating status of the image acquisition device.

[0011] In this implementation, the network status and image acquisition operation status can be used to determine whether the client's video upload status is abnormal, thus making the determination of the client's video upload status more accurate.

[0012] In this implementation, the network status and image acquisition operation status can be used to determine whether the client's video upload status is abnormal, thus making the determination of the client's video upload status more accurate.

[0013] In one possible implementation of the first aspect, the source of the replacement video includes: videos / images pre-uploaded by the user of the first client.

[0014] In this implementation, the user of the first client can upload a video or an image before the video conference begins. The replacement video can then be either directly used as the replacement video, or a video can be generated based on the image using video generation technology and used as the replacement video. Thus, a replacement video can be obtained from the first client.

[0015] In one possible implementation of the first aspect, the source of the replacement video includes: video data uploaded by the first client during the meeting.

[0016] In this implementation, users of the first client can continuously upload video data during the video conference. The cloud-based conferencing system then extracts and replaces the video from the uploaded data. This eliminates the need for pre-recording, and since users may have slightly different attire before and during the meeting, using the video data from the meeting further improves the smoothness of the visuals in the sub-windows related to the first client, thus enhancing the viewing experience for participants.

[0017] In one possible implementation of the first aspect, the source of the replacement video includes: video / images pre-uploaded by the host of the video conference.

[0018] In this implementation, the host of the second client can upload a video or an image related to the first client before the video conference begins. The replacement video can then be either directly used as the replacement video, or a video can be generated based on the image using video generation technology and used as the replacement video. Thus, a replacement video related to the first client can be obtained.

[0019] In one possible implementation of the first aspect, the method further includes: during the meeting, receiving a replacement video related to the first client uploaded by the first client or the second client, and generating a video replacement instruction, wherein the second client is any one of the participating clients.

[0020] For example, the host or user of the first client can upload a replacement video by dragging it to the first area, and the cloud conferencing system can distribute the uploaded replacement video based on the first area, which is associated with the target client.

[0021] In this implementation, the distribution of a replacement video, specified by the host or the user of the first client, can be controlled by either the host or the user of the first client. This can serve as a backup method when neither the host nor the user of the first client has set a replacement video. Furthermore, this implementation allows the host or the user of the first client to subjectively determine when to distribute the replacement video; for example, when a participant leaves their seat, the host can control the server to distribute the replacement video, thus ensuring the participation experience of other participants.

[0022] In one possible implementation of the first aspect, before generating the video replacement instruction, it further includes: determining that the first client has set up the black screen replacement function.

[0023] For example, when the host drags a replacement video to the area associated with the first client, the replacement video will not be distributed if the video upload status of the first client associated with that area is not abnormal. However, if the video upload status of the first client associated with that area is abnormal, a video replacement instruction will be generated, and the replacement video will be distributed.

[0024] Therefore, in this implementation, determining that the video upload status of the first client is abnormal can effectively prevent accidental operation by the host or the user of the first client, thereby avoiding the distribution of replacement videos due to accidental operation by the host or the user of the first client.

[0025] In one possible implementation of the first aspect, the first client includes a conference window, wherein the conference window is used to indicate sub-windows associated with each client, and the conference window also includes an auto-screen control, which is used to generate action modification instructions or voice generation instructions.

[0026] In this implementation, an automatic screen control is provided on the first client or the second client. The user of the first client or the host of the second client can generate modification commands (action modification commands or voice generation commands) by interacting with the automatic screen control.

[0027] In one possible implementation of the first aspect, the video includes the face of the user of the first client.

[0028] In this implementation, the replacement video includes the facial image of the relevant user, making the replaced image more natural.

[0029] In one possible implementation of the first aspect, the method further includes: receiving an action modification instruction from a first client or the host of a video conference; modifying the actions of the characters in the replacement video based on the action modification instruction; and sending the modified video to the participating clients.

[0030] In this implementation, the action modification command can change the actions of the characters in the replacement video, thereby generating a modified video. The modified video is then sent to the participating clients until the replacement video is distributed. This allows the replacement video to adjust its actions according to the action modification command during loop playback, resulting in smoother visuals in the sub-windows related to the first client and further enhancing the viewing experience for the participants.

[0031] In one possible implementation of the first aspect, the actions of the characters in the replacement video are modified based on the action modification instruction, including: extracting one or more video frames from the replacement video; and generating a new video based on the action modification instruction and one or more video frames, wherein the new video is the modified video.

[0032] In this implementation, one or more video frames are processed through action modification instructions, so that the new video retains the elements in the replacement video and makes the transition between the replacement video and the modified video more natural, thus ensuring the viewing experience of the participants.

[0033] In one possible implementation of the first aspect, sending the modified video to the participating client includes: sending the modified video to the participating client until the screen of the distributed replacement video reaches one or more video frames, and ceasing to send the replacement video to the participating client.

[0034] In this implementation, when the modified video is generated from one or more video frames, the beginning of the modified video is similar to the image of one or more video frames. Therefore, by using one or more video frames as the timing for distributing the modified video during the distribution of the replacement video, the transition between the modified video and the replacement video can be made more natural, further improving the viewing experience for the participants.

[0035] In one possible implementation of the first aspect, the method further includes: receiving a speech generation instruction from a first client or the host of a video conference; generating target speech based on the speech generation instruction; and distributing the target speech to participating clients.

[0036] For example, the target speech is a voice with a timbre that is essentially the same as that of the first client user.

[0037] In this implementation, the target speech is generated through a speech generation command. When the user of the first client is unable to speak, the target speech can be played through the control of the automatic screen controls by the first client or the host.

[0038] In one possible implementation of the first aspect, the method further includes: the cloud conferencing system detecting a video replacement instruction, including: obtaining the on / off state of the video capture function of the first client; detecting that the video capture function of the first client is in the on state, and generating a video replacement instruction.

[0039] For example, the first client has a video capture function. When the first client does not send a digital human request, and its video capture function is enabled, the first client uploads video data in real time, which is then distributed by the cloud conferencing system. After the first client sends a digital human request, and its video capture function is enabled, the cloud conferencing system generates a video replacement instruction, identifies a replacement video associated with the first client, and distributes the replacement video.

[0040] In this implementation, the digital human method allows users on the first client to conduct meetings even when it is inconvenient for them to video chat directly, thus improving the viewing experience for participants.

[0041] Secondly, this application provides a video conferencing control device, the device comprising: an application to a cloud conferencing system, the cloud conferencing system including a first client and at least one participating client, the first client and the participating client establishing a data connection through the cloud, the method comprising: in a multi-person video conferencing scenario, the cloud conferencing system detecting a video replacement instruction, determining a replacement video, wherein the video replacement instruction is used to instruct the replacement video to be used instead of distributing the video data obtained in real time from the first client to the participating client; sending the replacement video to the participating client, and no longer sending the video data obtained in real time from the first client to the participating client.

[0042] In one possible implementation of the second aspect, the processing module is further configured to: obtain the video upload status of the first client during the process of detecting a video replacement instruction in the cloud conferencing system; and generate a video replacement instruction upon detecting that the video upload status of the first client is abnormal.

[0043] In one possible implementation of the second aspect, an abnormal video upload status means that the network status of the uplink data channel of the first client is abnormal or the camera of the first client is turned off.

[0044] In one possible implementation of the second aspect, the source of the replacement video includes: videos / images pre-uploaded by the user of the first client.

[0045] In one possible implementation of the second aspect, the source of the replacement video includes: video data uploaded by the first client during the meeting.

[0046] In one possible implementation of the second aspect, the source of the alternative video includes: video / images pre-uploaded by the host of the video conference.

[0047] In one possible implementation of the second aspect, the processing module is further configured to receive, during the meeting, a replacement video related to the first client uploaded by the first client or the second client, and generate a video replacement instruction, wherein the second client is any one of the participating clients.

[0048] In one possible implementation of the second aspect, the processing module is further configured to determine that the first client has set up a black screen replacement function before generating the video replacement instruction.

[0049] In one possible implementation of the second aspect, the client includes a conference window, which is used to indicate sub-windows associated with each client; the conference window of the first client or the second client also includes an auto-screen control, which is used to generate action modification instructions or voice generation instructions.

[0050] In one possible implementation of the second aspect, the video includes the face of the user of the first client.

[0051] In one possible implementation of the second aspect, the processing module is further configured to receive action modification instructions from the first client or the host of the video conference; based on the action modification instructions, modify the actions of the characters in the replacement video, and send the modified video to the participating clients.

[0052] In one possible implementation of the second aspect, the processing module is further configured to: extract one or more video frames from the replacement video during the process of modifying the actions of a person in the replacement video based on the action modification instruction; and generate a new video based on the action modification instruction and one or more video frames, wherein the new video is the modified video.

[0053] In one possible implementation of the second aspect, the processing module is further configured to, during the process of sending the modified video to the participating client, send the modified video to the participating client until the screen of the distributed replacement video reaches one or more video frames, and then stop sending the replacement video to the participating client.

[0054] In one possible implementation of the second aspect, the processing module is further configured to receive a speech generation instruction from the first client or the host of the video conference; generate target speech based on the speech generation instruction; and distribute the target speech to the participating clients.

[0055] In one possible implementation of the second aspect, the processing module is further configured to: obtain the on / off status of the video capture function of the first client during the process of detecting a video replacement instruction in the cloud conferencing system; and generate a video replacement instruction upon detecting that the video capture function of the first client is enabled.

[0056] Thirdly, this application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the method described in the first aspect or any possible implementation of the first aspect.

[0057] Fourthly, this application provides a computer-readable storage medium including computer program instructions that, when executed by a cluster of computing devices, perform the method described in the first aspect or any possible implementation thereof. Exemplarily, the computing device cluster may include one or more computing devices.

[0058] Fifthly, this application provides a computer program product containing instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method described in the first aspect or any possible implementation thereof. Exemplarily, the cluster of computing devices may include one or more computing devices.

[0059] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0060] Figure 1 This is a schematic diagram illustrating an application scenario of a video conferencing control method provided in an embodiment of this application;

[0061] Figure 2 This is a schematic diagram of the interface of a meeting window under different situations provided in the embodiments of this application;

[0062] Figure 3 This is a schematic diagram illustrating the uploading and distribution of participants' videos according to an embodiment of this application;

[0063] Figure 4 This is a schematic diagram of a video conferencing control method provided in an embodiment of this application;

[0064] Figure 5 This is a comparative schematic diagram of different processing methods provided in the embodiments of this application;

[0065] Figure 6 This is a comparative diagram of meeting windows under different processing methods provided in the embodiments of this application;

[0066] Figure 7 This is a schematic diagram illustrating the setting of a replacement video in a video conferencing control method provided in an embodiment of this application;

[0067] Figure 8 This is a schematic diagram illustrating the setting of a replacement video in a video conferencing control method provided in an embodiment of this application;

[0068] Figure 9 This is a schematic diagram illustrating the extraction of replacement video in a video conferencing control method provided in an embodiment of this application;

[0069] Figure 10 This is a schematic diagram illustrating the extraction of replacement video in a video conferencing control method provided in an embodiment of this application;

[0070] Figure 11 This is a schematic diagram illustrating the setting of a replacement video in a video conferencing control method provided in an embodiment of this application;

[0071] Figure 12 This is a schematic diagram illustrating the setting of a replacement video in a video conferencing control method provided in an embodiment of this application;

[0072] Figure 13 This is a flowchart illustrating a video conferencing control method provided in an embodiment of this application;

[0073] Figure 14 This is a schematic diagram illustrating the setting of a replacement video in a video conferencing control method provided in an embodiment of this application;

[0074] Figure 15 This is a schematic diagram of the process of obtaining modified video in a video conferencing control method provided in an embodiment of this application;

[0075] Figure 16 This is a schematic diagram of the interactive interface for generating and modifying video in a video conferencing control method provided in an embodiment of this application;

[0076] Figure 17 This is a schematic diagram of the interactive interface for generating and modifying video in a video conferencing control method provided in an embodiment of this application;

[0077] Figure 18 This is a schematic diagram of the interactive interface for generating target audio in a video conferencing control method provided in an embodiment of this application;

[0078] Figure 19 This is a schematic diagram of the interactive interface in a method of using a digital human in a video conference provided in an embodiment of this application;

[0079] Figure 20This is a schematic diagram of the structure of a video conferencing control device provided in an embodiment of this application;

[0080] Figure 21 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;

[0081] Figure 22 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;

[0082] Figure 23 This is a schematic diagram of another computing device cluster structure provided in an embodiment of this application. Detailed Implementation

[0083] In this article, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The symbol " / " in this article indicates that the related objects are in an "or" relationship; for example, A / B means A or B.

[0084] The terms "first" and "second," etc., used in the specification and claims herein are used to distinguish different objects, not to describe a specific order of objects. For example, "first response message" and "second response message," etc., are used to distinguish different response messages, not to describe a specific order of response messages.

[0085] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0086] In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more, for example, multiple processing units means two or more processing units, multiple elements means two or more elements, etc.

[0087] First, the relevant technical terms involved in the technical solution provided in this application will be introduced.

[0088] Uplink and downlink channels: The uplink channel refers to the path through which the client transmits data to the cloud-based conferencing system. It is typically used to send data from the client to the cloud-based conferencing system, including audio and video streams, user commands (such as commands to mute or turn on the camera), shared files, or screen content. The downlink channel is the path through which the cloud-based conferencing system transmits data to the client. It is mainly used to receive data pushed from the cloud-based conferencing system to the client, including audio and video streams from other participants, shared screens or files, and meeting control information (such as the host's mute control).

[0089] Next, the technical solution provided in this application will be introduced.

[0090] For example, Figure 1 This illustration shows an application scenario for video conferencing provided by an embodiment of this application. For example... Figure 1 As shown, during a video conference, multiple participants can attend using their own electronic devices. Each participant's electronic device can be configured with a client for video conferencing, and each participant connects to the server through the client on their electronic device. In this embodiment, the client can be a desktop application, mobile application, web application, or web-based application, etc. The electronic device can be a personal computer (such as a desktop or laptop), smartphone, tablet, smartwatch, smart TV, in-vehicle system, etc. The server can be a local server or a cloud server, etc. The electronic devices used by the participants can be configured with audio acquisition devices (such as microphones) and image acquisition devices (such as cameras). The audio acquisition device is used to acquire the participants' audio; the image acquisition device is used to acquire the participants' video. For example, the audio acquisition device can be integrated into the electronic device or configured separately from the electronic device; the image acquisition device can be integrated into the electronic device or configured separately from the electronic device.

[0091] During the meeting, electronic devices acquire data from audio and video capture devices and upload it to the server via the client. The server then distributes the captured video data of the participants to each client, ensuring that each participant can view the video feeds of other participants on their electronic device. Alternatively, each participant can choose to view their own feed while simultaneously viewing the feeds of others.

[0092] The client also includes a meeting window. For example... Figure 2 As shown in (A), the conference window may contain one or more child windows, wherein a child window is associated with a client and is used to display the video feed uploaded by the associated client. In some embodiments, such as Figure 2As shown in (A), when none of the participants have enabled video capture, each sub-window can display only information about the client associated with that sub-window, such as the username or avatar used by the participant when logging into the client. When a participant enables video capture, the client uploads the video data captured by the participant's electronic device, and then the sub-window associated with that client plays the image content contained in the video data (hereinafter referred to as the video footage). Figure 2 As shown in (B), when all participants have enabled the video capture function, all sub-windows in the conference window will play video footage.

[0093] exist Figure 1 In the application scenario shown, the server can include a cloud-based conferencing system. The cloud-based conferencing system is primarily used to support real-time data transmission between multiple clients. For example... Figure 3 As shown, once a participant's client joins the meeting, the client establishes a remote connection with the cloud-based conferencing system. In some embodiments, when the client connects to the cloud-based conferencing system, an uplink and downlink channel are established between the client and the cloud-based conferencing system.

[0094] like Figure 3 As shown, when no participant has enabled video capture, the meeting window will appear as follows: Figure 3 As shown in Client 2-prev, when participant A activates the video capture function, the electronic device captures participant A's video data through an image capture device (such as a camera) and uploads it to the cloud conferencing system in real time via the uplink channel through Client 1. Subsequently, the cloud conferencing system distributes participant A's video data via the downlink channel. For example, the cloud conferencing system distributes participant A's video data to the clients of other participants via the downlink channel, allowing their client's meeting window to display the video data. The video data of participant A will then be played in the sub-window associated with Client 1 within the meeting windows of other participants. Figure 3 As shown in Client 2-now, after being distributed by the cloud conferencing system, participant A's video data is played in sub-window A. It's easy to imagine that when other participants enable video capture and upload their own video data through the client, it will also be distributed by the cloud conferencing system and played in the sub-window associated with that client.

[0095] However, the uplink channel between the client and the cloud conferencing system is not always stable. When a client fails to upload video data successfully, the screen in the child window associated with that client may go black or freeze. Figure 4As shown in (A), different participants occupy different sizes of sub-windows. Participant "Zhang San" occupies a different size sub-window. Figure 4 In (A), the largest sub-window occupies the largest display area. If the uplink channel corresponding to participant "Zhang San" cannot successfully upload video data, the sub-window related to participant "Zhang San" may become black. When this happens, the screens of participant "Wang Wu" and participant "Zhang San" in the conference window can be swapped. After that, as shown... Figure 4 As shown in (B), the largest display window shows the image of participant "Wang Wu". However, this method of comparison negatively impacts the meeting's effectiveness, especially the participants' viewing experience.

[0096] Therefore, this application proposes a video conferencing control method. In this method, a first client and multiple participating clients establish a data connection via the cloud. The method includes: in a multi-person video conferencing scenario, the cloud conferencing system detects a video replacement instruction and determines a replacement video. For distinction, the client associated with the video replacement instruction is referred to as the first client. This video replacement instruction instructs the participating clients to use a replacement video instead of the video data obtained in real-time from the first client. Subsequently, the cloud conferencing system sends the replacement video to the participating clients, no longer sending the video data obtained in real-time from the first client, so that the participating clients can play the replacement video.

[0097] In this embodiment, the video replacement command is used to instruct the cloud-based conferencing system to perform subsequent tasks such as retrieving and distributing a replacement video. The video replacement command can be triggered in various ways, such as when the video upload status is abnormal, or when a digital human request is received; in either case, a video replacement command can be generated.

[0098] The following section explains the video conferencing control method by using the video upload status to trigger a video replacement command.

[0099] During a meeting, the video upload status of each client is monitored. If an abnormal upload status is detected for a particular client, a video replacement command is generated. Next, a replacement video associated with that client is retrieved. This replacement video includes facial images of the relevant participants (also known as the users of that client). This replacement video is then distributed to play in child windows associated with that client. This allows the replacement video to replace the original video in a child window if playback is abnormal. Consequently, when the video upload status is abnormal, the child window will not display a black screen or other abnormalities, increasing the smoothness of the meeting.

[0100] To better understand video conferencing control methods, the following will be combined with Figure 5 and Figure 6 Please provide a detailed explanation.

[0101] like Figure 5 As shown, participant A joins the meeting through client 1, turns on video capture at time t0, and turns it off at time t3. It's easy to imagine that, under normal circumstances, child window A associated with client 1 will play the video uploaded by client 1 between time t0 and t3. Figure 5 As shown, assume that at time t1, client 1's video upload status is abnormal (e.g., the upload channel is stuck, and the video cannot be uploaded normally). After repair, at time t2, client 1's video upload status returns to normal. Then, without using a replacement video, during the time interval t0-t1, if... Figure 5 and Figure 6 As shown in (A), the sub-window associated with client 1 plays the video uploaded by client 1 normally. During the time interval t1-t2, because the cloud conferencing system fails to receive the video data uploaded by client 1, the video played in sub-window A associated with client 1 will experience black screens or severe stuttering. When using a replacement video, such as... Figure 5 and Figure 6 As shown in (B), during the time interval t1-t2, the cloud-based conferencing system can distribute the replacement video to client 1, causing the replacement video to play in the child window associated with client 1. This continues until the client's video upload status returns to normal (i.e., at time t2), at which point the system switches to distributing the video data uploaded by client 1, until participant A closes the video capture function. (Reference) Figure 6 It is known that the replaced video includes facial images of the participants. It should be noted that... Figure 5 In the middle, this refers to the replacement video technology used during the period when the video capture function is enabled; however,

[0102] like Figure 5 As shown, the replacement video can be a fixed-length video, such as a video of duration T. During the process where the video upload status associated with client 1 is in an abnormal state, the replacement video allows the child window associated with client 1 to loop the replacement video. This continues until the video upload status returns to normal, at which point the video data uploaded by client 1 is redistributed. For example, in... Figure 5 In the process, the replacement video plays for a duration of T / 2 during its third replay. When the video upload status returns to normal, the cloud conferencing system stops distributing the replacement video and distributes the video data uploaded from client 1, so that the sub-window associated with client 1 plays the video data uploaded by client 1.

[0103] In this embodiment, the replacement video is associated with the client. For example, one client can correspond to one or more replacement videos. When the upload status of the video corresponding to the client is abnormal, a video replacement instruction can be generated. Then, a replacement video is selected from the one or more replacement videos associated with that client and distributed. The association between the replacement video and the client ensures that the participant in the replacement video is the same participant as the participant in the video played by the client in a normal state. This makes the visuals in the sub-windows related to that client in other participants' meeting windows more natural when the video upload status of the relevant client is abnormal and a replacement video is used. The replacement video can be obtained through various methods. In some possible implementations, the replacement video can be a pre-recorded video uploaded by the participant through the client, or a video extracted during the meeting through the cloud conferencing system. Alternatively, the replacement video can be a video processed based on the aforementioned video; this is not limited here. The replacement video can be stored in the cloud conferencing system. Of course, the cloud conferencing system destroys all replacement videos stored in the meeting at the end of each meeting.

[0104] By setting up video replacement, a replacement video can be distributed to make the meeting run more smoothly when the video upload status on the client is abnormal.

[0105] The following section introduces the technical solutions for video conferencing control methods, step by step, before, during, and after the video conference.

[0106] 1) Before the video conference begins

[0107] Before the video conference begins, a moderator can be selected from all participants to lead the conference. During the conference, the current moderator can pass the moderation role to another participant. During the conference, the moderator can issue commands to all participants' clients or only to a subset of participants' clients.

[0108] Additionally, before the video conference begins, both participants and the moderator can configure whether to use a replacement video during the meeting. First, taking a participant as an example, in... Figure 7 In (A), the client can provide an interactive video replacement settings interface for participants. Participants can use this interface to select when to use a replacement video; that is, they can select the content included in the video upload status. For example... Figure 7 As shown in (A), participants can choose to attend meetings when the uplink network condition is poor. Figure 7In cases of poor signal or unexpected disconnection, a replacement video is used. Subsequently, when the cloud conferencing system detects poor network conditions on the uplink channel, it considers the video upload status abnormal and distributes the replacement video.

[0109] In some embodiments, participants can also configure how the replacement video is obtained. The replacement video can be obtained through various methods; below, from the participant's perspective, three technical concepts for obtaining the replacement video are introduced.

[0110] 1. Participants upload replacement videos via the client.

[0111] Participants can choose how the replacement video is generated by configuring the settings interface. For example, such as... Figure 7 As shown in (A), participants can upload videos located locally on their electronic devices or in the cloud to the cloud conferencing system via the client by setting preset options in the video settings interface (including the dark box with the plus sign).

[0112] For example, before the meeting begins, participants can open the video replacement settings interface on the client. Figure 7 As shown in (A), participants can operate the replacement video settings interface to select local upload as the method for obtaining the replacement video, and then select a video from the videos stored on their electronic device as the replacement video. Figure 7 In (B), participants can upload multiple videos in the video replacement settings interface and select one as the replacement video. After the participant confirms, the client can upload the selected replacement video to the cloud conferencing system.

[0113] Additionally, participants can choose to upload their replacement video as a live recording via the video replacement settings interface. For example, when a participant selects live recording as the upload method, the live recording function is enabled. In response to the participant using the live recording function, the image acquisition device will capture environmental images from certain directions. The client can acquire the video captured by the image acquisition device and upload it as the replacement video to the cloud-based conferencing system. After the live recording function is enabled, the client can provide a window for participants to select when to start recording and allow them to view the recorded information in real time.

[0114] 2. Extracting and replacing videos from cloud-based conferencing systems (methods for extraction during meetings)

[0115] like Figure 8As shown, participants can also choose to use the "extract from the meeting" method to obtain a replacement video through the replacement video settings interface. "Extract from the meeting" instructs the cloud conferencing system to use at least a portion of the video data uploaded by the participant's client during the meeting as the replacement video. The specific methods by which the cloud conferencing system extracts the video data uploaded by the client will be detailed in the subsequent description of the video conferencing process.

[0116] 3. Generate replacement video by replacing images

[0117] Video generation technology can be used to generate replacement videos based on replacement images. For example, the replacement image can be uploaded by participants using a method similar to uploading a replacement video via a client as described in 1), or it can be obtained by the cloud-based conferencing system using a method similar to extraction as described in 2). After being uploaded or extracted, the replacement image reaches the cloud-based conferencing system; this image contains the facial images of the relevant participants. Then, based on this replacement image, a corresponding replacement video is generated.

[0118] As explained earlier, the facilitator can also configure whether or not to use alternative videos during the meeting. The following explanation will focus on the facilitator's perspective.

[0119] When a moderator is present in a video conference, their client application can be configured with a master control interface for alternative videos. The moderator can use this interface to configure whether or not to use an alternative video during the conference. For example, such as... Figure 11 As shown, the video replacement control interface has a first control 1001. This first control 1001 is used to configure the black screen replacement function. The black screen replacement function instructs the cloud conferencing system whether to use a replacement video and under what circumstances it can be used. When the host activates the function of the first control 1001, the cloud control system receives this activation information and, if the video upload status with the specified client is abnormal, distributes the replacement video associated with that client. The specified client is one or more clients in the conference. The host can select which clients participating in the conference are designated as specified clients through the video replacement control interface.

[0120] In addition, the host can configure the method for obtaining replacement videos for each client. For example, the host can instruct the cloud conferencing system to generate replacement videos for specified clients using the extraction method described above, or upload replacement videos to specified clients through the master replacement video control interface (i.e., obtain replacement videos using local upload). The extraction method has been explained in detail in previous content and will not be repeated here. The following is a detailed introduction to how the host can upload replacement videos to specified clients through the master replacement video control interface.

[0121] For example, such as Figure 11 As shown, the video replacement control interface also includes a second control 1002. The second control 1002 is used to upload the replacement video and associate the uploaded replacement video with a specified client. For example, as... Figure 11 As shown, the display icon of the second control 1002 is a small window. There are four small windows (second control 1002) on the video replacement control interface, each associated with a participant's client. Through interaction between the host and the second control 1002, the replacement video can be associated with a specified client. The operation process is as follows... Figure 12 As shown, in Figure 12 In this context, one or more preset videos can be stored on the electronic device corresponding to the host's control terminal. A preset video is a pre-recorded video of a participant. Figure 12 In the process, the host can interact with the second control 1002 to select a replacement video. For example, they can click on the second control 1002 to bring up a file explorer. Then, they can select a preset video in the file explorer as the replacement video for the second control 1002; alternatively, they can drag the preset video to the display icon of the second control 1002 to use it as the replacement video. Since the second control 1002 is associated with the client, the replacement video associated with the second control 1002 will also be associated with the client. After the host selects a replacement video through the second control 1002, the replacement video, along with the identification information of the target client associated with the replacement video, will be uploaded to the cloud conferencing system.

[0122] As previously explained, the host can set the method for obtaining replacement videos for each client, and participants can also set their own replacement video acquisition method through their clients. In one possible scenario, both participants and the host have set their replacement video acquisition methods. In this case, the cloud conferencing system can prioritize using the replacement video acquisition method set by the participant. In another possible scenario, a participant uploads replacement video 1 through client 1, and the host uploads replacement video 2 related to client 1. In this case, if the upload status of the video related to client 1 is abnormal, the cloud conferencing system can distribute only the replacement video 1 uploaded by the participant, so that replacement video 1 can be played in the sub-window related to client 1.

[0123] 2) During the video conference

[0124] During a video conference, the cloud-based conferencing system is primarily responsible for: 1. Obtaining a replacement video based on the settings before the video conference begins; 2. Monitoring the video upload status and distributing the replacement video. These two aspects will then be described in detail.

[0125] 1. Obtain the replacement video according to the settings before the video conference starts.

[0126] There are three settings before the video conference starts: uploading a replacement video via the client, extracting the video during the conference, and generating a replacement video from a replacement image. The following explains how the cloud conferencing system obtains the replacement video when each of these three methods is selected.

[0127] A. When choosing to upload a replacement video via the client:

[0128] The client can upload alternative videos selected by participants or the host to the cloud-based conferencing system. The cloud-based conferencing system receives the uploaded alternative videos at the start of the meeting and associates them with the client.

[0129] B. When choosing to use the extraction method to generate the replacement video:

[0130] During a meeting, a cloud-based conferencing system can generate a replacement video associated with the client based on the video data uploaded by the client. Of course, when generating a replacement video using the meeting extraction method, participants must have had video capture enabled for at least a certain period and uploaded video data through their clients. For example, such as... Figure 9 As shown, the cloud-based conferencing system can extract a portion of the video data already uploaded by the client as a replacement video. When the cloud-based conferencing system retrieves the replacement video, it can check if the duration of the uploaded video data is greater than T1. If the duration is greater than T1, the cloud-based conferencing system can extract a segment of video with a duration of T1 from the uploaded video data and use this segment as the replacement video. T1 is a pre-set duration, such as 20 seconds. Additionally, the cloud-based conferencing system can also periodically update and replace the video during the meeting. Figure 9 In this cloud-based conferencing system, the replacement video is updated every T2 time interval. During each update, the system extracts video data from clients that have already uploaded the video at the time closest to the current update to obtain the replacement video.

[0131] Of course, cloud-based conferencing systems can also first determine which parts of the video data uploaded by the client meet the requirements for video replacement. For example... Figure 10As shown, the cloud-based conferencing system analyzes the video data uploaded by the client, identifies segments with high network latency, removes these segments, and then extracts a replacement video from the removed segments. Alternatively, it can perform facial recognition analysis on the uploaded video data, remove segments where no faces were detected, and extract a replacement video from the removed segments. For example, visual tracking algorithms and MTCNN (Multi-task Cascaded Convolutional Networks) algorithms can be used to determine which parts of the video frame contain faces and which do not. Those skilled in the art will understand from the above that other requirements can be applied to sequentially remove segments from the uploaded video data and extract a replacement video from the removed segments; these will not be elaborated upon here.

[0132] C. When choosing to generate a replacement video using a replacement image:

[0133] The cloud-based conferencing system will determine the replacement image and use machine learning-based video generation technology to generate a replacement video from it. This machine learning-based video generation technology can be other deep learning models such as First Order Motion Model (FOMM), DeepFaceLab, or GANimation.

[0134] As one possible implementation, the cloud-based conferencing system first extracts the subject from the replacement image using image matting technology. Then, the system uses facial recognition technology to obtain the skeletal and / or muscle node information of the subject. Next, using pre-set motion data, the system drives the obtained skeletal and / or muscle nodes to move dynamically, thereby generating an animated video of the participant. For example, the pre-set motion data could be the motion trajectories of skeletal and / or muscle nodes for actions such as clapping or turning the head. Based on the motion trajectories of the corresponding actions recorded in the pre-set motion data, the cloud-based conferencing system can further fit the replacement image to generate dynamically changing portrait images and then connect these dynamic images into a single replacement video. For example, the cloud-based conferencing system can generate multiple animated videos based on pre-set motion data for different actions, and then combine these animated videos to generate the final replacement video. For instance, it can acquire pre-set motion data related to blinking, smiling, and nodding, and generate animated videos of blinking, smiling, and nodding sequentially based on the pre-set motion data and the replacement image. These three animated videos are then sequentially stitched together to obtain the replacement video. Of course, the length of the generated animated videos can vary; for example, a blinking animated video could be 1 second long, while a smiling animated video could be 4 seconds long. By stitching together these animations, a replacement video can be obtained.

[0135] 2. Monitor video upload status and distribute replacement videos.

[0136] If the participants and the moderator have configured the previously mentioned black screen replacement function (which allows for the use of a replacement video during the meeting) for a specific client, the cloud-based conferencing system will monitor the video upload status of the specified client during the meeting and distribute the replacement video based on that upload status.

[0137] In some embodiments, the cloud-based conferencing system can identify the video upload status of each client and determine whether the video upload status is abnormal. The video upload status indicates the upload status of the client when uploading video. This status can be a direct indicator, such as network status, or an indirect indicator, such as when it is known that the image acquisition device is off, it can be inferred that no video is being uploaded, which indirectly indicates the video upload status.

[0138] For example, the video upload status can include the network status of the uplink channel when the client uploads video data. For example, the cloud conferencing system establishes a connection with the client via a network protocol and uses the communication interface defined in the network protocol to monitor the network status of the connected uplink channel in real time to assess the video upload status. For example, the network status can include metrics such as network latency, jitter, real-time data transmission volume, or packet loss rate between the client and the cloud conferencing system, thus reflecting the video upload status. For example, this communication interface can be an interface implemented based on the HTTP / 2 protocol, HTTPS API protocol, WebSocket protocol, or other network protocols. On the other hand, the video upload status can also include the on / off status of the image acquisition device. The client-side cloud conferencing system can monitor the operating status of the image acquisition device corresponding to the client when the client uploads video data. As one possible implementation, the client can report the operating status of its local image acquisition device by sending a specific signal to the cloud conferencing system. When the on / off status of the image acquisition device changes, the client can send the updated operating status of the image acquisition device to the cloud conferencing system via a separate signal. Thus, the cloud conferencing system can directly read the operating status of the image acquisition device recently reported by the client to determine the current working status of the image acquisition device.

[0139] Next, when determining whether the video upload status is abnormal, the network status of the uplink channel can be used. For example, metrics such as network latency, jitter, real-time data transmission volume, or packet loss rate can be compared to the thresholds set for each network parameter. When the value of any network parameter exceeds the preset threshold, the network status of the uplink channel is considered poor, and the video upload status is considered abnormal. For instance, the network latency can be compared to a preset latency threshold; if the network latency is greater than the preset threshold, the video upload status is abnormal, and vice versa. Alternatively, one or more quantities in the network status can be analyzed for judgment. For example, real-time data transmission volume can be analyzed. If the real-time data transmission volume of the uplink channel drops sharply to near zero, it may indicate a sudden malfunction and shutdown of the image acquisition device, and the video upload status can be considered abnormal. In another possible implementation, the on / off status of the image acquisition device obtained from the video acquisition section can be used to determine whether the video upload status is abnormal. For example, if the client's video capture function is enabled and the image capture device is switched off, then the video upload status is considered abnormal.

[0140] When the cloud-based conferencing system determines that a client's video upload status is abnormal, it determines whether to generate a video replacement instruction and distribute a replacement video based on whether the client or the host had enabled the black screen replacement function before the meeting started. If the client associated with the video upload status is the specified client, and a replacement video associated with that client can be found, a video replacement instruction is generated, and the replacement video is distributed.

[0141] For example, during the distribution of a replacement video, the cloud-based conferencing system can send the replacement video to other clients. Upon receiving the replacement video, each client determines the playback sub-window for it. In one possible implementation, the cloud-based conferencing system sends the target client's identification information along with the replacement video. That is, when a client receives the replacement video, it also receives the target client's identification information. The client then determines the sub-window within the conferencing window that corresponds to the target client's identification information based on that target client's identification information, and plays the video in that sub-window.

[0142] To better understand the video distribution process of a cloud-based conferencing system during a video conference, this section combines... Figure 13 Let's illustrate with examples.

[0143] For example, such as Figure 13As shown, the video distribution process of a cloud-based conferencing system mainly includes: video acquisition, anomaly detection, and video distribution. The video acquisition part is primarily used to acquire video data uploaded by each client during the meeting. Figure 13 Taking participant A as an example, participant A joins the meeting through client 1. After participant A activates the video capture function on client 1, client 1 uploads participant A's video data to the cloud conferencing system via the uplink channel. Subsequently, the cloud conferencing system receives the video data uploaded by participant A from client 1 via the uplink channel. Furthermore, while client 1 is uploading video data, the cloud conferencing system can monitor the video upload status related to client 1 and determine whether the video upload status is abnormal within the anomaly detection section.

[0144] The anomaly detection section is primarily used during the meeting to determine whether the video upload status, obtained from the video acquisition section and associated with each client, is abnormal. As one possible implementation, the cloud-based conferencing system is equipped with an anomaly detection module. This module can determine whether the video upload status is abnormal or normal using various specific methods. For example... Figure 13 As shown, the cloud-based conferencing system can obtain the on / off status of image acquisition devices (such as cameras) associated with client 1, and simultaneously detect whether the network status associated with client 1 is abnormal. When the video acquisition function of client 1 is enabled but the camera is not enabled, or when the network status is abnormal, the video upload status associated with client 1 is considered to be abnormal.

[0145] The video distribution section is primarily used to perform further operations based on the anomaly detection results of the video upload status obtained from the anomaly detection section. As one possible implementation, the cloud-based conferencing system is configured with a replacement video determination module. This module is used to obtain the replacement video associated with a client when the client's video upload status is abnormal, by leveraging the association between the replacement video and the client. Taking participant A as an example, when the video upload status associated with participant A is normal, participant A's video data is sent to other clients in the conference, causing the video data of participant A to play in the sub-windows associated with client 1 on those other clients. When the video upload status associated with participant A is abnormal, the replacement video determination module finds the replacement video A associated with client 1 and distributes replacement video A to other clients, causing replacement video A to play in the sub-windows associated with client 1 on those other clients.

[0146] During the meeting, the host can also specify a replacement video in the sub-window related to the participants in case of a black screen or other abnormal situation.

[0147] When the video upload status of the client is abnormal, if the corresponding participant has not set a replacement video recording method and the host has not set a replacement video, the cloud conferencing system will not have a replacement video related to the client in the abnormal state. In this case, the sub-window related to the client may be stuck or in a black screen state.

[0148] In the above scenario, the host or participant can drag a preset video into the first area of ​​the host's client or participant's client and release the preset video in the first area, so that the cloud conferencing system can distribute the preset video to other clients and play the preset video in the sub-windows related to the first area in the other clients.

[0149] Let's start by explaining from the perspective of the host. For example... Figure 14 As shown in (A), the host can see from the client's conference window that the child window related to client A (also known as the target client) is in a black screen state. Then, as... Figure 14 As shown in (B), the host's electronic device can store multiple preset videos. Responding to the host's selection operation, the electronic device selects a preset video and uses it as a replacement video, switching the replacement video from a deselected state to a selected state. The host then drags the selected replacement video to the first area of ​​the host's client and releases the replacement video in the first area. Releasing the replacement video means switching it from a selected state to a deselected state. The first area is associated with client A. Figure 14 As shown in (B), the first area can overlap with the sub-window associated with client A. In response to dragging the selected replacement video and releasing it within the preset first area in the host client, the host client uploads the replacement video and the identification information of client A associated with the first area to the cloud conferencing system. Afterwards, the cloud conferencing system obtains the replacement video and the identification information of client A, generates a video replacement instruction, and then distributes the replacement video so that it plays in the sub-window associated with client A. Finally, as... Figure 14 As shown in (C), a replacement video is played in a sub-window associated with client A. During this process, the target client's identification information serves as a unique identifier. Once the cloud conferencing system receives the target client's identification information, it can determine the target client based on this information. For example, when the first region is associated with client A, the identification information of the target client associated with the first region could be "client A".

[0150] When the host drags the selected replacement video to the first area of ​​the client and releases it, the host client determines the target client associated with the first area based on the release position.

[0151] For example, taking a Windows system as an example, the meeting window of the host client includes multiple child windows. Each child window can be set to have the attribute of receiving drag-and-drop files (in Windows, this could be the WS_EX_ACCEPTFILES attribute). Having the attribute of receiving drag-and-drop files indicates that the child window supports drag-and-drop file operations. For example, a child window can have the attribute of receiving drag-and-drop files only when the relevant video upload status is abnormal, or it can have this attribute normally. For ease of description, a child window with an abnormal video upload status will be referred to as child window A. The first area refers to the area associated with child window A; for example, if the size of the first area is approximately the same as the size of child window A, then the first area has the attribute of receiving drag-and-drop files.

[0152] When the video upload status of child window A is abnormal, child window A may be black or in another abnormal state. At this time, the host can switch a preset video from unselected to selected on their electronic device (hereafter referred to as replacement video A for ease of description). The host can then drag replacement video A to the first area and release it there. For example, when replacement video A is dragged, the Windows system can create an event message recording information such as the file path of replacement video A. Then, when replacement video A is dragged to the first area, the client can detect that a file has been dragged into the first area. When replacement video A is dragged to the first area, the client can view the event message to learn about the file's information, such as the file path and type. The client can decide whether to allow dragging and dropping based on the file type it sees. Of course, the client can also provide visual cues for allowing and disallowing dragging and dropping.

[0153] Subsequently, in response to the host's release operation of the replacement video A, the Windows system can generate a file message (which may be a WM_DROPFILES message in Windows) and send this file message to the host client. This file message typically contains the filename and storage path information of the replacement video A. Upon receiving this file message, the client can upload the replacement video A and the identification information of the target clients associated with the first region to the cloud conferencing system via the host client, based on the file message. After receiving the replacement video A and the identification information of the target clients associated with the first region, the cloud conferencing system distributes the replacement video A to the clients associated with the first region, enabling the replacement video A to play in child window A.

[0154] For example, on an electronic device with a mouse, you can first move the cursor to the icon of the replacement video A, then switch the replacement video A from a deselected state to a selected state by long-pressing the left mouse button, and then keep the left mouse button active (pressed) and move the mouse to drag the replacement video A; on another electronic device with a touch screen, you can also switch the replacement video A to a selected state by long-pressing the touch icon of the replacement video A, and move the touch point to make the replacement video A move with the touch point.

[0155] Furthermore, let's illustrate this from the participant's perspective. During the meeting, a participant can specify a replacement video if their associated sub-window is black or under other special circumstances. In this scenario, the participant's client's uplink network is normal, and the client can upload video through the uplink. At this time, the participant's client's video capture function may be enabled, but the corresponding image capture device may be disabled. In this situation, the participant drags the replacement video into the first area of ​​their client and releases it within that area, allowing the cloud conferencing system to distribute the replacement video to other clients and play it in the sub-windows associated with the first area on those other clients. The first area is associated with the participant's associated sub-window. The technical details in this case are largely the same as those provided above regarding the host specifying a replacement video when the participant's associated sub-window is black or under other abnormal circumstances, and will not be repeated here. In one possible implementation, the sub-window on the participant's client with the attribute of receiving drag-and-drop files can be different from the sub-window on the host's client with the same attribute. The host client can have one or more child windows that have the ability to receive drag-and-drop files, while the participant client can only have this ability in the child window associated with the participant. In other words, the participant client only supports participants uploading their selected replacement video as their own replacement video to the cloud conferencing system.

[0156] In one possible scenario, the replacement video might be detected as abnormal because it is playing in a loop.

[0157] To make the replacement video more realistic, further modifications can be made while it's playing in the child window to make it appear more natural. For example, while the replacement video is playing in the child window, the user can send modification instructions to the cloud-based conferencing system, which will then modify the replacement video accordingly. The user can be either the host or a participant related to the replacement video; there are no restrictions here.

[0158] The modification commands can be either commands that modify actions or commands that modify sounds. These will be explained in detail later.

[0159] 1. The modification command pertains to actions (the modification command is the action modification command).

[0160] like Figure 15 As shown in part (A), a method for generating and distributing modified videos is given, specifically including S1301 to S1304, as follows.

[0161] S1301: Receive action modification command.

[0162] The action modification instruction must include at least the action information indicated by the user. This action information instructs the cloud-based conferencing system to generate the modified video specified by the action information. The action information can be explicit, such as text like "wave" or other nouns related to the action, or it can be an identifier associated with the action, such as "ASDC," where the action associated with "ASDC" is "wave." Alternatively, the action information can be implicit, such as the cursor's trajectory when the user selects an action.

[0163] S1302: Extract one or more video frames from the replacement video according to the extraction rules.

[0164] A video frame can be understood as replacing a single image within a video. As one possible implementation, multiple video frames can be consecutive, which can also be understood as replacing a segment of the video. The cropping rule indicates the location where a video frame is cropped. The cropping rule is a pre-defined rule; it can be to crop the current video frame that is playing at the moment the action modification command is received, or a segment of video starting from the current video frame. Alternatively, it can be to crop the current video frame that is expected to play some time after the action modification command is received, or a segment of video starting from the current video frame. Cropping the expected video frame some time after the action modification command is received allows sufficient time for the subsequent generation of the modified video.

[0165] S1303: Generate modified video based on one or more video frames and motion modification instructions.

[0166] In this step, the duration of the modified video (also referred to as the new video or the modified video) can be determined based on the action information in the action modification instruction. For example, if the action information in the action modification instruction is "clapping," the default duration of the generated modified video might be 5 seconds; if the action information in the action modification instruction is "nodding," the default duration of the generated modified video might be 2 seconds. Additionally, the action modification instruction can specify the duration of the action, in which case the duration of the generated modified video will be the duration of that action. Those skilled in the art can generate modified videos of different durations for different actions based on the actual application, and this is not limited here.

[0167] The process of generating a modified video from video frames is similar to the process of generating a replacement video from a replacement image, as described earlier, and will not be repeated here.

[0168] S1304: Distribute modified videos.

[0169] For example, a cloud-based conferencing system can directly distribute the modified video after detecting that the modified video has been generated. Alternatively, the cloud-based conferencing system can distribute the modified video when the distributed replacement video reaches the captured video frame (one or more video frames). The distributed replacement video reaching the captured video frame means that the distributed replacement video meets a pre-set frame matching condition. For example, the pre-set frame matching condition could be whether the distributed replacement video belongs to a specified frame surrounding the first frame of the captured video frame. For instance, the frame matching condition could be whether the distributed replacement video is the frame preceding the first frame of the captured video frame. Therefore, when the distributed replacement video is the frame preceding the first frame of the captured video frame, it can be considered that the distributed replacement video has reached the captured video frame. The frame matching condition can usually be set based on the first frame of the captured video frame; however, those skilled in the art can easily conceive of other frame matching conditions based on the above techniques to determine whether the distributed replacement video has reached the captured video frame. Understandably, modifying a video will replace the currently playing replacement video, which will then play in the child window associated with that replacement video. Once the modified video distribution is complete, the replacement video can be distributed.

[0170] Then combine Figure 15 (B) details the methods for generating and distributing modified videos.

[0171] like Figure 15 As shown in (B), without any modification instructions, the replacement video is looping in a sub-window. At time t1, the cloud conferencing system receives the action modification instruction. Afterwards, a video frame can be selected to generate the modified video. Figure 15In (B), select the video frame at time t2 to replace. Then, use this video frame along with the motion modification instructions to generate the modified video. Alternatively, multiple video frames can be selected to generate the modified video. Figure 15 In (B), you can also choose to replace a segment of video after time t2 (a segment of video consists of multiple video frames). In this segment, the first frame is the video frame of the replacement video at time t2. Then, the modified video is generated based on the action modification instructions and this segment of video. Figure 15 In (B), time t2 is a time that is delayed by a period of T3 from time t1. That is, t2 = t1 + T3. In other words, at time t1, the cloud conferencing system receives the action modification instruction and begins generating the modified video. To allow sufficient time for generating the modified video, one or more video frames from time t2 are selected, and the modified video is generated based on the action modification instruction and the selected video frames. For example, the modified video may have already been generated before time t2, but the cloud conferencing system still distributes the replacement video until time t2, at which point it switches to distributing the modified video.

[0172] In one possible implementation, the replacement video can be modified using an automatic screen control 1401. The automatic screen control 1401 receives user input, generates action modification instructions, and sends these instructions to the cloud-based conferencing system.

[0173] For example, the client may also have an automatic screen control 1401. (e.g.) Figure 16 As shown in (A), the Auto Screen control 1401 can be displayed in the conference window. Assuming an action modification command is being sent to a participant, the participant can send the command through the Auto Screen control 1401 in the conference window. Figure 16 As shown in (B), participants can open a secondary menu of the automatic screen control 1401 to select generating a replacement video about the action. In some embodiments, such as Figure 16 As shown in (C), the automatic screen control 1401 can also have a three-level menu, which can contain pre-set actions. When a participant selects an action, the client creates an action modification instruction based on the selected action and sends the instruction to the cloud conferencing system. The cloud conferencing system then captures video frames and, based on the video frames and the action modification instruction, generates and distributes the modified video. For example, in... Figure 16 In step (C), the participant selects a waving gesture, and the cloud-based conferencing system then generates a modified video containing the waving gesture based on the video frames. The modified video is then distributed to the appropriate windows, such as... Figure 16 As shown in (D), the sub-window associated with the participant plays the generated modified video containing the waving gesture.

[0174] As explained earlier, cloud-based conferencing systems can contain preset motion data related to various actions. It's easy to understand that a preset motion data point is associated with a specific action. When the modification command is an action modification command, the cloud-based conferencing system can retrieve the preset motion data related to the action in that command. Then, it generates the modified video based on the preset motion data and the captured video frames.

[0175] For example, such as Figure 17 As shown, the automatic screen control 1401 can also obtain action modification instructions through the user's interaction trajectory in a specified area. The interaction trajectory indicates the trajectory record formed by the user's interaction within the specified area. Furthermore, preset trajectories can be stored in the cloud conferencing system, and the preset trajectories and actions can be associated. In use, as... Figure 17 As shown, the automatic screen control 1401 can provide an interactive canvas (i.e., a motion brush window interface). The user then creates a trajectory on the canvas using cursor or touch, thus obtaining the interaction trajectory. After obtaining the user's interaction trajectory, the client matches it with a preset trajectory. Based on the correlation between the preset trajectory and the action, the intent corresponding to the user's interaction trajectory is determined. Figure 17 As shown, after recognizing the user's interaction trajectory, a preset trajectory "Z" is obtained. The action related to "Z" is then identified as "waving," and a modification instruction is generated based on this action, which is then sent to the cloud conferencing system. Alternatively, the modified video can be generated based on the user's interaction trajectory and its location. In this case, the user can interact in a sub-window playing the replacement video, generating an interaction trajectory. The cloud conferencing system can then modify the replacement video based on this trajectory. For example, if the interaction trajectory occurs at the head position of a person in the replacement screen, and the trajectory is similar to "N," a corresponding modified video will be generated based on the interaction trajectory and its location. Besides the methods listed above, the automatic screen control 1401 can also use other methods to obtain the user's selected action. For example, the user's gesture can be associated with an action, and the selected action can be determined by recognizing the gesture. Alternatively, a shortcut key can be associated with an action, and the selected action can be determined through the shortcut key's command. Those skilled in the art can conceive of other ways to obtain the action modification instruction based on the above description, which are not limited here.

[0176] 2. The modification command is related to sound (the modification command is the speech generation command).

[0177] When the modification instruction concerns sound, the user can use voice-generated commands to instruct the cloud-based conferencing system to distribute the target audio. In some embodiments, the target audio can be obtained using machine learning-based audio generation techniques. For example, the audio generation technique could be a neural network model capable of customizing personalized voices, such as StarGAN-VC or VQ-VAE-2. Customizing the personalized voice can be achieved by training the audio generation technique with a small number of samples from a specific person, allowing it to mimic that person's voice output.

[0178] like Figure 18 As shown in (A), the automatic screen control 1401 can also generate voice generation instructions regarding sound. In some embodiments, audio can be generated by selecting a text-driven approach. Figure 18 As shown in (B), the automatic screen control 1401 can have a text conversion window through which the user can input text messages. The input text message is then used by the automatic screen control 1401 to generate a speech command, which is then uploaded to the cloud conferencing system via the client to generate target audio. This target audio is then distributed to the participants' clients by the cloud conferencing system. For example, in the cloud conferencing system, the target audio is combined with a replacement video, and the combined replacement video with the target audio is distributed to the participants' clients.

[0179] In other embodiments, the auto-screen control 1401 can also generate target audio by selecting a sound driver. Unlike text-driven audio, when the user selects a sound driver, the auto-screen control 1401 can also have a console window. Figure 18 As shown in (C), when sound-driven mode is selected, the automatic screen control 1401 can call the audio acquisition device of the current device to capture the user's voice and send the user's voice as a speech generation command to the cloud conferencing system to generate the target audio through audio generation technology. Of course, the user's voice can be sent to the cloud conferencing system in real time, or it can be sent to the cloud conferencing system after the user's entire voice is captured by manual confirmation.

[0180] In some embodiments, such as Figure 18 As shown in C, when the host selects the audio-driven method of generating target audio to replace the video, the host's client's audio capture function is disabled. Therefore, when the host is generating target audio for the replacement video, it prevents the host from simultaneously uploading their own audio data through the client. Otherwise, the cloud conferencing system might simultaneously distribute both the target audio and the host's audio data.

[0181] 3) After the meeting

[0182] When a meeting ends, all replacement videos associated with that meeting are destroyed. For example, when a meeting is created, a meeting ID may be included, and this meeting ID is destroyed when the meeting ends. In this example, when the meeting ID is destroyed, all replacement videos associated with that meeting ID are also destroyed. This way, each client needs to re-upload a replacement video each time a new meeting starts, instead of using the previous replacement video. For example, before meeting ID 123, participant A uploads replacement video A through client 1. The cloud conferencing system receives replacement video A and uses it as the replacement video for client 1. When an abnormal video upload status is detected for client 1, replacement video A is distributed to the corresponding sub-window. When meeting ID 123 ends, the cloud conferencing system destroys all replacement videos associated with that meeting ID (including at least replacement video A). Later, when participant A attends another meeting (ID 124) through client 1, participant A's appearance may have changed; for example, clothing, hairstyle, and jewelry may have changed compared to when attending meeting ID 123. Since the replacement video A has been destroyed, it will not be distributed in the meeting with ID 124, thus avoiding the phenomenon of the replacement video being distributed multiple times in meetings at different times to expose that the participants are using the replacement video.

[0183] On the other hand, the video conferencing control method is explained by using a digital human request to trigger a video replacement command.

[0184] In this method, before the video conference begins, either the participant on the first client or the host of the video conference can upload a replacement video / image in advance and turn on the digital human switch associated with that participant. When the digital human switch is on, it indicates that the participant will use the digital human function during the meeting. Figure 19 As shown, there is a digital human switch in the video replacement settings interface. Figure 19 The example states, "Use the replacement video directly after the camera is turned on." For instance, during a meeting, the cloud-based conferencing system monitors the participants' video capture capabilities. When a participant's video capture function is enabled, a video replacement instruction is generated. The cloud-based conferencing system can then distribute the replacement video to the clients within the meeting. Additionally, the cloud-based conferencing system can also distribute the replacement video to the clients associated with it; in other words, participants who uploaded the replacement video can also receive a distribution of their own video from the cloud-based conferencing system.

[0185] For example, when a participant turns on their digital human, the cloud-based conferencing system receives a replacement video / image uploaded from the participant's client or the host's client. If the cloud-based conferencing system receives a replacement image, it can convert it into a replacement video based on the replacement image and the previously mentioned video generation technology.

[0186] Of course, participants or the host can use modification commands to alter the replacement video being distributed by the cloud-based conferencing system to make the digital human more realistic. The technical details regarding modifying the digital human using specified modifications and modifying the replacement video using modification commands are quite similar and will not be elaborated upon here.

[0187] It is understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. In addition, the various embodiments described above can be combined according to actual conditions, and the combined solutions are still within the protection scope of this application.

[0188] Next, based on the methods in the above embodiments, a video conferencing control device 2000 provided in this application will be introduced.

[0189] For example, Figure 18 A schematic diagram of the structure of a video conferencing control device 2000 provided in an embodiment of this application is shown. Figure 18 As shown, the video conferencing control device 2000 includes: a detection module 2001 and a processing module 2002, wherein,

[0190] The detection module 2001 is used in a multi-person video conferencing scenario to detect a video replacement instruction and determine the replacement video. The video replacement instruction is used to instruct the replacement video to be used instead of the video data obtained in real time from the first client and distributed to the participating clients.

[0191] Processing module 2002 will send the replacement video to the participating clients, instead of sending the video data obtained in real time from the first client to the participating clients.

[0192] In some embodiments, the processing module 2002 is further configured to: obtain the video upload status of the first client during the process of detecting a video replacement instruction in the cloud conferencing system; and generate a video replacement instruction when the video upload status of the first client is detected to be abnormal.

[0193] In some embodiments, an abnormal video upload status means that the network status of the uplink data channel of the first client is abnormal or the camera of the first client is turned off.

[0194] In some embodiments, the video may include a view of the user of the first client.

[0195] In some embodiments, the source of the replacement video includes: videos / images pre-uploaded by a user of the first client.

[0196] In some embodiments, the source of the replacement video includes video data uploaded by the first client during the meeting.

[0197] In some embodiments, the source of the replacement video includes: a video / image pre-uploaded by the host of the video conference.

[0198] In some embodiments, the processing module 2002 is further configured to receive a replacement video related to the first client uploaded by the first client or the second client during the meeting, and generate a video replacement instruction, wherein the second client is any one of the participating clients.

[0199] In some embodiments, the processing module 2002 is further configured to determine that the first client has set the black screen replacement function before generating the video replacement instruction.

[0200] In some embodiments, the client includes a conference window, which is used to indicate sub-windows associated with each client; the conference window of the first client or the second client also includes an auto-screen control, which is used to generate action modification instructions or voice generation instructions.

[0201] In some embodiments, the processing module 2002 is further configured to receive an action modification instruction from a first client or the host of a video conference; modify the actions of the person in the replacement video based on the action modification instruction; and send the modified video to the participating clients.

[0202] In some embodiments, the processing module 2002 is further configured to: extract one or more video frames from the replacement video during the process of modifying the actions of a person in the replacement video based on the action modification instruction; and generate a new video based on the action modification instruction and one or more video frames, wherein the new video is the modified video.

[0203] In some embodiments, the processing module 2002, during the process of sending the modified video to the participating client, will send the modified video to the participating client once the screen of the distributed replacement video reaches one or more video frames, and will no longer send the replacement video to the participating client.

[0204] In some embodiments, the processing module 2002 is further configured to receive a voice generation instruction from a first client or the host of a video conference; generate a target voice based on the voice generation instruction; and distribute the target voice to participating clients.

[0205] In some embodiments, the processing module 2002, when the cloud conferencing system detects a video replacement instruction, includes: obtaining the on / off status of the video capture function of the first client; and generating a video replacement instruction when the video capture function of the first client is detected to be on.

[0206] In some embodiments, Figure 18 The detection module 2001 and processing modules 2002 shown can both be implemented in software or in hardware. For example, the implementation of the detection module 2001 will be described below. Similarly, the implementation of the processing module 2002 can refer to the implementation of the detection module 2001.

[0207] As an example of a software functional unit, the detection module 2001 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the detection module 2001 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0208] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0209] As an example of a hardware functional unit, the detection module 2001 may include at least one computing device, such as a server. Alternatively, the detection module 2001 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0210] The detection module 2001 includes multiple computing devices that can be distributed within the same region or in different regions. Similarly, the detection module 2001 can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the detection module 2001 can be distributed within the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0211] It should be noted that, in other embodiments, the detection module 2001 can be used to execute any step in the video conferencing control method described in the above embodiments, and the processing module 2002 can also be used to execute any step in the video conferencing control method described in the above embodiments. Furthermore, the steps implemented by the detection module 2001 and the processing module 2002 can also be specified as needed, and different steps in the video conferencing control method described in the above embodiments can be implemented by the detection module 2001 and the processing module 2002 respectively. Figure 18 The video conferencing control device 700 shown has all the functions of the device.

[0212] This application also provides a computing device 2100. For example... Figure 21 As shown, the computing device 2100 includes a bus 2102, a processor 2104, a memory 2106, and a communication interface 2108. The processor 2104, the memory 2106, and the communication interface 2108 communicate with each other via the bus 2102. The computing device 2100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 2100.

[0213] Bus 2102 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 21 The bus 2104 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 2104 may include a path for transmitting information between various components of the computing device 2100 (e.g., memory 2106, processor 2104, communication interface 2108).

[0214] Processor 2104 may include any one or more of the following processors: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP), or digital signal processor (DSP).

[0215] The memory 2106 may include volatile memory, such as random access memory (RAM). The processor 2104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0216] The memory 2106 stores executable program code, and the processor 2104 executes the executable program code to implement the aforementioned functions respectively. Figure 18 The detection module 2001 and processing module 2002 shown herein perform their functions to implement the video conferencing control method described in the above embodiments. That is, the memory 2106 stores instructions for executing the video conferencing control method described in the above embodiments.

[0217] Alternatively, the memory 2106 stores executable code, and the processor 2104 executes the executable code to implement the aforementioned functions respectively. Figure 18 The video conferencing control device 700 shown in the diagram performs the functions of the video conferencing control method described in the above embodiments. That is, the memory 2106 stores instructions for executing the video conferencing control method described in the above embodiments.

[0218] The communication interface 2108 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 2100 and other devices or communication networks.

[0219] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0220] like Figure 22 As shown, the computing device cluster includes at least one computing device 2100. The memory 2106 of one or more computing devices 2100 in the computing device cluster may store the same instructions for executing the video conferencing control method described in the above embodiments.

[0221] In some possible implementations, the memory 2106 of one or more computing devices 2100 in the computing device cluster may also store partial instructions for executing the video conferencing control method described in the above embodiments. In other words, a combination of one or more computing devices 2100 can jointly execute instructions for executing the video conferencing control method described in the above embodiments.

[0222] It should be noted that the memory 2106 in different computing devices 2100 within the computing device cluster can store different instructions, each used to execute the aforementioned instructions. Figure 18 The video conferencing control device 700 shown contains some of its functions. Specifically, the instructions stored in the memory 2106 of different computing devices 2100 can implement the functions of one or more modules in the detection module 2001 and processing module 2002.

[0223] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 23 One possible implementation is shown. For example... Figure 23 As shown, the two computing devices 2100A and 2100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this possible implementation, the memory 2106 in computing device 2100A stores instructions for executing the functions of the detection module 2001. Simultaneously, the memory 2106 in computing device 2100B stores instructions for executing the functions of the processing module 2002.

[0224] It should be understood that Figure 23The functions of computing device 2100A shown can also be performed by multiple computing devices 2100. Similarly, the functions of computing device 2100B can also be performed by multiple computing devices 2100.

[0225] This application also provides another computing device cluster. The connection relationships between the computing devices in this computing device cluster can be similarly referred to... Figure 22 and Figure 23 The connection method of the computing device cluster is different in that the memory 2106 of one or more computing devices 2100 in the computing device cluster can store the same instructions for executing the methods in the above embodiments.

[0226] In some possible implementations, the memory 2106 of one or more computing devices 2100 in the computing device cluster may also store partial instructions for executing the aforementioned video conferencing control method. In other words, a combination of one or more computing devices 2100 can jointly execute the instructions for executing the aforementioned video conferencing control method.

[0227] It should be understood that each step of the above method embodiments can be accomplished by hardware logic circuits or software instructions in a processor.

[0228] Based on the methods in the above embodiments, this application provides a computer-readable storage medium including computer program instructions. When executed by a cluster of computing devices including at least one computing device, the computer program instructions cause the cluster of computing devices to perform the methods in the above embodiments. Exemplarily, the computer-readable storage medium can be any available medium that the computing device can store, or a data storage device such as a data center containing one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives).

[0229] Based on the methods in the above embodiments, this application provides a computer program product containing instructions that, when executed by a cluster of computing devices containing at least one computing device, cause the cluster of computing devices to perform the methods in the above embodiments.

[0230] It is understood that the processor in the embodiments of this application can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.

[0231] The method steps in the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.

[0232] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0233] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application.

[0234] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.

Claims

1. A video conferencing control method, characterized in that, Applied to a cloud-based conferencing system, the cloud-based conferencing system includes a first client and at least one participating client, wherein the first client and the participating client establish a data connection through the cloud, the method includes: In a multi-person video conferencing scenario, the cloud conferencing system detects a video replacement instruction and determines to replace the video. The video replacement instruction is used to instruct the replacement video to replace the video data obtained in real time from the first client and distributed to the participating clients. The replacement video will be sent to the participating client instead of the video data obtained in real time from the first client.

2. The method according to claim 1, characterized in that, The cloud-based conferencing system detected a video replacement command, including: Get the video upload status of the first client; An abnormal video upload status was detected in the first client, and a video replacement instruction was generated.

3. The method according to claim 2, characterized in that, The video upload status being in an abnormal state means that the network status of the uplink data channel of the first client is abnormal or the camera of the first client is turned off.

4. The method according to any one of claims 1-3, characterized in that, The sources of the replacement video include: videos / images pre-uploaded by users of the first client.

5. The method according to any one of claims 1-3, characterized in that, The sources of the replacement video include: video data uploaded by the first client during the meeting.

6. The method according to any one of claims 1-3, characterized in that, The sources of the replacement video include: videos / images pre-uploaded by the host of the video conference.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: During the meeting, a replacement video related to the first client is received from either the first client or the second client, and a video replacement instruction is generated, wherein the second client can be any one of the participating clients.

8. The method according to claim 7, characterized in that, Before generating the video replacement instruction, it also includes: determining that the first client has set up the black screen replacement function.

9. The method according to any one of claims 1, characterized in that, The replacement video includes a facial image of the user of the first client.

10. The method according to any one of claims 1-9, characterized in that, The client includes a conference window, and the conference window of the first client or the second client also includes an automatic screen control, wherein the automatic screen control is used to generate action modification instructions or voice generation instructions, and the second client is any one of the participating clients.

11. The method according to any one of claims 1-10, characterized in that, The method further includes: Receive action modification instructions from the first client or the host of the video conference; Based on the action modification command, the actions of the characters in the replacement video are modified, and the modified video is sent to the participating client.

12. The method according to claim 11, characterized in that, The modification of the character's actions in the replacement video based on the action modification instruction includes: Extract one or more video frames from the replacement video; Based on the action modification instruction and the one or more video frames, a new video is generated, wherein the new video is the modified video.

13. The method according to claim 12, characterized in that, Sending the modified video to the participating client includes: Once the replaced video reaches one or more video frames, the modified video is sent to the participating client, and the replaced video is no longer sent to the participating client.

14. The method according to any one of claims 1-10, characterized in that, The method further includes: Receive voice generation instructions from the first client or the host of the video conference; Based on the speech generation instructions, the target speech is generated and distributed to the participating clients.

15. The method according to any one of claims 1-6, characterized in that, The cloud-based conferencing system detected a video replacement command, including: Get the on / off status of the video capture function on the first client; Upon detecting that the video capture function of the first client is enabled, a video replacement instruction is generated.

16. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-15.

17. A computer-readable storage medium, characterized in that, The method includes computer program instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method as described in any one of claims 1-15, wherein the cluster of computing devices includes at least one computing device.

18. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1-15, wherein the computing device cluster includes at least one computing device.