Network video conferencing processing methods, devices, electronic equipment and storage media
By simulating real meeting scenarios in online video conferencing and utilizing the collaboration of servers and terminals, virtual meeting scenarios are displayed and adapted to themes, solving the problems of limited terminal connection quantity and high power consumption, and enhancing the user's immersive experience and enjoyment.
Patent Information
- Application Number
- CN202111623524.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-08-12
- Filing Date
- 2021-12-28
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2041-12-28
AI Technical Summary
In existing network video conferencing systems, the number of terminal connections in multi-person video mode is limited, terminal power consumption is high, meeting scenarios are limited, user engagement is low, and the displayed meeting screen differs greatly from the actual scene, causing users to choose audio conferencing instead of video conferencing.
By simulating a real meeting scenario, the system utilizes the collaboration of servers and terminals to display a virtual meeting scene. The system displays the images of the participants in the virtual meeting scene at their respective corresponding positions. The positions are assigned according to the order of participants, their speaking order, and their identity information, adapting to the theme and scene. The system also uses masked images to remove the background and reduce the amount of computation on the terminal.
It achieves a good immersive experience, enhances user interest and engagement, reduces terminal power consumption, simulates a realistic meeting scenario, and increases user initiative.
Smart Images

Figure CN115706773B_ABST
Abstract
Description
[0001] This application claims priority to application number 202110924572.9, filed on August 12, 2021, entitled "Network Video Conferencing Processing Method, Apparatus, Electronic Device and Storage Medium". Technical Field
[0002] This application relates to the field of Internet technology, and in particular to a method, apparatus, electronic device and computer-readable storage medium for processing network video conferencing. Background Technology
[0003] The development of internet technology has accelerated the digital transformation of enterprises, with large-scale enterprises beginning to adopt online video conferencing to replace traditional meetings. Online video conferencing, based on cloud technology, greatly reduces the complexity of meeting organization and the barriers to participation, making it a more economical and flexible option for enterprises. It has a short preparation cycle, low cost, and is almost unrestricted by time and space.
[0004] However, the online video conferencing provided by related technologies is still mostly in a multi-person video mode, for example... Figure 1 This demonstrates a scenario where four users participating in a web-based video conference are engaged in a video call. Figure 1 It can be seen that the solutions provided by the relevant technologies simply piece together the images of people from different ends in the same frame, resulting in low user engagement and a significant discrepancy between the displayed meeting screen and the actual meeting scenario, thus reducing users' initiative to use online video conferencing. Summary of the Invention
[0005] This application provides a network video conferencing processing method, apparatus, electronic device, and computer-readable storage medium, which can simulate real meeting scenarios to achieve a good immersive experience during network video conferencing.
[0006] The technical solution of this application embodiment is implemented as follows:
[0007] This application provides a method for processing network video conferencing, including:
[0008] Receive video stream from a network video conference, wherein the video stream includes real-time images of multiple participants, and the multiple images of participants are obtained by separately capturing images of multiple participants in the network video conference;
[0009] Displaying a virtual meeting scene, and
[0010] In the virtual meeting scenario, the images of the multiple participants are displayed at positions corresponding to the multiple participants.
[0011] This application provides a network video conferencing processing device, including:
[0012] The receiving module is used to receive the video stream of the network video conference, wherein the video stream includes real-time images of multiple participants, and the multiple participants images are obtained by separately capturing images of multiple participants in the network video conference;
[0013] The display module is used to display the virtual meeting scene, and
[0014] In the virtual meeting scenario, the images of the multiple participants are displayed at positions corresponding to the multiple participants.
[0015] In the above scheme, the device further includes an allocation module, which is used to allocate corresponding positions to the multiple participants in the virtual meeting scene according to the order in which the multiple participants join the network video conference, wherein the position order of the multiple participants in the virtual meeting scene corresponds to the order in which they join; the display module is also used to display the images of the multiple participants according to the corresponding positions allocated to the multiple participants.
[0016] In the above scheme, the allocation module is further configured to allocate corresponding positions to the multiple participants in the virtual meeting scene according to the speaking order of the multiple participants in the online video conference, wherein the position order of the multiple participants in the virtual meeting scene corresponds to the speaking order; the display module is further configured to display the images of the multiple participants according to the corresponding positions allocated to the multiple participants respectively.
[0017] In the above scheme, the allocation module is further configured to sort the multiple participants according to their identity information in the online video conference, and assign corresponding positions to the multiple participants in the virtual conference scene, wherein the position sorting of the multiple participants in the virtual conference scene corresponds to the identity information sorting; the display module is further configured to display the images of the multiple participants according to the corresponding positions assigned to the multiple participants respectively.
[0018] In the above scheme, the display module is further configured to display a virtual meeting scene adapted to the theme of the online video conference based on the video stream; the device further includes an update module, configured to update the virtual meeting scene to adapt to the changed theme when it is determined from the video stream that the theme of the online video conference has changed.
[0019] In the above scheme, the device further includes a decoding module for decoding the video stream to obtain multiple video frames; the device further includes a topic recognition module for calling a topic recognition model to perform topic recognition processing on the multiple video frames to obtain topic recognition results for each video frame, and determining the topic recognition result with the highest repetition frequency as the topic of the network video conference; the display module is also used to display a virtual meeting scene adapted to the topic of the network video conference.
[0020] In the above scheme, the device further includes a moving module, used to move the image of the target participant from its originally assigned position to a specific position in the virtual meeting scene when it is identified that there is a target participant currently speaking among the plurality of participants, wherein the salience of the specific position is greater than that of the originally assigned position; the moving module is also used to move the image of the target participant from the specific position to the originally assigned position when it is identified that the target participant has finished speaking.
[0021] In the above scheme, the video stream also includes multiple mask images corresponding one-to-one with the multiple participant images, and the multiple mask images are obtained by object recognition of the multiple participant images respectively; the device also includes a masking module, which is used to perform the following processing for each participant image: masking the participant image based on the mask image corresponding to the participant image to obtain the participant image with the background removed; the display module is also used to display the multiple participant images with the background removed at the positions corresponding to the multiple participants in the virtual meeting scene.
[0022] In the above scheme, the range of pixel values of the mask image is smaller than the range of pixel values of the image of the participant; the device further includes a mapping module for mapping the mask image so that the range of pixel values of the mask image is consistent with the range of pixel values of the image of the participant.
[0023] In the above scheme, the receiving module is further configured to receive the video stream of the network video conference generated by the server in the following ways: when any participant among the multiple participants leaves the network video conference or the connection is abnormal, the participant is removed from the network video conference; the remaining participants of the network video conference receive the participant images, and the network video conference is generated based on the participant images of the remaining participants.
[0024] In the above scheme, the positions corresponding to the multiple participants are selected from the unoccupied positions in the virtual meeting scene; after removing any participant from the online video conference, the update module is also used to update the position corresponding to any participant in the virtual meeting scene from an occupied state to an unoccupied state.
[0025] In the above scheme, the video stream also includes an image of the virtual meeting scene; the decoding module is further used to decode the video stream to obtain video pixel data; the device also includes a rendering module, used to perform rendering processing based on the decoded video pixel data, so as to display the image of the virtual meeting scene in the human-computer interaction interface, and to display the images of the multiple participants in the image of the virtual meeting scene at positions corresponding to the multiple participants respectively.
[0026] In the above scheme, the device further includes an image segmentation module, used to perform image segmentation processing on each of the participant images to obtain a mask image corresponding to the participant image; the masking module is also used to perform masking processing on the corresponding participant image based on each mask image to obtain multiple participant images with background removed; the display module is also used to display the multiple participant images with background removed at the positions corresponding to the multiple participants in the image of the virtual meeting scene.
[0027] In the above scheme, the image segmentation module is further configured to perform the following processing for each of the participant images: calling an image segmentation model based on the participant image to identify the participants in the participant image, taking the area outside the participant as the background, and generating a mask image corresponding to the background; wherein, the image segmentation model is trained based on sample images and the objects labeled in the sample images.
[0028] In the above scheme, the receiving module is further configured to receive the video stream of the network video conference generated by the server in the following ways: acquiring multiple participant images obtained by image acquisition of multiple participants in the network video conference, and a mask image corresponding to each participant image; performing masking processing on the corresponding participant image based on each mask image to obtain multiple participant images with background removed; acquiring an image of a virtual meeting scene adapted to the theme of the network video conference; filling the positions corresponding to the multiple participants in the image of the virtual meeting scene with the multiple participant images with background removed to obtain a merged image; and performing encoding processing on the merged images corresponding to different times to obtain the video stream of the network video conference.
[0029] This application provides an electronic device, including:
[0030] Memory, used to store executable instructions;
[0031] The processor, when executing executable instructions stored in the memory, implements the network video conferencing processing method provided in the embodiments of this application.
[0032] This application provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement the network video conferencing processing method provided in this application.
[0033] This application provides a computer program product, which includes computer-executable instructions for implementing the network video conferencing processing method provided in this application when executed by a processor.
[0034] The embodiments of this application have the following beneficial effects:
[0035] By displaying a virtual meeting scene and showing multiple participant images at their respective positions within the virtual meeting scene, a realistic meeting scenario can be simulated, giving users a sense of being there and enhancing their immersive experience when using online video conferencing. Attached Figure Description
[0036] Figure 1 This is a schematic diagram illustrating the application scenarios of network video conferencing processing methods provided by related technologies;
[0037] Figure 2 This is a schematic diagram of the architecture of the network video conferencing processing system 100 provided in the embodiments of this application;
[0038] Figure 3 This is a schematic diagram of the structure of the electronic device 500 provided in the embodiments of this application;
[0039] Figure 4 This is a flowchart illustrating the network video conferencing processing method provided in an embodiment of this application;
[0040] Figures 5A to 5C This is a flowchart illustrating the network video conferencing processing method provided in an embodiment of this application;
[0041] Figures 6A to 6C This is a schematic diagram illustrating an application scenario of the network video conferencing processing method provided in the embodiments of this application;
[0042] Figure 7 This is a schematic diagram of the overall framework for the interaction between multiple terminals and the server provided in the embodiments of this application;
[0043] Figure 8This is a schematic diagram illustrating the interaction process between a single terminal and a server, provided in an embodiment of this application.
[0044] Figure 9 This is a schematic diagram illustrating a scenario of interaction between a single web conferencing client and a server, as provided in an embodiment of this application.
[0045] Figure 10 This is a schematic diagram of the precoding process for a human face mask provided in an embodiment of this application;
[0046] Figure 11 This is a schematic diagram illustrating a scenario of interaction between multiple web conferencing clients and a server, as provided in an embodiment of this application.
[0047] Figure 12 This is a schematic diagram of the post-decoding process for the decoded portrait mask image provided in an embodiment of this application;
[0048] Figure 13 This is a flowchart illustrating the network video conferencing processing method provided in an embodiment of this application;
[0049] Figure 14 This is a flowchart illustrating the network video conferencing processing method provided in an embodiment of this application;
[0050] Figure 15 This is a schematic diagram illustrating an application scenario of the network video conferencing processing method provided in the embodiments of this application;
[0051] Figure 16 This is a schematic diagram illustrating an application scenario of the network video conferencing processing method provided in the embodiments of this application. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0053] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0055] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0056] 1) Online video conferencing is an interactive mode that uses a network (such as the Internet or a local area network) as the communication medium. The multimedia data of any participant, such as voice and video, can be synchronized to other participants in real time, thereby breaking through the communication limitations of participants in terms of spatial distance.
[0057] 2) Video bitrate (data rate) refers to the amount of data a video file uses per unit of time. Also called bitrate, sampling rate, or stream rate, it's the most important part of image quality control in video encoding. The commonly used unit is kb / s or Mb / s. Generally, at the same resolution, the higher the bitrate of a video file, the lower the compression ratio and the higher the image quality. A higher bitrate means a higher sampling rate per unit of time, resulting in higher data flow precision. The processed file is closer to the original file, with better image quality and clearer picture. Of course, this also requires a higher decoding capability from the playback device.
[0058] 3) H.264, a commonly used data encoding algorithm, introduces a new concept at the system level: a conceptual division between the Video Coding Layer (VCL) and the Network Abstraction Layer (NAL). The former is the representation of the core compressed content of the video content, while the latter is the representation delivered through a specific type of network. This structure facilitates information encapsulation and better priority control of information.
[0059] 4) Virtual meeting scenarios are simulated environments used to accommodate participants, such as meeting rooms and classrooms.
[0060] Currently, most online video conferencing technologies still operate on a multi-person video model, simply piecing together images of people from different ends into the same frame. For example... Figure 1 This demonstrates a scenario where four users are simultaneously making a video call. Figure 1 It can be seen that the solutions provided by the relevant technologies are simply a patchwork of portrait images from various terminals. Further implementations would include virtual backgrounds and other effects on each terminal.
[0061] However, in implementing the embodiments of this application, the applicant discovered that the solutions provided by related technologies have the following obvious problems:
[0062] 1. The number of terminal connections is limited, usually within 10;
[0063] 2. If the virtual background is enabled, the terminal's algorithm will run entirely locally on the terminal. After prolonged use, the terminal will experience significant power consumption issues (such as overheating and excessive power consumption).
[0064] 3. The scenarios are relatively simple (just a simple patchwork of images), which makes it less interesting for users. The displayed meeting screen is also quite different from the actual meeting scene. In practical applications, most users are unwilling to turn on the video and choose to make an audio conference call instead.
[0065] To address the aforementioned technical issues, embodiments of this application provide a network video conferencing processing method, apparatus, electronic device, and computer-readable storage medium, which can simulate real meeting scenarios to achieve a good immersive experience during network video conferencing.
[0066] The following describes exemplary applications of the electronic devices provided in the embodiments of this application. The network video conferencing processing method provided in the embodiments of this application can be implemented by various electronic devices, such as laptops, tablets, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), and other types of user terminals. It can also be implemented collaboratively by a server and a terminal. The following description uses the implementation of the network video conferencing processing method provided in the embodiments of this application by a server and a terminal as an example.
[0067] See Figure 2 , Figure 2 This is a schematic diagram of the architecture of the network video conferencing processing system 100 provided in the embodiments of this application. In order to support an application that simulates a real meeting scenario, the network video conferencing processing system 100 includes: a server 200, a network 300, and terminals (terminals 400-1 to 400-N are shown as examples, where N is a positive integer constant greater than 1, for example, the value of N can be 2, 3, 5, etc.), which will be described below.
[0068] Server 200 is the backend server for clients 410-1 to 410-N, used to send the video stream of the online video conference to terminals 400-1 associated with participant 1 to terminals 400-N associated with participant N. The video stream sent by server 200 includes N participant images, which are obtained by capturing images of the N participants in the online video conference. For example, server 200 first receives participant images 1 to N sent by terminals 1 to N respectively. Participant image 1 is obtained by terminal 1 using its camera to capture images of participant 1, and participant image N is obtained by terminal N using its camera to capture images of participant N. Then, server 200 encodes the received participant images 1 to N to obtain the video stream for sending to terminals 1 to N.
[0069] Network 300 is used as a medium for communication between server 200 and terminals 400-1 to 400-N. Network 300 can be a wide area network or a local area network, or a combination of both.
[0070] Clients 410-1 to 410-N run on terminals 400-1 to 400-N respectively. Clients 410-1 to 410-N are of the same type, such as web conferencing clients or instant messaging clients. Taking client 410-1 as an example, after receiving the video stream of the web video conference sent by server 200, client 410-1 displays a virtual meeting scene on the human-computer interaction interface according to the received video stream, and displays the images of N participants in the virtual meeting scene at the positions corresponding to the N participants. In this way, compared with the related technology, which simply splices the images of participants collected by each terminal into the same screen, this embodiment of the application can simulate a real meeting scene by displaying a virtual meeting scene and displaying the corresponding participant image at the position corresponding to each participant in the virtual meeting scene, giving users an immersive meeting feeling, effectively improving the fun of web video conferencing, and thus achieving a good immersive experience in the process of web video conferencing.
[0071] In some embodiments, the embodiments of this application can be implemented using cloud technology, which refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to realize the computation, storage, processing, and sharing of data.
[0072] Cloud technology is a general term encompassing network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form resource pools, allowing for on-demand use with flexibility and convenience. Cloud computing technology will become a crucial support. For example, the service interaction functions between server 200 and terminals 400-1 to 400-N mentioned above can be realized through cloud technology.
[0073] Example, Figure 1 The server 200 shown can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Terminals 400-1 to 400-N can be smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, etc., but are not limited to these. Terminals 400-1 to 400-N can be directly or indirectly connected to server 200 via wired or wireless communication, which is not limited in this embodiment.
[0074] For example, the online video conferencing provided in this application embodiment can also be a cloud conferencing, which is an efficient, convenient, and low-cost form of conferencing based on cloud computing technology. Users only need to perform simple and easy-to-use operations through an internet interface to quickly and efficiently share voice, data files, and video with teams and customers around the world. The complex technologies such as data transmission and processing during the meeting are handled by the cloud conferencing service provider.
[0075] Currently, cloud conferencing in China mainly focuses on services based on the Software as a Service (SaaS) model, including telephone, internet, and video services. Video conferencing based on cloud computing is called cloud conferencing.
[0076] In the era of cloud conferencing, data transmission, processing, and storage are all handled by the computer resources of the video conferencing vendors. Users no longer need to purchase expensive hardware or install cumbersome software. They can simply open a browser, log in to the corresponding interface, and conduct efficient remote meetings.
[0077] In other embodiments, taking terminal 400-1 as an example, terminal 400-1 can also implement the network video conferencing processing method provided in the embodiments of this application by running a computer program. The computer program can be as follows: Figure 1The client 410-1 is shown. For example, a computer program can be a native program or software module in an operating system; it can be a native application (APP), i.e., a program that needs to be installed in the operating system to run, such as a web conferencing client or an instant messaging client; it can also be a small program, i.e., a program that only needs to be downloaded to a browser environment to run; or it can be a small program that can be embedded in any APP, wherein the small program can be controlled by the user to run or close. In short, the above-mentioned computer program can be any form of application, module, or plugin.
[0078] The structure of the electronic device provided in the embodiments of this application will be described below. Taking the electronic device as an example as a terminal, see [link to example]. Figure 3 , Figure 3 This is a schematic diagram of the structure of the electronic device 500 provided in the embodiments of this application. Figure 3 The illustrated electronic device 500 includes at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to implement communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 3 The general labeled all buses as Bus System 540.
[0079] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0080] User interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0081] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 550 may optionally include one or more storage devices physically located away from the processor 510.
[0082] The memory 550 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550 described in this application embodiment is intended to include any suitable type of memory.
[0083] In some embodiments, memory 550 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0084] Operating system 551 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;
[0085] The network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.
[0086] Presentation module 553 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 (e.g., a display screen, a speaker, etc.) associated with user interface 530;
[0087] The input processing module 554 is used to detect and translate one or more user inputs or interactions from one or more input devices 532.
[0088] In some embodiments, the network video conferencing processing apparatus provided in this application can be implemented in software. Figure 3 A network video conferencing processing device 555 stored in memory 550 is shown. This device can be software in the form of programs and plug-ins, including the following software modules: a receiving module 5551, a display module 5552, an allocation module 5553, an update module 5554, a decoding module 5555, a topic recognition module 5556, a movement module 5557, a masking module 5558, a mapping module 5559, a rendering module 55510, and an image segmentation module 55511. These modules are logically connected and can therefore be arbitrarily combined or further split according to the functions implemented. It should be noted that... Figure 3For ease of description, all the above modules are shown at once, but this should not be interpreted as excluding the implementation of the network video conferencing processing device 555 which may only include the receiving module 5551 and the display module 5552. The functions of each module will be described below.
[0089] In other embodiments, the network video conferencing processing device provided in this application can be implemented in hardware. As an example, the device provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the network video conferencing processing method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0090] The following will describe the network video conferencing processing method provided in this application embodiment, with reference to exemplary applications and implementations of the terminals provided in the embodiments of this application. For example, see... Figure 4 , Figure 4 This is a flowchart illustrating the network video conferencing processing method provided in the embodiments of this application, which will be combined with... Figure 4 The steps shown are explained.
[0091] It should be noted that, Figure 4 The method shown can be derived from Figure 1 The execution of various forms of computer programs running on the terminal 400-1 shown is not limited to the client 410-1 described above, but can also be the operating system 461, software modules and scripts described above. Therefore, the client described below should not be regarded as a limitation on the embodiments of this application.
[0092] In step S101, the video stream from the network video conference is received.
[0093] Here, the video stream may include real-time images of multiple participants, and these multiple participant images are obtained by capturing images of multiple participants in the online video conference.
[0094] For example, taking multiple participants in a network video conference as users A to D, the server first receives real-time facial images captured by the cameras of the terminals associated with users A to D (i.e., the server receives facial images of user A obtained by terminal A (associated with user A) using its camera, facial images of user B obtained by terminal B (associated with user B) using its camera, facial images of user C obtained by terminal C (associated with user C) using its camera, and facial images of user D obtained by terminal D (associated with user D) using its camera). Then, the server encodes the received facial images of users A to D to obtain video streams for the network video conference, which are sent to terminals A to D respectively. The video streams include the facial images of users A to D.
[0095] It should be noted that in practical applications, the terminal can receive a synchronous audio stream while receiving the video stream. Then, the audio stream is decoded to obtain audio sampling data. Subsequently, while rendering and displaying the video pixel data obtained from decoding the video stream, the audio sampling data synchronized with the video pixel data is played, thus achieving synchronized playback of video and audio.
[0096] In some embodiments, the above-mentioned receiving of video streams from network video conferences can be achieved by receiving video streams from network video conferences generated by the server in the following ways: removing any participant from the network video conference when any participant leaves the network video conference or when a connection error occurs; receiving participant images of the remaining participants in the network video conference, and generating video streams from the network video conference based on the participant images of the remaining participants.
[0097] For example, taking multiple participants in a video conference, users A through D, as an example, when user A leaves the video conference (e.g., user A clicks the "End Meeting" button displayed on the human-computer interaction interface) or a connection error occurs (e.g., the network connection of user A's associated terminal is abnormal, resulting in network outages or limited network connectivity), the server removes user A from the video conference. Then, the server receives images of user B's face captured by the camera of user B's associated terminal B, and images of user C captured by the camera of user C's associated terminal C. The server receives the facial images of users B through D, and the facial image of user D obtained by the camera of terminal D associated with user D. Then, the server generates a video stream for the online video conference based on the received facial images of users B through D (at this point, the video stream only includes the facial images of users B through D). Thus, when the terminal displays the virtual conference scene based on the video stream, the facial image of user A is removed from the virtual conference scene (i.e., user A's facial image is not displayed in the virtual conference scene), allowing the conference initiator to monitor the participants in real time, such as whether any users have left the online video conference.
[0098] In other embodiments, the positions corresponding to multiple participants are selected from unoccupied positions in the virtual meeting scene; then after removing any participant from the online video conference, the following process can also be performed: update the position corresponding to any participant in the virtual meeting scene from occupied to unoccupied.
[0099] For example, continuing from the previous text, the positions corresponding to users A through D (i.e., the positions used to display the face images of users A through D) are selected from unoccupied positions in the virtual meeting scene. Assume that user A's position is position 1 in the virtual meeting scene (at this time, position 1 is occupied). When the server removes user A from the video conference due to leaving or a connection error, it can update position 1 from occupied to unoccupied. At this time, other participants in the video conference (such as users B through D, or newly joined users E and F) can occupy position 1 in the virtual meeting scene. Thus, by reclaiming resources from positions in the virtual meeting scene that are inactive (i.e., the corresponding user has left the video conference or a connection error has occurred), server system resources are saved.
[0100] In other embodiments, the above-mentioned receiving of the video stream of the online video conference can also be achieved by the following method: receiving the video stream of the online video conference generated by the server in the following way: acquiring multiple participant images obtained by image acquisition of multiple participants in the online video conference, and a mask image corresponding to each participant image; performing masking processing on the corresponding participant image based on each mask image to obtain a participant image with the background removed; acquiring an image of a virtual meeting scene adapted to the theme of the online video conference; filling the positions corresponding to the multiple participants in the image of the virtual meeting scene with multiple participant images with the background removed to obtain a merged image; and encoding the merged images corresponding to different times to obtain the video stream of the online video conference.
[0101] For example, taking multiple participants in a network video conference, users A to D, as an example, the server first receives the face images of users A to D sent by the terminals associated with users A to D, as well as the corresponding mask images for each face image (that is, the server receives the face image and mask image A of user A sent by terminal A, the face image and mask image B of user B sent by terminal B, the face image and mask image C of user B sent by terminal C, and the face image and mask image D of user D sent by terminal D). Among them, mask image A is the face image of user A obtained by terminal A using the object recognition model. The images are obtained through object recognition. Mask image B is obtained by terminal B calling the object recognition model to perform object recognition on user B's face image. Mask image C is obtained by terminal C calling the object recognition model to perform object recognition on user C's face image. Mask image D is obtained by terminal D calling the object recognition model to perform object recognition on user D's face image. Then, the server performs masking processing on the corresponding face image based on each mask image to obtain the face image with the background removed. For example, the server performs masking processing on user A's face image based on mask image A to obtain user A's face image with the background removed. The server performs masking on user B's face image based on mask image B, resulting in a background-removed face image of user B. Similarly, it performs masking on user C's face image based on mask image C, resulting in a background-removed face image of user C. Finally, it performs masking on user D's face image based on mask image D, resulting in a background-removed face image of user D. The server then acquires an image of a virtual meeting scene that matches the theme of the current video conference (e.g., a background image selected by the host or automatically selected by the server based on the theme). After obtaining the virtual meeting scene image, the server fills the corresponding positions for users A through D with the background-removed face images of users A through D, resulting in a merged image. Finally, the server encodes the merged images corresponding to different times to obtain a video stream for sending to terminals A through D. Upon receiving the video stream from the server, terminals A through D can directly display the merged image. In other words, the terminals are only responsible for image acquisition and displaying the merged image, significantly reducing the computational load on the terminals and thus effectively lowering their power consumption.
[0102] In step S102, a virtual meeting scene is displayed.
[0103] In some embodiments, the above-described display of the virtual meeting scenario can be achieved by: displaying a virtual meeting scenario adapted to the topic of the online video conference based on the video stream; and updating the virtual meeting scenario to adapt to the changed topic when it is determined from the video stream that the topic of the online video conference has changed.
[0104] For example, the above-mentioned virtual meeting scene adapted to the theme of the online video conference can be implemented in the following way: decode the video stream to obtain multiple video frames; call the theme recognition model to perform theme recognition processing on the multiple video frames to obtain the theme recognition results of each video frame, and determine the theme recognition result with the highest repetition as the theme of the online video conference; obtain and display the virtual meeting scene adapted to the determined theme of the online video conference.
[0105] For example, suppose that after decoding the current video stream, 10 video frames are obtained. Then, the terminal calls the topic recognition model to perform topic recognition processing on these 10 video frames, and obtains the topic recognition results for each video frame. Subsequently, the topic recognition result with the highest repetition is determined as the topic of the current online video conference. For example, suppose that the topic recognition results of 6 video frames in these 10 video frames are academic conferences, then academic conferences are determined as the topic of the current online video conference. Finally, the terminal obtains and displays a virtual meeting scene (such as a virtual classroom scene) adapted to the academic conference.
[0106] In other embodiments, the virtual meeting scenario can also be manually set by the meeting initiator. For example, multiple candidate virtual meeting scenarios are presented in the human-computer interaction interface of the terminal associated with the meeting initiator for the meeting initiator to choose from. In this way, when the meeting initiator selects a virtual meeting scenario (e.g., virtual meeting scenario 1) from the multiple candidate virtual meeting scenarios, virtual meeting scenario 1 can be displayed in the human-computer interaction interface of the terminal associated with each participant in the online video conference.
[0107] It should be noted that in practical applications, the virtual meeting scenarios displayed on the human-computer interaction interfaces of the terminals associated with different participants in a network video conference can also be different. That is, multiple candidate virtual meeting scenarios can be displayed on the human-computer interaction interface of each participant's associated terminal for them to choose from. For example, if participant A selects virtual meeting scenario 1, then virtual meeting scenario 1 will be displayed on the human-computer interaction interface of participant A's associated terminal; while if participant B selects virtual meeting scenario 2, then virtual meeting scenario 2 will be displayed on the human-computer interaction interface of participant B's associated terminal. This can meet the personalized needs of users. Furthermore, it should be noted that the order of multiple participants remains consistent across different virtual meeting scenarios. That is, the human-computer interaction interfaces of the terminals associated with different participants simply display different types of virtual meeting scenarios; the order of multiple participants in a network video conference remains the same across different types of virtual meeting scenarios.
[0108] Furthermore, when the terminal determines that the topic of the online video conference has changed based on the video stream, it can automatically update the virtual meeting scene to adapt to the changed topic. For example, when the terminal calls the topic recognition model to perform topic recognition processing on multiple video frames obtained from the subsequent video stream decoding, and the topic recognition result with the highest repetition frequency changes to a seminar, the terminal automatically obtains and displays a virtual meeting scene (such as a virtual conference room scene) that is adapted to the seminar. In this way, the virtual meeting scene will automatically update as the topic of the online video conference changes, improving the user experience.
[0109] It should be noted that the topic recognition model mentioned above can be a neural network model (such as a convolutional neural network, a deep convolutional neural network, a fully connected neural network, etc.), a decision tree model, a gradient boosting tree, a multilayer perceptron, and a support vector machine, etc. The embodiments of this application do not specifically limit the type of topic recognition model.
[0110] In other embodiments, the topic of a video conference can also be determined based on the meeting schedule, shared files, etc. For example, the terminal can determine the topic of the video conference based on information such as the meeting name and meeting content carried in the meeting schedule. Of course, the terminal can also determine the topic of the video conference based on the files shared by the participants in the video conference, such as the file name and file content of the shared files.
[0111] In some embodiments, the virtual meeting scene can also be manually updated. For example, the meeting initiator may have the authority to update the virtual meeting scene. For instance, assuming the current virtual meeting scene is a virtual classroom scene, when the next agenda item is a more relaxed seminar, the meeting initiator may manually update the virtual meeting scene to a virtual meeting room scene in advance.
[0112] In step S103, multiple participant images are displayed at positions corresponding to the multiple participants in the virtual meeting scene.
[0113] In some embodiments, Figure 4 The illustrated step S103 can be achieved through Figure 5A Steps S1031A to S1032A shown are implemented, and will be combined with Figure 5A The steps shown are explained.
[0114] In step S1031A, based on the order in which multiple participants join the online video conference, corresponding positions are assigned to each participant in the virtual meeting scenario.
[0115] Here, the order in which multiple participants are positioned in the virtual meeting scenario corresponds to the order in which they join the online video conference.
[0116] In some embodiments, the above-mentioned method of assigning corresponding positions to multiple participants in a virtual meeting scenario based on their order of joining the online video conference can be achieved as follows: Multiple participants are sorted in descending order based on the time they join the online video conference; the first position in the virtual meeting scenario is designated as the position corresponding to the participant ranked first in the descending order (i.e., the participant who joined the online video conference earliest); the other positions after the first position in the virtual meeting scenario are sequentially designated as the positions corresponding to the other participants after the first participant in the descending order, for example, the second position in the virtual meeting scenario is designated as the position corresponding to the second participant in the descending order, and so on. In this way, by mapping the order of the positions of multiple participants in the virtual meeting scenario to the order in which they join the online video conference, the user's enthusiasm for joining the online video conference is increased.
[0117] In step S1032A, images of multiple participants are displayed according to the corresponding positions assigned to them.
[0118] In some embodiments, after assigning corresponding positions to multiple participants in the virtual meeting scene according to the order in which they join the online video conference, images of multiple participants can be displayed according to the corresponding positions assigned to each participant. For example, assuming that participant A is assigned position 1 in the virtual meeting scene, then participant A's image (e.g., participant A's face image) is displayed at position 1. If participant B is assigned position 2, then participant B's image (e.g., participant B's face image) is displayed at position 2, and so on.
[0119] It should be noted that in a virtual meeting scenario, the corresponding positions assigned to multiple participants can be consecutive or spaced out (for example, there is an empty position between two adjacent participants). As long as the position order corresponds to the order in which the multiple participants join the online video conference, this application embodiment does not impose specific limitations on this.
[0120] For example, see Figure 6A , Figure 6A This is a schematic diagram illustrating an application scenario of the network video conferencing processing method provided in the embodiments of this application, such as... Figure 6A As shown, in the virtual meeting scenario 601, multiple participant images (e.g., participant image 602, participant image 603, participant image 604, and participant image 605) are displayed in a consecutive order. The positional order of the multiple participant images corresponds to the order in which the participants joined the online video conference. For example, the participant corresponding to the leftmost participant image 602 (e.g., user A) joined the online video conference the earliest, the participant corresponding to the participant image 603 (e.g., user B) joined the online video conference the next earliest, the participant corresponding to the participant image 604 (e.g., user C) joined the online video conference the next earliest, and the participant corresponding to the rightmost participant image 605 (e.g., user D) joined the online video conference the latest. In this way, by corresponding the positional order to the order in which the participants joined the online video conference, the enthusiasm of users to join the online video conference can be improved.
[0121] For example, see Figure 6B , Figure 6B This is a schematic diagram illustrating an application scenario of the network video conferencing processing method provided in the embodiments of this application, such as... Figure 6BAs shown, multiple participant images (e.g., participant images 606, 607, 608, and 609) are displayed at intervals in the virtual meeting scene 610, and the order of their positions corresponds to the order in which they joined the online video conference. For example, the participant corresponding to the leftmost participant image 606 (e.g., user A) joined the online video conference the earliest, the participant corresponding to participant image 607 (e.g., user B) joined the online video conference the next earliest, the participant corresponding to participant image 608 (e.g., user C) joined the online video conference the next earliest, and the participant corresponding to the rightmost participant image 610 (e.g., user D) joined the online video conference the latest. In this way, by corresponding the positional order to the order in which they joined the online video conference, the user's enthusiasm for joining the online video conference can be improved.
[0122] In other embodiments, Figure 4 The step S103 shown can also be achieved through... Figure 5B Steps S1031B to S1032B shown are implemented, and will be combined with Figure 5B The steps shown are explained.
[0123] In step S1031B, based on the speaking order of multiple participants in the online video conference, corresponding positions are assigned to each participant in the virtual conference scenario.
[0124] Here, the order in which multiple participants are positioned in the virtual meeting scenario corresponds to their speaking order.
[0125] In some embodiments, the speaking order of multiple participants in a video conference can be obtained in advance. For example, the conference schedule (which records the conference content or the speaking time of each participant) can be obtained, and the speaking order of multiple participants in the video conference can be determined according to the conference schedule. Then, according to the speaking order of multiple participants in the video conference, corresponding positions are assigned to each participant in the virtual conference scene. For example, if participant A is the first to speak, the first position in the virtual conference scene can be used as the position corresponding to participant A. If participant B is the second to speak, the second position in the virtual conference scene can be used as the position corresponding to participant B, and so on. In this way, by mapping the position order to the speaking order of multiple participants, the efficiency of users using video conferencing for meetings is improved.
[0126] In step S1032B, images of multiple participants are displayed according to the corresponding positions assigned to them.
[0127] In some embodiments, after assigning corresponding positions to multiple participants in the virtual meeting scene according to their speaking order in the online video conference, images of multiple participants can be displayed according to the assigned positions. For example, assuming that participant A (the first participant to speak) is assigned position 1 in the virtual meeting scene, then participant A's image (e.g., participant A's face image) is displayed at position 1. If participant B (the second participant to speak) is assigned position 2, then participant B's image (e.g., participant B's face image) is displayed at position 2, and so on. In this way, multiple participant images are displayed according to their speaking order, improving the efficiency of the online video conference.
[0128] It should be noted that in a virtual meeting scenario, the corresponding positions assigned to multiple participants can be consecutive or spaced out (for example, there is an empty position between two adjacent participants), as long as the position order corresponds to the speaking order of the multiple participants in the online video conference. This application embodiment does not impose specific limitations on this.
[0129] In some embodiments, Figure 4 The step S103 shown can also be achieved through... Figure 5C Steps S1031C to S1032C shown are implemented, and will be combined with Figure 5C The steps shown are explained.
[0130] In step S1031C, based on the identity information of multiple participants in the online video conference, corresponding positions are assigned to each participant in the virtual conference scenario.
[0131] Here, the order of the positions of multiple participants in the virtual meeting scenario corresponds to the order of their identity information. The order of identity information can be based on the participant's role (e.g., host, speaker, audience), job title, account level, department, etc.
[0132] In some embodiments, while sending images of participants to the server, the terminal can also send the bound account to the server, so that the server can obtain the participant's identity information (such as participant role, position, account level, department, etc.) based on the account. Then, the server sorts the identity information of multiple participants and assigns corresponding positions to each participant in the virtual meeting scene according to the sorting of the participants' identity information in the online video conference. For example, taking the sorting of participant roles as an example, the host's position in the virtual meeting scene is the most forward (for example, the most forward position in the virtual meeting scene can be used as the host's position), the speaker's position in the virtual meeting scene is the next forward (for example, the middle part of the virtual meeting scene can be used as the speaker's position), and the audience's position in the virtual meeting scene is the most backward (for example, the back part of the virtual meeting scene can be used as the audience's position). In this way, by corresponding the position sorting with the sorting of the identity information of multiple participants, the management of online video conferences is facilitated.
[0133] It should be noted that in practical applications, a corresponding area can be pre-assigned in the virtual meeting scenario for each type of identity information. Different types of identity information correspond to different areas in the virtual meeting scenario. For example, the earlier the identity information is listed, the closer the corresponding area is to the center of the virtual meeting scenario.
[0134] In step S1032C, images of multiple participants are displayed according to the corresponding positions assigned to them.
[0135] In some embodiments, after assigning corresponding positions to multiple participants in the virtual meeting scene based on their identity information in the online video conference, images of multiple participants can be displayed according to the corresponding positions assigned to them. For example, taking the order of participant roles as an example, if the host of the online video conference is assigned position 1 in the first row in the virtual meeting scene, then the host's image (e.g., the host's face image) is displayed in position 1. If the speaker A is assigned position 2 in the second row, then the speaker A's image (e.g., the speaker A's face image) is displayed in position 2. If the audience B is assigned position 3 in the fifth row, then the audience B's image (e.g., the audience B's face image) is displayed in position 3. In this way, by corresponding the position order with the order of the participants' roles, the host can easily manage the online video conference.
[0136] In other embodiments, the terminal may also perform the following processing: when it is identified that there is a target participant who is currently speaking among multiple participants, the participant image corresponding to the target participant is moved from the originally assigned position to a specific position in the virtual meeting scene, wherein the significance of the specific position is greater than the original position assigned to the target participant; when it is identified that the target participant has finished speaking, the participant image corresponding to the target participant is moved back from the specific position to the originally assigned position.
[0137] For example, see Figure 6C , Figure 6C This is a schematic diagram illustrating an application scenario of the network video conferencing processing method provided in the embodiments of this application, such as... Figure 6C As shown, multiple participant images are displayed in the virtual meeting scene 611. When the terminal identifies a target participant (e.g., participant A) who is currently speaking among the multiple participants, the participant image 612 corresponding to participant A can be moved from its original assigned position 613 to a specific position 614 in the virtual meeting scene 611 (the specific position 614 can be the middle position in the first row of the virtual meeting scene 611). Then, when it is identified that participant A has finished speaking, the participant image 612 corresponding to participant A can be moved back from the specific position 614 to its original assigned position 613. In this way, by moving the participant image corresponding to the currently speaking participant to a specific position in the virtual meeting scene, the attention of other participants can be attracted, thus improving the user experience.
[0138] In other embodiments, the video stream may further include multiple mask images corresponding one-to-one with multiple participant images, and the multiple mask images are obtained by object recognition of the multiple participant images respectively; then the above-mentioned display of multiple participant images at positions corresponding to multiple participant images in the virtual meeting scene can be achieved in the following way: perform the following processing on each participant image: perform mask processing on the participant image based on the mask image corresponding to the participant image to obtain a participant image with the background removed; display the multiple participant images with the background removed at positions corresponding to multiple participant images in the virtual meeting scene.
[0139] For example, taking multiple participants in a network video conference, users A through D, as an example, the video stream sent by the server includes the face images of users A through D, as well as the corresponding mask image for each face image. After receiving the video stream, the terminal calls the decoder to perform decoding processing, obtaining the face images of users A through D, and the corresponding mask image for each face image. Then, the terminal can perform the following processing on each face image: masking the face image based on the mask image corresponding to the face image, obtaining a face image with the background removed. For example, taking user A's face image as an example, the mask image corresponding to user A's face image can be used to mask user A's face image, obtaining user A's face image with the background removed. Finally, the terminal can display the face images of users A through D with the background removed at the positions corresponding to users A through D in the virtual meeting scene. In this way, by masking the face images and removing the background of the face images, the meeting scene displayed on the human-computer interaction interface is more in line with the real meeting scene, improving the user experience.
[0140] In some embodiments, in order to reduce the bandwidth required for transmitting video bitstreams, the range of pixel values in the mask image may be smaller than the range of pixel values in the object image (i.e., the mask image is compressed). Before masking the object image based on the mask image corresponding to the object image, the following processing may also be performed: mapping the mask image so that the range of pixel values in the mapped mask image is consistent with the range of pixel values in the object image.
[0141] For example, assuming the pixel values of the decoded mask image range from [0, 128], while the pixel values of the participant image range from [0, 256], before masking the participant image using the mask image, the mask image first needs to be mapped to expand the pixel value range to [0, 256]. Then, the participant image is masked based on the mapped mask image to obtain a participant image with the background removed.
[0142] In other embodiments, the video stream may also include images of a virtual meeting scene. This can be achieved by displaying the virtual meeting scene based on the video stream and displaying multiple participant images at positions corresponding to the participants in the virtual meeting scene: The video stream is decoded to obtain video pixel data of video frames; the decoded video pixel data is then rendered to display the virtual meeting scene image in the human-computer interaction interface, and multiple participant images are displayed at positions corresponding to the participants in the virtual meeting scene image. Since the virtual meeting scene image is carried within the video stream, the terminal does not need to obtain an adapted virtual meeting scene image based on the topic of the network video conference, significantly reducing the local computational load on the terminal and thus effectively reducing power consumption.
[0143] In some embodiments, following the above examples, the display of multiple participant images at positions corresponding to multiple participants in the image of the virtual meeting scene can be achieved in the following way: Image segmentation is performed on each participant image to obtain a mask image corresponding to each participant image; based on each mask image, the corresponding participant image is masked to obtain multiple participant images with backgrounds removed; these multiple participant images with backgrounds removed are displayed at positions corresponding to multiple participants in the image of the virtual meeting scene. Thus, the participant images displayed in the virtual meeting scene have their backgrounds removed, which better reflects the real meeting scene and improves the user's visual experience.
[0144] For example, each participant image can be segmented to obtain a mask image corresponding to each participant image in the following way: Perform the following processing for each participant image: Call the image segmentation model based on the participant image to identify the participants in the participant image, take the area outside the participants as the background, and generate a mask image corresponding to the background. The image segmentation model is trained based on the sample image and the objects labeled in the sample image.
[0145] It should be noted that the image segmentation model mentioned above can be a neural network model (such as a convolutional neural network, a deep convolutional neural network, a fully connected neural network, etc.), a decision tree model, a gradient boosting tree, a multilayer perceptron, and a support vector machine, etc. The embodiments of this application do not specifically limit the type of image segmentation model.
[0146] The network video conferencing processing method provided in this application embodiment can simulate a real meeting scene by displaying a virtual meeting scene and displaying multiple participant images at positions corresponding to multiple participants in the virtual meeting scene, giving users an immersive meeting experience and achieving a good immersive experience during the network video conferencing process.
[0147] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.
[0148] This application provides a method for processing online video conferencing, offering a variety of engaging virtual meeting scenarios that enhance the enjoyment of online video conferencing. In virtual meeting scenarios that simulate real-world meetings, users can experience an immersive meeting atmosphere, thereby increasing their initiative in using online video conferencing.
[0149] Specifically, this application embodiment can aggregate the images of users holding different terminals into the same virtual meeting scene. The aggregated image, after rendering, is displayed on each user's respective terminal, creating an atmosphere similar to a real-world meeting. This narrows the gap between online video conferencing and in-person meetings, enhancing the overall experience of online video conferencing. Furthermore, the online video conferencing processing method provided in this application embodiment can support multiple (e.g., 100) terminals connecting simultaneously. Through a system design that separates the terminal from the online conferencing server, the terminal is only responsible for data collection and display, significantly reducing the computational load on the terminal itself and thus effectively lowering power consumption.
[0150] The following is a detailed description of the network video conferencing processing method provided in the embodiments of this application.
[0151] For example, see Figure 7 , Figure 7 This is a schematic diagram of the overall framework for the interaction between multiple terminals and the network conferencing server provided in the embodiments of this application, such as... Figure 7 As shown, each terminal (e.g., a mobile phone) collects data and reports the collected data to the web conferencing server, so that the web conferencing server processes the data collected by each terminal, merges and arranges the virtual meeting scene, and returns the processed data to each terminal for display.
[0152] For example, see Figure 8 , Figure 8 This is a schematic diagram illustrating the interaction process between a single terminal and a web conferencing server, as provided in an embodiment of this application. Figure 8As shown, the terminal uses the camera to capture images and segment faces of users (i.e., meeting participants), outputs the face image (i.e., the source image) and the segmented face mask image to the web conferencing server. The web conferencing server performs positional arrangement and image synthesis based on the source image and face mask image input by each terminal, and finally returns the synthesized meeting screen to the terminal for display.
[0153] The process of the terminal sending the H.264 bitstream obtained by encoding the source image and the human face mask image to the server is explained below.
[0154] For example, see Figure 9 , Figure 9 This is a schematic diagram illustrating a scenario of interaction between a single web conferencing client and a web conferencing server, as provided in an embodiment of this application. Figure 9 As shown, after the terminal (e.g., a mobile phone) captures a human image using its camera, it transmits the captured image to the web conferencing app (in...). Figure 9 In web conferencing (abbreviated as web conferencing), after receiving a human image, the web conferencing app calls the feedforward neural network inference engine (XNN) to process it with artificial intelligence algorithms, outputting a source image and a human image mask of a specified size. Then, the pre-encoding module performs range mapping on the human image mask, compressing the range of pixel values in the human image mask, so that less bitstream can be used for encoding in the subsequent encoding process (i.e., compression).
[0155] For example, see Figure 10 , Figure 10 This is a schematic diagram of the precoding process for a human face mask image provided in an embodiment of this application, such as... Figure 10 As shown, assuming the original pixel value range of the portrait mask output from XNN is [0, 255], the pixel value range of the portrait mask after normalization will become [0, 1]. Multiplying this by 128 and rounding down, it is mapped to the value range of [0, 128], thus obtaining the pre-encoded portrait mask. In this way, by compressing the pixel value range of the portrait mask from [0, 255] to [0, 128], less bitstream can be used to encode the portrait mask in the future, saving system resources.
[0156] See also Figure 9 After receiving the source image and the pre-encoded human face mask image, the encoder in the web conferencing app performs H.264 encoding on them to obtain the corresponding H.264 bitstream. Finally, the web conferencing app calls the sending module to send the encoded H.264 bitstream to the web conferencing server.
[0157] It should be noted that in practical applications, other video coding algorithms, such as H.265, H.266, and Advanced Video Coding (AVC), can also be used. This application does not specifically limit these algorithms.
[0158] The following section continues to explain the process by which the terminal receives the H.264 bitstream obtained by encoding the rendered image sent by the web conferencing server.
[0159] See also Figure 9 The web conferencing app uses a receiving module to receive the H.264 bitstream encoded from the rendered image from the web conferencing server. Then, it calls a decoder to decode the received H.264 bitstream to obtain the image rendered by the server. Finally, the web conferencing app displays the rendered image on the terminal screen.
[0160] The following describes the network video conferencing processing method provided in the embodiments of this application from the server side.
[0161] For example, see Figure 11 , Figure 11 This is a schematic diagram illustrating a scenario where multiple web conferencing clients interact with a web conferencing server, as provided in the embodiments of this application. Figure 11 As shown, the web conferencing server receives data from multiple web conferencing apps (in...) through the receiving module. Figure 11 In Chinese, this is often referred to as a web conferencing session. For example, web conferencing sessions 1 through n send H.264 streams (which could also be H.265, H.266, etc., depending on the video encoding algorithm used). The web conferencing server then calls a decoder to decode the H.264 streams sent by multiple web conferencing apps, obtaining multiple corresponding source images (e.g., source...). Figure 1 The web conferencing server can further map the pixel values of the multiple decoded image n and the image mask. For the multiple image masks obtained from decoding, the web conferencing server can also call the post-decoding module to perform further pixel value mapping on the multiple image masks, resulting in multiple mapped image masks (e.g., image mask). Figure 1 Let n be the image mask, where each source image corresponds to an image mask, for example, source n = n. Figure 1 Corresponding portrait mask Figure 1 (Source image n corresponds to the human face mask image n).
[0162] For example, see Figure 12 , Figure 12 This is a schematic diagram of the post-decoding process for the decoded portrait mask image provided in an embodiment of this application, such as... Figure 12 As shown, assuming the pixel values of the decoded portrait mask range from [0, 128], the following arithmetic operations are performed during the mask mapping:
[0163] Mask image pixel value / 128.0 * 255
[0164] This remaps the pixel values of the decoded portrait mask to the range of [0, 255].
[0165] In addition, the web conferencing app used by the meeting initiator or host can also display a background image selection interface, allowing the host to choose the desired background image from multiple candidate images. After the host selects a background image, the host's web conferencing app sends a corresponding notification message to the web conferencing server to inform it of the selected background image. Subsequently, the web conferencing server calls the layerout server to modify the layout of the participants' images based on the input information and adjusts their "seats" in the background image according to the order in which they join the meeting. After the adjustments are completed, the rendered image is output.
[0166] The process of joining a web video conference will be explained below.
[0167] For example, see Figure 13 , Figure 13 This is a flowchart illustrating the network video conferencing processing method provided in an embodiment of this application, as shown below. Figure 13 As shown, after the web conferencing app sends a membership request, the web conferencing server sends the source image and the person's mask image so that the screen layout server can check whether the input source image and the person's mask image are normal. If the screen layout server determines that the input source image and the person's mask image are abnormal, it will return an error code. If the screen layout server determines that the input source image and the person's mask image are normal, it will assign the coordinate position of the person's image in the background image according to the order of membership. Then, the screen layout server will merge the source image and the person's mask image into the specified position in the background image (that is, display the person's image with the background removed at the specified position). Finally, the screen layout server outputs the final merged image.
[0168] The following explains the process for exiting a web video conference.
[0169] For example, see Figure 14 , Figure 14 This is a flowchart illustrating the network video conferencing processing method provided in an embodiment of this application, as shown below. Figure 14 As shown, when a web conferencing app sends a request to end the meeting, or when the screen layout server detects an abnormal connection of a web conferencing app, it will call the algorithm to remove the image of the person in the corresponding position in the background image. At the same time, it will call the resource module to reclaim the resources in that position, and finally output the final image of the person whose image has been removed.
[0170] Finally, the web conferencing server can encode the rendered image and distribute the encoded H.264 stream to various web conferencing apps through the sending module.
[0171] For example, see Figure 15 , Figure 15 This is a schematic diagram illustrating an application scenario of the network video conferencing processing method provided in the embodiments of this application, such as... Figure 15 As shown, multiple portraits 1502 with their backgrounds removed are displayed in background image 1501 (e.g., a virtual classroom scene). The position sequence of the multiple portraits 1502 in background image 1501 is determined according to the order of joining the meeting. For example, the leftmost portrait 1502 is the first to join the online video conference, while the rightmost portrait 1502 is the last to join the online video conference.
[0172] For example, see Figure 16 , Figure 16 This is a schematic diagram illustrating an application scenario of the network video conferencing processing method provided in the embodiments of this application, such as... Figure 16 As shown, multiple portraits 1602 with their backgrounds removed are displayed in background image 1601 (e.g., a virtual symposium scene). The position sequence of the multiple portraits 1602 in background image 1601 is determined according to the order of joining the meeting. For example, the leftmost portrait 1602 is the first to join the online video conference, while the rightmost portrait 1602 is the last to join the online video conference.
[0173] The online video conferencing processing method provided in this application provides a variety of interesting virtual meeting scenarios, which can significantly enhance the fun of online video conferencing. In the virtual scenario that simulates a real-life meeting, it gives users an immersive meeting experience and increases users' initiative in using online video conferencing.
[0174] The following description continues to illustrate the exemplary structure of the network video conferencing processing device 555 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 3 As shown, the software modules stored in the network video conferencing processing device 555 in the memory 550 may include: a receiving module 5551 and a display module 5552.
[0175] The receiving module 5551 is used to receive the video stream of the network video conference, wherein the video stream includes real-time images of multiple participants, and the multiple participants images are obtained by separately capturing images of multiple participants in the network video conference; the display module 5552 is used to display the virtual conference scene, and display the multiple participants images at the positions corresponding to the multiple participants in the virtual conference scene.
[0176] In some embodiments, the network video conferencing processing device 555 further includes an allocation module 5553, which is used to allocate corresponding positions to multiple participants in a virtual meeting scene according to the order in which they join the network video conference, wherein the position order of the multiple participants in the virtual meeting scene corresponds to their order of joining; and a display module 5552, which is used to display images of multiple participants according to the corresponding positions allocated to the multiple participants.
[0177] In some embodiments, the allocation module 5553 is further configured to allocate corresponding positions to multiple participants in a virtual meeting scenario according to the speaking order of the multiple participants in the online video conference, wherein the position order of the multiple participants in the virtual meeting scenario corresponds to the speaking order; the display module 5552 is further configured to display images of multiple participants according to the corresponding positions allocated to the multiple participants respectively.
[0178] In some embodiments, the allocation module 5553 is further configured to allocate corresponding positions to the multiple participants in the virtual meeting scenario according to the sorting of the identity information of the multiple participants in the online video conference, wherein the sorting of the positions of the multiple participants in the virtual meeting scenario corresponds to the sorting of the identity information; the display module 5552 is further configured to display the images of the multiple participants according to the corresponding positions allocated to the multiple participants respectively.
[0179] In some embodiments, the display module 5552 is further configured to display a virtual meeting scene adapted to the theme of the online video conference according to the video bitstream; the online video conference processing device 555 further includes an update module 5554 configured to update the virtual meeting scene to adapt to the changed theme when it is determined from the video bitstream that the theme of the online video conference has changed.
[0180] In some embodiments, the network video conferencing processing device 555 further includes a decoding module 5555 for decoding the video stream to obtain multiple video frames; the network video conferencing processing device 555 also includes a topic recognition module 5556 for calling a topic recognition model to perform topic recognition processing on the multiple video frames to obtain topic recognition results for each video frame, and determining the topic recognition result with the highest repetition frequency as the topic of the network video conference; the display module 5552 is also used to display a virtual meeting scene adapted to the topic of the network video conference.
[0181] In some embodiments, the network video conferencing processing device 555 further includes a moving module 5557, which is used to move the image of the target participant from its original position to a specific position in the virtual meeting scene when it is identified that there is a target participant currently speaking among multiple participants, wherein the salience of the specific position is greater than that of the original position; the moving module 5557 is also used to move the image of the target participant from the specific position to the original position when it is identified that the target participant has finished speaking.
[0182] In some embodiments, the video stream further includes multiple mask images corresponding one-to-one with multiple participant images, wherein the multiple mask images are obtained by object recognition of the multiple participant images respectively; the network video conferencing processing device 555 further includes a mask module 5558, which is used to perform the following processing for each participant image: masking the participant image based on the mask image corresponding to the participant image to obtain a participant image with the background removed; the display module 5552 is further used to display the multiple participant images with the background removed at the positions corresponding to the multiple participants in the virtual meeting scene respectively.
[0183] In some embodiments, the range of pixel values in the mask image is smaller than the range of pixel values in the image of the participant; the network video conferencing processing device 555 further includes a mapping module 5559 for mapping the mask image so that the range of pixel values in the mask image is consistent with the range of pixel values in the image of the participant.
[0184] In some embodiments, the receiving module 5551 is further configured to receive the video stream of the network video conference generated by the server in the following manner: removing any participant from the network video conference when any participant leaves the network video conference or when the connection fails; receiving the participant images of the remaining participants in the network video conference, and generating the video stream of the network video conference based on the participant images of the remaining participants.
[0185] In some embodiments, the positions corresponding to the multiple participants are selected from the unoccupied positions in the virtual meeting scene; after any participant is removed from the network video conference, the update module 5554 is further used to update the position corresponding to any participant in the virtual meeting scene from the occupied state to the unoccupied state.
[0186] In some embodiments, the video stream further includes an image of a virtual meeting scene; the decoding module 5555 is further configured to decode the video stream to obtain video pixel data; the network video conferencing processing device 555 further includes a rendering module 55510, configured to perform rendering processing based on the decoded video pixel data to display the image of the virtual meeting scene in the human-computer interaction interface, and to display multiple images of participants at positions corresponding to multiple participants in the image of the virtual meeting scene.
[0187] In some embodiments, the network video conferencing processing device 555 further includes an image segmentation module 55511, used to perform image segmentation processing on each participant image to obtain a mask image corresponding to the participant image; a masking module 5558, used to perform masking processing on the corresponding participant image based on each mask image to obtain multiple participant images with background removed; and a display module 5552, used to display multiple participant images with background removed at positions corresponding to the multiple participants in the image of the virtual meeting scene.
[0188] In some embodiments, the image segmentation module 55511 is further configured to perform the following processing for each participant image: calling an image segmentation model based on the participant image to identify the participants in the participant image, using the area outside the participants as the background, and generating a mask image corresponding to the background; wherein, the image segmentation model is trained based on sample images and the objects labeled in the sample images.
[0189] In some embodiments, the receiving module 5551 is further configured to receive the video stream of the network video conference generated by the server in the following manner: acquiring multiple participant images obtained by image acquisition of multiple participants in the network video conference, and a mask image corresponding to each participant image; performing masking processing on the corresponding participant image based on each mask image to obtain multiple participant images with background removed; acquiring an image of a virtual meeting scene adapted to the theme of the network video conference; filling the positions corresponding to the multiple participants in the image of the virtual meeting scene with multiple participant images with background removed to obtain a merged image; and performing encoding processing on the merged images corresponding to different times to obtain the video stream of the network video conference.
[0190] It should be noted that the description of the device in the embodiments of this application is similar to the implementation of the network video conferencing processing method described above, and has similar beneficial effects, therefore it will not be repeated. For any technical details not covered in the network video conferencing processing device provided in the embodiments of this application, please refer to... Figure 4 , Figures 5A to 5C The meaning is understood in accordance with the description of any of the accompanying drawings.
[0191] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the network video conferencing processing method described in this application.
[0192] This application provides a computer-readable storage medium storing executable instructions. When these executable instructions are executed by a processor, they cause the processor to perform the method provided in this application, for example... Figure 4 ,or Figures 5A to 5C The network video conferencing processing method is illustrated in any of the attached figures.
[0193] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0194] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0195] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).
[0196] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.
[0197] In summary, the embodiments of this application, by displaying a virtual meeting scene and displaying multiple participant images at positions corresponding to multiple participants in the virtual meeting scene, can simulate a real meeting scene, giving users an immersive meeting experience, thereby enhancing users' initiative and enjoyment in using online video conferencing.
[0198] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A method for processing network video conferencing, characterized in that, The method includes: Receive video stream from a network video conference, wherein the video stream includes real-time images of multiple participants, and the multiple images of participants are obtained by separately capturing images of multiple participants in the network video conference; The virtual meeting scene is displayed, wherein the virtual meeting scene changes with the topic of the online video conference. The topic of the online video conference is determined by the following method: decoding the video stream to obtain multiple video frames; calling a topic recognition model to perform topic recognition processing on the multiple video frames to obtain the topic recognition result of each video frame; and determining the topic recognition result with the highest repetition frequency as the topic of the online video conference. In the virtual meeting scene, the images of the multiple participants, with their backgrounds removed, are displayed at their respective corresponding positions. These positions are selected from unoccupied locations within the virtual meeting scene and are determined based on any of the following methods: the order in which the participants joined the online video conference, the order in which they spoke in the online video conference, or the order in which their identity information was sorted within the online video conference. When any of the multiple participants leaves the online video conference or the connection fails, the participant will be removed from the online video conference, and the position of the participant in the virtual meeting scenario will be updated from occupied to unoccupied. When a target participant is identified among the multiple participants who are currently speaking, the participant image corresponding to the target participant is moved from its original position to a specific position in the virtual meeting scene, wherein the significance of the specific position is greater than that of the original position. When the target participant finishes speaking, the participant image corresponding to the target participant is moved from the specific position to the originally assigned position.
2. The method according to claim 1, characterized in that, The display of background-removed images of the multiple participants at their respective locations within the virtual meeting scene includes: Based on the order in which the multiple participants join the online video conference, corresponding positions are assigned to each of the multiple participants in the virtual conference scene, wherein the order of the positions of the multiple participants in the virtual conference scene corresponds to the order in which they join. Based on the corresponding positions assigned to the multiple participants, images of the multiple participants with the background removed are displayed.
3. The method according to claim 1, characterized in that, The display of background-removed images of the multiple participants at their respective locations within the virtual meeting scene includes: Based on the speaking order of the multiple participants in the online video conference, corresponding positions are assigned to each of the multiple participants in the virtual conference scene, wherein the order of the positions of the multiple participants in the virtual conference scene corresponds to the speaking order; Based on the corresponding positions assigned to the multiple participants, images of the multiple participants with the background removed are displayed.
4. The method according to claim 1, characterized in that, The display of background-removed images of the multiple participants at their respective locations within the virtual meeting scene includes: Based on the sorting of the identity information of the multiple participants in the online video conference, corresponding positions are assigned to the multiple participants in the virtual conference scene, wherein the sorting of the positions of the multiple participants in the virtual conference scene corresponds to the sorting of their identity information; Based on the corresponding positions assigned to the multiple participants, images of the multiple participants with the background removed are displayed.
5. The method according to claim 1, characterized in that, The virtual meeting scene display includes: The video stream displays a virtual meeting scene that matches the theme of the online video conference. When it is determined from the video stream that the topic of the online video conference has changed, the virtual conference scene is updated to adapt to the changed topic.
6. The method according to claim 1, characterized in that, The video stream also includes multiple mask images that correspond one-to-one with the multiple participant images. The multiple mask images are obtained by performing object recognition on the multiple participant images respectively. The display of background-removed images of the multiple participants at their respective locations within the virtual meeting scene includes: For each of the participant images, the following processing is performed: The image of the participant is masked based on the mask image corresponding to the image of the participant, and the background of the participant image is removed. In the virtual meeting scene, images of the multiple participants with their backgrounds removed are displayed at the positions corresponding to the multiple participants respectively.
7. The method according to claim 6, characterized in that, The range of pixel values in the mask image is smaller than the range of pixel values in the image of the participating object; Before masking the participant image based on the mask image corresponding to the participant image, the method further includes: The mask image is mapped so that the range of pixel values in the mask image is consistent with the range of pixel values in the image of the participating object.
8. The method according to claim 1, characterized in that, The method of receiving the video stream from the network video conference includes: The receiving server generates the video stream of the network video conference in the following manner: If any participant among the plurality of participants leaves the online video conference or if a connection error occurs, that participant will be removed from the online video conference. Receive the participant images of the remaining participants in the online video conference, and generate the video stream of the online video conference based on the participant images of the remaining participants.
9. The method according to claim 1, characterized in that, The video stream also includes images of the virtual meeting scene; The virtual meeting scene display includes: The video stream is decoded to obtain video pixel data; The video pixel data obtained from decoding is rendered to display an image of the virtual meeting scene in the human-computer interaction interface; The display of background-removed images of the multiple participants at their respective locations within the virtual meeting scene includes: In the image of the virtual meeting scene, the images of the multiple participants, with the background removed, are displayed at the positions corresponding to the multiple participants respectively.
10. The method according to claim 9, characterized in that, Displaying images of the multiple participants (with background removed) at their respective positions in the image of the virtual meeting scene includes: Perform image segmentation processing on each of the participating object images to obtain the mask image corresponding to the participating object image; Based on each of the mask images, the corresponding participant images are masked to obtain multiple participant images with the background removed. The images of the participants, with the background removed, are displayed at positions corresponding to the respective participants in the image of the virtual meeting scene.
11. The method according to claim 10, characterized in that, The step of performing image segmentation processing on each of the participant images to obtain the mask image corresponding to the participant image includes: For each of the participant images, the following processing is performed: Based on the image of the participants, an image segmentation model is invoked to identify the participants in the image of the participants, and the area outside the participants is used as the background to generate a mask image corresponding to the background. The image segmentation model is trained based on sample images and the objects labeled in the sample images.
12. The method according to claim 1, characterized in that, The method of receiving the video stream from the network video conference includes: The receiving server generates the video stream of the network video conference in the following manner: The system acquires multiple participant images obtained by capturing images of multiple participants in the online video conference, as well as a mask image corresponding to each participant image. Based on each of the mask images, the corresponding participant images are masked to obtain multiple participant images with the background removed. Acquire images of a virtual meeting scene that are adapted to the theme of the online video conference; In the image of the virtual meeting scene, the positions corresponding to the multiple participants are filled with the images of the multiple participants whose backgrounds have been removed, resulting in a merged image; The merged images corresponding to different times are encoded to obtain the video stream of the network video conference.
13. A network video conferencing processing device, characterized in that, The device includes: The receiving module is used to receive the video stream of the network video conference, wherein the video stream includes real-time images of multiple participants, and the multiple participants images are obtained by separately capturing images of multiple participants in the network video conference; The display module is used to display a virtual meeting scene, wherein the virtual meeting scene changes with the topic of the online video conference. The topic of the online video conference is determined by: decoding the video stream to obtain multiple video frames; calling a topic recognition model to perform topic recognition processing on the multiple video frames to obtain the topic recognition results of each video frame; and determining the topic recognition result with the highest repetition frequency as the topic of the online video conference. The display module is further configured to display images of the multiple participants (with background removed) at positions corresponding to the multiple participants in the virtual meeting scene, wherein the positions corresponding to the multiple participants are selected from unoccupied positions in the virtual meeting scene, and the positions corresponding to the multiple participants in the virtual meeting scene are determined according to any of the following methods: the order in which the multiple participants joined the online video conference, the speaking order of the multiple participants in the online video conference, or the sorting of the identity information of the multiple participants in the online video conference; The update module is used to remove any participant from the online video conference or update the position of any participant in the virtual meeting scene from occupied to unoccupied when any participant leaves the online video conference or the connection is abnormal. The moving module is used to move the image of the target participant from its original position to a specific position in the virtual meeting scene when it is identified that there is a target participant currently speaking among the multiple participants. The specific position is more significant than the original position. The moving module is also used to move the image of the target participant from the specific position to the originally assigned position when it is recognized that the target participant has finished speaking.
14. The apparatus according to claim 13, characterized in that, The receiving module is further configured to receive the video stream of the network video conference generated by the server in the following ways: acquiring multiple participant images obtained by image acquisition of multiple participants in the network video conference, and a mask image corresponding to each participant image; performing masking processing on the corresponding participant image based on each mask image to obtain multiple participant images with background removed; and acquiring an image of a virtual meeting scene adapted to the theme of the network video conference. In the image of the virtual meeting scene, the positions corresponding to the multiple participants are filled with the multiple participant images with the background removed, to obtain a merged image; the merged image corresponding to different times is encoded to obtain the video stream of the network video conference.
15. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the network video conferencing processing method according to any one of claims 1-12.
16. A computer-readable storage medium having executable instructions stored thereon, characterized in that, When the executable instructions are executed by the processor, they implement the network video conferencing processing method according to any one of claims 1-12.
17. A computer program product, characterized in that, The computer program product includes computer-executable instructions that, when executed by a processor, implement the network video conferencing processing method according to any one of claims 1-12.
Citation Information
Patent Citations
Method for replacing background, method for synthesizing virtual scene, as well as relevant system and equipment
CN101753851A
Meeting place creating method and system of video conference
CN104349111A
Image processing method and device
CN107529096A
Virtual scene generation method and device, electronic equipment and storage medium
CN110427227A
Method and apparatus for realizing virtual conference
KR1020180062045A